AI for Research and Data Organization
Devin Acheampong is a policy analyst at a state Department of Health. His director wants a memo by Friday: what are other states doing about rural ambulance "deserts," and what would it cost us to copy the best idea? Devin has 41 PDF reports, three legislative databases, a folder of news clippings, and 11,000 rows of county-level response-time data in a spreadsheet that someone exported badly. In the old world this was two weeks of reading and re-typing. Devin has four days. He turned to his agency's approved AI assistant and discovered it could either save his week or quietly wreck his credibility, depending entirely on how he used it. This lesson is the difference.
Why Evidence Access Is Uneven, and Why That Matters
Policy work requires research. You need to understand what other jurisdictions are doing, what the research says about a problem, and what approaches have been tried before. Historically that meant weeks of literature review, calls to other agencies, and database searching. Good policy is based on evidence, but access to evidence has never been evenly distributed. An analyst in a well-funded agency might spend weeks reviewing the literature. An analyst in a small county office, carrying several programs at once, might not have an afternoon.
This is the strongest honest case for AI in research work: it narrows that gap. It can help you quickly synthesize what is known about a problem, surface approaches you had not heard of, and point you toward evidence you would otherwise have missed. Better research leads to better policy, and better policy leads to better outcomes for citizens. What follows is how to capture that benefit without importing the failure mode that comes attached to it.
Research Is Not the Same as Retrieval
The single biggest mistake government analysts make with AI is treating it like a search engine. A search engine retrieves documents that exist. A large language model, software trained to predict plausible text, generates sentences that sound like answers, whether or not the underlying facts are real. Ask "which states passed rural ambulance funding laws in 2024?" with no source attached, and you may get five confident answers, two of them invented, including a bill number that does not exist.
So Devin reframed the whole job. He stopped asking the AI to know things and started asking it to work on things he gave it. Every productive use below shares that DNA: the analyst supplies the source, the AI organizes, compares, and surfaces, but it does not originate facts. Hold that sentence in mind through the next two sections, because the difference between the safe pattern and the dangerous one is often a single line in your prompt.
Four Research Jobs the AI Does Well
1. Categorizing a pile of documents. Devin pasted the executive summaries of his 41 reports, a few at a time, and asked the AI to tag each one: state, intervention type (subsidy, regional consolidation, telemedicine, volunteer EMS), and whether it reported cost figures. In twenty minutes he had a structured table that would have taken a day to build by hand. He spot-checked ten rows against the originals and all ten were right, which is genuine evidence that the tagging was working. It is not proof that the other rows were right, and extraction errors do happen even when the model is reading your own text, so the checking continued as he went.
2. Comparing and contrasting. "Here are three states' programs (pasted below). Build a comparison table: funding source, dollars per county, first-year results, and the main criticism each faced." The model is genuinely good at this when the raw material is in front of it.
3. Finding patterns and themes in unstructured text. Devin dumped 60 constituent comments from a public hearing transcript and asked for the recurring concerns, ranked by frequency, with a representative quote for each. Three themes emerged that he had not noticed reading them one by one.
4. Drafting the scaffolding. "Outline a five-section policy memo comparing these four approaches for a non-technical director." The AI built the skeleton; Devin filled it with verified facts.
The rule underneath all four: ask the AI to organize what you give it, never to remember what you did not.
Three Research Use Cases, With Their Trapdoors
The quick literature review. You are developing a policy on a topic and need to know what has been tried. The prompt looks like this: "What are the leading research-based approaches to [topic]? Include citations so I can find the original research." Then you go and read the citations, and you may find sources you did not know existed. The trapdoor is in the word "citations." A model asked for references from its own memory will produce some that are real and some that are not, formatted identically. Treat the list as a set of leads to check, not a bibliography. Anything you cannot locate in the actual literature does not go in your memo, no matter how plausible the title sounds.
Finding what other jurisdictions are doing. The question is how other cities, states or countries have approached a problem. Ask: "What approaches have [city/state/country] X, Y, Z taken to address [problem]? Where can I find more information?" The AI gives you an overview of different approaches; you research the most promising ones in detail; you adapt what fits your context. Done this way, benchmarking gets much faster and you see more examples than you could assemble alone. Done carelessly, it produces a memo full of programs that do not exist in the jurisdictions named. The overview is a search shortcut, and every named program still needs a primary source before it appears in your draft.
Identifying trends in your own data. This one is safer, because the material comes from you. Provide the data, or a de-identified summary of it, and ask: "What trends do you see in this data? What's surprising? What warrants further investigation?" The AI will flag patterns you might not see manually, and you then investigate what it flagged. Note the verb. It flags; you investigate. A pattern the model notices is a hypothesis with an origin, not a finding, and the next section explains why that distinction is the whole game once numbers are involved.
Organizing: Categorizing and Extracting at Volume
Beyond research, the second family of uses is organizing information you already hold. Categorizing comes first. You have hundreds of records, complaints, applications or inquiries, that need sorting. Define the categories clearly yourself rather than asking the AI to invent them. Give the AI a few examples of each category so it learns your boundaries, not its own. Ask it to categorize the rest. Spot-check its work. Then refine your instructions based on the errors you find, and run it again. That refinement loop is not optional polish; it is how the categories get sharp enough to trust.
Extraction is the second. You have unstructured text such as forms, letters or field notes, and you need specific fields out of it. Say you have 500 incident reports and you need date, location, type of incident, and severity from each. Provide sample reports, then ask: "For each report, extract: date, location, incident type, severity. Format as a table." Review what comes back, then use the extracted data for analysis. Keep the field list explicit and identical every time you run it, because a field you name loosely is a field the model will fill loosely.
Extraction is where unverified output does the most quiet damage. Nothing downstream will tell you that a severity column was misread; the analysis simply comes out wrong and looks fine. Review a sample of extracted rows against the source documents before anything is built on top of them, and review a fresh sample whenever the input format changes, because a new report template can break an extraction that has worked for months.
Organizing the Messy Spreadsheet
The 11,000-row response-time export was the kind of mess every analyst knows: inconsistent county names ("St. Clair" versus "Saint Clair Co."), times in three formats, blank cells, and a stray header row in the middle. AI assistants help here in two distinct ways, and confusing them is dangerous.
The safe way: ask the AI to write the cleanup instructions or formula, then you run it and check the result. "Give me a spreadsheet formula that standardizes county names against this reference list" produces a tool Devin can verify. The risky way: pasting all 11,000 rows in and asking the AI to "clean this and give me the averages." The model may silently drop rows, fabricate a plausible average, or mis-parse the times, and you would never see it happen.
Devin's rule for data: the AI may design the method; the numbers must come from the spreadsheet. He had the AI write the deduplication and date-parsing approach, applied it in the actual tool, and computed the averages himself. The AI accelerated the boring setup without touching the math the director would quote in a hearing.
The Verification Discipline
Devin's near-disaster: an early draft of his memo, AI-assisted, claimed "State X reduced average rural response time from 22 to 14 minutes." It was a clean, quotable, dangerous sentence. The AI had blended two real figures from different reports into one false claim. He caught it only because of a habit worth building into your bones: every load-bearing fact gets traced back to a named source before it enters the memo.
Practically, that means each statistic in his draft carried a bracketed source tag, [Report 12, p.4], until final formatting. Anything the AI produced without a traceable source got deleted, not filed under "probably fine." This is not bureaucratic caution. A wrong number in a memo that reaches a legislator is the kind of error that follows an analyst's name for years, and it is the kind an AI produces most fluently.
The same discipline applies to patterns, not just figures. If the model reports that response times worsened in counties that consolidated dispatch, that is a correlation it noticed in your data. Whether the relationship is causal, whether the difference is large enough to mean anything, and whether some third factor explains both, are questions the model has not answered and cannot answer from the table alone. Those answers come from human analysis, and a policy recommendation resting on an uninterrogated pattern is a recommendation resting on nothing.
A Usable Artifact: The Source-Grounding Research Log
Devin keeps a simple log for any AI-assisted research task. It makes the work fast and, just as important, defensible if anyone later asks "where did this come from?"
| Field | What to record | Why it matters |
|---|---|---|
| Task | What you asked the AI to do (categorize, compare, summarize, extract, draft). | Keeps the AI in "organize" mode rather than "originate" mode. |
| Source provided | The exact documents or data you pasted in. | If you gave no source, treat every output as unverified. |
| Output | What the AI returned, saved or screenshotted. | Reproducibility and records compliance. |
| Verified? | Which facts you spot-checked against originals, and the result. | Catches blended or invented figures before they spread. |
| Sensitive data? | Whether anything personal or protected was involved, and the approved tool used. | Privacy obligations apply to data too, not just letters. |
Worked Example: Building a Remote Work Policy
Your agency is developing a remote work policy. You need to understand two things: what other government agencies are doing, and what the research says about the effectiveness and risks of remote work. Here is the whole workflow, with the verification steps kept in rather than skipped.
- Ask the AI: "What remote work policies have federal government agencies implemented? Include links to policies so I can review them directly."
- Review the examples the AI provides, and read the actual policies from a few agencies. Any policy you cannot open and read yourself does not count as an example.
- Ask the AI: "What does research say about effectiveness of remote work for government employees? What are benefits and risks?"
- Review the research findings, and read the original research on the topics that matter most to your agency.
- Synthesize: "Based on this research and examples from other agencies, what remote work policy makes sense for our context?"
- Develop your policy, citing the research and examples you personally opened.
The result is much faster policy development, informed by evidence and by peer examples. Notice that steps two and four are not decoration. They are the steps that convert a plausible list into a defensible one, and they are the first steps to fall out when the deadline tightens. If you have to compress this workflow, compress the number of examples you chase, never the reading you do on the ones you keep.
The Payoff, Honestly Stated
Devin finished his memo Thursday afternoon. The AI had collapsed the categorization, comparison, and outlining from days into hours. It had not written a single fact he did not verify, and it had not touched a number the director would quote. That is the realistic win: not a robot researcher, but a tireless assistant that organizes your material at machine speed while you stay firmly in charge of what is true.
Anti-Patterns to Avoid
- Taking AI research at face value. The AI tells you about research on a topic and you cite it in your policy without reading the original. The risk is that the AI mischaracterized the research, or that the study does not exist. The original may say something quite different, and the citation is what a reviewer will check first.
- Over-relying on AI-identified trends. The AI identifies a trend in your data and you make a major decision on it without human analysis. The risk is that the pattern is a correlation with no causal link, or a difference too small to be meaningful. The model has no way to tell you which it is.
- Using unverified extracted data. You ask the AI to extract fields from records and use the output without checking it. The risk is extraction errors propagating silently into every analysis built on that table.
- Treating a clean spot-check as a clearance. Ten correct rows are evidence about those ten rows. They do not certify the rest of the batch, and they say nothing about the next batch or a changed input format. Sampling is a floor under your confidence, not a ceiling on the error rate.
- Asking the AI to be the source. Any prompt that begins "what are the leading" or "which states have" is asking the model to recall rather than organize. Use those prompts to generate leads if you like, but nothing that arrives that way belongs in a document until you have found it somewhere real.
Practice Prompts
- Take a research question you actually owe someone this month. Write the prompt you would use, then write, before you run it, the list of things you will have to verify afterward.
- Pick a folder of documents you already have. Ask the AI to tag each one on three fields you define, then check a sample of the rows against the originals and record how many were right.
- Take a messy spreadsheet and ask the AI for the cleanup formula rather than the cleaned data. Run the formula yourself and compare the result with what you expected.
- Ask the AI for the leading research-based approaches to a topic you know well. Try to locate each source it names. Count how many you could actually find.
Reflection
- Think of a research question relevant to your work. How would you use AI to help answer it, and what would the next steps be after the AI provided initial findings?
- Do you have data in your role that you would like to understand better? How might AI help you identify trends or patterns, and who would judge whether a pattern means anything?
- When you use AI for research, how would you verify that the information is accurate? Could you reconstruct that verification for an auditor long after the fact?
Glossary
- Literature Review: A comprehensive examination of published research on a topic.
- Benchmarking: Comparing your approach or performance to others.
- Trend: A general direction or tendency in data over time.
- Pattern Recognition: Identifying repeated or related elements in data.
- Data Extraction: Pulling specific information from larger bodies of text or data.
- Large language model: Software trained to predict plausible text. It generates sentences that sound like answers, which is not the same as retrieving documents that exist.
- Blended fact: A false claim the model assembles by fusing two real figures from different sources into one sentence. It reads perfectly and is the hardest error to catch.
Related Lessons
- Your Agency's Approved AI Tools determines which tool you may paste agency data into at all.
- Prompt Engineering Basics covers the prompt structure the research and extraction prompts here depend on.
- Evaluating AI Outputs gives you the systematic pass for checking what comes back.
- AI for Government Tasks: Summarization and Drafting applies the same source-grounding discipline to correspondence and briefing work.
- AI Confidence and Hallucination explains why an invented citation arrives with exactly as much confidence as a real one.
Closing
Research and data organization are practical applications where AI genuinely helps government work go faster and go better. Use these tools. But remain skeptical, verify findings, and do not rely on AI for final answers. Rely on it for starting points and for pattern identification, then do the analytical work that turns a starting point into something you would put your name on. Devin got his memo done in four days because he knew which half of the job the machine could hold.
Key Takeaways
- Organize, don't recall. The AI generates plausible text; ask it to work on sources you provide, never to remember facts on its own.
- Categorizing, comparing, theming, and outlining are the sweet spot. These tasks turn piles of your own material into structure in minutes.
- AI can accelerate research, but human verification is essential. Always read the original if you are citing research, and drop anything you cannot locate.
- Not all patterns are meaningful. The model can flag a trend but cannot tell you whether it is causal or significant. That judgment stays human.
- For data, let the AI design the method, not produce the numbers. Have it write formulas and cleanup steps you run and verify; compute quoted figures yourself.
- Extraction saves time and requires checking. Spot-check extracted fields against the source documents before anything is built on them, and check again when the input format changes.
- Watch for blended facts. The model can fuse two real figures into one false claim that reads perfectly. It is the most dangerous error in this whole lesson.
- Trace every load-bearing fact to a named source, and keep a research log. Task, source, output, verification: fast now, defensible later.
Frequently Asked Questions
Can I ask an AI tool for citations to back up a policy memo? You can ask, and you must then find every one of them in the real literature. Models produce plausible references and real references in the same format. A citation you could not open is not a citation, and it does not go in the memo.
What is the safest way to use AI on a spreadsheet? Ask it to write the cleanup method, formula or parsing approach, then run that yourself in the actual tool and check the result. Pasting thousands of rows in and asking for averages invites silently dropped rows and fabricated numbers.
How many rows should I spot-check after an AI categorization or extraction? The lesson does not set a number, and any number you pick is a floor rather than a guarantee. What matters is that you check against the original documents, that you keep checking after the process feels settled, and that you check afresh when the input format changes.
The AI found a trend in our data. Can I put it in the recommendation? Not on its own. It has surfaced a pattern for you to investigate. Establish whether the relationship is causal, whether the effect is big enough to act on, and what else might explain it, before the pattern carries any weight in a decision.
Is it safe to paste agency data into the tool at all? That depends on the data and on which tool your agency has approved for it. Privacy obligations cover datasets exactly as they cover letters, and a de-identified summary is often enough for the trend question you actually want answered.
Why keep a research log if the work turned out fine? Because "where did this come from?" is asked months later, usually by someone with authority, and usually about the one number that ended up in a press release. The log is how you answer in a minute instead of a week.
Skill.re