The AI Decision Framework: When to Use and When Not To
Priyanka Wojciechowski managed the constituent services unit for a mid-sized city's Department of Human Services, and she had just been handed a budget surplus and a mandate: "find places to use AI." Within a week she had a stack of proposals. A chatbot to answer benefits questions. An automated tool to prioritize housing-assistance applications. A system to summarize the 200-page inspection reports her caseworkers waded through. A model to predict which families were "at risk" so the department could intervene early. Each proposal sounded plausible. Each had an eager sponsor. What Priyanka lacked was not enthusiasm but a way to decide: a repeatable test that would tell her which of these should go forward, which should be redesigned, and which should be politely killed before they damaged the people her department existed to serve.
This lesson gives you that test: a decision framework for when AI is appropriate for a government task and when it is not. The goal is not to maximize AI adoption. The goal is to deploy AI exactly where it helps and to recognize, without embarrassment, the many places where it does not. Sometimes AI is the right tool. Sometimes it is a bad idea. Sometimes it is workable but not worth the implementation cost. In the public sector, "we chose not to automate this" is frequently the correct and defensible answer.
Start With the Right Question
The instinct "find places to use AI" inverts the proper order of inquiry. AI is a means, not an end. The right opening question is never "where can we use AI?" but "what problem are we trying to solve, and is AI the appropriate tool for it?" A tool in search of a problem produces expensive systems that solve nothing and erode public trust when they fail. You will recognize the moment it happens: a process is slow, frustrating or expensive, someone says "we should use AI to fix this," and everyone nods because the idea feels modern and technological.
Priyanka rewrote each proposal to lead with the problem rather than the technology. The chatbot became "constituents wait 40 minutes on hold for routine benefits questions." The risk model became "we want to identify families needing early intervention." Stated as problems, the proposals could be evaluated. Stated as AI projects, they could only be approved or denied on enthusiasm. That rewrite costs an afternoon and it is the highest-leverage thing you can do before any of the analysis below begins.
What a Badly Made AI Decision Costs
Agencies are under pressure to modernize. Budgets are constrained, backlogs are deep, and someone eventually asks whether AI can help. The honest answer is that it sometimes can, and that sometimes the costs and risks outweigh the benefits. The reason it is worth slowing down long enough to tell those cases apart is that a bad choice here is not a neutral experiment. It spends real money, it spends public trust, and in the wrong domain it spends other people's stability.
- Bad AI decisions waste resources. Building, maintaining and ultimately replacing a poorly chosen AI system costs time and money that could have been spent elsewhere.
- Bad AI decisions damage trust. When a system makes wrong decisions that affect citizens, and people learn that the decision to use AI was never seriously examined, confidence in government AI and in government generally declines.
- Bad AI decisions can cause harm. A system that systematically disadvantages certain groups, or that errs in high-stakes situations, does real damage to real people.
- Good AI decisions are invisible. When you use AI where it is appropriate and avoid it where it is not, the benefits compound quietly. Decisions improve, processes get more efficient, citizens get better service, and nobody writes a story about it.
The Decision Gates
The framework runs each candidate use case through a sequence of gates. Walk them in order. A use case must clear every gate to proceed, and failing any single gate is a stop or a redesign, not a point to be averaged against the others. If you hit a "no" or a high risk you cannot mitigate, the answer is probably "do not use AI" or "reconsider this approach." The gates are ordered so that the cheapest questions come first: there is no point costing out an explainability review for a use case that has no definable task.
Gate 1: Is there a clear, well-defined task, and data to learn from?
AI performs well on bounded tasks with a definable notion of "correct" and poorly on open-ended judgment. Well-defined problems look like "identify which benefit applications have high risk of fraud," "predict which neighborhoods have the highest infrastructure maintenance need," or "summarize a long policy document into key points": each has a clear outcome that can be evaluated against data you have. Poorly defined problems look like "improve service delivery," which is too vague, "make hiring better," which is broad and deeply subjective, or "streamline everything," which is not specific enough to build against.
Quality of data belongs in this same gate. AI requires examples to learn from, and if you have only a few hundred of them, or your records are full of errors, the system will not work well no matter how well specified the task is. Priyanka's report-summarization proposal cleared this gate easily. Her "at-risk family" prediction did not: the underlying question was not really predictable, it was a contested judgment about what risk even means, dressed up as a forecast.
Gate 2: Do we have the right to use that data?
Having data is not the same as being allowed to use it. Three sub-questions decide this. Does the data exist at sufficient quality? Is it representative of the population the system will serve, or does it overrepresent some groups and miss others? And do you have legal authority to use it this way, considering the Privacy Act of 1974, any applicable state privacy law, and the original purpose for which the data was collected? Data gathered to administer one program often cannot lawfully be repurposed to train a model for another.
Priyanka's chatbot drew only on published policy documents, which was clean. The risk model would have required combining child-welfare, benefits and police data in ways that raised immediate Privacy Act and ethical concerns. Notice that this gate can fail on its own even when the task is perfectly well defined and the technical performance would have been excellent. Authority is not a technical property, and no amount of model quality substitutes for it.
Gate 3: How sensitive is the data the system needs?
Sensitivity sets the height of the bar before anything else is negotiated. Work out precisely what data the system requires, not what data you happen to hold, then locate that requirement on the ladder below. The more sensitive the data, the higher the bar for proceeding, and the top rung is a stop rather than a hurdle. Answer this gate honestly at the design stage, because a use case that quietly needs personal data will discover that fact later, in front of a privacy office, at a point when a redesign is expensive and a pilot is already running.
| Data the system needs | Posture | What that requires |
|---|---|---|
| Classified information | Stop. Do not do it | Unless you have a classified AI system, which is rare |
| PII or PHI | Proceed with extreme caution | Legal authority, security review, privacy office approval, strong data handling and security measures |
| CUI | Moderate caution | Legal authority, security review, data minimization |
| Unclassified data only | Lower caution | Still maintain good security |
| No personal data at all | Lowest caution | Ordinary diligence |
Gate 4: What is the citizen impact, and who bears it?
This is the gate that most distinguishes government from private-sector decision-making, and it is critical: high-stakes decisions require more caution. High-stakes uses include determining eligibility for vital services such as food, housing, healthcare and income support; criminal justice applications such as who to investigate, predicting dangerousness, and bail decisions; employment decisions; immigration decisions; and anything at all that affects freedom or access to essentials. Low-stakes uses include suggesting which form someone might need, summarizing documents, categorizing incoming inquiries, and finding good times for service appointments.
For high-stakes decisions you need higher accuracy requirements, more extensive testing and validation, more human oversight, more transparency and explainability, and clear fallback and appeal processes. For low-stakes decisions you can move faster with lower accuracy requirements. Plot the use case on two axes, data sensitivity from public information to highly sensitive personal data, and citizen impact from convenience to consequential decisions about benefits, liberty or safety. Where the two axes land together sets the regime.
- Low sensitivity, low impact (answering published-policy questions): AI is appropriate with standard oversight.
- High sensitivity or high impact, but AI only assists a human who decides (summarizing a report a caseworker still reads): appropriate with strong human-in-the-loop controls and explainability.
- High sensitivity and high impact, with AI making or effectively driving the decision (auto-prioritizing who gets scarce housing assistance): the bar is very high, the system is rights-impacting under federal frameworks, and the default leans toward not automating the decision itself.
Gate 5: What error rate is acceptable, and what does each error cost?
Every system makes errors, so the question is never whether errors will happen but what they cost and who absorbs them. Start from the requirement rather than the capability. If your system has to be 99.9% accurate to be useful, you may simply not be able to get there. If you need 85% accuracy and the system measures 90%, that is a workable position, though a measured figure is an estimate of past performance on the cases you tested, not a promise about the cases you have not seen yet.
Then split the error types, because they land on different people. What is the cost of a false positive, where the system incorrectly flags something? What is the cost of a false negative, where the system misses something? And can humans catch the errors the AI makes? For benefits determination, if the system denies someone a benefit they should receive, they go without food or housing: a severe cost demanding high accuracy. For document summarization, if the AI leaves out one detail a human can catch it, so the cost is lower and you can tolerate less accuracy.
Priyanka held on to the distributional half of this question. A chatbot that occasionally gives a wrong answer about office hours is a minor, recoverable inconvenience. A prioritization model that wrongly deprioritizes a family in crisis produces irreversible harm borne by the most vulnerable. When the cost of error is high, irreversible, and concentrated on people with the least power to contest it, the framework demands either far stronger safeguards or a decision not to proceed at all.
Gate 6: Can we explain the decision, and can a citizen contest it?
Government decisions affecting individuals carry due-process obligations. If the system denies someone a benefit, you have to be able to tell them why. "The AI decided you don't qualify" is not acceptable. "The AI identified these factors as indicating you're outside the eligibility range" is better. If the system cannot produce an explanation a caseworker understands and a citizen can challenge, it is unsuitable for any decision affecting rights or benefits, regardless of how accurate it is.
Some architectures are more explainable than others. Decision trees and linear models are transparent; deep neural networks can be opaque. For high-stakes decisions you should prefer the more explainable system even when it is less accurate, over the black-box system even when it is more accurate. That trade runs against engineering instinct, which is exactly why it has to be a stated rule rather than a judgment call made under deadline. Accuracy does not substitute for explainability when someone's benefits are at stake.
Gate 7: Can we mitigate the risks we have found?
By this point you have named specific risks: data sensitivity, accuracy shortfalls, explainability limits. The gate asks whether each one has a mitigation you can actually staff and fund. The standard moves are implementing strong security for sensitive data, requiring human review before high-stakes decisions, testing for bias and fairness, providing transparency and explanation to affected people, building in fallback and appeal processes, and monitoring performance over time. If you cannot mitigate the risks, do not proceed. An unmitigated risk you have merely documented is still an unmitigated risk.
Gate 8: Is there an existing solution that is better?
Maybe you do not need AI. Maybe traditional software is cheaper and more appropriate. Maybe the right answer is to hire more people, or to redesign the process so the work disappears. Before committing, ask what the alternatives are and how their pros and cons compare with AI, on cost, on time to deliver, and on how the failure modes land on citizens. Sometimes AI wins that comparison and sometimes it does not, but a comparison you never ran cannot be defended to an inspector general.
Gate 9: What are the approval and accountability requirements?
Before building, identify what governance the use case triggers. A rights-impacting or safety-impacting system in a federal agency invokes obligations under OMB Memorandum M-24-10: inventory, impact assessment, minimum risk-management practices, and Chief AI Officer review. State and local governments increasingly have parallel requirements. If you cannot name the approval pathway, you are not ready to start, and discovering the pathway after a pilot is running is the most expensive possible order in which to learn it.
Three Use Cases Run Through the Gates
The framework earns its keep on cases that are genuinely arguable, not on the ones that were obviously fine or obviously reckless. Here are three, worked end to end through the gates, with the recommendation each one produces at the bottom. Note how differently they come out despite all three being sensible-sounding proposals that a reasonable manager might have championed, and note that none of them ends in a plain yes or a plain no. Each ends in a scope, a condition, or a substitute.
Should we use AI for hiring?
The problem: the hiring process is slow, the agency gets 1,000 applications per position, and reviewing them all takes months. On the task gate the answer is only "somewhat": years of hiring data exist, but "good candidate" is subjective and hiring success depends on many factors nobody measures well, so the risk is moderate. Impact is high-stakes, since this determines access to government employment. Data sensitivity is moderate, covering educational background, work history and possibly demographic information. Accuracy requirements are very high, because systematically excluding qualified candidates is not an acceptable failure mode.
Explainability is a high risk too: you need to be able to tell candidates why they were not selected. Mitigations exist on paper, requiring human review of every candidate the AI recommends, testing for bias before deployment, and providing explanations to candidates, but the accuracy and review requirements together might make the process not much faster than manual review, which was the entire point. The alternatives are traditional software such as a resume database and keyword search for initial screening followed by human review, hiring a recruiter to do initial screening, or a simpler AI that ranks candidates by keywords rather than a complex model.
The recommendation: do not use a complex AI system for hiring. Use a simple tool to help humans with initial screening, and keep humans in control. This is a case where the framework does not produce a flat no, it produces a scoped-down yes, and the difference matters because the underlying problem, the months of review, is real and still needs solving.
Should we use AI for document summarization?
The problem: the agency receives thousands of citizen letters per year and staff spend days summarizing them for leadership briefings. The task gate clears easily, because summarization is a clear task with many examples available. Impact is low-stakes, since a summary that is incomplete or slightly inaccurate will not cause harm. Data sensitivity is low: some letters contain personal information, but only the summaries are used and the original data is not stored in the AI system. Accuracy needs are moderate; the summary must hit the main points and does not need to be perfect.
Explainability is not critical for a summary. You can explain how the system works, that it looks at key phrases and extracts main points, without explaining every word choice. The mitigation is straightforward: have a human review summaries before they go to leadership, and do not automate the summaries without human review. The alternatives are hiring someone to summarize, using templates to guide summarization, or not summarizing at all and sending full letters to leadership. The recommendation is that AI for document summarization is a good fit here. Use an approved AI system, have humans review the summaries, and deploy.
Should we automate responses to routine inquiries?
The problem: the agency receives thousands of inquiries per year, many of them routine, such as checking the status of an application or asking general questions about eligibility, and answering them consumes significant staff time. The task gate clears: years of inquiry data exist, routine and complex cases can be classified, and example responses are on hand. The impact gate is where the design decision lives. If the AI is providing general information, the use is low-stakes. If it is making determinations about eligibility or benefits, it is high-stakes. The recommendation follows directly: use AI only for information provision, and route determinations to humans.
Data sensitivity is low to moderate, since some inquiries contain personal information. The design response is to build the system so it does not store PII and uses the inquiries only to categorize and route to appropriate responses. Be precise about what that buys you: not storing PII limits retention, it does not mean PII never reaches the system, so the input path still needs the controls Gate 3 requires. Accuracy needs are moderate, because a wrong routing means a human fixes it, which is annoying rather than disastrous, and explainability matters little for routine information.
Mitigations are to have humans review automated responses before they go to citizens and to build an escalation pathway for uncertain cases. Alternatives include hiring more staff, redesigning the inquiry process to be self-service, improving the published FAQ, or not responding to routine inquiries at all. The recommendation: AI for routine inquiry response is appropriate. Use it to augment human staff rather than replace them, have humans review responses, and monitor performance after launch rather than assuming the launch-day behavior persists.
When the Answer Is No
A framework that never says no is not a framework. The honest "do not use AI here" cases include decisions that are fundamentally about contested values rather than prediction; situations where the available data is biased or unrepresentative and cannot be fixed; decisions whose errors are irreversible and fall on the vulnerable; and tasks where a simpler, more transparent tool, a clear rule, a better form, an additional staff member, would serve citizens better. Choosing a non-AI solution is not a failure of ambition. It is often the most accountable decision available.
Priyanka's Dispositions
Run through the gates, Priyanka's four proposals sorted cleanly. The benefits chatbot cleared every gate, bounded task, public data, low impact, explainable, low cost of error, and proceeded with a human-handoff path for anything beyond routine questions. The report summarizer cleared the gates as an assistive tool because a caseworker still read every report and the AI only drafted a summary; it advanced with a mandatory human-review step. The housing prioritization tool was redesigned: rather than ranking applicants, it was scoped to surface missing-document flags so caseworkers could help applicants complete files, removing AI from the allocation decision itself.
The at-risk family prediction was declined. It failed the clear-task, data-rights and cost-of-error gates simultaneously, and the department recorded the decision and its reasoning in a one-page memo so the question would not simply resurface with a new sponsor next quarter. That memo mattered as much as the deployments. Documenting why a use case was rejected is itself a governance artifact: it shows that the agency exercised judgment rather than reflex, and it protects the next analyst from relitigating a settled question.
Anti-Patterns to Avoid
- "AI sounds good, so let's do it." An organization falls in love with the idea and skips careful evaluation, until "AI is innovative" is the only rationale on the page. The risk is that they build a system for a task where it is not appropriate, waste resources, and cause harm.
- Ignoring the sensitivity and stakes combination. An organization commits AI to a high-stakes decision such as benefit determination without seriously working through data sensitivity and accuracy requirements. The risk is a system that does not work well enough, making wrong decisions that affect people's lives.
- Assuming AI will be more accurate than it is. An organization overestimates the system and designs the process as though AI makes final decisions. The risk is that the errors arrive anyway, land on people, and bring liability and lost trust with them.
- Not considering alternatives. An organization fixates on "should we use AI?" and never asks "what is the best solution to this problem?" The risk is a complex, expensive AI build where something simpler would have served better.
- Treating a benchmark number as a guarantee. "We need 85% and it tests at 90%, so we are covered" reads a measurement of past cases as a promise about future ones. A test result is evidence that the system cleared the bar on the data you had; it does not certify performance on the population you are about to point it at. Keep measuring after launch.
- Reading "we do not store PII" as "no PII is involved." A design that discards personal data after routing still receives it, transmits it and processes it. Retention is one control among several, and the sensitivity gate applies to what enters the system, not only to what survives in it.
- Letting human review become a rubber stamp. Human-in-the-loop is a real mitigation only when the reviewer has the time, the information and the standing to overturn the AI. A queue of recommendations approved at a rate nobody could actually read is an automated decision wearing a signature.
Practice Prompts
Take an AI initiative your agency is considering, or one you know about, and answer these in writing. The gaps you cannot answer are the finding.
- What problem is it trying to solve?
- How high-stakes is that problem?
- What data does it need?
- What accuracy does it need to achieve?
- Can failures in the AI system be caught and corrected by humans?
- What are the risks if the system is wrong?
- Can those risks be mitigated?
- What would the alternative, non-AI solution be?
- Is AI actually better than the alternative?
- Would you recommend proceeding with this AI initiative, and why or why not?
Reflection
- Think of a process in your agency that is slow or inefficient. Would AI be a good solution? Apply the decision framework. What is your recommendation?
- Think of an AI system you know about, in your agency or elsewhere. Evaluate it using this framework. Does it seem like a good fit for the task? If you were asked to defend the decision to use AI for that task, what would you say?
- What is the most high-stakes decision your agency makes? Should AI be involved in that decision, and why or why not?
- In your agency, who makes the decision about whether to pursue an AI solution? Is that person familiar with the factors discussed in this framework?
Glossary
- High-stakes decision. A decision that significantly affects a person's rights, freedom, access to essential services, or opportunities.
- False positive. An error where the system incorrectly identifies or predicts something, such as incorrectly flagging someone as high-risk.
- False negative. An error where the system fails to identify something, such as failing to identify someone who actually is high-risk.
- Data sensitivity. The degree of harm that would result from unauthorized disclosure of data.
- Explainability. The ability to understand and explain why an AI system made a particular decision.
- Mitigation. Actions taken to reduce or eliminate a risk.
- Well-defined problem. A problem with a clear goal, measurable outcomes, and available data.
Related Lessons
- Government AI Policy Landscape sets out the policy environment the approval gate points into.
- Data Sensitivity and Classification goes deeper on the ladder used in Gate 3, from unclassified through CUI to classified material.
- PII and AI: The Bright Red Lines covers the personal-data limits that decide Gate 2 and Gate 3 for most citizen-facing use cases.
- Understanding AI Bias explains why unrepresentative training data is a stop condition rather than a tuning problem.
- Reporting AI Concerns in Your Agency is what you use when a system already in production starts failing gates it once cleared.
- Risk Classification: Safety-Impacting vs. Rights-Impacting formalizes the impact judgment made in Gate 4.
Closing
You now have a framework for making good AI decisions, and the only thing that makes a framework worth anything is using it when it is inconvenient. When your agency is considering an AI initiative, walk the gates. Push back if the decision seems premature or insufficiently considered. Ask the hard questions early, when the answer is a redesign, rather than late, when the answer is a public failure. A well-decided no is better than a poorly decided yes, and it is far easier to defend.
That is how you ensure the AI initiatives in your agency are thoughtful, well designed and appropriate to the work. If you cannot clearly articulate why AI is right for a task, that is the finding, and the correct response is to say so out loud in the meeting rather than to hope the question gets answered later by someone with more authority. The framework does not make you the person who blocks things. It makes you the person who can explain, on the record and in a paragraph, why a given decision went the way it did.
Key Takeaways
- Lead with the problem, not the technology. "Where can we use AI?" is the wrong question. "What problem are we solving, and is AI the right tool?" is the right one. Not every problem is an AI problem.
- Run every use case through every gate. Clear task and usable data, lawful authority, data sensitivity, citizen impact, error cost, explainability, mitigation, alternatives, and approval requirements. Failing one gate is a stop, not an average.
- Data sensitivity times citizen impact sets the bar. Classified data is a stop; PII and PHI demand extreme caution; high-sensitivity, high-impact decisions made by AI face the strictest scrutiny and often should not be automated at all.
- Split the error types and ask who absorbs them. False positives and false negatives cost different people different amounts, and a benchmark accuracy figure is evidence about past cases rather than a guarantee about future ones.
- Explainability is non-negotiable for rights-affecting decisions. Due process requires an outcome a citizen can understand and contest, and you should prefer an explainable system even when a black-box one scores higher.
- Always consider alternatives, and be willing to say no. Contested-value decisions, biased data, and irreversible harms concentrated on the vulnerable are reasons to choose a simpler, more transparent solution.
- Document the rejections. A short memo explaining why a use case was declined is a governance artifact that prevents the question from resurfacing on enthusiasm alone.
Frequently Asked Questions
Do I have to run all the gates for a small, low-stakes tool? Run them, but expect most to clear in a sentence each. The gates are cheap when the answers are easy; the ordering exists so that a low-sensitivity, low-impact use case like answering published-policy questions reaches "appropriate with standard oversight" quickly. The cost of the framework is concentrated exactly where the risk is.
What if we can meet the accuracy requirement but not the explainability one? Then the use case fails for any decision that affects rights or benefits. For high-stakes decisions you should prefer more explainable systems even if they are less accurate over black-box systems even if they are more accurate, because "the AI decided you don't qualify" is not something a citizen can contest.
Our data exists and is high quality. Is that enough to clear the data gate? No. Quality is one of three questions. The data also has to be representative of the population the system will serve, and you need legal authority to use it for this purpose, considering the Privacy Act of 1974, applicable state privacy law, and the purpose for which it was originally collected. Data collected to administer one program often cannot lawfully train a model for another.
Is human review always enough of a mitigation? Only when the review is real. Requiring a human to approve every AI recommendation is a standard mitigation, but in the hiring example the review requirement was heavy enough to erase most of the time savings that motivated the project. Cost the review honestly, and confirm the reviewer can actually overturn the recommendation.
Who should own the decision to proceed? Someone who knows the approval pathway the use case triggers. A rights-impacting or safety-impacting federal system invokes obligations under OMB Memorandum M-24-10 including inventory, impact assessment, minimum risk-management practices and Chief AI Officer review, with parallel requirements increasingly common at state and local level. If nobody can name the pathway, the project is not ready to start.
Skill.re