←
AI for Government
Capable · M33 · lesson 33 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Requirements Gathering for AI
📖
now learning

Requirements Gathering for AI

15 min

Rosalind Achterberg-Mbeki had run requirements gathering for a dozen government IT projects, and she was good at it. She knew how to interview stakeholders, write user stories, and pin down acceptance criteria before a single line of code was written. So when she was assigned to gather requirements for the labor department's first AI project, a tool to help caseworkers match unemployed workers to job-training programs, she ran her usual playbook. Three weeks later, she had a tidy requirements document, and an uneasy feeling that it described the wrong kind of system. Her document said things like "the system shall match workers to programs." But it could not answer the questions the data scientists kept asking: matched how well? Compared to what? What counts as a good match versus a bad one? How will we know if it is getting worse? Her traditional requirements were precise about what the system should do and silent about the things that actually determine whether an AI system succeeds or fails.

This lesson is about requirements gathering for AI: how it differs from traditional IT requirements, what new questions you must ask, and how to produce a requirements document that sets an AI project up to succeed rather than to disappoint. The core insight is that traditional requirements describe deterministic behavior ("when the user clicks submit, the form is saved"), while AI systems are probabilistic. They are right most of the time and wrong some of the time, and the entire game is defining "most," "some," and "acceptable." If your requirements do not pin those down, your project has no definition of done.

Why AI Requirements Are Different

Traditional software does what you tell it. You specify behavior; the system either implements it correctly or has a bug. AI software does what the data teaches it, approximately. You cannot specify "match every worker to the perfect program," because no system achieves that and "perfect" is not even well defined. Instead you specify a target, a measurable level of performance that counts as good enough, and you specify the conditions under which the system must hold that target.

This changes the requirements interview. Rosalind's old questions ("what should the system do?") had to be joined by new ones: how do we measure whether it did it well? what error rate is acceptable, and what is the cost of each kind of error? how must performance hold across different groups of workers? and how will we detect when it degrades? These are the questions that determine whether an AI project delivers value or quietly fails.

A traditional IT requirements document typically specifies four things: what the system will do, how it will perform on speed, capacity and uptime, what it will integrate with, and what it will cost. Those four still apply. AI requirements have to carry eight more on top of them: what success actually looks like in accuracy, fairness and equity outcomes; what data the system needs and where that data comes from; what the system cannot be expected to do, stated as explicit limitations and failure modes; how the team will validate that it works; how it will be monitored in production; what happens when it fails or must be overridden; the fairness and bias considerations specific to the domain; and how it will be updated as conditions change.

Traditional IT requirements specifyAI requirements must also specify
What the system will do (features and functions)What success actually looks like (accuracy, fairness, equity outcomes)
How it will perform (speed, capacity, uptime)What data the system needs and where it comes from
What it will integrate with (data sources, other systems)What the system cannot be expected to do (limitations, failure modes)
What it will costHow validation will be performed before deployment
How the system will be monitored and alerted on in production
What happens when the system fails or must be overridden
Fairness and bias considerations specific to the domain
How the system will be updated and maintained as conditions change

A Problem Statement Is Not a Requirement

You will walk into a requirements meeting and hear something like "we want to use AI to detect fraud in our tax filings" or "we need AI to improve our hiring process." Those sound like requirements. They are problem statements. A real requirement answers much harder questions: what counts as fraud that the system should catch? What does "improve" actually mean, more diversity, faster processing, better accuracy? Who decides whether the system is working? And what happens when the AI gets it wrong? Until those are answered, nobody can build to the statement and nobody can test against it.

So the first move is to refuse the shortcut. Do not accept "we want to use AI" as a problem statement; dig into what is failing in the current process, how the team knows it is failing and which metrics show it, what success would look like, who experiences the current failure (staff, citizens, other agencies), and what has already been tried. The difference is stark. Instead of "we want AI to improve our hiring," the useful version reads: "we are getting complaints that hiring decisions take too long and candidates from certain universities are rarely selected. We want to reduce time-to-hire from 90 days to 30 days and increase diversity in our selection."

Success Criteria You Can Actually Test

Every AI requirements document must define how success is measured and what level of performance is acceptable before deployment. Vague goals like "improve matching" are not requirements; they cannot be tested. Rosalind reworked the goal into a measurable target negotiated with the program office: the tool's recommended programs should result in program completion at a rate at least as high as current caseworker matching, measured over a defined evaluation window, with a pre-registered threshold so success could not be redefined after the fact. The act of setting that number forced a conversation the program office had been avoiding: what does "better" actually mean here?

Pinning a target down means answering a specific set of questions with the stakeholders in the room. What metrics will we use? What is the baseline, meaning current performance without AI? What improvement are we actually trying to achieve? Who defines success, given that technical teams and business teams routinely hold different criteria and often do not discover the difference until testing? And how will we know in production whether success is being achieved? A criterion that cannot be measured with data the agency will actually hold is not a criterion; it is an aspiration wearing a number.

Good criteria look concrete enough to argue about. Examples from real government settings: "reduce fraud detection time from 40 hours of analyst review to 10 hours per case"; "identify 85% of high-risk maintenance items with a false positive rate below 5%"; "achieve demographic parity in approval rates across protected classes within 2% tolerance"; "reduce decision variance for similar cases from 15% to less than 5%." Each of those can be checked against evidence and each can fail. That is what makes them requirements rather than intentions.

Error costs and the asymmetry of mistakes

AI systems make two kinds of mistakes, and they rarely cost the same. Recommending a worker to a program they are not suited for wastes a training slot and the worker's time. Failing to recommend a program a worker would have thrived in costs that worker an opportunity they will never know they missed. Requirements must state which error is worse and by how much, because that asymmetry shapes how the system is tuned. Rosalind's document specified that under-recommendation (missing a good match) was the costlier error, which told the data team to tune toward recall over precision, a design decision that lives or dies on a requirement.

Fairness and subgroup performance

An AI requirement that specifies only overall performance permits a system that works well on average while failing a specific group. For a government system, that is unacceptable and likely unlawful. Requirements must state that performance be measured and meet the target across demographic and geographic subgroups, not just in aggregate. Rosalind added an explicit requirement: matching quality must hold within defined tolerance across race, gender, age band, and rural-versus-urban location, with disaggregated reporting. This single requirement is what later let the team catch that early model versions performed markedly worse for rural workers.

Data Requirements Decide What Is Possible

AI systems have requirements traditional systems do not: what data is needed, in what quantity and quality, with what labels, and with what legal authority to use it. Rosalind's document specified the training data sources, the minimum data quality, the labeling approach (what counts as a "successful" historical match), and, crucially, confirmation that using historical caseworker outcomes to train the model was consistent with the data's original collection purpose and the Privacy Act. Data determines what is possible in an AI system, and that constraint is non-negotiable.

The data conversation has its own question set, and it belongs in requirements rather than in a later technical phase. What data sources are available? How complete is the data, including missing values and coverage gaps? How recent is it, and is historical data stale for the purpose? Is it representative of the population the system will be applied to? Are there protected class fields present, such as demographic data or disability status? What biases or known problems sit in the historical record? Who owns the data, and can the agency actually use it for this purpose? What cleaning or preparation is needed, and what data governance constraints apply, including privacy regulations and data protection requirements?

The reason to spend requirements time here is blunt: the system can only be as good as the data it learns from. Bad data produces a bad system no matter how good the algorithm is. A requirements document that specifies a demanding accuracy target on data nobody has assessed has specified a wish, and the assessment that would have revealed the problem costs far less at requirements time than after a build.

Designing the Human Oversight Model

Because the tool informs decisions about people, requirements must state how the system explains its recommendations to the caseworker and how the human stays meaningfully in control. Rosalind specified that every recommendation be accompanied by the top factors driving it, in plain language, and that the caseworker retain and routinely exercise the ability to override, with overrides logged and reviewed, so "human oversight" was real rather than nominal. Oversight that is never exercised is not evidence that the system is right; it is evidence that nobody is looking.

Oversight is a governance decision made in requirements, not a detail settled later. Ask whether the system will be fully automated or human-reviewed. If humans review, what percentage of decisions do they see, and what triggers review: all decisions below a confidence threshold, all adverse decisions, random sampling? How much time does a reviewer actually have per decision? What training will reviewers need? Can reviewers override the AI recommendation, and here the answer is not open: they should be able to. And how will you monitor whether the reviewers are providing value rather than clicking through?

Governance modelWhat it meansWhere it fits
Full automationSystem makes final decisions with no human reviewVery limited use in government, only appropriate for low-stakes decisions
High-risk human reviewHumans review all decisions, especially adverse decisionsAppropriate for rights-impacting systems
Sampling reviewHumans review a random sample of decisions to check qualityUseful for monitoring but inadequate for full validation
Alert-based reviewSystem flags potentially problematic decisions for human reviewAppropriate for fraud detection and outlier identification
Assistance modelSystem provides recommendations that humans use to inform their own decisionsAppropriate for many government processes

Limits, Failure Modes, and Degradation

Every system has limitations, and writing them down early is what prevents deployment failures. Requirements should state what the system is not expected to do, what would constitute a failure of the system, what happens if it is wrong including the relative cost of false positives and false negatives, whether there are populations or case types where it will inevitably perform worse, what the operational constraints are on speed, availability and cost, what happens when the system cannot process something (does it escalate, or default to accept or reject?), and how often it will need updating as the environment changes.

Traditional software does not get worse on its own; AI does, as the world drifts from its training data. Requirements must specify what is monitored after deployment, how often, and what performance drop triggers retraining or rollback. Rosalind's document named the monitoring metrics, a monthly cadence, and a defined threshold that would automatically trigger a model review. That threshold was a management commitment the agency set for itself, not a legal test and not a line that separates a lawful system from an unlawful one. Without something like it, though, an AI system degrades silently until someone complains.

The AI Requirements Template

A structured template is how you stop relying on memory. Rosalind's final artifact was an AI requirements specification organized in two parts: the familiar functional requirements (what the system does and how users interact with it) and a new AI performance specification. The second part was what made the project buildable and evaluable, and it was the part her original document had entirely lacked. The template below is the checklist form of that second part, and each block is a set of prompts to fill in with the stakeholders present rather than a form to complete alone at a desk.

  • Business context. Name of system; target users and stakeholders; the current process before AI; why AI is being considered.
  • Problem statement. What specific problem are we solving; how the problem is measured today, meaning current metrics and baseline; who experiences the problem; what has already been tried.
  • Success criteria. Performance metrics such as accuracy and speed; business outcome metrics; fairness and equity metrics where applicable; acceptable performance ranges; how success will be measured in production.
  • Scope and scale. How many decisions or cases per month; which departments and offices will use it; which populations will be affected; whether scope will expand, and planning for that in advance.
  • Data requirements. Sources available; completeness and quality issues; recency and staleness concerns; demographic composition against the population served; protected class fields and fairness concerns; data governance and privacy constraints.
  • System functionality. What inputs the system receives; what outputs it produces; what confidence level is needed to act on a recommendation; edge cases and special scenarios; what situations trigger human escalation.
  • Validation and testing. How accuracy will be tested; how fairness and bias will be tested; which edge cases must be tested; who validates before deployment; what validation documentation is required.
  • Human oversight design. Automated or human-reviewed; if reviewed, what percentage, which decisions and what process; whether reviewers can override AI decisions, and they must be able to; what training reviewers need; how reviewer performance is monitored.
  • Monitoring and maintenance. What metrics are monitored in production; how often performance is reviewed; what triggers escalation or investigation; how often the model is retrained; how performance degradation is handled.
  • Risks and constraints. What could cause the system to fail; which populations might be affected differently; privacy and security considerations; operational constraints on speed, cost and availability; the contingency if the system does not work.
  • Governance and compliance. Which policies apply, including civil rights and federal acquisition; who needs to approve the system; what documentation is required; how incidents and failures will be handled; what training stakeholders need.

Widening the Room: Stakeholder Mapping

Gathering these requirements means widening the room. Rosalind's traditional interviews included program staff and end users. Her AI requirements interviews added the data team (to ground targets in what the data could support), the privacy office (data authority and impact-assessment implications), the civil rights office (subgroup fairness), and legal (notice and appeal obligations). The interview is no longer just "what do you want the system to do?" It is a structured negotiation that produces measurable, testable, fairness-aware targets that every stakeholder has agreed to in advance.

The typical cast in a government AI project is larger than most requirements leads expect. The business sponsor wants business value in efficiency, accuracy and cost savings. End users care about usability, time impact and reliability. The compliance and legal team worries about policy compliance and legal exposure. The civil rights or EEO team worries about discrimination and fairness. IT and security care about data handling, security and architecture. Audit and oversight care about documentation and traceability. And affected communities care about fairness, accuracy and the impacts they will live with. Each group holds a legitimate concern that should shape the requirements.

The hard part is that these concerns conflict. Business teams want fast systems; civil rights teams want careful validation. IT teams want simple data pipelines; fairness teams want detailed demographic tracking. Your job as requirements lead is to gather from all stakeholders, make conflicts explicit rather than burying them, help teams negotiate and prioritize, document what was committed to and for whom, and mark clearly where a requirement represents a tradeoff. A worked example of that last move reads: "to meet business timeline requirements (3-month deployment), we are accepting demographic performance gaps up to 5% rather than pursuing perfect parity. Legal team and civil rights team have signed off on this tradeoff. This will require enhanced monitoring in production."

Validating Requirements and Getting Sign-Off

Do not just gather requirements and move on. Validate them while everyone is still in the room. Do all stakeholders understand the success criteria the same way? Are those criteria actually measurable with the data you have? Are there conflicting requirements still to be negotiated? Have you identified the real constraints, given that constraints often hide assumptions? Is the scope realistic for the timeline and budget? Have you stated what the system is not expected to do? Do you have commitment from data owners to provide the required data, and from the business to implement human oversight as designed? Have fairness and bias concerns been addressed?

Then get formal sign-off: from the business sponsor that this solves the business problem; from legal and compliance that it can be done lawfully; from civil rights that the fairness approach is appropriate; from the data owner that the necessary data can be provided; and from IT that this can be built and maintained. Each signer should initial against the specific requirements they are committing to, not the document as a whole. Record the close: requirements finalized on a stated date, signed off by named people, with any change requiring a written amendment. A signature does not make a requirement achievable, but it does make a later dispute resolvable against a record rather than against memory.

Anti-Patterns

  • Vague success criteria. Requirements say "the system should improve decision quality" or "should work well" without defining metrics. Teams discover during testing or deployment that they disagree about whether the system succeeded. Every success criterion must be testable and measurable. "Improve" is not a criterion. "Reduce false positive rate to below 3%" is. Make people commit to numbers.
  • Ignoring data quality. Requirements focus on what the AI should do without assessing whether the data is good enough to do it. The system is built, tested on good data, deployed, and fails because production data is different. Spend real effort on data quality during requirements, and get data owners into the room.
  • Unrealistic timelines and scope. Requirements are gathered in a compressed window with limited stakeholder input; the team learns later they are incomplete. Scope expands during development, timelines slip, stress rises. Allocate enough time, involve everyone, document scope and timeline assumptions, and build in discovery time rather than locking requirements too early.
  • Fairness as an afterthought. Requirements focus on accuracy and business outcomes, and fairness is bolted on late, by which point the architecture cannot support fairness monitoring. Ask from the first meeting: who could be harmed if this system is unfair, and what fairness metrics matter? Then require fairness validation inside the success criteria.
  • Missing failure modes. Requirements never say what happens if the system fails or makes a bad decision, so when something goes wrong in production the team improvises. Ask explicitly what the worst case is and what you would do if the system had to be turned off, and document fallback and override procedures.
  • Treating the template as coverage. A completed template is evidence that a set of known questions was asked. It cannot prompt the question nobody thought of, and a filled-in field is not a validated answer. Use the template to structure the conversation, then ask the stakeholders directly what it failed to ask about their situation.
  • Selling requirements as a harm preventer. Requirements gathering is a governance control, and a strong one, but calling it harm prevention overstates it. It reduces the harms someone anticipated and wrote down. It does nothing about the harm nobody raised, which is why the civil rights and affected-community voices are load-bearing rather than decorative.

Practice Prompts

  • Requirements template application. Take a government AI system you know, or one from your agency, and apply the AI requirements template. Complete each block as if you were gathering requirements for it. Identify which blocks you can fill with real information and which are gaps, and note what additional stakeholder conversations would be needed to close each gap.
  • Stakeholder perspective mapping. Choose a hypothetical government AI system, such as predictive maintenance for federal buildings, benefits eligibility screening, or hiring assistance. List the likely stakeholders and their perspectives. What does each care about? Where might their interests conflict? How would you negotiate those conflicts, and what would each stakeholder sign off on?
  • Success criteria definition. Write detailed success criteria for one of these: an AI system that flags suspicious purchases in federal procurement; an AI system that assists in staffing decisions for federal offices; or an AI system that predicts infrastructure maintenance needs. For your chosen system, define performance metrics, business outcome metrics, fairness metrics, and acceptable ranges, and say how each will be measured.
  • Break your own criteria. Take one success criterion you wrote in the previous prompt and describe a system that would satisfy it while still being unacceptable to deploy. Then write the additional requirement that closes that gap. This is the fastest way to find the difference between a testable criterion and a sufficient one.

Reflection

Reflect on a complex government process you know well. What would it take to use AI to improve it? What would you need to know about the people affected by its decisions? What would you have to measure to know whether the AI was actually helping rather than simply producing output faster? What could go wrong, and who would find out first? Use this reflection to build intuition for the depth of requirements work an AI project needs, and notice how many of your answers require someone outside your own office to supply them.

Glossary

  • Requirements. Specifications of what a system must do, including functional requirements (what it does), performance requirements (how well it does it), and constraints (what it cannot do, or limits on how it operates).
  • Success criteria. Measurable, specific standards by which the success or failure of a system will be evaluated, including performance metrics, fairness metrics, and business outcome metrics.
  • Stakeholder. Any person or group with an interest in, or affected by, an AI system, including business sponsors, end users, compliance teams, and affected populations.
  • Fairness metric. A measurable standard for assessing whether an AI system treats different populations equitably, such as demographic parity, equalized odds, or calibration.
  • Data quality assessment. Evaluation of whether available data is suitable for the intended AI application, considering completeness, recency, representativeness, and bias.
  • Baseline. Current performance of the process without AI, measured before the project starts, against which any claimed improvement is judged.
  • Error cost asymmetry. The difference in real-world consequence between a false positive and a false negative for a given system, which determines how the system should be tuned.

Closing

Requirements gathering is where you establish the foundation for the entire AI project. Bad requirements lead to failed projects, governance failures, and wasted resources; good requirements lead to systems that solve real problems and earn trust. The key insight for government is that AI systems are not just technical systems. They are governance systems. They affect people's access to services, their opportunities, and their treatment by their government, and requirements gathering is where those impacts get understood and encoded rather than discovered later.

You will return to this document repeatedly. During testing you will check the system against the success criteria you defined. During deployment you will implement the human oversight model you designed. During monitoring you will track the fairness metrics you included. Requirements discipline early is what prevents chaos later, which is exactly why Rosalind's instinct that something was missing turned out to be correct: the second half of her specification, the AI performance half, was the half that made the project buildable at all.

Key Takeaways

  • AI is probabilistic; requirements must define "good enough." Traditional requirements describe deterministic behavior; AI requirements must specify measurable performance targets and the conditions under which they hold.
  • A problem statement is not a requirement. "We want AI to improve hiring" becomes a requirement only when someone states the current failure, the baseline, the target, and who decides whether it was met.
  • Set measurable success criteria with a pre-registered threshold. Criteria must be specific enough to fail, and negotiated before deployment so success cannot be redefined afterward.
  • Specify the asymmetry of error costs. The two kinds of AI mistakes rarely cost the same; stating which is worse drives how the system is tuned.
  • Require subgroup fairness, not just average performance. Performance must meet the target across demographic and geographic groups, with disaggregated reporting, or a system can fail a group while looking fine on average.
  • Treat data and legal authority as first-class requirements. The system can only be as good as its data. Specify sources, quality, labels, and confirmation that the intended use is lawful under the Privacy Act and consistent with the original collection purpose.
  • Design the oversight model explicitly. Choose among full automation, full human review, sampling, alert-based review, and assistance, and remember that sampling is useful for monitoring but inadequate as full validation.
  • Document what the system is not expected to do. Limitations, failure modes, and contingencies belong in requirements, not in the incident report.
  • Make conflicts explicit and get initialled sign-off. Stakeholders bring legitimate but competing requirements; your job is to surface tradeoffs, document them, and have each signer initial the specific commitments they are making.

Frequently Asked Questions

How is this different from just writing good traditional requirements? Traditional requirements are complete when they describe behavior precisely. AI requirements are complete only when they also describe a measurable performance target, the acceptable error profile, how performance must hold across groups, what data makes it possible, who oversees it, and what triggers action when it degrades. The functional half of the document does not change much. The other half does not exist in traditional practice.

Our stakeholders refuse to commit to a number. What do we do? Treat the refusal as the finding. If nobody will say what performance counts as acceptable, then nobody can say later whether the system worked, and every post-deployment argument becomes unresolvable. Bring the baseline to the table first, because people commit to a target far more readily once they can see current performance measured.

Who owns fairness requirements, the data team or the civil rights office? The civil rights or EEO office defines what fairness means for the system and what disparities are tolerable; the data team determines what is measurable with the available data and reports the disaggregated results. Neither can do it alone, and the requirements document is where those two answers get reconciled in writing rather than in a dispute after launch.

Can a template guarantee we covered everything? No. A template captures the questions someone previously thought to ask, which is genuinely valuable and still not coverage. Use it to structure the conversation, then ask each stakeholder what the template failed to ask about their situation. The gap between a filled-in field and a validated answer is where most requirements failures live.

How firm should the requirements be, given that AI projects learn as they go? Firm on success criteria, fairness tolerances, data authority, and oversight design; flexible on implementation approach. Lock the definition of done and leave room for discovery about how to reach it. Locking everything too early produces a specification the team quietly abandons, and locking nothing produces a project with no way to tell success from motion.

What if the requirements reveal the project should not proceed? That is a successful requirements exercise, not a failed one. Learning at requirements time that the data cannot support the target, that the legal authority is absent, or that no acceptable oversight model fits the workflow is the cheapest possible moment to learn it. Document the finding and the reasoning; the next team to propose something similar will need it.