Quality Assurance for AI Work Products
Priya Nair, an analyst at a state environmental agency, used an AI assistant to draft a summary of public comments on a new water permit. It was a 200-page comment file and the deadline was Thursday. The AI summary was clean, organised, and read beautifully. It also invented two comments that no one had submitted, and merged a citizen's objection with a developer's support so that they read as a single view. Priya's supervisor caught it in the meeting, in front of the regional director. The summary had to be redone, the meeting rescheduled, and Priya spent a week rebuilding trust she had built over two years.
The lesson Priya learned the hard way: AI does not produce work products. It produces drafts. The work product is what you sign off on after you have checked it. Quality assurance is the professional discipline that turns an AI draft into something safe to put your name on, and at agency scale it is the discipline that prevents public harm before it reaches a citizen. This lesson gives you a repeatable process for a single document and the structure of a quality assurance plan for a system that thousands of people depend on.
Why AI output needs a different kind of check
When a junior colleague drafts something, their errors are usually visible. They hedge, they leave gaps, they say "I am not sure about this part." AI does the opposite. It is fluent and confident even when it is wrong, and it will state a fabricated fact in the same smooth tone as a true one. That is the central quality assurance challenge, and it is the reason ordinary proofreading does not transfer. Proofreading hunts for typos and awkward phrasing, which these systems rarely produce. Checking AI work means hunting for confident falsehoods: invented facts, wrong numbers, citations to documents that do not exist, subtle misreadings of a source, and bias baked into the framing. The polish is exactly what makes the errors dangerous, because it lowers your guard.
The pattern repeats across professions. A federal attorney drafting with a generative tool may receive citations to cases that do not exist, as happened in the 2023 Avianca matter in federal court. A caseworker drafting benefits correspondence may receive plausible but incorrect regulatory references, which look right to everyone in the review chain who is not holding the regulation open. A policy analyst may receive numeric figures that sound authoritative and come from no verified source at all. Priya's fabricated comments belong to the same family. In every case the output was not garbled; it was well formed and wrong, which is the only kind of error that survives a casual read.
Verification, validation, and why you need both
Two questions sit underneath every quality assurance activity, and confusing them is how programs end up well tested and still harmful. Verification asks whether the system was built correctly. Validation asks whether the right system was built. Verification covers unit tests, integration tests and regression suites, and it is familiar territory for any agency quality assurance team. Validation is the harder half: subject matter expert review, ground-truth accuracy testing on representative samples, calibration checks, subgroup fairness analysis, and real-world piloting in shadow mode before anything reaches production.
Shadow mode deserves particular emphasis because it is cheap and routinely skipped. The system runs against live inputs, produces its outputs, and those outputs go nowhere except into a comparison log against what the humans actually decided. You get real distribution, real edge cases and real volume without a single citizen being affected by a mistake. The instruments around this are already published: the MEASURE function of the NIST AI Risk Management Framework provides the governance anchor, OMB Memorandum M-24-10 requires pre-deployment testing and ongoing monitoring for rights-impacting and safety-impacting AI, the GAO AI Accountability Framework organises performance into data, models, operations and outcomes, and ISO/IEC 42001 supplies management system clauses that tie the four together.
The five-layer check for a single work product
Run any AI-assisted document, analysis or recommendation through five layers, in order. Each layer catches a different class of failure, and the order matters because a completeness problem is easier to see once the facts are settled. Stop and fix before moving on. Be honest with yourself about what the five layers do: they catch five known classes of error, and a disciplined pass through them is strong evidence of care rather than a guarantee that the product is correct.
Layer 1: Accuracy, is every fact true?
- Trace every number, name, date and dollar figure back to a source you trust. If you cannot find it, it does not go in.
- Check every citation, regulation reference and quotation against the real document. These systems invent plausible-looking citations routinely, and a citation that resolves to a real document may still not say what the draft claims it says.
- Re-read the source that was summarised. Was it represented faithfully, or were points merged, dropped or invented, as they were in Priya's summary?
Layer 2: Completeness, is anything missing?
- Did the system quietly skip an important comment, exception or counterargument because it did not fit a clean narrative? Omission leaves no trace in the text, which is what makes it the hardest layer.
- Are the required sections, disclosures and caveats present, including any statement that the product was AI-assisted?
Layer 3: Bias and fairness, who is helped or harmed?
- Does the framing favour one group, vendor or outcome? Read it as the affected member of the public would read it, not as the author.
- Did the system flatten disagreement into a false consensus, which is precisely what merging an objection with a supporting comment accomplishes?
Layer 4: Compliance and sensitivity, is it allowed?
- Does it contain personal information or other sensitive data that should not be there, and would its presence be a disclosure if the document were released?
- Is it accessible under Section 508 if it will be published, with alt text, real headings and readable structure?
- Does it meet your agency's records and disclosure rules, given that it may later be produced under a records request?
Layer 5: Judgment, does the recommendation make sense?
- If the product carries a recommendation, would a reasonable expert reach the same one from the same facts?
- Does it overstate certainty? These systems rarely write "we do not have enough data to say," and often that is the honest sentence.
Designing the pre-deployment evaluation
The five layers govern a document. A deployed system needs an evaluation plan written before launch, because an evaluation designed afterwards tends to measure whatever the system happens to be good at. The plan starts with a ground-truth dataset: real cases with correct answers established independently of the system, drawn to reflect the population the system will actually serve rather than the cases that were easiest to collect. Everything downstream depends on this, and a ground-truth set assembled carelessly will make a weak system look strong in a way no later analysis can undo.
On top of that sit four further components. Subgroup stratification reports performance separately for defined groups, so an aggregate figure cannot conceal a disparity, and it is the practical expression of the equity principle in the Blueprint for an AI Bill of Rights, a non-binding federal framework rather than a statute. Calibration assessment asks whether stated confidence matches observed accuracy, because a system that is confidently wrong routes past every review threshold you set. Uncertainty quantification gives reviewers a signal about which outputs deserve their attention. Adversarial robustness testing establishes what happens under inputs designed to break it. For rights-impacting and safety-impacting uses, M-24-10 attaches minimum practices to this work rather than leaving it to professional judgement.
Setting thresholds you can defend
Quality assurance runs on numbers you choose: the share of outputs you sample for review, the accuracy floor below which the system pauses, the subgroup gap that triggers an investigation, the confidence level below which an output must reach a human. Every one of those is a threshold the agency sets for itself. Say so plainly in the plan. A self-set accuracy floor is a management commitment, not a legal test, and it is not a line separating lawful from unlawful conduct. Presenting it as though it were invites two failures at once: staff below the line assume they are safe, and staff above it assume they are exposed.
Two rules make self-set thresholds useful rather than decorative. Record them in advance, with the reasoning, before you have seen results that might tempt you to move them. And define what happens when one is crossed, naming the person who is notified and the action that follows, because a threshold with no consequence attached is a number in a document. A disparity trigger works the same way: it is the point at which the agency has committed to look, not the point at which harm begins. Harm can occur below any threshold you set, and your plan should say that in as many words so nobody reads the number as a floor of acceptability.
Citation verification and generative output
Generative systems introduce their own failure catalogue: hallucination, fabricated citations, overconfident wrong answers and prompt-injection vulnerabilities. Quality assurance for these tools therefore needs mechanisms rather than exhortations. Retrieval-augmented generation grounds outputs against an authoritative agency knowledge base instead of the model's own recall. Source logging records which document each assertion came from. Structured reviewer protocols tell a reviewer exactly which elements to verify rather than leaving them to read for plausibility. And clear disclosure tells the end user that the output was AI-assisted and human-reviewed.
Be precise about what retrieval buys you. Grounding a system in real documents reduces the rate at which it invents sources; it does not guarantee that a retrieved source says what the draft claims, that the right source was retrieved, or that the passage was read in context. The verification step stays with the human either way, which is why the reviewer protocol matters more than the architecture. The Notice and Explanation principle in the Blueprint for an AI Bill of Rights supports the disclosure practice, and the human-on-the-loop expectations in M-24-10 reinforce the review obligation for rights-impacting uses. Disclosure that the product was AI-assisted is a strict practice and worth applying more broadly than any single instrument strictly demands.
Peer review, and the independent review above it
Your own check catches many errors. A colleague's check catches a different set, the ones your own assumptions hide from you. For anything that leaves the building or informs a real decision, build in a peer review, and keep it light but real.
- Tell the reviewer it is AI-assisted. They should check facts, not just read for flow. This is the single most useful instruction you can give.
- Give them the source material. A reviewer cannot catch a misrepresented source without the source in front of them.
- Use a tiered rule. Low-stakes internal notes get a self-check. Public-facing or decision-driving products get a second reviewer, every time.
Above peer review sits independent review, and it is a different instrument with a different purpose. The Inspector General, GAO, agency civil rights officers, the Senior Agency Official for Privacy, the Chief Information Security Officer and bargaining units such as AFGE and NTEU each bring a question your program will not ask itself. Agencies that invite these parties in early reduce the risk of a surprise finding later, and the Houston Federation of Teachers case against HISD showed that courts themselves become independent reviewers of opaque systems when nobody else has, with due process as the accountability anchor. The Allegheny County Department of Human Services offers the positive model: published validation studies by external researchers, community advisory input, and ongoing monitoring with public reporting.
Red teaming as structured adversarial probing
Red teaming is the deliberate attempt to make a system fail, run by people whose job is to succeed at that. For government AI it should cover prompt injection, jailbreaks, data exfiltration attempts, harmful output categories including defamatory and discriminatory language, and scenario-based probing for the specific rights-impacting decisions your system touches. The CISA Guidelines for Secure AI System Development and the NIST Generative AI Profile provide reference material, and CISA and GSA have hosted federal red team exercises that agencies can learn from rather than starting cold.
Two design choices decide whether a red team programme is real. Build it into the development lifecycle rather than bolting it on before launch, because findings that arrive after the architecture is frozen become risk acceptances rather than fixes. And attach clear remediation pathways with follow-up testing, so that a finding is closed by evidence that the behaviour changed rather than by a note that it was reviewed. A red team report with no retest is an inventory of known weaknesses, which is worse than useless in a records request because it documents that you knew.
Monitoring and the pause-and-fix posture
Ongoing monitoring is not optional, because models drift, data changes and adversaries adapt. Agencies must watch performance, fairness, incident rates and user satisfaction continuously, and each of those catches something the others miss: user satisfaction surfaces harms your metrics never defined, and incident rates surface the failures that reached a person. Understand the limit as well: monitoring reports on what you chose to instrument, so the choice of signals is itself a quality assurance decision that deserves to be reviewed.
What separates a mature programme from a vulnerable one is what happens when monitoring finds something. The agency must have both the authority and the institutional posture to pause, investigate and remediate. That authority has to be granted in advance, in writing, naming who may exercise it, because in the moment the pressure will run entirely the other way. Agencies that promise leadership they will never need to pause are setting up the next Michigan MIDAS, where tens of thousands of false fraud determinations accumulated because inadequate quality assurance was never allowed to interrupt operations.
Documentation and the paper trail
Government work gets audited, challenged and requested under public records laws. You need to be able to show how a product was made, and three habits make that painless at the desk level.
- Mark AI-assisted drafts as drafts. Use clear file names such as
water-permit-summary_AIdraft_v1.docx, then_reviewed_v2, then_final, so that anyone can see the lineage at a glance. - Keep the source and the prompt. Save the source documents and a note of what was asked. If a number is questioned later, you can reconstruct where it came from instead of reasoning about it.
- Record who checked what. A one-line note naming the person who verified the facts and the person who peer-reviewed it makes the product defensible. It does not make the product correct, and the note should never be read as though it did.
At system level the same instinct produces five standing artifacts. Model cards document intended use, limitations, training data, evaluation methodology and known risks. Datasheets document dataset provenance, collection processes and preprocessing. Test reports document pre-deployment evaluation results. Incident reports document failures and the corrective actions taken. Operator checklists translate abstract requirements into the concrete steps a reviewer performs daily. Together these support responses under the Freedom of Information Act, the Privacy Act, the Administrative Procedure Act and congressional oversight, which is to say they are produced whether you prepared them or not.
Priya's redo, done right
The second time, Priya kept the comment file open beside the AI draft and checked every summarised point against the original, which is Layer 1. She found the two fabricated comments and the merged objection in fifteen minutes. She gave Tom the draft and the source for peer review rather than the draft alone. She named the files by version and noted who had verified what. The redo took ninety minutes in total, against a first attempt that cost a week of recovery. Quality assurance is not the slow path; it is the cheap one, and the cost comparison is the argument that persuades a supervisor who thinks checking is optional.
Your quality assurance plan
The exercise this lesson builds toward is a written quality assurance plan for one rights-impacting AI system in your own agency. It covers pre-deployment evaluation, ongoing monitoring, red teaming cadence, independent review engagement and incident response, and it should name people rather than offices wherever a decision has to be made. Have a peer review the plan itself, on the same principle that governs every other product here: the gaps you cannot see are the ones somebody else finds immediately.
Then commit to implementing it within thirty days, which is the source's own cadence and is chosen to be short enough that the plan does not become a document about intentions. Thirty days is long enough to establish the ground-truth dataset, write the reviewer protocol and get the pause authority signed, and short enough that the people who agreed to it still remember agreeing. What you should not do is treat the finished plan as the accomplishment. Michigan had procedures. The Houston evaluation model had documentation. Each of the failures in this lesson traces back to a gap in quality assurance, and each was preventable, which is the uncomfortable point of studying them.
Anti-Patterns to Avoid
- Treating the completed checklist as the assurance. A completed checklist has never once made a false statement true. It is evidence that a defined set of checks was run, which is valuable and is not the same as evidence that the product is right.
- Letting peer review become a rubber stamp. A reviewer who is not told the product is AI-assisted and is not given the source material will read for flow and sign. That signature then transfers accountability without transferring any actual checking.
- Reading fluency as reliability. The smoothest paragraph in the draft is the one most likely to slip through, because nothing about it triggers suspicion. Polish should raise your guard.
- Trusting retrieval to solve fabrication. Grounding a system in authoritative documents lowers the invention rate. It does not guarantee the retrieved source supports the claim, and a real citation attached to a wrong assertion is harder to catch than an invented one.
- Presenting a self-set threshold as a legal standard. Your sampling rate and accuracy floor are commitments the agency makes to itself. Harm is possible on the safe side of any of them, and staff who believe otherwise will stop looking.
- Running a red team with no retest. An unretested finding is a documented weakness you knew about, which is a worse position than not having tested, particularly in a records request or an inspector general review.
- Monitoring only what is easy to instrument. Accuracy against available labels is the cheapest signal and the least likely to surface a novel harm. User complaints and incident rates catch what your metrics never defined.
- Promising leadership that the system will never need to pause. The promise is what makes pausing unthinkable at the moment it is needed, and it converts a recoverable incident into an escalating one.
- Deferring independent review until it is compelled. Inspectors general, civil rights officers and bargaining units cost less as early participants than as sources of a surprise finding, and courts are the most expensive independent reviewer of all.
Practice Prompts
These are checking aids, not drafting aids, and none of them replaces reading the source yourself. Use them only on material cleared for the tool you are using, and treat every output as a claim to verify rather than a result.
- "Here is a summary and here is the source document it was drawn from. List every assertion in the summary that does not appear in the source, and every point in the source that the summary omits."
- "Extract every factual claim, number, citation and quotation from this draft into a table with a column for the source I will fill in myself. Do not populate the source column."
- "Read this draft as the member of the public it is about. Where does the framing favour one party, and where has disagreement been presented as consensus?"
- "List the reviewer instructions someone would need to check this specific product, naming which elements must be verified against a primary source and which may be read for judgement."
- "Here is my draft quality assurance plan. Identify every threshold in it that is stated without saying who set it, what happens when it is crossed, and who is notified."
Reflection Questions
- Think of the last AI-assisted product you signed off on. Which of the five layers did you actually run, and which did you assume?
- If a fabricated fact reached a published agency document through your team, how would you find out, and how long would that take?
- Which thresholds does your programme use today? Were they recorded in advance, and does anyone downstream believe they mark the boundary of acceptable harm?
- Who could pause your AI system tomorrow, and is that authority written down anywhere a person could point to under pressure?
- Which independent reviewer would find your programme's weakest point fastest, and what stops you from inviting them now?
Glossary
- Verification. Checking that the system was built correctly, through unit, integration and regression testing.
- Validation. Checking that the right system was built, through expert review, ground-truth accuracy testing, calibration, subgroup analysis and shadow-mode piloting.
- Hallucination. A fluent, confident output that is not grounded in any real source, including invented facts and citations to documents that do not exist.
- Ground-truth dataset. Real cases with correct answers established independently of the system, drawn to represent the population the system will serve.
- Calibration. The correspondence between a system's stated confidence and its observed accuracy; poor calibration lets wrong answers pass confidence-based review gates.
- Shadow mode. Running a system against live inputs while its outputs affect nothing, purely for comparison against human decisions.
- Retrieval-augmented generation. Grounding generated output in retrieved documents from an authoritative knowledge base rather than in the model's own recall.
- Red teaming. Structured adversarial probing covering prompt injection, jailbreaks, exfiltration attempts and harmful output categories.
- Model card. A document stating a model's intended use, limitations, training data, evaluation methodology and known risks.
- Datasheet. A document stating a dataset's provenance, collection process and preprocessing, so that later users know what they are working with.
- Pause-and-fix authority. The pre-granted, written authority to suspend a system pending investigation when monitoring shows degradation.
Related Lessons
- Evaluating AI Outputs works through the judgement layer in more depth on individual outputs.
- Bias Detection Tools and Methods supplies the subgroup analysis techniques behind the fairness layer.
- Human-in-the-Loop: Design and Implementation covers how review points are designed so that they are real rather than nominal.
- Advanced Adversarial Testing extends the red teaming material to a full programme.
- Testing and Validating AI Systems develops the pre-deployment evaluation design.
- AI Audit Methodology shows what an independent reviewer will look for in your documentation.
- Building an AI Quality Culture addresses the part no checklist reaches: whether people feel able to raise a problem.
Closing Thoughts
Priya's mistake was not using AI to summarise a long comment file. That was a reasonable use of a tool for a real deadline. Her mistake was treating the output as a work product rather than a draft, and the correction cost fifteen minutes the second time. Almost everything in this lesson scales that same substitution upward: replace an impression that the output looks right with a specific, recorded check against a specific source, performed by a named person, at a point decided in advance.
What none of this delivers is certainty. The five layers catch five classes of error, evaluation catches what the ground-truth set represents, red teaming catches what the red team thought to try, and monitoring catches what you instrumented. A quality assurance programme is a set of deliberate, documented bets about where the failures will be, revisited as you learn you bet wrong. Agencies that hold it that way keep improving. Agencies that treat the completed plan as the accomplishment discover the gap the way Michigan did, in public, years later, at somebody else's expense.
Key Takeaways
- AI gives you drafts, not work products. The work product is what you verify and sign; accountability for accuracy never transfers to the tool.
- Fluency is the trap. These systems state false facts as confidently as true ones, so polish should raise your guard rather than lower it.
- Run the five layers in order. Accuracy, completeness, bias, compliance and judgement each catch a different failure, and together they are strong evidence of care rather than a guarantee of correctness.
- Verification and validation are different questions. Built correctly is not the same as built right; validation needs ground truth, calibration, subgroup analysis and shadow-mode piloting.
- Trace every fact and citation to a real source. Retrieval reduces fabrication but does not confirm that a retrieved source supports the claim attached to it.
- Say out loud that your thresholds are self-set. A sampling rate or accuracy floor is a management commitment recorded in advance, not a legal test and not the point where harm begins.
- Match review depth to stakes, and tell the reviewer. Self-check internal notes; require a second reviewer with the source material for anything public-facing or decision-driving.
- Red teaming without retest is an inventory of known weaknesses. Build it into the lifecycle and close findings with evidence that behaviour changed.
- Pause authority must exist before you need it. Write down who may suspend the system, because the pressure at the moment of decision runs entirely the other way.
- Keep the paper trail. Versioned drafts, saved prompts and sources, model cards, datasheets, test reports and operator checklists are produced under records law whether or not you prepared them.
Frequently Asked Questions
Does a peer review make an AI-assisted product safe to publish?
It makes it better checked, which is not the same thing. A peer review catches errors your own assumptions hide, and it does so only if the reviewer is told the product is AI-assisted and is given the source material. Without both, the reviewer reads for flow and the signature adds accountability without adding scrutiny. Treat the review as one layer among several, and keep the accuracy check that precedes it rather than delegating it.
How much of the output should we sample for review?
That is a threshold your agency sets, and this lesson deliberately does not supply a number, because a rate imported from another programme carries none of the reasoning that justified it. Set it from the severity of the worst error your system can make and the volume you can genuinely review, record it with that reasoning before you see results, and define what happens when the review finds a problem. Then revisit it when monitoring shows the error profile changing.
Is retrieval-augmented generation enough to stop hallucinated citations?
It reduces the rate at which sources are invented, because the system is drawing from a real corpus rather than from recall. It does not establish that the retrieved document supports the assertion, that the right document was retrieved, or that a passage was read in context. Citation verification against the primary document stays a human step, which is why the reviewer protocol matters more than the retrieval architecture.
Who should be involved in independent review?
The Inspector General, GAO, agency civil rights officers, the Senior Agency Official for Privacy, the Chief Information Security Officer and bargaining units such as AFGE and NTEU all bring questions the programme will not ask itself. Engaging them early converts what could become an adversarial finding into shared work on the fix. The Allegheny County example adds external researchers and community advisory input with public reporting, which is the fuller version of the same posture.
What do we do when monitoring shows the system degrading?
Use the pause-and-fix authority, which is why it has to be written down and assigned before the incident. Suspend or throttle the system, investigate, remediate, retest, and document the whole sequence in an incident report that feeds the next evaluation cycle. The failure mode to avoid is a programme that can detect degradation and cannot act on it, because the detection then only documents how long the harm continued.
Do these practices apply to low-stakes internal work?
Scale them rather than skipping them. A low-stakes internal note gets a self-check across the five layers and a versioned file name. A rights-impacting system gets the full apparatus: ground-truth evaluation, subgroup stratification, calibration, red teaming, independent review and monitoring with pause authority. The trigger for the heavier treatment is not the size of the project but whether the output materially affects a person's rights, benefits or safety.
Skill.re