←
AI for Government
Capable · M12 · lesson 12 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Capstone Lab: End-to-End AI Integration Project

15 min

Sofia Reyes, a program analyst at a city housing authority, had taken every AI lesson her agency offered and could not point to one thing she had actually shipped. She knew prompts. She knew the risks. She had never carried a single real task from "this is a problem" all the way to "this is running and validated." Her manager finally said the thing that stung: "You have learned about swimming. Now get in the pool." This lab is the pool. By the end you will have taken one real government task from framing through data, build, validation, deployment and monitoring, and you will hold a finished artifact and a completed project record with your name on it.

Sofia's task: her team spends roughly 12 hours a week reading inbound tenant complaint emails and sorting each into one of six categories (maintenance, billing, safety, noise, discrimination, other) so they reach the right desk. It is dull, error-prone, and a good first AI integration. Pick your own equivalent now, a real, bounded, repetitive task you actually own. You will build it as you read. Alongside Sofia we will run a second, larger case at federal scale, so that you can see the same five stages carry from a two-person pilot to a program that a Chief AI Officer would have to defend to an inspector general.

The federal case is the one the capstone is graded against. You are a program manager at a USDA Rural Development grants office processing 5,000 permit applications per month. Current processing time averages 60 days. The success rate, meaning approved without rework, is 65 percent. Three GS-13 analysts spend the equivalent of 1.8 FTE on processing, at an annual cost of roughly $500,000. The problem statement is that inconsistent rule application delays qualified applicants and denies claims that should succeed. Your task is to plan an AI-assisted system that reduces processing time to under 20 days, raises first-pass success to 85 percent, and produces legally defensible decisions.

What a capstone has to prove

A capstone is not a quiz on definitions. It is the point where everything you have studied, risk frameworks, workflow analysis, use-case development, prompt engineering, vendor management, testing, change management and measurement, has to become a single plan somebody could actually run. Reviewers judge the work on six criteria: completeness, appropriateness, specificity, integration, feasibility and governance. That is six criteria counted from the source list, and specificity is the one that decides most outcomes. Generic plans read fine and fail in practice.

The bar to aim for is concrete: a good capstone reads like something a reviewer at the Office of Management and Budget or the Government Accountability Office could hold in their hands and understand without asking you three clarifying questions. Connect it to your real agency context wherever you can. Document your assumptions rather than hiding them. Show your decision-making, including the options you rejected. Integrate concepts across lessons instead of treating each as a separate box to tick. This is the difference between an AI-ready practitioner and an AI-curious observer.

Stage 1: Frame the problem

Most AI projects fail here, by being vague. "Use AI for complaints" is not a task; it is a wish. A framed problem is specific, bounded and measurable. Write yours as a single sentence carrying a baseline and a target. Sofia's reads: classify each inbound tenant complaint email into one of six categories, currently done manually in about 12 hours a week with an estimated 1-in-10 misrouting rate, with the target being an automatic draft classification that a human confirms. Notice what that sentence contains and what it refuses to contain.

It names the exact input and the exact output. It states the current cost, 12 hours a week, and the current error rate, roughly one in ten, so improvement can be proved rather than asserted. And it says draft for human confirmation, not decide, which is the right posture for any government output touching a member of the public. The federal case follows the same shape at larger scale: 5,000 applications a month, a 60-day average and a 65 percent first-pass rate become the baseline, and under 20 days plus 85 percent first-pass become the target. If you cannot write your task in one sentence with a current cost and a target, you are not ready to build it, because the fuzziness will become the failure.

The screening decision: should AI do this at all?

Before building, screen the task against three questions. Is it repetitive and loose enough in its rules that AI helps, yet consistent enough to be checkable? Is the consequence of an error recoverable with a human in the loop? And is the data you would feed it allowed to enter the tool you are using? Sofia's task passes: misrouting is annoying but caught and fixed, and because the emails contain personal information she confirms she is using an approved internal tool rather than pasting tenant data into a public chatbot. If your task fails the data question, stop and fix that before you write a single prompt.

Classify the risk before you build

Screening tells you whether to proceed. Risk classification tells you what obligations follow. Under OMB Memorandum M-24-10, issued in 2024, rights-impacting AI is AI whose output materially affects a person's access to, or loss of, opportunities, rights or benefits. Federal grant decisions affecting individuals or small businesses are rights-impacting. The classification is not a label you apply privately; you document it in your agency's AI use-case inventory with the identifier the memorandum's Section 5 requires, and the classification then triggers the minimum practices in Section 5(c).

The source lists five of those minimum practices, counted from its own enumeration: pre-deployment testing; documentation of intended purpose and reasonably foreseeable misuse; human review of consequential outcomes; consultation with affected communities; and publication of the intended use. Read them as a set rather than a checklist, because each one exists to catch a failure the others miss. This is not bureaucracy for its own sake. It is the lesson Michigan learned expensively with MIDAS, where unemployment fraud allegations were issued by an automated system without adequate human review, resulting in a $20 million settlement and restored benefits to thousands of claimants.

Working the classification on real examples is faster than arguing about the definition. The grid below reproduces the worked classifications the source supplies. Note that a single system can land in more than one category, and that landing in a category is a statement about the effect on a person, not about how sophisticated the model is.

SystemClassification
Michigan MIDAS unemployment fraud detectionRights-impacting
CMS eligibility predictionRights-impacting
HHS vaccine-appointment schedulingLimited-risk
CBP Traveler VerificationRights-impacting and safety-impacting
USDA grant approval (this capstone)Rights-impacting

Stage 2: Prepare the data

You cannot validate what you cannot compare against. Build a small gold set: 25 to 50 real examples you have hand-labeled correctly. Sofia pulled 40 past complaint emails and wrote down the correct category for each. This is the most skipped and most valuable step in the whole lab, because the gold set is the only thing standing between you and a subjective impression that the AI "seems pretty good." Strip or mask anything sensitive that does not need to be there. Sofia removed tenant names and unit numbers, since the category does not depend on who sent the complaint.

At program scale, the same discipline becomes a written data requirement. For the grants case that means naming the application fields, the supporting documents, the historical disposition data and the legal rules that will be encoded as features, and saying where each comes from. Government data is almost always secondhand: it was collected for a different purpose, under a different notice, sometimes under a different authority. Complete or update the privacy impact assessment for the system before production, and confirm that the records you intend to use may lawfully be used this way rather than assuming that possession implies permission.

Write the requirements document

The AI Requirements Document is the artifact your reviewers will actually read, and it is where a vague plan gets caught. It carries the business context, meaning why the agency is doing this, which citizens it serves, and why now. It carries the quantified problem statement, which for the grants case is the 60-day average, the 65 percent first-pass rate, the 5,000 applications a month and the roughly $500,000 annual cost. It carries success criteria: under 20 days, 85 percent first-pass, fairness equal to or better than the current process across demographic groups, and a demonstrable reduction in arbitrary decisions.

It then carries the parts people skip. Data requirements, covering application fields, supporting documents, historical disposition data and legal rules as features. System functionality, covering document extraction, rule application, confidence scoring and human handoff. A validation approach: pre-deployment evaluation on held-out data, fairness tests, adversarial tests and user acceptance. A human-oversight design in which every denial is reviewed by a human, all edge cases route to a human, and a human can override approvals. A monitoring plan for drift, fairness, accuracy and user feedback. And a stakeholder list with each stakeholder's requirements attached.

Name the stakeholders explicitly rather than writing "the usual parties." For the grants case the source names applicants, USDA program staff, the USDA Office of Inspector General, the USDA privacy officer, GAO, OMB and potentially Congress. Each of those has a different question. Applicants want to know how to contest a decision. The privacy officer wants to know what data moved and under what notice. GAO and the inspector general want to know how you would prove any of this after the fact. Writing their requirements down early is cheaper than discovering them during an audit.

Stage 3: Build the solution

Now the part everyone wants to start with, which is genuinely the easy part once framing and data are done. Write a structured prompt that does one job clearly. A reliable classification prompt has four parts, and each part removes a specific failure mode rather than being decoration.

  • Role and task. "You are a housing-authority assistant. Classify the complaint below into exactly one category."
  • The fixed options. List the six categories with one line each defining them, so the model has no room to invent a seventh.
  • The output format. "Reply with only the category name and one sentence of justification." Constrained output is checkable output.
  • A fallback. "If the complaint is ambiguous or fits none well, reply NEEDS HUMAN and say why." This is the safety valve that routes hard cases to a person instead of forcing a wrong guess.

Run the prompt against your gold set one example at a time, and record the model's answer next to your known-correct answer. Sofia did all 40. At federal scale the equivalent step is the internal dry run: in the grants case, the team runs the system against 200 historical applications and compares the output to ground truth before a single live application touches it. Same idea, different blast radius. In both cases the point is that you learn what the system does on your data, not on a vendor's demo data.

Stage 4: Validate honestly

This is where amateurs declare victory and professionals do the work. Compare the model's answers to your gold set and count three separate things, because a single headline number will hide the thing that hurts you.

  • Accuracy. How many did it get exactly right? Sofia got 35 of 40, which is 87.5 percent, against a manual baseline whose stated misrouting rate of one in ten implies about 90 percent. The model started slightly below the humans, not above them.
  • Error pattern. Of the misses, is there a theme? Sofia found that 4 of her 5 errors were safety complaints misfiled as maintenance: a clear, fixable pattern rather than random noise.
  • Error severity. Are the errors harmless or dangerous? A safety complaint routed to maintenance is the worst kind for a housing authority, because safety issues are time-sensitive. The error rate looked ordinary; the error type was alarming.

Severity matters more than raw accuracy, and this is the single most transferable lesson in the lab. Sofia fixed her prompt by adding one line: any mention of fire, gas, mold, structural damage or threat of harm is safety, even if maintenance is also implied. On the re-run, safety errors dropped to zero. Measure, find the pattern, fix, re-measure. That loop, not the prompt itself, is the skill. Note also that her first result did not clear the human baseline, which is exactly the sort of finding a capstone should report rather than bury.

The testing strategy at program scale

For a rights-impacting system, gold-set accuracy is the first test of several. Accuracy testing runs on a representative sample of historical applications, deliberately including the legally difficult cases rather than the clean ones. Fairness testing runs across applicant demographics where those are collected, with attention to disparate-impact patterns. Edge-case testing covers applications with unusual field combinations, missing documentation or foreign-language material. Adversarial testing probes robustness to manipulation. Drift monitoring exists to catch silent degradation after launch, when nobody is watching closely any more.

Anchor the strategy to published references rather than to your own judgement alone. The source directs the testing plan to cite the NIST AI Risk Management Framework Playbook subcategories MEASURE-2.1 through MEASURE-2.11, and, where the system is vendor-supplied, the EU AI Act Article 15 robustness obligations. The framework is voluntary guidance rather than binding law in the United States, which is precisely why citing specific subcategories matters: it turns "we tested it" into a claim a reviewer can check. The classification work earlier in the lab maps to the Playbook's MAP-1.1 through MAP-5.2 subcategories in the same way.

Let real cases shape the fairness tests instead of inventing hypotheticals. COMPAS produced recidivism risk scores with disparate fairness across groups. Houston HISD's EVAAS teacher evaluation was overturned on due-process grounds because its outputs could not be meaningfully challenged. A San Francisco district attorney tool experimented with race redaction and produced useful lessons about what redaction does and does not remove. Each of those tells you which subgroup comparison to run and which explanation you will need to produce when somebody contests a decision.

The 26-week project plan

A capstone plan that says "then we deploy" is not a plan. The source sets out a 26-week federal pilot with five phases, and the phase boundaries are decision points rather than calendar decorations. The weeks below sum to 26: four for discovery, six for the MVP, six for pilot, six for scale-up and four for the steady-state transition.

WeeksPhaseWhat it produces
1 to 4DiscoveryStakeholder interviews, privacy impact assessment, FISMA security categorization, data-sharing agreements where needed, confirmed use-case inventory entry, initial documentation
5 to 10MVP buildModel developed or procured, human-review interface built, logging instrumented, internal dry run on 200 historical applications compared to ground truth
11 to 16Pilot20 percent of incoming applications routed through the AI-assisted flow, strict human review on every denial, weekly fairness and accuracy reports to the Chief AI Officer
17 to 22Scale-upGradual expansion toward 100 percent of applications, human review preserved for denials and edge cases, drift monitored
23 to 26Steady-state transitionOperating procedures finalized, training updated, results published, lessons documented, pilot budget closed out

Staff the plan with named roles, not with a box marked "the team." The source names the Chief AI Officer, the program manager, the vendor lead, the privacy officer, the inspector general liaison and the contracting officer. Government-wide digital service teams have historically supplied surge design and acquisition support to projects like this, and their published patterns remain useful, but do not plan a delivery dependency on a support organisation without first confirming it currently exists and is available to you. A capstone that assumes an unavailable partner is a capstone that fails on feasibility.

Stage 5: Deploy with a human in the loop, and monitor

You do not flip a switch and walk away. Deploy in the safest posture: the AI drafts, a human confirms. Sofia's rollout had the tool suggest a category and her team click to accept or override in about a second instead of reading and sorting in about ninety. The human stays accountable and the AI removes the drudgery. The federal equivalent is stricter, because the stakes are higher: every denial is reviewed by a human, every edge case routes to a human, and a human can override an approval as well as a denial.

Then monitor, because models drift and inputs change. Sofia kept a simple log: each week, how often did staff override the AI, and on what? A rising override rate is an early warning that something has shifted. Set a reminder to re-run your gold set on a schedule; if accuracy slips, you catch it before a member of the public does. Be clear-eyed about the limit of this: monitoring surfaces only what you chose to watch. An override-rate log tells you nothing about a harm your staff never noticed and never overrode.

The measurement framework

Decide the measures before launch, because measures chosen afterwards tend to be the ones that flatter you. The source sets out these key indicators for the grants case: processing time against the under-20-day target; first-pass success against the 85 percent target; human-review rate, set at 100 percent of denials, 100 percent of edge cases and a sample of approvals; applicant satisfaction scores; appeal rates; and cost per processed application. Take a baseline measurement of the current process before deployment, because a baseline reconstructed later is an argument rather than a measurement.

Attribution is the part most pilots skip. The source recommends a quasi-experimental design comparing the pilot cohort to a control cohort where feasible, analysed as a difference in differences, so that you can distinguish the effect of the system from a seasonal swing in application volume. Post-deployment reporting then runs on a cadence: a monthly dashboard to the Chief AI Officer, a quarterly report to OMB, and an annual public-facing report. Publishing on a schedule is itself a control, because a number that will be published in ninety days gets checked more carefully than one that will not.

Brief the decision maker

The plan is worthless if it cannot survive fifteen minutes with an executive. Structure the briefing around six moves: the problem, the proposed solution, the expected outcomes stated quantitatively, the investment and return, the risks and their mitigations, and the decision you are requesting. For an Under Secretary or a congressional staffer this is not a technical presentation. It is a decision brief that answers why we should do this, what could go wrong, and how we will know when we are done.

Three patterns from federal digital service practice make the difference. Show a short live demonstration if you possibly can, because five seconds of the real thing beats five slides describing it. Present one worst-case scenario and its mitigation, rather than hoping nobody asks; the executive who hears the bad case from you trusts the rest of the brief more. And name the decision requested explicitly, in one sentence, so the meeting ends with a decision instead of a follow-up meeting.

Change management and vendor terms

A technically sound system still fails if the people who must use it were not consulted. Plan stakeholder engagement, training and, where employees are affected, coordination with union representatives, early enough that it can change the design rather than merely announcing it. Where the change requires public notice or rulemaking, the applicable transparency and privacy authorities are 5 USC 552, the Freedom of Information Act, and 5 USC 552a, the Privacy Act. Design deliberately for appeal and contestation: a person told no by a system they cannot question is the shape most government AI litigation takes.

On the vendor side, four terms carry most of the weight: FedRAMP authorization for the hosting environment, alignment with ISO/IEC 42001 for the vendor's AI management system, contract performance clauses that tie payment to the accuracy and fairness measures you defined, and an exit strategy that says what happens to your data, your fine-tuned artifacts and your service continuity if the relationship ends. Ask for an evidence package rather than a claim. A vendor tool accepted as a black box is a risk you have taken on your agency's behalf without measuring it.

Validation and explainability are prerequisites

The final piece of the plan is the standing validation apparatus: the QA checklist applied to every release, confidence calibration that maps model confidence to human-review thresholds, bias detection pipelines, monitoring dashboards, the human-in-the-loop pattern, and an incident response path naming who is paged, within what time and with what playbook. Idaho's Medicaid case shows why. An opaque algorithm reduced benefit allocations without explainable justification, litigation followed successfully, and the system was rolled back. Validation and explainability are not nice-to-haves; they are the conditions on which the system is allowed to keep running.

Your deliverable

Complete this one-page project record for your own task. It is the artifact that proves you did the lab and the template you will reuse for every future AI integration. Keep it to a page: the discipline of fitting it on one page is what forces the specificity reviewers are looking for.

  1. Problem statement. One sentence: input, output, current cost, current error rate, target.
  2. Risk classification. Rights-impacting, safety-impacting or neither, with the reason and the inventory entry.
  3. Suitability check. Recoverable errors? Data allowed in this tool? Human stays accountable?
  4. Gold set. Number of hand-labeled real examples; sensitive data masked.
  5. The prompt or system spec. Role, fixed options, constrained output, and a NEEDS HUMAN fallback.
  6. Validation results. Accuracy against baseline; the dominant error pattern; the most severe error type; the subgroup breakdown.
  7. The fix. What you changed and the re-measured result.
  8. Deployment posture. AI drafts, human confirms; who is accountable by name.
  9. Monitoring plan. The signal you watch, your re-test cadence, and who reads the result.
  10. Sunset criteria. The conditions under which you would turn it off rather than tune it.

Sofia's finished record fit on one page. It turned 12 hours a week of sorting into a few minutes of confirming, eliminated the dangerous safety-misrouting error, and gave her something she had never had before: a real, shipped, validated AI integration with her name on it and a record of how it was tested. That is the difference between learning about swimming and getting in the pool.

Anti-Patterns to Avoid

  • Treating a completed plan as containment. A finished capstone document, like a completed checklist, has never once made a wrong output right. The plan is evidence that you thought about the failure modes; it is not proof that you covered them.
  • Declaring the pilot a success and assuming scale-up follows. Pilot-scale success and production-scale success are different claims. A 20 percent pilot with strict human review on every denial says almost nothing about behaviour at 100 percent volume with review capacity stretched.
  • Underestimating privacy and civil-rights review time. Teams routinely budget weeks for model work and days for review, then discover the sequence is reversed. Discovery in this plan is four weeks partly because the reviews are the long pole.
  • Skipping engagement with union representatives where staff are affected. Consultation that arrives after the design is frozen is an announcement, and it converts the people who must operate the system into people who resent it.
  • Failing to design for appeal and contestation. If a denied applicant cannot get a human explanation and a route to challenge it, the system's legal exposure is unbounded regardless of its accuracy.
  • Accepting vendor tools as black boxes. A vendor claim without an evidence package is marketing. Ask for the test data, the subgroup results and the methodology, and treat refusal as a finding.
  • Neglecting drift monitoring. Systems degrade quietly, and the absence of complaints is not the presence of accuracy. Monitoring catches only what you chose to watch, so choose deliberately and revisit the choice.
  • Producing generic documentation. A plan that does not engage with the specific statutory authority governing your program and the specific population you serve will read as competent and land as useless.
  • Judging a model on headline accuracy alone. Sofia's 35 of 40 looked respectable and hid the one error type that could have hurt somebody. Always ask which misses would cause harm, not just how many there were.

Practice Prompts

Work these against your own chosen task rather than Sofia's. The first three are drafting aids; the last two are checks you run on your own output. Never paste real personal information into a tool that is not approved for it.

  • "Here is my task description. Rewrite it as a single sentence containing the exact input, the exact output, the current time cost, the current error rate, and the target. Tell me which of those five I have failed to supply."
  • "Given this task, list the categories or output values the system may produce, with a one-line definition for each, and add an explicit fallback value for cases that fit none of them."
  • "Draft a stakeholder table for this system. For each stakeholder, state the one question they will ask about it and the artifact that answers that question."
  • "Here are my validation results: the counts, the error list and the baseline. Group the errors by type, tell me which type would cause the most harm to a member of the public, and tell me what my numbers do not tell me."
  • "Read this project record as a skeptical inspector general. List every claim in it that is asserted rather than evidenced, and every number whose source is not stated."

Reflection Questions

  • Write your task in one sentence with a baseline and a target. Which part was hardest to fill in, and what does that gap tell you about how well the current manual process is measured?
  • Is your task rights-impacting, safety-impacting or neither? Write the reason in two sentences, as if for the inventory entry, and name the person who would review that judgement.
  • What is the most severe error your system could make? Not the most likely, the most severe. Who would be harmed and how would they find out?
  • If your first validation run came in below the human baseline, as Sofia's did, would your organisation let you report that honestly? What would you do if it would not?
  • Name the condition under which you would turn this system off. If you cannot name one, you have not finished the plan.

Glossary

  • Gold set. A small collection of real examples you have hand-labeled with the correct answer, used as ground truth to measure a system against.
  • Rights-impacting AI. Under OMB M-24-10, AI whose output materially affects a person's access to, or loss of, opportunities, rights or benefits.
  • Safety-impacting AI. AI whose output bears on physical safety; a system can be both rights-impacting and safety-impacting, as with CBP Traveler Verification.
  • Minimum practices. The obligations M-24-10 Section 5(c) attaches to rights-impacting and safety-impacting uses, including pre-deployment testing, documentation, human review, community consultation and publication of intended use.
  • AI use-case inventory. The agency register in which each AI use case is recorded with an identifier, as required by M-24-10 Section 5.
  • Drift. Silent degradation of performance over time as real-world inputs move away from what the system was built and tested on.
  • Difference in differences. An attribution method comparing the change in a pilot cohort against the change in a control cohort, so that outside trends are not credited to the system.
  • Confidence calibration. The mapping from a model's stated confidence to the human-review threshold, so that low-confidence outputs reliably reach a person.
  • Evidence package. The test data, methodology and results a vendor supplies to substantiate a performance or fairness claim, as distinct from the claim itself.

Closing Thoughts

The gap this lab closes is not a knowledge gap. Sofia knew the frameworks before she started; what she lacked was the experience of carrying one thing all the way through, including the uncomfortable middle where the first result came in below the human baseline and she had to decide whether to report it. That decision, more than the prompt or the plan, is what a capstone tests. Everything else in this lesson is scaffolding around it.

Run the five stages on a real task, at whatever scale you actually control. If that scale is 40 emails, run it on 40 emails; the discipline transfers upward far better than enthusiasm transfers downward. Write the one-page record, including the sunset criteria and the results you would rather not show anyone. When the next AI proposal arrives on your desk, you will read it as someone who has been through it once, which is the only reliable defence against a good demonstration.

Key Takeaways

  • Framing is where projects live or die. Write the task as one sentence with a current cost and a target; if you cannot, the fuzziness will become the failure.
  • Classify the risk before you build. Under OMB M-24-10, a rights-impacting classification triggers pre-deployment testing, documentation, human review, community consultation and publication of intended use.
  • Build the gold set before the prompt. Twenty-five to fifty hand-labeled real examples are your ground truth; without them you are guessing.
  • Constrain the output and add a fallback. Fixed options, a required format and a NEEDS HUMAN escape route make the work checkable and route hard cases to a person.
  • Severity beats accuracy. A respectable error rate can hide a dangerous error type; Sofia's 35 of 40 was 87.5 percent against a roughly 90 percent human baseline, and the type of the five misses mattered more than the count.
  • The real skill is the validation loop. Measure, find the pattern, fix, re-measure; that cycle, not the prompt, is what the lab teaches.
  • Plan in phases with decision points. The 26-week structure of discovery, MVP, pilot, scale-up and steady state exists so that each boundary is a go or no-go, not a Gantt milestone.
  • Deploy as draft-and-confirm and monitor deliberately. A named human approves, and monitoring surfaces only what you chose to watch, so choose the signals and the cadence in advance.
  • Specificity is the grading criterion. Completeness, appropriateness, specificity, integration, feasibility and governance are the six criteria; generic plans fail in practice.

Frequently Asked Questions

How big should my gold set be?

The lab uses 25 to 50 hand-labeled real examples, and Sofia used 40. Smaller than that and a single misclassification swings your percentage wildly; larger and the labeling effort starts to compete with the work you are trying to save. Whatever size you choose, the examples must be real cases drawn from your actual queue, and the labels must be ones you would defend, because the gold set is the standard everything else is measured against.

My first validation run scored below the human baseline. Do I abandon the project?

Not automatically, and Sofia's case shows why. Her 35 of 40, or 87.5 percent, came in under the roughly 90 percent implied by a one-in-ten manual misrouting rate. What made the project continue was that the errors clustered in one identifiable pattern that a single prompt line addressed. A below-baseline result with random errors is a different situation from a below-baseline result with a structured, fixable pattern. Report the number either way.

Does the 26-week timeline apply to a small local project?

The durations are drawn from a federal pilot at a grants office processing 5,000 applications a month, so do not transplant the week counts onto a smaller effort. What does transplant is the phase structure and the meaning of each boundary: discovery ends when you know what is feasible, the MVP ends when it has been run against historical cases, the pilot ends when a limited share of live volume has been through it under strict review, and scale-up ends when human review is still holding at full volume.

Who has to review outputs, and how many?

In the federal case the source sets human review at 100 percent of denials, 100 percent of edge cases and a sample of approvals. That is a posture the agency commits to and writes into its own plan, not a universal rule. Your own review rates should be set in advance, recorded, and justified by the severity of the worst error the system can make, then revisited when monitoring shows the error profile changing.

Can I use a public AI tool for the build if I mask the data?

Answer the question at the policy level before the technical one. Sofia confirmed she was using an approved internal tool because the complaints contained personal information, and she masked names and unit numbers in her test set as well. Masking reduces exposure; it does not by itself make a tool approved for the data, and a masking scheme that leaves enough detail to identify someone in a small population has not accomplished what its name suggests.

What makes the difference between a passing capstone and a strong one?

Specificity, according to the six evaluation criteria the source lists. A strong capstone names the statute governing the program, the population served, the individual stakeholders and their questions, the exact measures and their baselines, and the conditions under which the system would be switched off. A weak one describes the same project in language that would fit any agency, which is precisely what makes it useless to the reviewer who has to decide whether it will work here.