←
AI for Government
Visionary · M8 · lesson 8 of 46 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI in Education
📖
now learning

AI in Education

15 min

Dr. Renata Acevedo runs the office of teaching and learning for a state department of education that oversees 1.1 million students. In August, a vendor offered her district a "personalized learning AI" that promised to lift reading scores by adapting to each child. By October, a parent in one of her smallest rural districts filed a complaint: her dyslexic son had been routed into a remedial loop the system never let him exit, and no teacher could explain why. Renata had signed the contract. Now she had to explain it to the state board, the parent, and a reporter, all in the same week.

That sequence is the real shape of AI in public education. The promise is genuine and so is the risk, and the two arrive on the same purchase order. As an agency leader, your job is not to be excited or afraid. It is to know which uses are ready, which are dangerous, and how to put guardrails around both before a parent, an auditor, or a court does it for you. The complaint is coming either way. What you control is whether you can answer it.

Where AI actually earns its place in education

Strip away the marketing and AI does four useful things in a school system. It personalizes practice. It removes administrative load from teachers. It widens access for students with disabilities. And it gives workforce programs a faster path from skill gap to credential. Each carries a different risk profile, and you should fund them in the order that risk suggests rather than the order the market pushes. The market pushes the highest-risk category first, because it is the one that sells.

K-12 personalization, done honestly

Adaptive practice tools, the kind that adjust math problem difficulty in real time, are mature. They work best as a supplement a teacher controls, not a track a child gets locked into. The failure in Renata's district was not the algorithm. It was the absence of an off-ramp. A defensible deployment gives every teacher a one-click override and shows, in plain language, why a student was placed where they were placed.

Both halves of that sentence need scrutiny before you accept them as protections. An override that teachers have not been trained on, are not permitted to use without approval, or are quietly discouraged from using is not an off-ramp. It is a button. Log how often it is exercised and treat a district with zero overrides as a warning rather than a success. Likewise, an explanation shown on a screen is only useful if it reflects what actually drove the placement. Ask the vendor to demonstrate that by changing one input and showing you the explanation change with it.

Watch the equity math. If a tool needs reliable home internet to do homework, it will quietly widen the gap between a student with fiber and one using a phone hotspot. Before you scale, run the numbers on who actually has access, by district and by income band. Do that before procurement rather than after, because once the tool is embedded in daily instruction the finding that many of your students cannot use it at home becomes a political problem rather than a design input.

Higher education efficiency

At the campus level, the biggest wins are administrative: drafting financial-aid award letters, triaging the 40,000 routine emails an admissions office gets each cycle, flagging students who have stopped logging into the learning system so an advisor can call before they drop out. A mid-size public university that automates first-draft transcript evaluations can cut processing from three weeks to three days. That is a real service improvement with low civil-rights exposure, which is exactly why it should come first.

One caution inside that category. Flagging a student who has stopped logging in is administrative only as long as the flag reaches a human who calls. The moment it feeds an automated status change, a funding consequence, or a note in a permanent record, it has crossed into a different risk class and needs the safeguards that class carries. Draw that line explicitly in the requirements, because vendors will happily extend a low-risk tool into a high-risk one as a roadmap item and describe it as the same product.

Special education, the highest-stakes use

AI can transcribe meetings for an Individualized Education Program, draft goal language, and translate documents for families. These save overworked case managers real hours, and the hours are not trivial when caseloads are already past what the staffing model assumed. But an Individualized Education Program is a legal document under federal disability law. AI may draft; a credentialed human must own every word. Never let a model make a placement or eligibility decision. That is the line that turned Renata's pilot into a complaint.

Treat that rule as absolute rather than as a default with exceptions, because the exceptions are where the harm lives. The pressure to relax it will not arrive as a proposal to let the model decide. It will arrive as a workflow in which the model's recommendation is pre-filled, the human signature is the last field on the form, and the caseload makes disagreement expensive. Design against that pressure by requiring the human to enter the determination rather than confirm it, and by making the model's draft visible as a draft rather than as a completed document awaiting approval.

Workforce development

For state labor and workforce agencies, AI can match a laid-off worker's existing skills to in-demand local jobs and recommend the shortest credential to close the gap. The risk is a model trained on historical hiring data that steers women toward lower-wage tracks, reproducing the pattern in the training data and presenting it as a recommendation. Test for that before launch, not after. The test is not complicated: look at what the system recommends for otherwise similar workers who differ on a protected characteristic, and look at whether the recommended credentials differ in wage outcome.

Workforce matching has a second failure that is easy to miss because it looks like accuracy. A model trained on which credentials people actually completed will favor the credentials that are easiest to complete, which correlates with cost, schedule, and proximity rather than with earning power. It will therefore recommend the short cheap certificate to the worker who most needs the longer one, and it will be right about completion and wrong about the outcome the program exists to produce. Decide which outcome you are optimizing for and check that the system agrees with you.

The frameworks already governing your decision

You are not writing rules from scratch. Three bodies of guidance already apply, and an auditor will expect you to know them. Knowing them is also the fastest way to get a vendor conversation onto solid ground, because a vendor who has never heard of any of them is telling you how much public sector work they have actually done.

The NIST AI Risk Management Framework, the federal standards body's voluntary playbook, gives you four plain-language functions: Govern, Map, Measure, Manage. In school terms: assign who is accountable, list where the tool touches a student's rights, decide how you will measure harm, and set how you will respond when it appears. It is voluntary and non-binding, which people sometimes read as optional. Read it instead as the structure an auditor will use to organize their questions whether or not you adopted it.

If you receive federal funds, the Office of Management and Budget memo M-24-10, issued in 2024 on advancing governance, innovation, and risk management for agency use of artificial intelligence, matters. It calls any AI that affects a person's access to education or benefits rights-impacting and requires specific safeguards: an impact assessment, human fallback, and the ability to opt out. A student-placement engine is rights-impacting. A tool that summarizes faculty meeting notes is not. Sort your tools into those two buckets first, in writing, before you sort anything else.

Finally, Section 508 of the Rehabilitation Act requires that technology you buy be accessible to people with disabilities. An AI tutor a blind student cannot use with a screen reader is not a smaller problem to fix later. It is a procurement violation. Note what this means for evidence: a vendor accessibility report is the vendor's own assessment of their own product. It is a starting point, not a finding. Test the tool with the assistive technology your students actually use, with people who actually use it, before you accept the report as an answer.

Doing the sort before you do anything else

Everything downstream depends on one act of classification, so do it deliberately and put it on paper. For each tool, name the decision it touches, name the person that decision lands on, and state whether the output changes that person's access to education, services, or opportunity. If the answer is yes, the tool is rights-impacting and carries the heavier obligations. If the answer is no, it is administrative and can move faster. The sort is not difficult. What makes it fail is that nobody is assigned to do it and it therefore happens implicitly, at purchase, by whoever wrote the requisition.

Two habits make the sort durable. First, classify the tool as deployed rather than as described, because the same product configured two ways lands in two different buckets. An early-alert system that emails an advisor is administrative; the identical model wired to a registration hold is not. Second, reclassify on change. Vendors extend products, integrations get added, and a tool that was correctly classified in January can be doing something different by September without anyone re-examining it. Put reclassification on the same schedule as your drift review so the two questions get asked together.

The output of the sort is not a spreadsheet for its own sake. It is the document you hand an auditor, a board member, or a reporter when they ask which of your systems make decisions about children. Having that document ready is a substantially different position from assembling it in the week a complaint lands, and the difference is visible in every conversation that follows.

Student data, minors, and the records question

Every tool in this lesson runs on data about children, and much of it sits in or near the education record. The specific obligations that attach to that data in your jurisdiction come from your counsel and your records officer, and this lesson deliberately states none of them. What it can give you is the procurement posture that keeps you out of trouble regardless of which specific rules apply: know exactly what data leaves your systems, know exactly who can see it, and know exactly what the vendor is permitted to do with it once it is theirs.

The contract term to insist on is the one in the checklist below, barring secondary use of student data. Be clear-eyed about what a contract term is, though. It is a remedy, not a control. It gives you standing to act after a disclosure, and it does nothing mechanical to prevent one. Pair it with technical limits on what is transmitted in the first place, because the data you never sent is the only data that cannot be misused. When a vendor argues that a broad data feed is necessary for model quality, treat that as a negotiation about your exposure, not as a technical fact.

A readiness checklist before you sign anything

This is the artifact Renata wished she had used in August. Run every proposed education AI tool through its eight questions and require written answers before procurement, not after. Written matters. A verbal assurance from a sales engineer in a demo has no owner three months later, and the person who gave it will not be on the account.

QuestionWhy it mattersPass condition
Is this tool rights-impacting?Triggers the OMB M-24-10 safeguardsClassified in writing; impact assessment if yes
Can a teacher override every decision?Prevents the locked-loop failureOne-click override, logged, with override rates reviewed
Can it explain a placement in plain language?You must answer parents and auditorsPer-student reason on demand, demonstrably tied to the inputs
Is it accessible under Section 508?Legal requirement and equity issueTested with real assistive technology, not only a vendor report
Who can the data be sold or shared with?Student privacy law (FERPA)Contract bars secondary use, plus limits on what is transmitted
What is the equity baseline?Tools can widen access gapsAccess and outcome data by income and disability status
How will we monitor for drift?Accuracy degrades over a school yearQuarterly outcome review scheduled and owned
What is the human fallback if it fails?A non-AI path must existDocumented manual process that staff have actually run

Run a pilot that can fail safely

The biggest mistake leaders make is the district-wide launch. Pilot in two or three sites for one semester, with a named educator owner, a control group, and a pre-defined stopping threshold. Decide in advance what result ends the pilot: for example, if any subgroup of students shows worse outcomes than the control group by more than a set margin, you stop. Write that margin down before you start, so the decision is not a political fight later when a vendor relationship and a superintendent's reputation are attached to the answer.

Writing it down is necessary and not sufficient. Pre-committed thresholds get renegotiated at exactly the moment they bind, which is why the second half of the design matters as much as the first: name who has the authority to stop, put that authority somewhere other than the office that bought the tool, and require that any change to the threshold be documented with a reason. A threshold that only the purchasing office can invoke or revise is a threshold that will be revised.

Budget the unglamorous parts. A district that spends $400,000 on licenses and $0 on teacher training and monitoring has bought a liability, not a tool. A useful rule of thumb: for every dollar of license cost, plan for at least an equal amount in training, integration, and oversight in year one. That ratio is a planning heuristic rather than a measured constant, and the reason to hold it is not precision. It is that the line item you underfund is always the one that would have caught the problem.

What your monitoring will and will not catch

The checklist asks how you will monitor for drift, and the honest answer in most districts is that nobody has decided. Accuracy degrades over a school year for ordinary reasons: the students change, the curriculum changes, a state assessment changes, a teacher changes how the tool is used in one building. None of these announce themselves. A model that was well calibrated in September can be quietly mismatched by March while producing outputs that look exactly as confident as they did before, because confidence is a property of the output rather than a measurement of correctness.

Be precise about the limits of monitoring rather than treating it as a general safety net. Monitoring detects the things you chose to watch, at the interval you chose to watch them, in the populations large enough to show a signal. It is structurally poor at catching harm concentrated in a small subgroup, which is exactly where the harm in education tends to concentrate. A remedial loop trapping a handful of dyslexic students in one rural district will not move a district-wide accuracy metric at all. That is not a flaw in your dashboard. It is the reason the complaint channel and the teacher override log are part of your monitoring rather than adjacent to it.

Design the review to answer three questions on a fixed schedule: has the outcome gap between groups moved, are overrides rising or falling and where, and what did the complaints say. Assign the review to a person, put it on the calendar, and give the reviewer somewhere to escalate. A quarterly review with no owner and no escalation path is a line in a contract, not a control, and it will be discovered to have lapsed at the least convenient moment.

What Renata does differently now

The complaint resolved, eventually, in the way these things do: the student was moved, the district apologized, and the contract was amended. What changed permanently was the sequence. Renata's office now sorts every proposed tool into rights-impacting or not before it reaches procurement, funds the administrative and accessibility categories first, and requires the eight written answers above from any vendor who wants a pilot. The dyslexic student in the remedial loop is the reason, and she says so out loud in every vendor meeting, because a specific child is harder to argue with than a principle.

Anti-patterns to watch for

Buying personalization first because it demos best. Adaptive learning is the category vendors lead with and the category with the highest civil-rights exposure. Administrative automation and accessibility work saves more staff time per dollar, carries less legal risk, and generates the operational track record that makes the harder deployments defensible later. Funding in demo order rather than risk order is how a district ends up explaining an algorithm to a reporter before it has ever successfully run one.

Treating the override as the safeguard. An override button satisfies a requirement on paper and protects nobody unless teachers know it exists, are permitted to use it without seeking approval, and face no professional cost for disagreeing with the system. Measure usage. If overrides are near zero across a large deployment, the most likely explanation is not that the model is right every time.

Accepting a generated explanation as a reason. Systems can produce fluent, plausible accounts of a decision that have no reliable relationship to what actually drove it. That is worse than no explanation, because it manufactures confidence in front of a parent who has no way to check it. Require that explanations be demonstrably derived from the decision, and test that by changing an input and watching whether the explanation follows.

Letting the model pre-fill a legal determination. Nobody proposes letting AI decide eligibility. What actually happens is a workflow where the recommendation is already in the field, the human step is a signature, and the caseload makes disagreement costly. The determination should be entered by the person accountable for it, with the model's draft visible as a draft. Anything else converts a credentialed judgment into a confirmation click.

Piloting without a stopping rule you do not control. A pilot with no pre-defined failure threshold always succeeds, because success gets defined after the results are in. A pilot whose threshold can be revised by the office that bought the tool always succeeds too. Put the stopping authority somewhere else and require a written reason for any change to it.

Confusing a signed data clause with data protection. Contract language gives you recourse after a disclosure and prevents nothing at the moment of transmission. Pair every clause with a technical limit on what actually leaves your systems, and treat a vendor's claim that a broad feed is required for model quality as a negotiating position about your exposure.

Practice prompts

  1. Take every AI tool currently in use or under consideration in your agency and sort each one, in writing, into rights-impacting or not. For any tool you cannot confidently place, write down what you would need to know. That list is your first procurement conversation.
  2. Pick one deployed tool and run the eight checklist questions against it retroactively. Note every question you cannot answer from existing documentation, and treat those as open findings rather than as history.
  3. Draft the equity baseline for one proposed tool: which access and outcome measures you will collect, broken down how, and what result would cause you to stop.
  4. Write the stopping rule for a pilot you are planning. Name the threshold, name the person who can invoke it, name where that person sits, and specify what documentation a change to the threshold requires.
  5. Script the three questions you will ask the next vendor about accessibility, and decide in advance which answers you will accept as evidence and which you will treat as a self-assessment.

Reflection

These are worth answering honestly rather than quickly, and most of them have an uncomfortable answer that is more useful than the comfortable one. Take them to the people who run the tools day to day rather than answering from the office, because the gap between how a deployment is described upward and how it actually works in a classroom is where most of the risk in this lesson lives.

  • Which tools in your system are currently making or shaping decisions about individual students without anyone having classified them as rights-impacting?
  • If a parent called tomorrow and asked why their child was placed where they were placed, who in your organization could answer, and how long would it take?
  • Where in your current deployments would a teacher who disagreed with the system face more friction than a teacher who went along with it?
  • What percentage of your AI spending this year went to training, monitoring, and oversight rather than to licenses, and what does that ratio say about what you expect to happen?
  • Which category are you funding first, and is that because of risk or because of what demonstrated well?

Glossary

  • Adaptive practice. Software that adjusts exercise difficulty or sequence from a student's responses. Mature as a teacher-controlled supplement; risky as a track a student cannot leave.
  • Rights-impacting AI. A system whose output affects a person's access to education, benefits, or opportunity, triggering heightened safeguards including impact assessment, human fallback, and opt-out.
  • Human fallback. A documented non-AI path to the same outcome that staff have actually practiced.
  • Individualized Education Program. A legal document governing services for a student with a disability. AI may assist in drafting; a credentialed human owns the content and the determination.
  • Equity baseline. Access and outcome measures, disaggregated by group, collected before deployment so later comparisons mean something.
  • Drift. Degradation of a model's accuracy as students, curricula, and conditions change away from its training data.
  • Stopping rule. A failure threshold defined before a pilot begins, together with the named person authorized to invoke it and the documentation required to change it.
  • Secondary use. Any vendor use of student data beyond delivering the contracted service, including model training, product development, or transfer onward.

Closing

The honest summary of AI in public education is that the safest uses are the least exciting and the most exciting uses are the ones that produce complaints. That is not an argument against the exciting ones. It is an argument for earning them: build the operational record on transcript evaluation and accessibility and administrative triage, learn what your own monitoring actually catches, and arrive at student-facing personalization with a governance apparatus that already works rather than one you are assembling under pressure.

Renata's October is available to anyone who signs in August without the eight answers. What separates the leaders who avoid it is not caution and not technical depth. It is the willingness to do the sorting before the buying, to fund the oversight alongside the license, and to keep a specific child in mind while reading a contract. The student in the remedial loop was not a statistical outcome. He was the point.

Key takeaways

  • Sequence by risk, not by hype. Fund administrative automation and accessibility first; treat placement and eligibility tools as the highest-stakes, last-to-scale category.
  • Every decision needs a human off-ramp that people can actually use. An override nobody is trained on or permitted to use freely is a button, not a safeguard; watch override rates and treat zero as a warning.
  • Classify each tool as rights-impacting or not. That single written sort determines which safeguards, impact assessments, and opt-outs you owe.
  • AI drafts, credentialed humans own. For an Individualized Education Program and any legal document, the model may write a first draft, but a licensed professional enters the determination and is accountable for it.
  • Test explanations and accessibility rather than accepting reports. A generated explanation can be a plausible reconstruction, and a vendor accessibility report is a self-assessment; verify both against reality.
  • Run the equity and access math before scaling. Tools that assume home internet or fail screen readers quietly widen the gaps you are paid to close.
  • Pre-commit to a stopping rule and put it out of reach. Define the failure threshold before the pilot starts and give the authority to invoke it to someone other than the office that bought the tool.
  • Budget for oversight, not just licenses. Training, monitoring, and integration are the cost of doing this safely; the line you underfund is the one that would have caught the problem.

Frequently Asked Questions

How do I decide whether a tool is rights-impacting?

Ask whether the output affects an individual student's access to education, services, or opportunity. A placement engine, an eligibility screen, or anything feeding a permanent record is rights-impacting. A meeting summarizer or email triage tool is not. Boundary cases are usually tools that start administrative and acquire consequences later, so classify as deployed and reclassify when outputs start driving a status change.

Can AI write an Individualized Education Program?

It can draft language, transcribe meetings, and translate documents for families, all of which save case managers real time. It cannot own the content and it must never make a placement or eligibility determination. Build the workflow so the accountable human enters the determination rather than confirms a pre-filled one, and keep the model's output visibly a draft. The failure mode is not a policy decision to automate; it is a form design that makes disagreement expensive.

Is a vendor's accessibility report enough evidence for Section 508?

No. It is the vendor's own assessment of the vendor's own product, which makes it a starting point for a conversation rather than a finding you can rely on. Test the tool with the assistive technology your students actually use and with people who actually use it. Accessibility failures tend to surface in workflows rather than in features, which is exactly what a feature-level self-assessment will miss.

What should a pilot's stopping rule look like?

It needs four parts: a specific measured outcome, a threshold set before the pilot begins, a named person authorized to invoke it, and a documentation requirement for any change. The fourth is usually missing and decides whether the rule survives contact with a vendor relationship. Put the stopping authority outside the purchasing office.

How much should I budget beyond licenses?

The rule of thumb in this lesson is at least an equal amount for training, integration, and oversight in year one. Treat that as a planning heuristic rather than a measured constant; your actual figure depends on how much integration your systems need and how much of your monitoring is manual. The reason to hold roughly to it is that a district spending on licenses and nothing on oversight has bought a liability rather than a tool.

What do I do about student data leaving my systems?

Know what leaves, who can see it, and what the vendor may do with it, and get all three in writing. The specific legal obligations attaching to student records in your jurisdiction come from your counsel and records officer, and nothing in a training lesson substitutes for them. Operationally, remember that a contract clause is a remedy rather than a control, so pair it with a technical limit on what is transmitted in the first place.