←
AI for Government
Strategic · M24 · lesson 24 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Emerging AI Technologies Assessment
📖
now learning

Emerging AI Technologies Assessment

15 min

Alejandra Montes-Reyes chairs the technology evaluation committee for a mid-sized state department of corrections: 14,000 staff, 48 facilities, and a fiscal year 2026 IT budget of $62 million. Every quarter her committee receives an average of nine vendor demonstrations for AI-based products, including predictive risk assessment tools, automated grievance routing, AI-assisted reentry planning, computer vision for facility monitoring, and natural language systems for translating staff reports. Every vendor calls its product "proven," "state-of-the-art," and "specifically designed for corrections environments." Her committee has approved exactly two tools in the past eighteen months, both through a structured evaluation process. This lesson is about how to make those calls reliably.

Assessing emerging AI technology is a distinct skill from assessing AI risk or managing AI vendors. It asks a narrower question: given a capability that did not exist in usable form two years ago, is it real yet, is it real for us, and is now the moment? Four practices answer that question, and this lesson works through each in turn: evaluation frameworks for new capabilities, hype cycle analysis, adoption timing decisions, and horizon scanning.

The evaluation problem in government

Government agencies are not early adopters by design. Procurement rules, budget cycles, and accountability structures all favor proven solutions over experimental ones. That is mostly appropriate. But it creates a particular difficulty with AI: the field moves faster than the traditional evaluation cycle, vendor claims routinely outpace demonstrated performance, and the agencies with the most to gain from AI adoption, those with the highest caseloads, the most legacy system debt, and the fewest technology staff, are also the ones least equipped to evaluate unfamiliar technology claims.

The result is a predictable pattern with two failure modes. Agencies either adopt too early, deploying tools that are technically functional but operationally unproven in government contexts, and then absorb failures that generate inspector general reports and media coverage. Or they adopt too late, waiting for a certainty that never fully arrives, while operational problems AI could have addressed continue accruing cost. Neither pattern is inevitable. Both are the result of having no structured evaluation framework, which means each decision is argued from scratch by whoever is in the room.

The structural point worth naming early is that a framework does more than improve individual decisions. It makes decisions defensible. A committee that can explain why it said yes to one tool and no to another, using the same criteria for both, can withstand a vendor's escalation to a legislator, a program director's frustration, and an auditor's later question about due diligence. A committee working case by case on instinct cannot, even when its instincts were correct.

Hype cycle analysis: where is this technology, really?

The Gartner Hype Cycle is a widely used framework describing a trajectory that new technologies commonly follow through five phases: an innovation trigger, a peak of inflated expectations, a trough of disillusionment, a slope of enlightenment, and a plateau of productivity. It is a descriptive model rather than a law of nature, and the value of using it is not prediction. It is that placing a specific capability on the curve forces you to separate what the technology can do from how loudly it is currently being discussed, which are the two things a vendor demonstration is designed to blur together.

AI capabilities in 2026 sit at radically different points on this curve depending on the specific capability, which is why "AI" is far too coarse a unit of evaluation. Large language models for text summarization are at or near the productivity plateau for many government use cases. Predictive risk assessment in corrections sits in the trough of disillusionment following a wave of academic criticism and state-level moratoriums, with genuine open questions about whether the plateau will ever be reached for high-stakes applications at all. That last hedge matters and should be preserved when you brief leadership: not every technology completes the curve, and some should not.

Position on the curve translates directly into how much skepticism a vendor's claims warrant. A tool at the peak of inflated expectations is being sold on trajectory rather than evidence, so the burden of proof on performance claims should be at its highest. A tool at the plateau has an accumulated body of deployment experience you can go and check, so the interesting question shifts from "does it work" to "does it work here." This is the first filter in Alejandra's process, and it costs nothing to apply. It happens before anyone schedules a demonstration.

Evaluating the three kinds of vendor claim

Vendors make three categories of claims, and government evaluators need a different kind of evidence for each. Performance claims assert a capability level, as in "97 percent accuracy." Deployment claims assert adoption, as in "proven in 40 government agencies." Context claims assert fit, as in "designed specifically for corrections environments." Committees get into trouble by accepting evidence for one category as though it settled another, which is exactly what a well-built demonstration encourages. A polished demo is a context claim dressed as a performance claim.

Performance claims require evaluation results produced on data comparable to your agency's, not the vendor's curated demonstration dataset. Deployment claims require verifiable references at comparable agencies, with permission to discuss specific performance metrics rather than general satisfaction, because "we are happy with it" survives contact with almost any product. Context claims require documentation of training data composition: if a tool was trained primarily on data from large urban facilities, its performance in rural or specialty facilities may be substantially lower, and no amount of configuration will close a gap that originates in the training distribution.

Claim typeWhat the vendor saysEvidence that actually answers itWhat does not answer it
PerformanceA stated accuracy, precision, or throughput levelEvaluation results on data comparable to yours, with the evaluation method describedResults on a vendor-curated demonstration set, or a headline metric with no denominator
DeploymentIn production at a stated number of agenciesNamed reference agencies of comparable type and size, with permission to discuss specific metricsLogos on a slide, or references who may only confirm general satisfaction
ContextPurpose-built for your domainDocumentation of training data composition and the settings it was drawn fromA domain-specific user interface, or a configuration option named after your agency type

Alejandra's committee turns this into one artifact: a standard evidence request form sent to every vendor before any demonstration is scheduled. It asks for third-party evaluation reports, comparable deployment references with contact permission, and training data documentation. Vendors who cannot or will not supply those do not advance to the demonstration stage. Of the nine submissions in a typical quarter, three to four survive the evidence request. The form does most of the committee's filtering work before anyone spends an hour in a conference room, which is the entire point of putting it first.

One caution about third-party evaluation, because it is the strongest single piece of evidence on the list and therefore the easiest to over-trust. A report is third-party in the sense that the vendor did not write it. It is usually not independent in the sense that the vendor did not choose the evaluator, define the scope, or decide whether the result got published. Read what was measured and on what data before you read the conclusion, and treat an evaluation whose scope you cannot see as weaker evidence than one whose scope you can, however impressive its letterhead.

Adoption timing: is your agency ready?

A tool that performs well in a well-resourced agency with clean, structured data and a mature AI governance function may perform poorly in an agency with fragmented legacy systems, inconsistent data quality, and no designated AI oversight. Technology readiness and organizational readiness are separate questions, and both have to be answered affirmatively before adoption makes sense. The common error is answering the first and assuming the second, usually because the first is the one the vendor is prepared to discuss.

The most frequent version of that error is evaluating a tool in isolation from the data environment it will run in. A reentry planning system trained on structured case management data will produce meaningless output if your case management data is 40 percent incomplete or inconsistently coded, which is the actual state of many corrections case management systems. Before evaluating any AI tool, assess whether your underlying data is in a condition where the tool could plausibly work. If it is not, the honest answer may be to invest in data quality first and revisit the tool in 18 to 24 months.

Organizational readiness is harder to assess because there is no vendor to ask. The questions worth writing down before any go decision are concrete: who will own this system after the implementation team disperses, what happens to it when that person leaves, which existing process does it replace rather than sit alongside, who has authority to pause it, and what does the agency do on the days it is unavailable. An adoption decision that cannot answer those is not a technology decision yet. It is an aspiration with a purchase order attached.

Higher-stakes applications warrant a longer bar. Risk assessment in corrections, benefit eligibility in social services, and enforcement prioritization in regulatory agencies all involve decisions where a wrong output has consequences a reversal cannot fully undo. For those, the standard Alejandra applies is meaningful deployment history in comparable contexts, on the order of 18 to 24 months, before adoption is considered at all. The cost of premature deployment in those settings is institutional, legal, and human, in roughly that order of visibility and the reverse order of importance.

Horizon scanning: what is 12 to 18 months away?

Adoption timing is not only a question about current readiness. It is also a question about what is coming. A tool that is best in class today may have significantly better successors in development, and an agency that commits to a multi-year contract for a specific capability without understanding the near-term development pipeline can find itself locked into an inferior tool 14 months into a 36-month contract, with two years left to run and no clean exit.

Horizon scanning does not require a research team. It requires a structured process: designate a function responsible for monitoring published research, peer agency deployments, and vendor roadmaps for the capability categories most relevant to your agency. Report to leadership quarterly. Flag capabilities approaching deployment readiness in contexts like yours. Done consistently, this can give procurement planning 12 to 18 months of lead time, so that contract timing follows technology maturity instead of following whichever vendor happened to demonstrate most recently.

Two disciplines keep horizon scanning honest. The first is writing down what you expected, because a scanning function that only records what happened cannot tell whether it is any good at reading the field. The second is treating vendor roadmaps as marketing documents rather than forecasts. A roadmap tells you what a company wants you to believe is coming and when, which is genuinely useful information about the company's intentions and close to worthless as a delivery date. Weight peer agency deployments and published evaluation results far more heavily, because those describe what has already happened rather than what is planned.

A usable artifact: the emerging technology assessment scorecard

The four practices work best as an ordered gate sequence rather than a set of considerations weighed together, because the order determines how much of the committee's time each submission consumes. Cheap filters go first. Placing a capability on the maturity curve costs a reading afternoon; a structured demonstration and reference check costs a week of several people's attention. Running them in the wrong order means spending the expensive resource on submissions the cheap one would have removed. Score each row in sequence and stop at the first gate a submission cannot pass.

GateQuestionPass conditionIf it fails
1. Maturity positionWhere does this specific capability sit on the curve?Placed, with the evidence for the placement written downNot a decline; a raised evidence burden, and a note in the horizon scanning file
2. Evidence requestDid the vendor supply evaluation results, references, and training data documentation?All three supplied and inspectableDoes not advance to demonstration; written decline naming the missing item
3. Data readinessCould this tool plausibly work on our data as it stands?The data it depends on is complete and consistent enough to support the functionInvest in data quality first and revisit later; the tool is not the problem
4. Organizational readinessWho owns, pauses, and covers for this system?Named owner, defined pause authority, and a documented downtime procedureDefer until the operating model exists, regardless of how good the tool is
5. Stakes tierWhat happens to a person if the output is wrong?Deployment history in comparable contexts proportionate to the stakesDecline for now, and record what evidence would change the answer
6. Horizon checkWhat is likely to be available before this contract ends?Contract length and exit terms fit the expected pace of the categoryRenegotiate term and exit rights before award, not after

The sixth gate is the one committees skip most often, because by the time a submission has cleared five gates everyone in the room wants to say yes. It is also the only gate that asks about time rather than about the product, and it is the one that determines whether a correct decision today is still a correct decision in the third year of the contract. Ask it out loud, and record the answer, because it is the gate a future auditor is most likely to ask whether you considered.

When the honest answer is "not yet"

A structured process produces a third outcome besides yes and no, and it is the one agencies handle worst. "Not yet" means the capability is real, the fit is plausible, and something on your side is not ready. It is a genuinely different decision from a decline, and it needs to be recorded differently, because the thing that makes it useful is that it names a condition. A no with no condition attached simply recurs next quarter with a slightly different slide deck.

The most common form is a data verdict. If the records a tool depends on are substantially incomplete or inconsistently coded, no configuration will fix it, and buying the tool first in the hope that it will drive the data cleanup inverts the dependency. Write the deferral as a condition instead: the capability is approved in principle once the underlying data reaches a stated state, with a named owner for that remediation and a revisit date. That converts a frustrating decline into a funded project with a clear finish line, which is a much easier conversation with the program office that wanted the tool.

The second form is a stakes verdict. A capability may be perfectly mature for a low-stakes use and clearly premature for a high-stakes one, and the same vendor will often propose both in the same meeting. Splitting the decision is legitimate and usually the strongest available answer: approve the routing, the translation, or the summarization, and defer the risk scoring. It gives the agency real operational benefit, gives the vendor a foothold that motivates continued engagement, and keeps the consequential decision on the far side of a much higher evidence bar.

The third form is an organizational verdict, and it is the hardest to say out loud because it is about the agency rather than the product. There is no owner for the system after implementation, or no one with authority to pause it, or no plan for the days it is down. Naming that plainly is more useful than any technical objection, because it identifies work the agency can actually do. It is also the deferral most likely to be overturned by a leader who wants the tool, which is exactly why it belongs in writing, with the specific missing role named, before the meeting where that pressure arrives.

The two tools that got approved

Of Alejandra's two approved tools, one was an AI-assisted translation system for staff incident reports: a mature capability with well-documented performance, no high-stakes decision authority, and a clear fallback in human translation for any report flagged as ambiguous. The other was an automated grievance routing system that classified incoming grievances by topic and directed them to the appropriate response team, with human review required before any routing decision was finalized. Neither tool decides anything about a person. Both accelerate a step in front of a decision a person still makes.

Both cleared the evidence request stage with third-party evaluation data. Both had verifiable references at comparable state corrections agencies. Both had training data that included corrections-specific documentation. Both were deployed in facilities where the underlying data was structured enough to support the tool's function. Neither was adopted at the peak of vendor enthusiasm; both were evaluated only after the technology had been operating in comparable contexts for at least 18 months. That is not a slow evaluation process. It is a disciplined one, and both tools remain in service with no reported incident.

"No reported incident" is the right phrase and it is deliberately narrower than "no problems." An agency sees the failures its monitoring was built to surface. A translation system that quietly degrades on a dialect the reporting staff use, or a routing system that systematically misroutes one grievance category, will not announce itself. The follow-through obligation on any approved tool is to define, before launch, what failure would look like in the data, and then to actually look. Approval discipline at the front door means very little without observation discipline afterward.

Running the committee so the answers hold up

The evaluation process is itself a thing that has to be designed, and three choices carry most of the weight. First, the criteria are written down and applied in the same order to every submission, so that a decision can be reconstructed later by someone who was not in the room. Second, the evidence request precedes the demonstration rather than following it, because a committee that has already watched a compelling demonstration is negotiating against its own impression when it asks for proof. Third, declines are explained in writing to the vendor, in terms of the specific evidence that was missing.

That last practice looks like courtesy and is actually self-defense. A written decline that names the gap gives the vendor a route back with the evidence rather than a route around the committee to a program director or a legislator. It also builds the record. Over several quarters, the accumulated declines describe exactly what the market is failing to supply to agencies like yours, which is the most useful market intelligence a committee produces and the part most often thrown away.

Approval discipline earns credibility with two audiences at once. Staff who watched a committee decline a well-marketed product will take its approvals seriously, which is most of what determines whether an adopted tool actually gets used. Oversight bodies that can see consistent criteria applied consistently treat the committee as a control rather than as a formality. A committee that approves whatever arrives has neither audience, and a committee that approves nothing eventually loses the mandate to review anything.

Anti-Patterns

  • Evaluating "AI" instead of a capability. Treating every AI product as occupying the same position on the maturity curve means either blanket caution or blanket enthusiasm, and both are wrong most of the time. Name the specific capability, place that capability on the curve, and evaluate against its own maturity rather than the field's reputation.
  • Scheduling the demonstration first. Once a committee has seen a polished demonstration it is arguing against its own impression when it later asks for evidence. Put the evidence request ahead of the demonstration so the filtering happens on documents rather than on presentation quality.
  • Treating a third-party report as independence. The vendor did not write it, but the vendor may well have selected the evaluator, scoped the test, and decided whether to publish the result. Read the method and the dataset before the conclusion, and downgrade any evaluation whose scope you cannot inspect.
  • Evaluating the tool and not the data. A capable system fed fragmented, incomplete, or inconsistently coded records produces confident output that means nothing. Assess data readiness first, and be willing to conclude that the correct next investment is data quality rather than a product.
  • Mistaking a pilot's success for readiness to scale. A pilot runs in the facility with the best data, the most engaged staff, and the implementation team on site. None of those conditions survives expansion. Ask which of the pilot's advantages are structural and which were temporary before treating the result as generalizable.
  • Declining without a written reason. An unexplained no reads to a vendor as an invitation to route around the committee, and it leaves the agency with no record of what the market failed to provide. Write the gap down, send it, and keep it.
  • Approving and then not watching. A tool that has produced no reported incident may be working, or may be failing in a way nobody instrumented for. Define what failure would look like in the data before launch, and schedule someone to look.

Practice Prompts

  • Place three capabilities on the curve. Take three specific AI capabilities your agency is being offered. For each, state where you believe it sits between innovation trigger and productivity plateau, and what evidence would move your placement. Note which of the three you cannot place, and what you would need to read to place it.
  • Draft the evidence request. Write the standard form your agency would send to every vendor before scheduling a demonstration. Specify what performance evidence, what deployment references, and what training data documentation you require, and decide in advance what happens when a vendor supplies two of the three.
  • Assess the data first. Pick one AI tool your agency is considering. Before evaluating the tool at all, describe the completeness, consistency, and structure of the data it would depend on. Then state honestly whether the tool could plausibly work on that data as it stands today.
  • Test organizational readiness. For the same tool, answer in writing: who owns it after implementation, what happens when that person leaves, which existing process it replaces, who may pause it, and what the agency does on the days it is unavailable.
  • Stand up horizon scanning. Name the function that will monitor research, peer deployments, and vendor roadmaps for your three most relevant capability categories, define the reporting cadence to leadership, and write down what you currently expect for each category so that you can check yourself later.

Reflection

Take twenty minutes. Think about the last AI product your agency declined, or the last one it adopted. Could you reconstruct the reasoning today from what was written down at the time, or would you be reconstructing it from memory and inference? If it is the second, that is the gap this lesson is really about, and it is a documentation gap rather than a judgment gap.

Then consider the pressure that runs in the other direction. Evaluation discipline is expensive in the currency agencies have least of, which is leadership attention and elapsed time, and its benefits are almost entirely invisible: the failed deployment that never happened, the contract you did not sign, the news story that did not run. How would you make that invisible value legible to a secretary or a legislator who is asking why the agency is behind its peers on AI adoption, and what would you say if the honest answer is that being behind was the correct decision?

Glossary

  • Hype cycle. A descriptive model of the trajectory many technologies follow through an innovation trigger, a peak of inflated expectations, a trough of disillusionment, a slope of enlightenment, and a plateau of productivity. Useful for placing a capability, not for predicting one.
  • Performance claim. A vendor assertion about capability level, such as an accuracy or throughput figure. Answered only by evaluation results produced on data comparable to yours.
  • Deployment claim. A vendor assertion about adoption, such as a count of agencies in production. Answered only by verifiable references at comparable agencies with permission to discuss specifics.
  • Context claim. A vendor assertion of fit for your domain. Answered only by documentation of what the system was trained on and in what settings.
  • Evidence request. A standard form sent to every vendor before any demonstration is scheduled, specifying the evaluation reports, references, and training data documentation required to advance.
  • Technology readiness. Whether the capability itself is mature enough to work reliably in settings like yours.
  • Organizational readiness. Whether your agency's data, staffing, ownership, governance, and processes can support the tool once the implementation team leaves. A separate question from technology readiness, and independently capable of sinking a deployment.
  • Horizon scanning. A standing function monitoring published research, peer agency deployments, and vendor roadmaps for relevant capability categories, reporting on a fixed cadence so procurement timing can follow technology maturity.

Closing

The committee Alejandra chairs is not an obstacle to innovation, though it is described that way roughly once per quarter by someone whose product it declined. It is the mechanism by which her department can adopt AI at all, because it is the only thing standing between a compelling demonstration and a multi-year contract for a capability that has not been shown to work in a facility like hers, on data like hers, for people like the ones in her custody.

The four practices in this lesson are not complicated. Place the capability on the maturity curve before you look at the product. Demand the specific evidence that answers each specific kind of claim, and demand it before the demonstration. Assess your own data and your own organization as seriously as you assess the tool. And keep a standing eye on what is a year or so out, so that contract timing serves technology maturity rather than the other way around.

What makes them work is applying them consistently, including to the products you want to say yes to. A framework used only on tools the committee is already skeptical of is not a framework. It is a justification, and everyone in the room will know the difference within about two quarters.

Key Takeaways

  • Evaluate capabilities, not "AI." Different AI capabilities sit at radically different points of maturity, and some may never reach a productivity plateau for high-stakes public sector uses. Name the capability before evaluating the product.
  • Hype cycle position is the first filter. A tool at the peak of inflated expectations is being sold on trajectory; a tool at the plateau has deployment history you can go and check. Position sets the burden of proof.
  • Match evidence to claim type. Performance claims need results on comparable data, deployment claims need references with permission to discuss specifics, context claims need training data documentation. A demonstration answers none of the three.
  • Put the evidence request before the demonstration. Filtering on documents is cheaper and more honest than filtering on presentation quality, and it protects the committee from arguing against its own impression.
  • Third-party is not the same as independent. Read the method and the dataset before the conclusion, and discount any evaluation whose scope, evaluator selection, or publication decision you cannot see.
  • Evaluate your data, not just the tool. A technically capable system will fail on fragmented, incomplete, or inconsistently coded records. Sometimes the correct next investment is data quality, with the tool revisited later.
  • Organizational readiness is a separate go or no-go. Ownership after implementation, process replacement, pause authority, and downtime handling all have to be answered before adoption, and no vendor can answer them for you.
  • High-stakes uses need a longer bar. Risk assessment, benefit eligibility, and enforcement prioritization warrant substantial deployment history in comparable contexts before adoption is considered at all.
  • Horizon scanning buys procurement lead time. A modest standing function watching research, peer deployments, and roadmaps lets contract timing follow technology maturity instead of vendor scheduling.
  • No reported incident is not the same as no harm. Define what failure would look like in the data before launch, and assign someone to look, or approval discipline stops at the front door.

Frequently Asked Questions

How do we place a capability on the hype cycle without a research budget?

You do not need original research; you need three readable signals. Look for published evaluations conducted by someone other than the vendors, look for deployments at agencies comparable to yours that have been running long enough to have had problems, and look for whether the serious criticism of the capability is about implementation details or about whether the approach works at all. Criticism of the second kind is the clearest indicator you are looking at a trough rather than a plateau, and it is usually easy to find once you look for it deliberately.

What if a vendor refuses to provide training data documentation on confidentiality grounds?

Treat it as an answer rather than an obstacle. You are not asking for the data; you are asking for a description of what it contained and which settings it came from, which is disclosable without exposing anything proprietary. A vendor who cannot describe the composition of its training data at that level either does not know or does not want you to know, and both make its context claim unverifiable. Record the refusal in your written decline so the reasoning is on file.

Is it ever right to adopt something at the peak of inflated expectations?

Occasionally, and the conditions are specific. It is defensible when the use is low stakes, the decision authority remains with a person, the failure mode is visible rather than silent, there is a working fallback, and the commitment is short enough that being wrong is survivable. Alejandra's two approvals both meet the low-stakes and fallback conditions. What is rarely defensible is meeting none of those conditions and adopting anyway because the demonstration was impressive and a peer agency announced something similar.

Our leadership says we are falling behind peer agencies. How do we respond?

Ask which specific capability, at which specific agency, in production for how long, and with what published result. That question is not rhetorical; it frequently reveals that the comparison is to an announcement rather than a deployment. Where the comparison turns out to be real, it is genuinely useful, because a peer agency running the capability is exactly the deployment reference your evidence request asks vendors for and cannot always get. Either way, the answer improves your position.

How does this differ from ordinary vendor evaluation?

Ordinary vendor evaluation asks whether this supplier can deliver this product on these terms, and assumes the product category is established. Emerging technology assessment asks a prior question: is this category real enough to buy from at all, and is now the moment for us specifically. Skipping that prior question is how agencies run a rigorous procurement, select the strongest vendor by every criterion, and still end up with a tool that could not have worked for anyone, because the whole category was not ready and no vendor in it would have said so.