Measuring Organizational AI Maturity
The question that stopped Andres Morales cold came from a junior analyst on his team during a budget presentation: "What level are we at?" Andres is the Deputy CIO at a large county health department with 2,700 staff, a $480 million annual budget, and responsibility for AI systems that now touch immunization records, disease surveillance, care management referrals, and public health emergency response. He had been presenting a slide deck to the county board that described the department's AI journey in terms of tools deployed, pilots completed, and staff trained. The analyst's question was pointing at something different: not what the department had done, but where it stood on a meaningful scale of organizational capability. Andres did not have an answer. The board did not push back on the omission, because they did not know the right question to ask either. But the analyst's question stayed with him, and three months later Andres conducted his department's first formal AI maturity assessment.
This lesson is about how to run that assessment, what it can and cannot tell you, and how to use the result to advance rather than merely to measure. The distinction matters more than it sounds. Most maturity assessments in government are commissioned as reassurance and consumed as a slide, and the whole value of the exercise sits in what happens after the report is filed.
What AI Maturity Actually Measures
An AI maturity model is a structured framework for assessing an organization's current capabilities across the dimensions that determine whether it can deploy AI responsibly and effectively at scale. Maturity is not the same as ambition, activity, or investment. An agency can spend $10 million on AI tools and still be immature if the governance infrastructure, data quality, and workforce capability to use those tools responsibly are absent. A smaller agency with solid data governance, trained staff, and a working oversight structure may be more mature at a fraction of the cost.
It is equally important to be clear about what a maturity score does not establish. A tier describes the state of your practices. It is not a finding about any particular system. An organization assessed at a high tier can still be running a model that is inaccurate, poorly matched to its task, or producing disparate outcomes, because maturity measures whether you have the machinery to notice such things, not whether you have noticed them. Treat the assessment as a statement about the department's capacity to govern AI, and keep the per-system evidence separate. Leaders who conflate the two use a strong maturity result to justify skipping a specific system review, which is exactly backwards.
The Dimensions That Carry the Weight
The dimensions that matter most in a government AI maturity model are typically five, and they are chosen because weakness in any one of them can invalidate strength in the others.
- Data governance. The quality, documentation, lineage and access controls on the data that AI systems consume. This is the dimension that determines whether anything downstream can be trusted.
- AI governance. The policies, oversight structures, decision rights and accountability mechanisms that govern how AI is approved, deployed, monitored and retired.
- Technical infrastructure. The systems and architectural patterns that make AI integration possible, including how models reach data and how their outputs reach the people who act on them.
- Workforce capability. AI literacy and skill across staff functions, and specifically the gap between the people who select AI systems and the people who use them daily.
- Equity and accountability. The mechanisms for monitoring, auditing and correcting disparate impacts, and for explaining a decision to the person it was made about.
Each dimension is scored on a scale running from initial, meaning ad hoc and undocumented, to optimizing, meaning systematically improved on the basis of measured outcomes. The overall maturity profile, which is to say the shape of which dimensions are advanced and which are lagging, is what determines where to invest next. The average across dimensions is the least useful number the assessment produces, because averaging is exactly the operation that hides the gap you most need to see.
Maturity Assessment Instruments
Several assessment frameworks exist for government contexts, and they are not interchangeable. Choosing one is a decision about what you want the result to be comparable to.
The NIST AI Risk Management Framework tiers. The framework describes implementation tiers that function as a rough maturity scale: partial, where risk management is not formalized; risk informed, where risk practices exist but are not integrated organizationally; repeatable, where practices are consistent across functions and updated regularly; and adaptive, where the organization responds dynamically to emerging risks and learns from operational experience. The framework is voluntary and non-binding, and the tiers provide a quick orientation rather than a granular basis for investment decisions. Their advantage is that everyone in federal work recognizes the vocabulary.
The GAO AI accountability framework. The Government Accountability Office's framework for evaluating federal agency AI accountability is organized around four principles: governance, data, performance and monitoring. Each principle carries specific attributes that can be assessed as present, partially present or absent. This framework matters particularly for federal agencies and for state and local bodies administering federal funds, because GAO uses it in its own audits. Understanding it before a GAO engagement is both preparation and substantive good practice. Note that it is an accountability framework rather than a maturity model, so it tells you whether the evidence exists rather than how developed the capability is.
Custom agency assessments. Many large agencies adapt an existing framework to their own context, adding domain-specific dimensions relevant to their programs. Andres's department built a custom assessment that added a public health emergency response dimension, measuring how well the department's AI systems could be activated, modified or suspended during a declared emergency. No general-purpose maturity model covers that, and for a health department it is central to the mission. The cost of customizing is comparability: the more you tailor, the harder it becomes to benchmark against anyone else. Andres's compromise was to keep the standard dimensions scored in the standard way and add his own dimension alongside rather than folding it into the others.
Conducting the Assessment
A maturity assessment is only useful when it is conducted rigorously, as a structured evaluation with evidence rather than as a self-certification exercise. Andres's department used a three-part process. First, a document review covering existing AI policies, data governance documentation, training records and audit logs. Second, interviews with staff across functions, deliberately including program managers, data analysts, IT staff and frontline workers rather than only the leadership who commissioned the work. Third, a structured comparison of documented practices against the assessment framework's criteria, scored against the evidence gathered rather than against what people believed was true.
That third step is where most assessments quietly fail. When the people accountable for a score are the people assigning it, the exercise stops being a measurement and becomes a negotiation, and the reliable symptom is a report with no bad news in it. The structural fix is not complicated: separate who gathers the evidence from who is accountable for the result, require a documented artifact for every score above the lowest tier, and make an unevidenced claim score as absent rather than as partially present. Interviewing frontline staff matters for the same reason. The distance between the documented process and the actual practice is usually the finding.
The assessment took eight weeks and produced a 22-page report. The overall finding placed the department in the risk informed tier, with significant variation across dimensions. Data governance was the strongest area, reflecting three years of sustained investment in data quality and documentation. AI governance was the weakest: there was no formal AI oversight structure, no incident reporting protocol, and no equity monitoring program for any of the four deployed AI systems. Workforce capability sat in between, and it split in a revealing way. Senior staff had reasonable AI literacy, while frontline staff had received no training at all on the AI systems they used every day. That gap, between the people who choose the systems and the people who operate them, is one of the most common findings in government assessments and one of the cheapest to close.
It is worth sitting with what one of those findings actually meant. The department had four AI systems in operation, covering immunization records, disease surveillance, care management referrals and emergency response, and not one of them had equity monitoring attached. That is not a documentation gap. It means that for four systems touching health services in a diverse county, nobody could say whether outcomes differed across the populations served, in either direction. The department was not managing a known small disparity; it was operating without the instrument that would detect a large one. Findings of that shape are the reason a maturity assessment is worth eight weeks of staff time.
The same reading applies to the missing incident reporting protocol. Its absence does not mean the department had no incidents. It means that if one occurred, there was no defined path by which it would reach the people able to suspend a system, no record that would exist afterwards, and no way to distinguish a one-off from a pattern. An assessment that records this as a governance score of two is being polite. Stated plainly, it is a finding that the department's ability to respond to an AI failure depended on whoever happened to notice it and on their judgment about whom to tell.
Reading the Profile Rather Than the Score
The single number is the part of the report everyone quotes and the part that carries the least information. Andres's department scored 2.4 on a five-point scale, and by itself that number cannot tell you whether the department is in reasonable shape with one serious hole or in uniformly mediocre condition throughout. Those two profiles average to the same figure and demand completely different responses.
Read the shape instead. Ask which dimension is lowest and whether that dimension gates the others, because data governance and AI governance both act as gates: weak data undermines every model built on it, and weak AI governance means nobody is positioned to detect the failure. Ask where the variance sits inside a dimension, since a department can have excellent governance for one flagship system and none for the other three. And ask which findings would change if a different person had been interviewed, which is the fastest test of whether a score reflects the organization or the respondent.
Benchmarking Against Peers
A maturity score is most useful in context. Knowing that your department scores 2.4 on a five-point scale is far less informative than knowing that peer agencies of comparable size and budget score between 1.8 and 2.6, with the leading agencies at 3.1. Benchmarking supplies that context, and it converts an abstract number into a budget argument that a board can act on.
Formal benchmarking data for government AI maturity is limited but growing. The National Association of State Chief Information Officers publishes annual surveys that include some AI maturity indicators, and state and local government associations increasingly include AI maturity questions in their own technology surveys. Andres's department participated in a voluntary county health department benchmarking consortium of twelve departments across six states, which shared anonymized assessment results. The consortium found that his department's data governance score sat above the consortium median while its AI governance score was the second lowest in the group. That single comparison shaped the investment priorities in the department's next budget request more effectively than any internal argument had.
Two cautions about benchmarks. Being above the median tells you about the peer group, not about adequacy; in a field this young, an above-median score can still describe a department with no incident reporting protocol. And comparisons only hold if everyone scored the same way, which is why consortium arrangements that publish their scoring guidance are worth far more than surveys that simply collect self-reported tiers.
Advancement: Investing in the Right Gaps
A maturity assessment is not a report card. It is a prioritization tool, and the right investment is rarely in the dimension where the agency is already strongest. It is in the dimension where weakness creates the highest risk to mission performance, equity or accountability. The instinct runs the other way, because the strong dimension has an owner who can describe what more money would buy, while the weak dimension usually has no owner at all. That absence is the finding, not an obstacle to acting on it.
For Andres's department, the weakest dimension was also the highest-risk gap. Without an incident reporting protocol, a significant failure in the disease surveillance system would surface first as a public health crisis rather than as a managed operational incident. Without equity monitoring, the care management referral system's potential for disparate impact was unknown and unmanaged, which is a materially different position from knowing that the impact is small. The first post-assessment investments were therefore not in data infrastructure, where the department already sat above the peer median, but in an AI oversight function staffed at one full-time equivalent costing $120,000 annually, and an equity monitoring protocol developed in partnership with the county's civil rights office, which took 90 days and $18,000 in staff time.
Note the shape of those two investments. Both were small, both were specific, and both created a capability the department did not previously have rather than improving one it already had. That is what a maturity assessment is good at producing: not a transformation programme, but a short list of missing pieces that can be funded inside a normal budget cycle. A department that emerges from an assessment with a multiyear roadmap and no immediate action has misused the instrument.
Tracking Progress Over Time
Maturity is not a destination. It is a position on a moving scale, with both the agency and the field advancing at the same time, which means a department can improve its practices and lose ground relative to its peers in the same year. Annual reassessment using the same instrument, so that year-over-year comparison means something, is the mechanism that turns maturity measurement from a one-time exercise into a management tool.
Andres's department committed to annual cycles beginning the year after the initial baseline. At the first reassessment, overall maturity had advanced from 2.4 to 2.9, driven by the AI governance investments, and the equity dimension had advanced from 1.2 to 2.8, the largest single-dimension gain. The pattern there is worth naming, because it recurs: a structured investment in a well-defined gap produces rapid movement, precisely because scoring near the bottom of a scale usually means the capability is absent rather than weak, and building something where nothing existed moves the score further than improving something that already works.
Guard the comparison itself. The same instrument scored by different people, or by the same people with a year more understanding of what the criteria mean, is not a clean comparison, and interpretation drift can produce movement that looks like progress. Keep the scoring guidance written down, keep at least one assessor consistent across cycles, and record the evidence behind each score so that next year's assessor is arguing with an artifact rather than with a number.
Making the Cycle Survive Its Sponsor
The final risk is organizational rather than analytical. A maturity programme that depends on one interested deputy CIO ends when that person moves, and government tenure being what it is, the second assessment is the one most likely never to happen. Andres's protection against that was to bind the cycle to things that persist: the assessment date sits ahead of the budget submission so its findings feed the request, the oversight function he funded owns the instrument rather than his personal office, and the consortium participation creates an external commitment that a successor would have to actively withdraw from.
None of that is elaborate, and that is the point. Institutionalizing a measurement cycle is mostly a matter of attaching it to a process that already recurs, naming an owner whose job description includes it, and leaving behind enough written scoring guidance that a stranger could run the next cycle. The alternative is what the analyst's question exposed in the first place: a department describing its progress in tools deployed and pilots completed, because nobody had ever established a scale on which to describe it any other way.
Anti-Patterns to Avoid
Scoring yourself. When the officials accountable for a dimension assign its score, the assessment becomes a negotiation with a predictable outcome. Separate evidence gathering from accountability, require an artifact for every score above the bottom tier, and treat an unevidenced claim as absent rather than partial.
Treating the tier as a safety finding. Maturity describes whether you have the machinery to detect problems. It does not establish that any deployed model is accurate, appropriate or equitable. A strong maturity result is never a reason to shorten a system-level review, and leadership will make that leap unless you close it explicitly.
Managing the average. The overall score is the number that gets quoted and the one that hides the gap. Two departments with identical averages can face completely different risks. Report the profile, and put the lowest gating dimension at the top of the summary.
Benchmarking as reassurance. Above the median in a young field can still mean no incident reporting and no equity monitoring. A benchmark tells you about the comparison group; it says nothing about adequacy, and it should never be used to close a gap discussion.
Assessing once. A single assessment is a baseline and nothing more. Without a committed second cycle, the report becomes a historical document, and the effort spent producing it bought a slide rather than a management instrument.
Investing where you are already strong. The strong dimension has an articulate owner with a funding proposal ready; the weak dimension often has no owner at all. If the budget follows the loudest advocate, the assessment will have identified the right gap and funded the wrong one.
Practice Prompts
1. Score one dimension honestly. Pick the dimension you believe is your organization's weakest. Write the score you would defend publicly, then list the specific artifacts that support it. Where an artifact does not exist, mark the criterion absent rather than partial and see what the score becomes.
2. Find the frontline gap. Identify one AI system your organization operates. Ask a person who uses it daily what training they received and what they do when they disagree with its output. Compare their answer with the training records and the documented procedure.
3. Design your custom dimension. Name the capability that is central to your mission and absent from general-purpose maturity models, as emergency activation and suspension was for Andres's department. Draft the scoring criteria for its lowest and highest tiers, and decide whether it will be scored alongside the standard dimensions or folded into them.
4. Build the investment case. Take your lowest-scoring dimension and write the smallest specific investment that would create a capability you do not currently have. State its cost, who would own it, and what evidence would show at the next reassessment that it worked.
5. Stress-test the second cycle. Assume you leave your role next month. Write down who would run the next assessment, what written guidance they would use, when it would happen, and what would stop it from quietly not happening. Fix whichever of those four answers is weakest.
Reflection
If someone asked your leadership team what level your organization is at with AI, what would they say, and on what scale? Consider which of the five dimensions you would be least comfortable having assessed by an outsider with access to your documentation and your frontline staff, and what that discomfort tells you about where the real gap sits. Then consider the harder question that Andres faced: if the assessment had come back reassuringly, would you have run it again the following year, or would a good result have quietly ended the practice that produced it?
Glossary
- AI maturity model. A structured framework for assessing an organization's capabilities across the dimensions that determine whether it can deploy AI responsibly and effectively at scale.
- Maturity profile. The pattern of scores across dimensions, as distinct from the overall average. The profile, not the average, indicates where to invest.
- Implementation tier. A named position on a maturity scale, running in the NIST framework from partial through risk informed and repeatable to adaptive.
- Baseline assessment. The first assessment in a series, whose primary value is the comparison it enables rather than the score it produces.
- Benchmarking consortium. A voluntary group of comparable organizations that share anonymized assessment results against agreed scoring guidance.
- Gating dimension. A dimension whose weakness undermines strength elsewhere, such as data governance, where poor data quality compromises every model built on it.
- Equity monitoring. The ongoing measurement of whether an AI system produces disparate outcomes across populations, and the mechanism for acting when it does.
- Interpretation drift. Change in how scoring criteria are understood between assessment cycles, which produces apparent movement in scores without any change in practice.
Related Lessons
This lesson sits at the point where measurement meets governance. AI Maturity Assessment covers the assessment mechanics at a more introductory level and is a useful companion for the staff who will do the evidence gathering. Your Agency's AI Governance Structure and Establishing an AI Governance Board address the dimension that scores lowest most often, and they describe what an oversight function like the one Andres funded actually does. Data Governance for AI covers the gating dimension in depth, Workforce Planning for AI addresses the frontline capability gap, and GAO AI Accountability: Four Principles in Practice takes the accountability framework further. AI Audit Preparation is the natural next step for any organization whose assessment revealed thin evidence.
Closing
The analyst's question was better than the presentation it interrupted. Tools deployed and pilots completed describe activity; they do not describe capability, and they give a board no way to judge whether the department is ready for what it is already doing. A maturity assessment supplies that scale, and its value is entirely in what follows: a profile read for its shape rather than its average, a benchmark that provides context without providing comfort, one or two specific investments in the gaps that gate everything else, and a second cycle that actually happens. Andres's department moved from 2.4 to 2.9 in a year, but the durable result was not the number. It was that the department now had a way to answer the question, and an owner responsible for answering it again.
Key Takeaways
- Maturity is not the same as investment or activity. A department that has spent heavily on AI tools but lacks governance infrastructure may be less mature than a smaller one with solid oversight structures and trained staff.
- A tier describes your machinery, not your systems. A strong maturity result says you are positioned to detect problems. It is never evidence that a particular deployed model is accurate, appropriate or equitable.
- Assess five dimensions and read the profile. Data governance, AI governance, technical infrastructure, workforce capability, and equity and accountability. The variation across them, not the average, tells you where to invest.
- Use evidence, not self-certification. Document review, cross-functional interviews including frontline staff, and scoring against artifacts. When the accountable official assigns the score, the assessment becomes a negotiation.
- Benchmark for context, not for comfort. Peer comparison converts a score into a budget argument, but being above the median in a young field can still describe a department with no incident reporting protocol.
- Invest in the weakest gating dimension, not the strongest. The highest-value investment creates a capability you do not have, and it is usually small, specific and fundable inside one budget cycle.
- Commit to the second cycle before you finish the first. Annual reassessment with the same instrument, consistent scoring guidance and at least one consistent assessor is what makes a comparison meaningful and guards against interpretation drift.
- Know the GAO accountability framework before your first audit. Its four principles of governance, data, performance and monitoring are the headings your evidence will be judged under, and mapping to them in advance is preparation rather than protection.
Frequently Asked Questions
Which maturity instrument should we use? Choose based on what you want the result to be comparable to. If your peers and your auditors speak in the NIST tiers, use those, because a score nobody else recognizes is hard to act on. If you need to know whether your evidence would survive an audit, work against the GAO accountability principles instead, remembering that it assesses whether evidence exists rather than how developed a capability is. Most organizations end up with a standard core plus one or two custom dimensions for mission-specific capabilities, which is a reasonable compromise as long as the standard dimensions stay scored in the standard way.
How long should an assessment take, and who should run it? Andres's department completed its assessment in eight weeks and produced a 22-page report, which is a reasonable shape for a first baseline in an organization of that size. What matters more than the duration is the separation: the people gathering evidence should not be the people accountable for the scores. That can mean an internal team from a different part of the organization, a peer reviewer from a consortium partner, or an external assessor. What does not work is the leadership team scoring its own dimensions in a workshop.
What if our score comes back low and leadership reacts badly? Frame the assessment as a prioritization tool before it runs, not after the result arrives. A low score in a young field is the expected finding, and the useful comparison is against peers rather than against an abstract ideal. It also helps to arrive with the score and the first proposed investment in the same conversation, so the discussion moves immediately to what a small, specific, fundable fix would create. A result that produces a defensive reaction and no action means the assessment was positioned as a judgment rather than as a plan.
Can we compare this year's score to last year's if we changed the instrument? Not reliably, and it is better to say so than to publish a movement you cannot defend. If you must change the instrument, score the new cycle both ways for one year so you have a bridge, or state plainly that the series restarts. The same caution applies to changes of assessor and to drift in how criteria are interpreted, which is why the scoring guidance and the evidence behind each score need to be written down and kept.
How does this apply to a small agency with only one or two AI systems? The dimensions still apply; the depth changes. A small organization is unlikely to need a dedicated oversight function, but it still needs to know who approves an AI deployment, what happens when the system fails, whether anyone is checking for disparate outcomes, and whether the people using the system were trained on it. Score those questions honestly and invest in whichever answer is missing entirely.
Skill.re