AI Maturity Assessment
Constance Leblanc-Moreau had been the Chief Data Officer for the Louisiana Department of Children and Family Services for eleven months when her new secretary asked her to present the department's AI maturity level at a legislative briefing. Constance drafted a one-pager describing the department's data warehouse, its two in-production predictive risk models, and its governance policy. She felt reasonably confident going in. The first question from a committee staffer stopped her cold: "What does your maturity score actually mean? Can you define the scale?" Constance could not. She had described activities, not maturity. She had listed what the department owned, not what it could reliably do. She spent the week after the briefing learning the difference, and learning how to use three assessment frameworks that would have prevented that moment.
You have built a strategy. Now you need to assess honestly where you actually stand. Not where you hope to be. Not where you claimed to be last year. Where you are right now. Many agencies skip this step because they assume they are starting from a blank slate and will build the AI program from scratch. That is rarely accurate. You probably have scattered AI work already happening, partial data systems, people with some expertise, and institutional constraints you have not fully acknowledged. This lesson is about seeing all of that clearly enough to plan against it.
What Maturity Means and Why It Is Not Activity
Maturity is not a list of things your agency has done. It is a measure of what your agency can reliably do, repeatedly and consistently at operational scale, across the dimensions that determine whether AI succeeds in practice. An agency that has deployed two AI pilots is not necessarily more mature than an agency that has deployed none, if the pilots are undocumented, ungoverned, and unmonitored. The pilots are activities. Maturity is the underlying capability that makes activities sustainable, and the distinction is the whole point of the exercise.
Think of it like a hospital's surgical capability. A hospital that has performed 20 open-heart surgeries with excellent outcomes has demonstrated capability. A hospital that has performed 3 surgeries with inconsistent documentation, no infection-control protocol, and no outcomes tracking has demonstrated activity. The maturity question is whether you can do this reliably at scale, with consistent governance, and with warranted confidence that the outcomes will hold when the conditions change.
For government AI, maturity spans strategy and governance, data and technical infrastructure, organizational capability, responsible AI practices, and outcomes measurement. An agency can be strong in one and weak in another. Mature overall capability requires reasonable strength across all of them: not perfection in each, but no critical gap in any. Be clear about what a level is, though. It is a description an agency produces about itself using a published scale. No body certifies an agency at Level 3, and a level is only as good as the evidence behind the claim.
Why Maturity Assessment Matters
Maturity frameworks exist because organizations do not all start from the same place. A federal agency with 50 data scientists is in a completely different position from a local government with 2 people and legacy databases. A ministry with pilots running for two years faces different constraints from one starting at zero. Knowing where you actually stand lets you set roadmaps that do not assume you can move faster than you can, identify which capability gaps are blocking you most, benchmark against organizations of similar size, allocate resources to the highest-impact gaps, and avoid spending effort on capability you do not need yet.
Government organizations face maturity assessment challenges the private sector does not. You have multiple stakeholders with genuinely different views of what maturity means: IT leadership cares about infrastructure, program staff care about usability, executives care about outcomes, compliance cares about governance. You have legacy systems and decades-old data infrastructure that are partly broken and expensive to replace. You have limited ability to hire quickly, constrained by budgets and pay scales. You operate in a highly regulated environment where moves must be documented, auditable, and defensible. And AI competes with many other priorities for the same resources.
This means the assessment has to be brutally honest about constraints, and it means you cannot simply assess yourself against private-sector scales. Assess relative to what government organizations actually look like. An agency that scores itself poorly against a commercial benchmark and concludes it is failing has learned nothing useful; an agency that identifies which of its own constraints is binding has learned exactly what to do next.
Three Frameworks You Need to Know
Three frameworks are most commonly referenced in government AI maturity assessment. They are complementary rather than competing, and using them together gives a more complete picture than any one alone. Start with the GSA model if you are a government organization, since it was designed for that context. Supplement with MITRE for granular governance assessment. Use Gartner when the question you most need answered is whether your strategy and your ability to execute it are actually connected.
The MITRE AI Maturity Model
MITRE's model, developed with federal agency input, spans four domains: AI leadership and culture, data and technology, talent and workforce, and governance and accountability. Each domain is scored on five levels. Its governance treatment is more detailed than most competing frameworks, which is why it is worth running even when another model is your primary instrument. The five levels are named Initial, Managed, Repeatable, Optimized, and Advanced, and the descriptors differ by domain.
| Domain | Level 1 Initial | Level 2 Managed | Level 3 Repeatable | Level 4 Optimized | Level 5 Advanced |
|---|---|---|---|---|---|
| Governance and accountability | Informal AI decision-making; no governance structures; pilots scatter with no coordination | Basic governance established; steering committee exists; decisions get documented | Processes defined and followed; AI project review standards exist | Processes optimized based on data; continuous improvement happens | Governance is predictive; the organization learns from patterns and adjusts proactively |
| Data | Data scattered; no governance; quality unknown; access ad hoc | Some data inventory exists; basic governance principles documented; quality issues identified | Data governance implemented; master data managed; quality standards exist and are measured | Data treated as an asset; pipelines optimized; lineage tracked | Data continuously optimized; predictive quality management; data serves all needs efficiently |
| Technology | Limited technical capability; no ML platforms; ad hoc AI work | Pilot-stage infrastructure; some tools adopted; limited integration | Standardized ML platforms exist; repeatable development processes | Enterprise ML platforms in place; continuous integration and deployment for models | Advanced platforms enable real-time learning; infrastructure fully optimized |
| Talent and culture | Limited awareness; no training; skepticism is high | Basic awareness; some training; early champions exist | AI literacy growing; training programs in place; change management works | Organization understands AI broadly; career paths exist; culture is adaptive | Organization proactively adopts AI; continuous learning embedded; culture is innovation-focused |
Read across a row and the pattern becomes obvious: the difference between Level 2 and Level 3 is almost never "more" of something. A department at Level 2 in data has established some quality standards but applies them inconsistently. At Level 3 it has documented standards, applies them across most programs, and can measure compliance. The difference is consistent process and measurement, not volume. That distinction matters enormously for planning, because a system built on Level 2 data infrastructure will produce Level 2 results: sometimes good, sometimes not, with limited ability to predict which in advance.
The Gartner AI Maturity Framework
Gartner's model, widely used across sectors, has five levels. Awareness means leadership understands AI exists and is relevant. Active means initial projects are underway. Operational means AI is used routinely in some areas with repeatable processes. Systemic means AI is integrated across multiple functions with organization-level governance. Transformational means AI is a core component of how the organization operates and improves. Those five levels correspond roughly to the vocabulary government executives already use when discussing AI progress, which makes the model effective for executive briefings and legislative testimony.
Underneath the levels, Gartner assesses three dimensions: AI strategy and governance, meaning whether you have a clear strategy and whether governance actually supports it; data and AI foundation, meaning whether you can execute projects with the data and technical foundation you have; and talent and culture, meaning whether you have the people, skills, and organizational culture to sustain the work. The framework's particular usefulness is that it forces the connection between strategy and execution. You cannot be mature on strategy while your foundation and talent lag, because the strategy will not survive contact with them.
The GSA AI Capability Maturity Model
The General Services Administration's AI Capability Maturity Model was designed specifically for civilian federal agencies and has been widely adapted by state agencies. It is explicitly built for the government context, accounting for budget constraints, acquisition processes, federal compliance obligations, and typical government organizational structure. Its dimensions cover vision and strategy, governance, data strategy, technical infrastructure, workforce and skills, operations and management, and continuous learning. Counting that enumeration gives seven dimensions, which is the widest coverage of the three models and the reason it works well as a primary instrument.
Two features distinguish it. First, it separates capability, meaning what the agency can do, from deployment, meaning what the agency has actually done, and treats them independently. An agency can hold strong capability it has not yet deployed, or weak capability it has deployed recklessly, and both facts matter to a planner. Second, it includes a readiness-for-responsible-AI scoring dimension that tests specifically whether governance practices meet minimum standards for rights-impacting and safety-impacting applications, which maps directly onto the obligations government agencies already carry.
Running Your Own Assessment
A maturity assessment is not an annual report. It is a candid internal exercise whose value depends entirely on its honesty. Constance ran her department's assessment over four weeks using a structured questionnaire drawn from all three frameworks. The sequence below combines the mechanics of scoring with the discipline that keeps the scoring truthful, and the first step matters more than any of the scoring rules.
- Assemble a cross-functional team. Do not let IT assess alone. Include program leadership, data owners, compliance and audit, operations staff, and at least one skeptic who thinks AI claims are inflated. The skeptic is not decoration. An assessment run entirely by the people whose work is being scored produces a predictable number.
- Choose your frameworks deliberately. Start with GSA for a government organization, add MITRE for granular governance detail, and use Gartner where the strategy-to-execution connection is the live question. Decide before you start rather than picking the model that flatters the result.
- Inventory what exists. List every AI system or AI-assisted tool currently in use, including the informal ones: staff using commercial AI tools without a procurement, AI-assisted features inside software platforms you already own. For each, document what it does, what data it uses, who owns it, when it was last reviewed, and whether it has a governance record. Constance found 9 systems she knew about and 4 she did not, including a document summarizer two program managers had adopted independently on personal accounts with a commercial platform and no data-handling agreement.
- Score each dimension on evidence. Rate on the framework's scale, based on what you can show rather than what you intend. "We have a data governance policy" scores differently from "we have a data governance policy that is consistently applied and audited quarterly." For each dimension ask what evidence shows you are at this level, and then ask the harder question: what is the strongest evidence you are not more mature than you claim?
- Expect unevenness and record it. Most organizations are uneven, perhaps Level 3 on technical infrastructure and Level 1 on governance. Resist the urge to average. The unevenness is the finding.
- Write the assessment report. For each dimension document the current level, the evidence supporting it, the major gaps or risks, and specifically what prevents movement to the next level. That last field is what turns the document into a plan.
- Identify the binding constraint. Your AI readiness is limited by your lowest maturity dimension in critical areas. If governance is Level 1 while technical infrastructure is Level 3, governance is the binding constraint and no amount of platform work will move you.
- Prioritize capability investments against real applications. Not all gaps are equally consequential. Prioritize gaps that block specific high-value planned applications over gaps that are theoretical with no near-term deployment implication. Constance's department prioritized data quality and monitoring over governance documentation, which was adequate for current applications, and over technical infrastructure, which was already strong.
- Build the maturity target into the AI strategy. The assessment is not separate from the strategy; it is the foundation of the strategy's sequencing. Applications in years one and two should match current maturity. Applications in years three through five should be sequenced after the capability investments that will raise maturity to support them.
A gap worth naming is a dimension where the current score sits two or more levels below what a specific planned application requires. Constance's department planned to expand its predictive risk model from two county offices to all 64 parishes. That expansion required Level 3 data quality, meaning consistent standards applied statewide, and the assessment showed the department at Level 2, inconsistent and county by county. The gap did not mean the expansion should be abandoned. It meant the expansion should be sequenced after a 12-month data quality improvement project.
Binding Constraints in Practice
The binding-constraint idea is easy to nod at and hard to act on, because the constraint is usually in the dimension nobody in the room owns. Three worked assessments show the pattern. In the first, a federal agency believed it was at Level 3 across the board: the CIO had implemented ML tools, hired data scientists, and had several pilots running. An honest cross-functional assessment found governance at Level 2, with pilots that existed but were uncoordinated and no clear approval process; data management at Level 1, scattered across legacy systems with no data governance; technical infrastructure at Level 3, with genuinely good platforms; and organizational readiness at Level 1, with frontline staff who did not understand AI and minimal training.
The finding was that data management, not technical capability, was the binding constraint. The agency had good tools it could not use effectively because the data was too fragmented to feed them. The roadmap was refocused on data governance rather than more AI tooling, and eighteen months later, with better data governance in place, the agency could scale its AI projects. Note how easily this could have gone the other way: the visible, fundable, satisfying investment was more platform, and the platform was already the strongest dimension.
In the second, a mid-sized city produced a sobering but clear picture: vision and strategy at Level 1 with scattered interest and no clear strategy, governance at Level 1 with no formal structures, technical infrastructure at Level 1 with old systems and limited cloud capability, workforce and skills at Level 1 with a single person holding a data science background, operations at Level 1 with no AI-specific processes, and a data strategy the assessment placed below the bottom of the scale, with no strategy at all and data living in department silos. Rather than pretend it could jump to Level 3, the city built a realistic three-year roadmap: year one on strategy and foundational data governance, year two adding basic technical infrastructure, year three starting actual pilots. The honest assessment prevented spending on projects it could not have sustained.
In the third, a health ministry was highly uneven: AI strategy at Level 3 with clear direction and good mission alignment, governance at Level 2 with structures that existed but were not optimized, data management at Level 2 with a framework applied inconsistently, technical at Level 2 with platforms that existed but were not optimized, and organizational readiness at Level 1 with skeptical staff and weak change management. Organizational readiness was the binding constraint. Despite good strategy and reasonable infrastructure, adoption was slow because frontline staff did not understand why AI mattered. Redirecting significant resources to change management and training accelerated adoption markedly within a year.
The Three Gaps That Recur
Three gaps appear most frequently in government AI maturity assessments, and all three are organizationally simple rather than technically hard, which is exactly why they persist.
Monitoring and performance measurement. Many agencies deploy AI systems with no ongoing monitoring plan. Performance is measured at deployment and then not again until something goes wrong. The gap shows up as a sentence nobody wants to say out loud: we do not know whether this system still performs the way it did when we tested it. Closing it requires a monitoring schedule with performance checks at least quarterly, a named person responsible for running them, and a defined threshold that triggers escalation when performance degrades. Note the limit of that control: monitoring surfaces drift only in the dimensions you chose to watch, so the choice of metric is part of the design.
Data documentation. AI systems built on undocumented data are audit risks. When an inspector general or a legislative auditor asks what data the system was trained on and what quality checks were performed, the answer has to be documented rather than reconstructed from memory. Closing the gap means maintaining a data lineage record for every AI system: where the training data came from, what preprocessing was done, what quality checks were applied, and when the dataset was last updated. Make it a deployment prerequisite rather than an after-the-fact documentation project, because reconstructed lineage is the least reliable kind.
Staff AI literacy. AI tools deployed to staff who do not understand how to use them appropriately can produce worse outcomes than no AI at all. Staff who do not grasp that a fraud-detection system produces probabilities rather than determinations will either over-rely on high-probability flags and skip the human review, or ignore the flags entirely and eliminate the value. Closing this gap requires role-appropriate training before deployment, not a general AI awareness course delivered years later. This is also the dimension most often assumed rather than verified, which is why it shows up as Level 1 in assessment after assessment.
What the Roadmap Looks Like at Each Level
Different organizations need different roadmaps, and applying a Level 3 roadmap to a Level 1 organization is one of the more expensive mistakes available. At Level 1, in the emerging or awareness stage, the focus is establishing basic governance, identifying one high-value use case, and beginning data inventory work. Do not build sophisticated ML infrastructure yet. The key investments are defining an AI strategy aligned to mission, establishing a governance committee, inventorying your data, building internal AI literacy through training, and selecting a single pilot use case. Plan on roughly 12 to 18 months to reach Level 2.
At Level 2, with basic infrastructure and governance in place, the focus shifts to repeatable processes, better data governance, and a sustainable data platform. The investments are standardizing the AI project approval process, implementing data quality standards, building or expanding data warehousing capability, developing technical standards for model development, and expanding training and change management. Plan on roughly 18 to 24 months to reach Level 3.
At Level 3, with working governance, decent data, and real technical capability, the focus becomes optimization and depth: tuning governance for speed without sacrificing oversight, implementing continuous improvement processes, integrating AI into broader organizational processes, expanding talent into more specialized roles, and developing advanced analytics capability. Advancement beyond this point generally takes 24 months or more. Those timelines are planning ranges rather than schedules, and the constraint that determines whether you hit them is usually the binding one you identified during the assessment.
Reporting Maturity Externally
When Constance returned to the legislative briefing six months later, she used the Gartner five-level framework because it matches the vocabulary legislators and staff already use. Her sentence was: we are at Level 2, Active, in most dimensions, with targeted investments in data quality and monitoring bringing us toward Level 3 in those specific areas over the next eighteen months. That statement is honest, specific, and actionable. It tells the committee what the department can do now, what it cannot do yet, and what investment is closing the gap.
Two things make that sentence work, and both are worth copying. It names the framework, so "Level 2" has a definition the audience can check rather than being a number that sounds like a grade. And it scopes the claim to specific dimensions rather than asserting a single agency-wide score, which is the form most likely to be wrong. Used this way, the assessment is also a powerful advocacy tool. Saying "we are Level 2 in data governance, which is why we cannot move faster on AI projects" explains reality in a form leadership can act on, which a request for more funding on its own does not.
Anti-Patterns
- Inflated self-assessment. You want to believe the organization is more mature than it is, so you claim Level 3 governance because a committee exists, even though the committee rubber-stamps decisions and oversees nothing. Inflated assessments produce roadmaps that do not match reality; you plan for Level 3 execution at Level 1 capability, projects fail, and the blame lands on technology or talent rather than on having moved faster than you could. Avoid it by including skeptics, documenting evidence for every claim, and asking what would convince an auditor. If you cannot articulate the evidence, you are overstating.
- Treating a maturity level as a certification. A level is an assertion an agency makes about itself against a published scale. No framework body awards it, no assessor confers it, and citing a model by name does not transfer the model's authority to your score. Present levels as claims with evidence attached, and be explicit in briefings that the assessment is a self-assessment. A level presented as a measurement invites an audience to rely on it as one.
- Ignoring the binding constraint. You find yourself weak in governance, then invest heavily in technical infrastructure where you are already strong. Actual capability is limited by the weakest dimension, so strengthening a strength changes nothing. Once you have assessed every dimension, name the binding constraint explicitly and aim the next 12 to 18 months at it. You are not abandoning the other dimensions, you are prioritizing what is preventing progress.
- Static assessment. You assess once, conclude you are Level 2, and plan against that finding for three years. Maturity moves as you invest and as the environment changes, so a stale baseline means you miss chances to advance faster or fail to notice you are slipping. Reassess annually; it is a light process after the first time. Track which dimensions are improving and which have stalled, and adjust the roadmap accordingly.
- Maturity theater. You write an assessment that looks good, assign a Level 3 rating, and use it to justify a funding request. It will not match what people inside the organization experience, frontline staff will know, and your credibility erodes; when projects fail, people conclude you were dishonest about capability rather than unlucky. Treat the assessment as a genuine learning tool, share the results including the weak areas, and use the weak areas to argue for resources instead of hiding them.
- Scoring maturity from activity. Counting deployments, tools purchased, or pilots launched and reading the total as a level is the error Constance made in her first briefing. Activity is what you have done; maturity is what you can repeat. An agency with many undocumented deployments may be less mature than one with none, and its inventory is a liability rather than a score.
- Averaging away the unevenness. Producing a single agency-wide number by averaging dimensions hides the one finding that matters. An agency strong on three dimensions and at Level 1 on the fourth does not behave like an agency sitting midway up the scale. It behaves like an agency stopped by its Level 1.
Practice Prompts
- Framework selection. Which maturity model is most useful for your organization? GSA is designed for government, MITRE is more detailed, Gartner connects strategy to execution. Which one addresses the biggest question you actually have, and what would you use the other two for?
- Honest assessment of one dimension. Take a single dimension, governance for instance, and assess your organization's current maturity. What evidence supports that assessment? What is the strongest evidence against it? What specifically would make you more mature?
- Find your binding constraint. Assess all four MITRE domains. Which is lowest? What makes it hardest to improve? What would have to happen, and who would have to act, for it to advance one level?
- Roadmap implications. Based on your current level, what should your organization focus on over the next 12 to 18 months? How does that differ from what you are currently doing, and what would you have to stop doing to make room?
- Organizational readiness. Organizational readiness is often the binding constraint in government. How mature is your organization on AI literacy, change management, and culture? What would advancing it enable that is currently blocked?
- Write the briefing. Draft the presentation you would give your agency head six months from now. Where do you stand today, what progress will you show, what will still be your binding constraint, and what resources would you ask for to advance it? Write it now and keep it as the reference point for the next phase of work.
Reflection
Think about the last time someone asked your agency how mature its AI capability was. What did the answer describe: activities or capability? If a committee staffer asked you today to define the scale you were using, could you? Which of your dimensions would a skeptic on your own staff score lower than you would, and what evidence would they cite? If your binding constraint is the one nobody in the room owns, who needs to be in the room next time? And if the honest assessment says a planned application should be sequenced eighteen months later than currently promised, who has to hear that, and how soon does the conversation get harder if you wait?
Glossary
- Maturity model. A framework for assessing organizational capability across dimensions using levels, where the bottom level is initial or ad hoc and the top is optimized or advanced. Enables comparison and benchmarking; the rating is self-assigned.
- Binding constraint. The lowest maturity dimension, which limits overall organizational capability. Improving other dimensions does not help while the binding constraint is unchanged.
- Data governance. Formal structures, processes, and accountability for managing data as an asset, including quality standards, data ownership, access controls, and lineage tracking.
- Governance and accountability. The decision-making frameworks for overseeing AI initiatives, including project approval, risk review, and compliance oversight.
- Technical infrastructure. The platforms, tools, and systems that enable AI work, including ML platforms, cloud infrastructure, data pipelines, and model deployment systems.
- Organizational readiness. The capability of an organization's people and culture to adopt and use AI, including AI literacy, change management, and openness to innovation.
- Capability versus deployment. The distinction the GSA model makes between what an agency is able to do and what it has actually put into operation. The two can diverge in either direction.
- Data lineage. The documented record of where training data came from, what preprocessing it received, what quality checks were applied, and when it was last updated.
Related Lessons
This assessment is the honest floor under the work in Developing an Organizational AI Strategy, and it feeds directly into AI Roadmap Development, where the sequencing decisions it implies get written down. Prioritization Frameworks for Government AI helps choose among the applications your current level can actually support, and Building the AI Business Case turns an identified capability gap into a funding request. Measuring Organizational AI Maturity extends the assessment work at enterprise scale, Data Governance for AI addresses the dimension most often found to be the binding constraint, and Change Management for AI Adoption addresses the second most common one.
Closing
Maturity assessment is where strategy meets honesty. It is easy to write aspirational strategies that assume perfect conditions and unlimited capability. It is harder to build realistic roadmaps that account for where you actually are, what is actually blocking progress, and what genuine advancement would take. Organizations that succeed with AI do this work thoroughly: they know where they stand, they know what is stopping them, and they plan within their real constraints rather than around them.
The next step is doing the assessment rather than reading about it: forming the team, choosing the framework, and getting honest about the score. If you are at Level 1, your next eighteen months look nothing like a Level 3 organization's. If your binding constraint is governance, you know where the money goes. If it is data, you know what is blocking you. Constance could not answer a single question about her scale at her first briefing. Six months later she could describe her department's position precisely enough that a legislative committee could act on it, and none of what changed was the technology.
Key Takeaways
- Maturity is what you can reliably do, not what you have done. A list of deployments is not a score. Maturity measures consistent process, governance, and outcomes tracking across strategy, data infrastructure, organizational capability, responsible AI practices, and performance measurement.
- A level is a self-assessment, not a certification. No framework body awards it. Present every level as a claim with evidence attached, and say so plainly in external briefings.
- Use the three frameworks together. GSA is built for government and the widest in coverage, MITRE is strongest on governance detail, and Gartner's five levels translate best to executive and legislative audiences while forcing the strategy-to-execution connection.
- Score on evidence, with a skeptic in the room. "We have a policy" scores differently from "we have a policy that is consistently applied and audited." Ask what would convince an auditor, and ask what the strongest evidence against your own claim would be.
- Your capability is limited by your weakest critical dimension. Identify the binding constraint explicitly and aim the next 12 to 18 months at it. Do not average dimensions into a single number, and do not invest in a strength because the investment is easier to fund.
- Gaps of two or more levels block specific planned applications. An expansion requiring Level 3 data quality at a Level 2 agency should be sequenced after a data quality project, not abandoned and not attempted on schedule.
- Three gaps recur: monitoring, data documentation, and staff literacy. Monitoring is neglected because it requires effort after the satisfying moment of deployment, documentation is skipped because it feels bureaucratic, and literacy is assumed rather than verified. All three are audit risks and operational risks.
- Reassess annually and put the target in the strategy. Years one and two should match current maturity; years three through five should follow the capability investments that raise it. A strategy disconnected from its maturity baseline is a wish list, and a baseline three years old is not a baseline.
Frequently Asked Questions
Which framework should we actually use? If you are a government organization, start with the GSA model, which was designed for your constraints and covers the widest set of dimensions. Add MITRE when you need granular detail on governance, which is where its treatment is strongest. Use Gartner when you are briefing executives or legislators, because its five named levels match the vocabulary those audiences already use. Running all three is not duplication; each answers a question the others answer less well.
Can we compare our score to another agency's? Cautiously, and only if you know how the other agency scored itself. These are self-assessments against published scales, so two agencies claiming Level 3 may have applied very different evidence thresholds. Benchmarking against organizations of similar size and type is genuinely useful for setting expectations, but treat a published level as a starting question rather than a comparable measurement.
What if our assessment comes back much lower than leadership expects? That is the assessment doing its job, and it is easier to deliver before a failed project than after one. Lead with the binding constraint rather than the low scores: the useful sentence is not "we are Level 1 in governance" but "governance is what is preventing the expansion we promised, and here is what advancing it requires." A low honest score is an argument for resources. An inflated one is a debt that comes due when the roadmap fails.
How often should we reassess? Annually. The first assessment is the expensive one; afterward it is a light process because the instrument, the team, and the evidence base already exist. Track which dimensions improved and which stalled rather than only the headline levels, since a stalled dimension that has quietly become your binding constraint is exactly what an annual cadence is meant to catch.
We are a small local government with almost nothing in place. Is this exercise worth the time? Yes, and the mid-sized city example shows why. Assessing at Level 1 across nearly every dimension is not a humiliation; it produced a three-year sequence, strategy and data governance first, infrastructure second, pilots third, that prevented spending on projects the city could not have sustained. The alternative is not staying at Level 1; it is buying a system at Level 1 and discovering the gap after the contract is signed.
Our infrastructure is strong but nothing gets adopted. What does that mean? Most likely that organizational readiness is your binding constraint, which is the most common one in government. The health ministry example had good strategy and reasonable platforms and still saw slow adoption because frontline staff did not understand why AI mattered. The remedy is change management and training rather than more platform, and it moves faster than most infrastructure work does.
Skill.re