Assessing Your Reporting AI Readiness
Three weeks before the board sign-off, the disclosure lead opened the Scope 3 file and found a single emission factor doing the work of EUR 40 million in reported purchased-goods emissions. Nobody could say which version of the factor library it came from, who chose it, or when. The CSO had already told the audit committee that the function was "ready to bring AI into the CSRD process next year." Sitting in front of an unsourced factor that one analyst remembered typing in sometime last spring, that promise suddenly looked less like a roadmap and more like a liability. The uncomfortable truth in that room is the truth in most reporting functions: the board believes you are far more ready than you are, and an honest assessment has to come before any tool, any vendor, and any roadmap.
This lesson gives you that honest assessment. Not a maturity-model poster for the wall, but a scorecard you run on your own function across three axes that actually determine whether reporting AI will help you or quietly manufacture assurance findings. You will score each axis from 1 to 5, sum the result, and interpret the band. More importantly, you will learn why the average score is a trap and the weakest axis is the only number that matters. By the end you should be able to sit across from your assurer, your controller, and your CSO and say, with evidence, exactly what your function can and cannot safely automate yet.
Why Readiness Comes Before The Roadmap
There is a predictable sequence in how reporting AI goes wrong. A senior executive reads that a peer is "using AI for sustainability reporting." The instruction comes down: find out how we do the same. A vendor demo follows, the slides are persuasive, a pilot is scoped, and twelve months later the assurance team raises a finding because nobody can reconstruct how a number was produced. The roadmap was built on a foundation that was never assessed. That is the failure this lesson is designed to prevent.
The regulatory environment makes the stakes concrete. The Corporate Sustainability Reporting Directive survived the 2026 Omnibus simplification: Directive (EU) 2026/470 was published on 26 February 2026 and entered into force on 18 March 2026. Large undertakings remain in scope where they exceed 1,000 employees and EUR 450 million in turnover; listed SMEs were exempted; member states must transpose by 19 March 2027; and EFRAG released its simplified ESRS draft on 3 December 2025. Treat every one of those figures as something to verify against the official text rather than to repeat from memory, because thresholds and dates move and your assurer will check them. The point for readiness is simpler: the disclosure obligation did not go away, the scrutiny did not soften, and external assurance is now the norm rather than the exception.
That assurance backdrop is the reason readiness is not optional. Roughly 73 percent of large global companies now obtain external assurance on at least some sustainability disclosures, up from about 51 percent in 2019, and GHG emissions are the most-assured category. Most engagements are limited assurance today, but the trajectory points toward reasonable assurance. When AI sits inside a number that an external assurer will test, the question is no longer "does the tool work" but "can we evidence every figure it touched." A function that cannot answer that for its current manual process will not magically answer it once AI is layered on top.
AI does not fix an unready reporting function. It industrializes whatever that function already is. Point fast automation at governed data and you scale quality; point it at scattered spreadsheets and unsourced factors and you scale the mess, faster and with more authority than before.
Readiness is therefore the gate, not a formality you clear on the way to the interesting work. The three axes that follow are not arbitrary. Each maps to a specific way that reporting AI fails in practice, and each is something a senior disclosure leader can assess without being a data scientist.
Axis One: Data Maturity
The first axis asks a deceptively simple question: when you need the activity data behind a disclosed number, where is it, and can you trace it to its source? Activity data is the raw input to every emissions calculation: litres of fuel, kilowatt-hours of electricity, tonnes of purchased steel, kilometres flown, spend by supplier category. The maturity of this data is the single largest determinant of whether AI helps or harms, because AI amplifies whatever it sits on. Point a capable model at structured, centralized, source-linked data and it accelerates genuine work. Point the same model at activity data scattered across spreadsheets, shared drives, and email attachments and it produces fast garbage with a confident tone.
What you are actually scoring
Three concrete properties separate mature data from immature data. The first is structure and centralization: is activity data captured in systems with consistent fields and units, or reassembled each cycle from whatever colleagues send in? The second is traceability to source: can you start from a disclosed figure and walk back to the meter reading, the invoice, or the supplier submission that produced it, without relying on someone's memory? The third is the primary-versus-secondary balance: how much of your inventory rests on measured primary data versus modelled or spend-based secondary estimates, and do you know the split?
Scope 3 is where this axis bites hardest. Scope 3 typically represents around 75 percent of total emissions across the fifteen GHG Protocol categories, and it depends on data you do not own. Industry surveys make the pain explicit: roughly 62 percent of reporters cite internal data quality and about 79 percent cite supplier-data availability as their top barriers (Sphera 2025). If most of your footprint sits in a category where the data is both the largest and the least controlled, then your data-maturity score is effectively your Scope 3 data-maturity score. An AI that drafts beautiful narrative around a Scope 3 number built on unverifiable supplier estimates has not solved your problem; it has dressed it up.
The reconstructability test
There is one practical test that cuts through self-assessment optimism: can you reconstruct last year's inventory from source today, without the analyst who built it? If the prior-year number depends on tribal knowledge, undocumented judgment calls, and files only one person can locate, your data is not mature regardless of how polished the final report looked. Assurers love this test because it is exactly what they do. If your own team cannot pass it internally, an external assurer will not pass it for you, and an AI trained to be helpful will happily fill the gaps with plausible reconstructions that are not evidence.
Axis Two: Emission-Factor Governance
The second axis is the one most functions score themselves too high on, because it is invisible until something goes wrong. Activity data tells you how much of something happened; an emission factor converts that activity into a quantity of greenhouse gas. The governance question is whether your factors live in a single, named, dated, version-controlled library, or whether they live in analysts' heads, last year's spreadsheets, and a shared folder nobody fully trusts.
Why this axis is the hallucination gate
Emission-factor governance is the specific control that closes the most dangerous reporting AI failure mode: the hallucinated factor. When you ask a model to "find the right emission factor" for an activity, it can return a number that looks entirely plausible, carries no provenance, and is wrong. If your function has a governed factor library, you can constrain the AI to retrieve only from that authorized, versioned source and reject anything else. The library becomes the ground truth the AI must cite. If your function has no such library, there is nothing trustworthy to retrieve against, and the hallucinated-factor failure mode is wide open. The AI will invent a factor, and an unready function has no control to catch it.
This is why factor governance, not data maturity, is often the binding constraint for the very first AI use case that functions reach for, which is factor lookup and selection. It feels like a low-risk, high-value place to start. It is only low-risk if the retrieval target is governed. Without governance, the most attractive entry point is also the most dangerous.
What a governed factor library actually has
A governed library is not a spreadsheet of numbers. It has, at minimum: a single authoritative location; a named owner accountable for it; version control so you can say which factor set applied to which reporting period; documented provenance for each factor (which database, which publication, which year); a change log recording who updated what and why; and a defined review cadence. The test is whether you can answer, for any factor in your current report, the five questions an assurer will ask: which factor did you use, where did it come from, which version, who approved it, and why was it appropriate for this activity. If those answers live only in an analyst's head, your factor governance is immature no matter how confident the analyst sounds.
The cardinal rule sits at the centre of this axis and the next. Every figure must trace to evidence, and "the AI estimated it" is not evidence. A governed factor library is what lets an AI-assisted lookup produce a traceable answer instead of an authoritative-sounding guess. Without it, automation simply makes the guessing faster.
Axis Three: The Assurance Relationship
The third axis is not about your data or your factors at all. It is about a relationship, and it is the axis that disclosure leaders most often forget to score because it does not feel like infrastructure. The question is straightforward: do you talk to your assurer before you build, or only after? Have you walked them through where AI would sit in your process, what it touches, and what controls surround it? Or will the first time your assurer learns that AI drafted a disclosure be the moment they find it in the file?
An assurer surprised by AI is a finding waiting to happen
Assurance is built on understanding the process that produced the numbers. When an assurer encounters a process element they were not briefed on, especially one as consequential as AI generating or selecting figures, their professional reaction is not curiosity, it is caution. Undisclosed AI in the file reads as undisclosed risk. It invites additional testing, additional questions, and in the worst case a qualification or a finding that lands in front of the same audit committee the CSO reassured. Readiness on this axis means the assurer is never surprised, because they helped you decide where AI could and could not sit.
The assurer as design partner
The mature posture treats the assurer as a design partner rather than an examiner you meet at the end. Practically, that means engaging them early about your intended AI use cases, showing them the controls (the governed factor library, the audit trail, the human review gates), and getting their reaction while the design is still changeable. This is not about seeking their approval as a rubber stamp, and it is certainly not about transferring responsibility to them. Accountability stays human and stays with you. It is about removing surprise from a relationship where surprise is expensive.
Note carefully what readiness on this axis is not. It is not the assurer endorsing a vendor. The same caution that applies to factors applies to tools: a platform category exists, but obligations do not transfer to the platform. For orientation only and never as endorsement, the market includes ESG and CSRD reporting platforms and carbon-accounting tools (the category includes names such as Watershed, Persefoni, Sweep, Workiva, Position Green, Sphera, IBM Envizi, SAP Sustainability, and Salesforce Net Zero Cloud). Whichever you use, your assurer is assuring your disclosures, not the vendor's software, and your accountability is unchanged.
Why this axis interacts with regulation
The assurance relationship is also where converging regulation shows up. ISSB standards IFRS S1 and S2 are adopted or planned across more than thirty jurisdictions representing over half of global GDP, with around twenty-one jurisdictions having adopted as of 1 January 2026 and roughly sixteen more planning, and targeted amendments to IFRS S2 issued in December 2025. The CBAM definitive phase went live on 1 January 2026, authorised-declarant applications were due 31 March 2026, and the first certificate surrender falls in 2027, covering cement, iron and steel, aluminium, fertilisers, hydrogen, and electricity, with a 50-tonne de minimis and a choice between actual and default values for embedded emissions. Every one of those regimes lands on the same assurer-facing evidence chain. An assurer who already understands where AI sits in your CSRD process is an assurer prepared for the next regime; one meeting AI for the first time is starting from behind on all of them. As always, verify these figures against primary sources rather than repeating them blindly.
The Readiness Scorecard: A Worked Example
Now we make this operational. The scorecard scores each of the three axes from 1 to 5, sums to a total out of 15, and reads the band. The discipline is in the level descriptions: score honestly against the described level, not against where you wish you were.
Scoring each axis 1 to 5
| Score | Data Maturity | Emission-Factor Governance | Assurance Relationship |
|---|---|---|---|
| 1 | Activity data scattered across spreadsheets, inboxes, and drives; rebuilt from scratch each cycle; no traceability. | No library; factors live in analysts' heads and old spreadsheets; provenance unknown. | No contact until the file is submitted; assurer meets your process at the end. |
| 2 | Some data centralized but inconsistent units and fields; tracing a number to source needs the analyst who built it. | A spreadsheet of factors exists but is unversioned, unowned, and partly undocumented. | Annual transactional contact only; no discussion of process design or change. |
| 3 | Core activity data in systems with consistent fields; primary-versus-secondary split known; Scope 3 still patchy. | A named factor source with some documentation, but version control and change log are weak or informal. | Regular contact and a walkthrough of the manual process, but AI plans not yet discussed. |
| 4 | Centralized, source-linked data; prior-year inventory reconstructable without the original analyst for most categories. | Single named, dated, version-controlled library with documented provenance and a review cadence. | Assurer engaged early on process changes and briefed on intended AI use cases. |
| 5 | Fully structured, source-traceable data including audited Scope 3 primary data; reconstructable end to end. | Governed library with full change log, approval workflow, and AI-retrieval constrained to it. | Assurer is a design partner; controls co-reviewed; no surprises in the file by design. |
A realistic company scores itself
Consider a large manufacturer, in scope for CSRD, that the board believes is "ready for AI." Run the scorecard honestly.
- Data maturity: 3. Core energy and fuel data sits in proper systems with consistent units, and the primary-versus-secondary split is known. But Scope 3, which dominates the footprint, still depends on supplier estimates that are patchy and slow to arrive. Solid middle, not better.
- Emission-factor governance: 2. There is a factor spreadsheet, and a capable analyst maintains it. But it has no version control, no named accountable owner beyond that analyst, and provenance is documented for only some entries. The EUR 40 million factor from the opening scene lives here.
- Assurance relationship: 3. The assurer runs a limited-assurance engagement, contact is regular, and they have walked the manual process. But nobody has told them AI is on next year's plan, and no controls have been co-reviewed.
The total is 8 out of 15. The board hears "more than half, basically ready." That reading is wrong, and understanding why is the whole point of the scorecard.
Interpreting the bands
| Total score | Band | What it means for AI |
|---|---|---|
| 3 to 6 | Not ready | Foundational gaps. AI will industrialize existing weaknesses. Fix the foundation before any AI project; a roadmap here produces assurance findings. |
| 7 to 10 | Low-risk pilots only | Enough maturity to attempt narrow, well-bounded pilots with heavy human review, but not to scale. The weakest axis dictates which pilots are even safe. |
| 11 to 15 | Ready to scale | Governed foundation across all three axes. AI can be extended to higher-value use cases under existing controls and assurer awareness. |
The weakest axis governs, not the average
At 8 of 15, our manufacturer lands in the "low-risk pilots only" band on the total. But the total is the less important number. The factor-governance score of 2 is the binding constraint, and here is why it overrides everything else. The most attractive first use case for this function is exactly AI-assisted factor lookup. That use case retrieves against the factor library. A factor library scoring 2 has nothing trustworthy to retrieve against, which means the single most tempting pilot is also the one most exposed to the hallucinated-factor failure mode. The data being a 3 does not rescue it. The assurance relationship being a 3 does not rescue it. The weakest axis caps what AI can safely do, because the weakest axis is precisely where automation will break first.
The correct conclusion is not "we scored 8, start a pilot." It is "fix factor governance from a 2 to at least a 4 before any AI factor-lookup project, and in the meantime brief the assurer that this is the plan." Stand up a single named, dated, version-controlled library with documented provenance, then revisit the scorecard. This is what an honest assessment buys you: it redirects the program from the seductive-but-unsafe first step to the unglamorous foundational fix that has to come first.
Using the average would have hidden the risk
Had the manufacturer averaged the three axes (about 2.7 out of 5) and rounded to "roughly a 3, middle of the pack," it would have lost the most important signal entirely. Averaging launders the weakest axis into a comfortable middle. The whole value of scoring axes separately and reading the minimum is that it refuses to let a strong axis hide a fatal one. Report the three scores and the minimum to your CSO. Never report only the average.
Running The Assessment Honestly
A scorecard is only as good as the honesty behind it, and the structural temptation is to score high. The CSO told the board the function was ready. The team is proud of last year's report. Nobody wants to write "2" next to factor governance in a document the executive will read. Building the honesty in is a design problem, not a willpower problem.
Score against evidence, not against pride
Require evidence for every score above a 2. A 3 in data maturity means showing the system and the known primary-secondary split, not asserting it. A 4 in factor governance means producing the version-controlled library and its change log, not describing the spreadsheet you intend to upgrade. Tie each level to an artifact you could hand an assurer. If the artifact does not exist, the score does not apply. This single rule defeats most optimism, because it converts "I think we are mature" into "show me the file."
Score with the people who do the work
The analyst who maintains the factor spreadsheet knows it is a 2 even if the org chart wants it to be a 4. The people closest to the data give you honest scores precisely because they are not the ones who promised the board. Run the assessment with them in the room, and protect them from being punished for honest low scores. A readiness assessment that punishes honesty will reliably produce dishonest scores and a roadmap built on a fiction.
Re-score on a cadence and after change
Readiness is not a one-time gate. Score at least annually, and re-score whenever something material changes: a new ERP, an acquisition that adds ungoverned data, a change of assurer, a regulatory expansion such as CBAM bringing new embedded-emissions data into scope. The scorecard is a living instrument that tells you, at any moment, the ceiling on what AI can safely do for your function and the one axis to invest in next.
Key Takeaways
- Assess readiness before you build a roadmap. AI does not fix an unready function; it industrializes whatever the function already is, so an honest assessment is the gate, not a formality.
- Score three axes that map to real failure modes: data maturity (does activity data trace to source), emission-factor governance (is there one named, dated, version-controlled library), and the assurance relationship (is your assurer a design partner or an examiner you meet at the end).
- Score each axis 1 to 5 and read the bands: 3 to 6 is not ready, 7 to 10 is low-risk pilots only, 11 to 15 is ready to scale. The worked manufacturer scored data 3, factors 2, assurance 3, for a total of 8 of 15.
- The weakest axis governs, never the average. The manufacturer's factor-governance 2 is the binding constraint because the most tempting first use case (AI factor lookup) retrieves against that very library; fix it to a 4 before any such pilot.
- Factor governance is the hallucination gate. A governed library gives AI a trustworthy retrieval target; without it, the hallucinated-factor failure mode is wide open and automation just makes the guessing faster.
- Apply the cardinal rule everywhere: every figure must trace to evidence, "the AI estimated it" is not evidence, and accountability stays human regardless of which platform you use, because obligations do not transfer to the vendor.
- Build honesty into the assessment: require an artifact for every score above 2, run it with the people who do the work, and protect honest low scores. Report the three axis scores and the minimum to leadership, never just the average.
- Re-score annually and after material change (new ERP, acquisition, new assurer, CBAM or other regulatory expansion). Treat the scorecard as a living ceiling on what AI can safely do, and verify all regulatory figures against primary sources rather than repeating them blindly.
Skill.re