Training Sustainability and Finance Staff
The controller has never opened the GHG Protocol in her life, and next quarter her signature goes under a Scope 3 figure that is three-quarters of the company's footprint. Down the hall, the carbon accountant who built that figure has never been through a financial close, does not speak the language of controls and materiality thresholds the way finance does, and has just been handed an AI copilot that promises to "do the estimation for him." Two people, two blind spots, one report that an external assurer will test line by line. This is the training problem that most reporting AI programs never even name: the tool is new, but the harder gap is that sustainability and finance now co-own a disclosure neither of them was trained to produce together. A training program is not a nice-to-have wrapped around the software. It is the thing that decides whether your numbers trace to evidence and survive an assurer, or whether they quietly become the finding that ends the cycle in a restatement.
Why Finance Staff Are Suddenly in Scope
For a decade, non-financial reporting lived in a sustainability team that operated at arm's length from the finance function. The CSO's team gathered the data, wrote the narrative, and published a glossy report that few people audited and fewer people litigated. That world is gone. Under CSRD, the sustainability statement sits inside the management report, is subject to external assurance, and is signed off with the same seriousness as the financial statements. Under ISSB's IFRS S1 and S2, sustainability-related financial disclosures are explicitly connectable to the financial statements and are adopted or planned across more than thirty jurisdictions. The practical consequence is blunt: the controller, the reporting accountant, and the internal audit function are now co-owners of numbers they used to ignore, and the sustainability team is now producing figures that must survive the discipline finance has always lived under.
This is not a reporting-line reshuffle. It is a collision of two professional cultures that were never trained on each other. Finance staff know controls, materiality, sign-off hierarchies, restatement discipline, and what an assurer actually tests, but most cannot tell you what a Scope 3 category is, why a spend-based estimate is weaker than an activity-based one, or what an emission factor's provenance means. Sustainability staff know the GHG Protocol, the fifteen Scope 3 categories, the ESRS datapoints, and the science, but many have never operated inside a controlled close, have never been on the receiving end of a limited-assurance procedure, and instinctively treat an estimate as good enough because "the exact number is impossible." Drop an AI tool into that gap and you get the worst of both: finance signing figures it cannot interrogate, and sustainability accepting AI outputs it cannot defend. The training program exists to close that gap deliberately, before the assurer does it for you at the engagement table.
There is a second reason finance's involvement matters for AI specifically. Finance is where the instinct for auditability already lives. The controller who asks "where does this number come from and who checked it" is asking exactly the question that saves an AI-assisted disclosure. Bring finance in as a trained co-owner rather than a nervous signatory, and you import the assurance mindset the sustainability team needs. Leave finance untrained, and you get a signatory who either rubber-stamps figures they do not understand or blocks the whole cycle out of fear. Neither is acceptable when a failed disclosure at a company above the CSRD threshold, more than 1,000 employees and more than EUR 450M turnover, is a board-level event.
Design the Program Around Doing, Not Knowing
The instinct of most training is to teach knowledge: here is the GHG Protocol, here is what a hallucination is, here is the ESRS taxonomy, now take a quiz. That produces people who can recognize the right answer on a multiple-choice test and still cannot produce a defensible figure under deadline. A gold-standard program is built around capability, not comprehension. The organizing question for every module is not "does the learner know this" but "can the learner do this, and can they produce an artifact that proves they did it right." The difference is the difference between a course and a competent team.
Concretely, this means every track culminates not in a score but in an artifact: a real or realistic disclosure output that the learner produced, that traces to evidence, and that a reviewer could put in front of an assurer. A carbon accountant does not pass by defining primary versus secondary data; they pass by producing a Scope 3 category estimate with the AI-drafted extraction, the named and dated emission factor from the approved library, the human decision, and the estimate-versus-measured label all captured. A finance controller does not pass by reciting the definition of limited assurance; they pass by running a sample of AI-assisted figures back to source and writing the review note they would actually sign. The artifact is the assessment, because the artifact is the job.
Train people to produce a file that survives an assurer, not to pass a quiz an assurer will never see.
This reframing also fixes the most common training failure: the learner who finishes the course and reverts to the old habit the moment a deadline hits. When the training output is the actual work product, there is no gap between the exercise and the job. The learner practices the exact move they will make in the cycle: take the AI draft, ground it, decide, and capture the proof. Repetition of that move under realistic pressure is what builds the reflex that holds when the assurer is at the door and the deadline is tomorrow.
Role-Based Tracks: What Each Person Must Be Able to Do
A single generic "AI for ESG" course fails everyone, because the carbon accountant, the disclosure lead, the controller, the procurement analyst, and the CSO face completely different AI risks and need completely different capabilities. The program must fork into role-based tracks, each defined by what that role must be able to do, not by a shared body of trivia. What follows is the spine of five tracks. The verbs matter more than the nouns: each track is a list of things the learner must be able to perform and prove.
The Analyst and Carbon Accountant Track
This is the person at the coalface of the numbers, and the person most exposed to AI's most dangerous failure mode: the hallucinated emission factor and the fabricated activity datum. Their track is the deepest on the mechanics of grounding. They must be able to take an AI-extracted activity datum and verify it against the source document; require and check that any AI-proposed emission factor returns a named, dated, versioned source from the approved factor library rather than a plausible number the model invented; distinguish spend-based from activity-based method and choose defensibly between them; label every output as primary or secondary and as estimated or measured; and reconcile a current-period figure against the prior period so an implausible jump is caught before it ships. The artifact they produce to pass is a completed Scope 3 category with full provenance that a reviewer signs. The competency being built is suspicion of a confident number, which is the single most valuable instinct an AI-assisted analyst can have.
The Disclosure-Lead Track
The disclosure lead owns the narrative and the mapping from one fact base to many frameworks. Their AI risk is different: not a fake number, but a softened negative impact, an invented target the company never set, or an ESRS datapoint that the model quietly answered without evidence. Their track must build the ability to use AI to draft narrative datapoints and materiality-assessment inputs while checking every claim against the underlying evidence; to detect where an AI draft has hedged a genuinely negative impact into something anodyne, which is a greenwashing risk with legal teeth; to trace an AI-mapped ESRS or ISSB datapoint back to its source and confirm it was not fabricated to fill a gap; and to keep a clean line between defensible estimation, labeled and disclosed, and a claim dressed up as fact. Their pass artifact is a narrative disclosure section with a claim-to-evidence trace for every material statement. The competency is editorial integrity under an assurer's eye.
The Finance and Controller Track
This track exists precisely because finance is now a co-owner. It is not a lightweight awareness course; it is a translation program that gives finance staff enough of the GHG Protocol, ESRS, and assurance-for-sustainability to interrogate the numbers they now sign, plus enough AI literacy to know what the tool can and cannot be trusted to do. They must be able to read a Scope 3 inventory and ask the right control questions; distinguish measured primary data from a defensible estimate from a fabrication laundered as measurement; apply financial-grade controls, sign-off hierarchy, segregation of duties, and restatement discipline, to non-financial figures; and treat an AI-assisted figure with the same "trace it to evidence and name the human who decided" rigor they apply to a journal entry. Their pass artifact is a controller's review file on a sample of AI-assisted disclosures, complete with the exceptions they would raise. The competency is importing the audit reflex into the sustainability statement.
The Procurement and Supplier-Data Track
Scope 3 averages around 75% of the footprint and lives or dies on supplier data, yet 79% of reporters cite supplier-data availability and 62% cite internal data quality as their top barriers. The procurement and supplier-data role sits on that fault line, and it is the place where the temptation to let AI "fill the gap" with an industry average is strongest and most dangerous. Their track must build the ability to use AI to triage and validate supplier questionnaire responses at scale without treating the AI's cleanup as new primary data; to distinguish a genuine supplier-reported primary figure from a modeled secondary estimate and label each correctly; to recognize when an AI-proposed proxy is defensible and when it is fabrication; and to build the supplier-data trail that lets a reviewer see exactly which figures came from suppliers and which were estimated. Their pass artifact is a supplier-data set with clear primary-versus-secondary labeling and the provenance of every estimate. The competency is refusing to let a gap become a fabrication.
The CSO and Executive Awareness Track
The people who sign the statement do not need to operate the tools, but they need to know what they are signing off on and what questions to ask. This track is short, sharp, and strategic. Executives must be able to articulate the cardinal rule and hold the organization to it: AI assists, the human decides, the file proves it. They must be able to ask the three questions that expose a weak program, does every figure trace to evidence, is every estimate labeled, and can any number be reconstructed on demand; understand that obligations never transfer to the AI vendor no matter what the contract says; recognize the career-ending failure modes well enough to smell them in a status report; and read a program's health from assurance findings and adoption rather than from a speed dashboard. Their pass artifact is a signed governance attestation grounded in evidence they actually reviewed, not a certificate of attendance.
The L1-to-L5 Progression as the Spine
Role-based tracks tell you what each person needs; a maturity spine tells you how deep each person goes. The program threads the same L1-to-L5 progression through every track, so a carbon accountant and a controller can be at different levels of the same ladder and everyone speaks a shared language of competence. Think of it as depth, not just breadth.
Level 1, Awareness: the learner knows what reporting AI is, what it can and cannot be trusted to do, and the cardinal rule. This is the floor for everyone, including the CSO. Level 2, Practitioner: the learner can operate the AI-assisted workflow for their role under supervision, producing outputs a reviewer checks. Level 3, Independent: the learner produces defensible, fully-provenanced artifacts without hand-holding and can catch the common AI failure modes in their own domain. Level 4, Strategist: the learner can design the workflow, set the controls, and defend the approach to an assurer, which is where disclosure leads and controllers need to reach. Level 5, Leader: the learner can build and govern the whole program, train others, and own the relationship with the assurance provider. The point of the spine is that "trained" is not binary. A team is a distribution across these levels, and the program's job is to move each role to the level its risk demands: a carbon accountant touching Scope 3 must reach at least Level 3, while a junior data-gatherer might sit safely at Level 2 under review.
Embed Training in the Reporting Cycle, Not Beside It
The most expensive mistake in training design is to run it as an event: a two-day workshop in the quiet season, a certificate, and then nothing until next year, by which point the deadline pressure has erased every good habit. A gold-standard program embeds training into the actual reporting cycle so the learning happens on the real work at the real moment of need. This is how a capability survives contact with a deadline.
In practice this means mapping the curriculum to the cycle. Before data collection opens, the analyst and procurement tracks run their grounding and labeling refreshers on this year's actual sources. As the inventory is built, reviewers apply the controller-track review on live figures, catching problems while they can still be fixed rather than in the assurance post-mortem. Before the assurer arrives, the disclosure lead and controller rehearse the reconstruction of sampled figures on the real file. New hires are onboarded onto the workflow on their first real task, not on a sandbox that bears no resemblance to the pressure they will feel. The training calendar and the reporting calendar become the same calendar. This also solves the reinforcement problem: the moves are practiced every cycle on live stakes, so they become the way work is done rather than a memory of a workshop.
Embedding also lets the program turn every assurance finding into a training input. When the assurer raises an exception, that exact scenario becomes next cycle's exercise for the relevant track. The program learns from its own scars, and the curriculum stays anchored to the failures that actually happen at your company rather than generic ones from a slide deck.
Measure Success by Assurance Outcomes and Adoption
How you measure a training program determines what it becomes. Measure seat-time and completion rates, and you will optimize for people sitting in chairs and clicking "next," which correlates with nothing that matters. A program built to produce assurable practitioners measures the two things that actually prove it worked: assurance outcomes and genuine adoption.
On the assurance side, the real metrics are the ones an assurer would recognize. Are the number of assurance findings and exceptions related to traceability and unsupported figures falling cycle over cycle? Are sampled AI-assisted figures reconstructable on demand, and what share reconstruct cleanly? Is the engagement scope holding steady or narrowing rather than ballooning as AI use grows? Are estimates consistently labeled with method and uncertainty in the live file, not just in training? These are lagging indicators of a healthy program and the only ones the assurer and the board should care about.
On the adoption side, the metric is not "did people attend" but "are trained practitioners actually using the grounded workflow instead of routing around it." The classic failure is the analyst who completes the course and then pastes supplier spend into an open chatbot the moment the deadline bites, because the fast wrong way is easier than the trained right way. Real adoption shows up as the share of figures produced through the governed workflow with provenance captured at creation, and as a falling rate of free-text, ungrounded shortcuts. A program can have 100% completion and 0% real adoption; only the second number tells you whether your numbers will survive.
A Worked Example: Two Ways to Train a Controller
Consider the same finance controller, newly made co-owner of the sustainability statement, trained two ways, and watch what each does at the assurance engagement.
In the seat-time version, the controller attends a generic half-day "AI and ESG awareness" session. She watches slides on the GHG Protocol, sees a definition of hallucination, learns that CSRD requires assurance, and receives a completion certificate. She retains almost none of it, because none of it touched her actual work. When the sustainability team sends her the Scope 3 numbers to sign, she has no framework to interrogate them. She either signs on trust, because she does not know which questions expose a weak figure, or she panics and demands the team re-explain everything from scratch, delaying the close. When the assurer samples a purchased-goods figure and asks for the basis, the controller cannot answer, because she never learned that the answer should exist. The finding lands, the scope widens, and the training is revealed as theater.
In the capability version, the controller's track was embedded in the cycle and built around an artifact. She learned just enough GHG Protocol and ESRS to ask the right control questions, and she practiced on this year's real Scope 3 inventory. Her pass artifact was a controller's review file: she took a sample of AI-assisted figures, traced each back to its source, confirmed the emission factor was named and dated rather than invented, checked that every estimate was labeled as an estimate with its method, and wrote the review note flagging two figures where the provenance was thin. Those two flags got fixed before the assurer ever arrived. When the assurer samples a figure and asks for the basis, the controller hands over the reconstruction immediately, because she has done exactly this move on live data. Same person, same tool, opposite outcome at the engagement table. The only difference was that one program taught her to know about AI and the other taught her to do the job and prove she did it. The proof, as always, is the file: AI assisted, the human decided, and the file shows it.
Key Takeaways
- Finance staff are now co-owners of non-financial reporting under CSRD and ISSB, which means the training gap is not just the AI tool but two professional cultures, sustainability and finance, that were never trained on each other. The program must close both gaps at once.
- Design every module around what the learner must be able to do and prove, not what they must know. The assessment is a traceable, assurable artifact the learner produced, not a quiz an assurer will never see.
- Fork the program into role-based tracks: analyst and carbon accountant, disclosure lead, finance and controller, procurement and supplier-data, and CSO and executive awareness. Each faces different AI risks and needs different capabilities defined by verbs, not nouns.
- Each track targets its role's signature failure mode: the hallucinated factor for the analyst, the softened impact or invented target for the disclosure lead, the unaudited sign-off for the controller, the gap-as-fabrication for procurement, and the ungoverned attestation for the executive.
- Thread an L1-to-L5 maturity spine through every track so "trained" is a level, not a binary. Move each role to the level its risk demands: a carbon accountant on Scope 3 needs at least Level 3, while a junior gatherer can sit at Level 2 under review.
- Embed training in the actual reporting cycle, mapped to data collection, review, and the assurance walkthrough, so learning happens on real work at the moment of need and every assurance finding becomes next cycle's exercise. Training beside the cycle evaporates under deadline.
- Measure success by assurance outcomes and real adoption, never seat-time. Falling traceability findings, figures that reconstruct on demand, a scope that holds as AI use grows, and a rising share of work done through the governed workflow are the only metrics that prove the training worked.
- The whole program serves one cardinal rule at the level of every practitioner and every figure: AI assists, the human decides, the file proves it. A team that can live that sentence under deadline is what the training is for.
Skill.re