Evaluating ESG AI Vendors for Traceability
The vendor demo is gorgeous. A sales engineer pastes a supplier invoice into the platform, and three seconds later a fully categorized Scope 3 line appears with an emission factor, a tidy total, and a confidence badge that glows green. The room nods. Then the head of sustainability, who has sat through one too many assurance engagements, asks a single question: "Show me the source of that factor, the version, and the date, and then export the audit trail for this number to a file my assurer can read offline." The sales engineer clicks around for a while and eventually says the export is "on the roadmap." That moment, not the demo, is the evaluation. This lesson is the set of questions that separate an assurable platform from a glossy one, asked before you sign rather than discovered after.
Why the Demo Always Looks Better Than the Engagement
A reporting-AI vendor demo is engineered to show speed, because speed is what sells. It compresses months of value-chain data wrangling into a clean animation, and that compression is genuinely valuable. But the demo is built on a curated dataset with known answers, and it ends at the moment a number appears on screen. Your job does not end there. It ends nine months later when an external assurer, under a limited- or reasonable-assurance engagement, points at that same number and asks you to reconstruct it from raw evidence. 73% of large global companies now obtain external assurance on at least some sustainability disclosures, up from 51% in 2019, and emissions are the most-assured category. So the right question is never "how fast does it produce a number" but "what does it leave behind that I can defend."
This reframes the entire buying decision. You are not buying a calculation engine. You are buying an evidence-production system that happens to calculate. A platform can be fast, beautiful, well-funded, and used by impressive logos, and still be unassurable, because its speed comes from hiding the very lineage the assurer needs to see. The traceability evaluation exists to find that out in the sales cycle, when you still have leverage, rather than in the engagement, when you have a deadline and a signed contract.
It helps to picture the two timelines the platform has to serve, because vendors design only for the first. The first timeline is the buying cycle: a procurement window of a few weeks in which the platform must look capable, and where speed, polish, and a confident answer to "can you do X" win the deal. The second timeline is the reporting cycle: a year of real data flowing in, gaps appearing, factors changing, and at the end an assurance engagement where the same platform's output is interrogated by a professional whose entire job is to doubt a number until it is proven. A platform optimized only for the first timeline will dazzle in March and collapse the following February. The traceability evaluation is the discipline of testing the platform against the second timeline while you are still living in the first. That is the only point at which the test is cheap, because afterward the cost of being wrong is a live, public, assured misstatement rather than a vendor you declined to sign.
You are not buying a number generator. You are buying an evidence-production system. If it cannot hand the evidence to your assurer without you in the room, the speed is a liability dressed as a feature.
The Four Questions That Separate Assurable From Glossy
Strip the evaluation down to four load-bearing questions. Each maps to a thing an assurer will actually do to you. A platform that cannot pass all four is not disqualified from your toolkit, but it is disqualified from being the system of record for any number you will publish and assure.
One: Can It Show Provenance on Every Factor?
An emission factor is the coefficient that turns activity data into emissions, kilograms of CO2 equivalent per unit of activity. Provenance is the named, dated, authoritative source behind that coefficient: which database, which version, which identifier, published when. The single most common silent failure in reporting AI is the hallucinated factor, a confident, plausible coefficient with no source behind it, which then corrupts every line it touches. So the first test is mechanical: pick any factor the platform applied and ask it to show you the source on screen, named and dated, with a click-through. Not "we use recognized databases" in a brochure. The actual factor, the actual database, the actual version, on the actual line. If the platform can only tell you it used "industry-standard factors" but cannot show you which factor came from which dated source for a specific line, it is laundering unsourced numbers into your inventory, and you will own that in the engagement.
Be specific about why the version and the date matter, because vendors will offer you a database name and hope you stop there. Factor libraries are revised; a coefficient for purchased aluminium published in one release can differ from the next, and an assurer comparing two reporting years will ask which version applied to which year and why the number moved if it did. A platform that names the database but cannot pin the version and date for a given line has given you provenance that is too coarse to defend. The same applies to the identifier: a named, dated database with a specific factor identifier lets you and the assurer go to the exact entry and confirm it says what the platform claims. Provenance is not a logo on a slide; it is a coordinate precise enough that a stranger can stand on it. Anything less is a story about where the number came from, and an assurer does not assure stories.
Two: Can It Reconstruct a Number?
Reconstructability is the assurer's core move. They take a published figure and walk it backward to raw evidence: this Scope 3 total decomposes into these category lines, this line is this activity datum times this factor, this activity datum came from this supplier document at this location on this date, this factor came from this dated database. The test is to take one number the platform produced and ask it to walk you back, step by step, to the source documents and the factor sources, with the arithmetic visible at every level. A platform that shows you a total and a confidence score but cannot decompose the total into traceable lines has given you a result you cannot defend. The number may even be correct. Correct and indefensible are different states, and only one of them survives assurance.
Three: Can It Export the Audit Trail?
Your assurer does not live inside the vendor's platform and will not be issued a login to poke around your live environment. They work from an assurance file, an offline, portable evidence package they can read, annotate, and retain. So the platform must export the audit trail: the activity data with sources, the factors with provenance, the methods, the human decisions and overrides, the version history, all in a format that stands on its own outside the tool. The test is to ask for a sample export of one disclosure's full lineage and hand it to someone who has never seen the platform, then ask them to follow a number from published figure to raw source using only the export. If they can, the platform is assurance-ready. If the lineage only exists inside the live dashboard and evaporates on export, you have a system that cannot feed an assurance engagement, however good it looks logged in.
Four: Where Does It Estimate, and Does It Say So?
This is the question vendors least want asked, and the one that matters most for Scope 3, which averages roughly 75% of a footprint across the fifteen GHG Protocol categories and where 79% of reporters cite supplier-data availability as a top barrier. No platform gets primary data for everything; it fills gaps with estimates. Defensible estimation is fine, even necessary. Fabrication dressed as measurement is fatal. So the test is: show me every place this platform estimated rather than measured, show me the method (spend-based, average-data), show me the uncertainty, and show me that an estimated figure is labeled secondary and never looks identical to a primary, supplier-reported figure in the output. A platform that quietly backfills missing supplier data with industry averages and presents the result as a clean measured number is the single most dangerous tool you can buy, because it manufactures exactly the assurance finding that becomes a restatement and a greenwashing headline.
The reason this test outranks the others in practical danger is that the failure is invisible until the assurer finds it. A hallucinated factor can sometimes be caught by a reviewer who knows the right range; a missing export announces itself the day you try to export. But a silently backfilled estimate looks exactly like good data. It sits in the inventory wearing the costume of a measured supplier figure, it sums into a clean total, and nothing on the surface betrays it. That is precisely why it is the move that fails assurance: the assurer's job is to lift the costume, and when they ask "is this line a supplier-reported figure or an estimate," the platform that cannot answer, or that answered wrong by labeling it primary, has converted a routine question into a finding. The right behavior is the opposite of invisible: an honest platform makes its estimates loud. It flags the gap, names the method, shows the uncertainty, and tags the line secondary, so that the estimate is disclosed rather than disguised. When you watch a platform handle a gap, you are watching the single most diagnostic thing it will ever show you.
The Vendor Landscape, as Orientation Only
It helps to know the category map, purely to orient yourself, never as an endorsement. The market splits loosely into broad ESG and CSRD reporting platforms that handle disclosure assembly and tagging across frameworks, and carbon-accounting and footprinting tools focused on the GHG inventory. Examples of the category, named only so you recognize the shape of the market and not because any of them is recommended, include the reporting-platform cluster and the footprinting cluster that any procurement search will surface. The point is the opposite of a shortlist: no platform's output is trustworthy because of the platform's name. A factor is not sourced because a well-funded tool produced it. A number is not assurable because a famous logo computed it. The four tests apply identically to every vendor regardless of brand, funding, or customer roster, and a smaller tool that passes them beats a market leader that does not.
The deeper point, which the next lessons in this chapter develop, is that the obligation does not transfer to the platform. When you publish a number, you publish it. The assurer holds you accountable, the regulator holds you accountable, and "the platform calculated it" is no more a defense than "the AI estimated it." You are buying a tool that helps you discharge an obligation you keep. Evaluating for traceability is how you make sure the tool actually helps, rather than handing you speed and quietly keeping the evidence to itself.
A Traceability Scorecard You Can Take Into the Room
Turn the four tests into a scorecard you score live during the evaluation, on the vendor's own data and ideally on a slice of yours. Score each dimension 0 to 2: 0 means the capability is absent or "on the roadmap," 1 means partial or manual, 2 means demonstrated on screen end to end. Anything that scores on a slide rather than in the product scores 0.
| Dimension | The question you ask, live | 0 | 1 | 2 |
|---|---|---|---|---|
| Factor provenance | "Show me the named, dated source and version for this specific factor on this line." | Brochure claim only | Source named but not dated or versioned | Named, dated, versioned, click-through on the line |
| Reconstructability | "Walk this published number back to raw source, step by step, with the arithmetic visible." | Total plus a confidence score only | Partial decomposition, gaps in the chain | Full walk-back to source documents and factor sources |
| Audit-trail export | "Export this disclosure's full lineage to a file my assurer can read offline." | No export; lives only in dashboard | Export exists but is incomplete or unreadable standalone | Complete, standalone, followable by an outsider |
| Estimation transparency | "Show me everywhere you estimated, the method, the uncertainty, and the primary/secondary label." | Estimates hidden or unlabeled | Labeled but method or uncertainty missing | Every estimate labeled secondary, method and uncertainty shown |
| Data ownership and exit | "If we leave, do we keep the data, the lineage, and the provenance in usable form?" | Lineage lost on exit | Raw data exportable, lineage degraded | Full data and lineage portable on exit |
A fifth row, data ownership and exit, is added because a platform that holds your provenance hostage fails you the day you switch tools or the day an assurer wants the history from a system you no longer use. Set a threshold before the demo, not after, so the score is honest. A reasonable bar: no dimension below 1, and the three core traceability dimensions (provenance, reconstructability, export) at 2, before any number from this platform may reach a published, assured disclosure.
A Worked Example: Two Vendors, One Scorecard
Watch the scorecard do its work on two plausible vendors evaluated for the same Scope 3 program.
Vendor A, the market leader. The demo is flawless. It ingests invoices, categorizes spend across the fifteen categories, and produces a total in minutes, with a slick confidence dashboard. On the scorecard, asked to show the source of a specific freight factor, the platform displays "DEFRA-based factor" with no version and no date: provenance scores 1. Asked to reconstruct a category total, it decomposes one level to spend lines but cannot tie each line to a source document, because much of the activity data was inferred from spend: reconstructability scores 1. Asked to export an offline audit trail, the sales engineer offers a PDF summary that lists totals but not lineage: export scores 0. Asked where it estimated, the platform admits that missing supplier data is backfilled with sector averages, but these are not separately labeled in the standard output: estimation scores 1. Data ownership on exit: raw inputs export, lineage does not: 1. Total: 4 of 10, and two of the three core dimensions miss the bar. Beautiful, fast, unassurable as a system of record.
Vendor B, the smaller specialist. The demo is plainer and slower, and the sales engineer is a former assurance practitioner. Asked for a factor source, the platform shows the named database, the version, the publication date, and a click-through to the underlying entry on the line: provenance scores 2. Asked to reconstruct, it walks the total down to category, to line, to the supplier document and page the activity datum came from, to the dated factor: reconstructability scores 2. Asked to export, it produces a structured file that a colleague who has never seen the tool follows from published number to raw source: export scores 2. Estimation: every estimated line is tagged secondary, with method and an uncertainty range, visually distinct from primary lines: 2. Data ownership: full data and lineage portable on exit: 2. Total: 10 of 10. Slower in the demo, faster in the engagement, because every number arrives defensible and nothing has to be reconstructed under deadline.
The lesson of the comparison is not "buy the small one." It is that the demo ranked the vendors in exactly the wrong order, and only the scorecard, scored live on the four questions, surfaced the truth before a signature. Vendor A might still earn a place for low-risk drafting and triage. But the number that reaches your assured disclosure should come from the tool that can show its work, and the scorecard is how you tell them apart while you still have the leverage of an unsigned contract.
Phrasing the Questions So They Cannot Be Dodged
Vendors answer vague questions with vague reassurance, so make every question demand an on-screen demonstration on a specific line, not a capability statement. "Do you support audit trails?" gets a yes. "Export the full lineage of this number to a file and let my colleague follow it offline" gets the truth. Insist the evaluation runs on a slice of your real data, not the vendor's curated demo set, because the gap between the two is exactly where traceability quietly fails. And put the four tests and their pass thresholds into the procurement requirements and, where you can, the contract, so that "on the roadmap" during the demo becomes a contractual commitment with a date, rather than a promise that evaporates after signature.
There is a tell worth watching for in how a vendor reacts to these questions, because the reaction is itself information. A vendor whose platform is genuinely assurable tends to welcome the provenance and reconstruction questions, because answering them is how they win against glossier competitors, and a former assurance practitioner on the vendor's side will often pre-empt them. A vendor whose speed depends on hiding lineage tends to redirect: to the confidence dashboard, to the customer logos, to the roadmap, to the breadth of features, anywhere but the specific line you asked about. When the answer to "show me the source of this factor" becomes a tour of everything except the source of that factor, you have learned what you needed to learn. The questions are not hostile; they are the same questions your assurer will ask, asked early enough to matter. A vendor who cannot stand them in a sales meeting will not survive them in your engagement, and it is far cheaper to discover that now.
Key Takeaways
- You are not buying a number generator but an evidence-production system; the demo shows speed on a curated dataset and ends where your real job, defending the number to an assurer, begins.
- Four questions separate assurable from glossy: can it show provenance on every factor, can it reconstruct a number to raw source, can it export the audit trail offline, and where does it estimate and does it say so.
- Factor provenance means a named, dated, versioned database with a click-through on the actual line, not a brochure claim of "recognized factors"; the alternative is the hallucinated factor laundered into your inventory.
- Reconstructability is the assurer's core move: a number you cannot walk back to source documents and dated factors is correct-but-indefensible, and only defensible survives the engagement.
- The audit trail must export to a standalone file an outsider can follow offline, because your assurer works from a portable assurance file, not a login to your live dashboard.
- Estimation transparency matters most for Scope 3 (around 75% of a footprint, 79% supplier-data barrier): every estimate must be labeled secondary with method and uncertainty, never dressed up as a measured primary figure.
- The vendor category map (reporting platforms, footprinting tools) is orientation only; no output is trustworthy because of a brand, funding, or logo, and the obligation never transfers to the platform.
- Score the four tests plus data-ownership-on-exit live, 0 to 2, on a slice of your real data, set the pass threshold before the demo, and write the tests into procurement and contract while you still hold the leverage of an unsigned deal.
Skill.re