←
AI for ESG & Sustainability Reporting
Strategic · M16 · lesson 16 of 23 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
PoC Design with an Assurance Gate
📖
now learning

PoC Design with an Assurance Gate

15 min

Six weeks into a proof of concept, the vendor's reporting AI has done something that looks like a miracle. Fed a tidy folder of fifty supplier emissions reports, it parsed every one, mapped them to the right Scope 3 categories, applied factors, and produced a category total in an afternoon. The pilot team is thrilled and the slide deck writes itself: "92% accuracy, weeks of work in hours, ready to scale." Then the assurance partner, invited late as a courtesy, asks to see the file behind one number. There is no file. The fifty reports were the clean ones, hand-picked because they were machine-readable; the real supplier population is messy, partial, and silent. The PoC proved the tool could fly in still air. It proved nothing about whether the number it produced could survive an assurance engagement. This lesson is how to design a proof of concept that proves the right thing.

What Most Proofs of Concept Actually Prove

A badly designed PoC is a demo you ran yourself. It uses a curated dataset, measures speed and a rough notion of accuracy, and declares success when a number appears. The trouble is that none of those are the thing that matters for a regulated discloser. The number that emerges from a PoC will, if you scale the tool, eventually become a published figure under an external assurance engagement, and 73% of large global companies now obtain assurance on at least some sustainability disclosures. So the only success criterion that counts is whether the PoC's output could survive that engagement. A PoC that proves speed and ignores assurability has answered the easy question and left the expensive one untested.

The deeper error is the dataset. A PoC run on the cleanest, most machine-readable slice of your data measures the tool under conditions you will never actually face. Real value-chain data is the opposite of clean: 79% of reporters cite supplier-data availability and 62% cite internal data quality as top barriers. The categories that are hard are hard precisely because the data is partial, inconsistent, or absent, and those are the categories where the tool either earns its keep or quietly fabricates. A PoC on the easy data tells you how the tool behaves where you do not need help. The whole point is to test it where you do.

Notice the perverse incentive at work. The pilot team wants the PoC to succeed, because a successful PoC justifies the project, the budget, and the months already invested in choosing a direction. The vendor wants the PoC to succeed because it closes the sale. Both sides, without any bad faith, are pulled toward the curated dataset, the generous success criterion, and the early declaration of victory, because all three make the PoC look good. The assurance gate exists to counter that gravity. It is a deliberately uncomfortable criterion, set by the one party who is paid to be skeptical, that the friendly incentives cannot soften. Without it, a PoC becomes a ceremony that confirms a decision already made rather than a test that could change it, and a test that cannot fail is not a test. The whole value of a proof of concept is its capacity to tell you no while saying no is still cheap.

A proof of concept on the clean data proves the tool can fly in still air. The assurance gate asks whether it can land in the weather you actually report in.

The Assurance Gate: The One Criterion That Matters

An assurance gate is a single, explicit pass/fail criterion placed at the end of the PoC: the output must be reconstructable and defensible against a real assurer's expectation, not a demo's. The gate is not a score and not an average. It is a binary the PoC must clear before anyone is allowed to say the word "scale." Frame it as a sentence the whole pilot team agrees to before the PoC begins: at the end of this PoC, an assurer must be able to take a number the tool produced and walk it back to raw evidence, see what is primary and what is estimated, and find the audit trail in a file they can read, or the PoC has failed regardless of how fast or accurate it looked.

Two properties make the gate real. First, it is defined before the PoC starts, so success cannot be redefined to fit whatever the tool happened to do. Second, the standard is a real assurer's expectation, ideally tested by an actual assurance practitioner, not the pilot team's guess at what an assurer might want. The cheapest way to make a PoC honest is to invite your assurance contact into the design phase and ask one question: "If we scaled this, what would you need to see for a number it produces to be defensible?" Their answer is your gate. Everything else the PoC measures, speed, cost, accuracy, is secondary to clearing it.

The Two Things the Gate Tests

The gate decomposes into two concrete tests, the same two that define assurability everywhere in this program. The first is provenance: for a number the PoC produced, can you show the named, dated source of every input, the activity data back to a supplier document and the emission factor back to a named, dated database? A tool that produces a number with no traceable inputs has failed, however close to the right answer it landed. The second is reconstructability: can someone who was not in the room take the PoC's output and rebuild the number from raw evidence, following the chain from published figure to source? If the only person who can explain the number is the analyst who ran the tool, the number is not reconstructable, and an assurance engagement assumes that analyst is unavailable.

A third property rides alongside both: the honest treatment of estimation. Because real data is partial, the PoC will hit gaps, and how the tool handles them is the most diagnostic thing you will observe. A tool that flags the gap and labels its fill as a secondary estimate with a method passes; a tool that silently backfills the gap with an average and presents a clean primary-looking number fails the gate outright, because that is the exact behavior that produces an assurance finding at scale.

Designing the PoC So the Gate Means Something

Four design choices turn a PoC from a demo into a real test. Get these right and the gate's verdict is trustworthy; get them wrong and a passing PoC tells you nothing.

Use a Representative Data Slice, Including the Ugly Parts

Pick a slice of real data that includes the hard cases deliberately: the supplier who sent a scanned PDF, the one who reported in the wrong units, the one who never replied, the spend line with no activity data behind it. If the PoC dataset has no gaps and no mess, it cannot test how the tool handles gaps and mess, which is the whole question. A good PoC slice is not the easiest fifty reports; it is fifty reports chosen to mirror the real population's difficulty, so that clearing the gate on this slice predicts clearing it at scale.

Test Against a Real Assurer Expectation

Bring an assurance perspective into the PoC, not after it. If you can, have your external assurer or an internal audit colleague review a sample of the PoC's output against what a real engagement would demand. The difference between "we think this looks defensible" and "our assurer confirmed this would be defensible" is the difference between a PoC that de-risks scaling and one that merely postpones the discovery of risk to the most expensive moment.

Seed Known-Answer Cases to Catch Fabrication

Plant a few cases where you already know the correct answer and the correct source, including at least one where the right behavior is to refuse or flag a gap rather than produce a number. A tool that confidently produces a number for the case where the honest answer is "no data" has revealed its most dangerous failure mode under controlled conditions, which is exactly what a PoC is for. Catching one hallucinated factor in a PoC is worth more than any accuracy percentage on the slide.

Require the Tool to Produce the Assurance Artifact, Not Just the Number

Make the PoC's deliverable the evidence package, not the headline figure. The tool must output, for the slice, the numbers plus their lineage: activity data with sources, factors with provenance, estimates labeled with method and uncertainty, and a log of decisions, all in an exportable, standalone form. If the tool can produce the number but not the artifact, the PoC has surfaced the exact deficiency that would sink you at scale, while it is still cheap to walk away.

Designing the PoC Dataset to Mirror Real Assurance Conditions

The single highest-leverage decision in a proof of concept is the one made before any tool runs: what data goes in. A demo dataset and an assurance-grade dataset can be the same size and produce the same headline number, yet only one of them predicts what happens at scale. The discipline is to build a slice that reproduces, in miniature, the exact conditions an assurer will encounter when the tool is live. That means the slice must carry the same texture of difficulty as the real value-chain population, not a sanitized excerpt of it.

Start from a simple principle: the assurer does not test your tool on your best day. An assurance engagement samples across the population, and the sample will land, sooner or later, on the supplier who never replied, the file that arrived as a photographed invoice, the line item where activity data simply does not exist. If your PoC slice contains none of those, the PoC has quietly excluded the very cases that decide whether the number is defensible. A representative slice is therefore not a random sample; it is a deliberately stratified one, engineered to contain the hard cases in roughly the proportion they occur in reality, so that clearing the gate on the slice is evidence, not luck.

What a Mirror Dataset Contains

Compose the slice so that it stresses the tool where reality will. A workable composition for a Scope 3 category slice includes clean machine-readable supplier reports in the same rough proportion they exist in your data, several files in awkward formats (scanned PDFs, images, spreadsheets with merged cells), at least one report in non-standard units, a handful of suppliers who never responded so the tool must confront a genuine gap, at least one spend line with no activity data behind it, and one or two records with an internal inconsistency such as a total that does not match its components. The point is not volume. Fifty well-chosen lines that mirror the population's difficulty tell you more than five hundred clean ones.

Recall the anchor numbers that make this non-negotiable: Scope 3 averages roughly 75% of a company's total footprint, and reporters cite 79% supplier-data availability and 62% internal data quality as their top barriers. Those barriers are not edge cases; they are the median condition. A PoC dataset that does not contain them is testing the tool on the 25% of the problem that was never hard.

Pre-Label the Slice So You Can Grade the Tool

A mirror dataset is only useful if you know the truth about it in advance. Before the tool runs, record for each record what the correct treatment is: what the right factor and source would be, whether the honest answer is a number or a flagged gap, and which records should be labeled secondary rather than primary. This pre-labeled key is what lets you grade the tool's behavior objectively rather than admiring an output you cannot check. Without it, you are back to trusting the number because it looks plausible, which is the failure the whole exercise exists to prevent.

The Exact Assurance-Gate Criteria

An assurance gate is only as good as the specificity of its criteria. "Could this survive assurance" is a direction, not a test. To make the gate operate as a real pass/fail, decompose it into four concrete, observable criteria, each of which the tool's output either demonstrably meets or does not. These four are the practical expression of the provenance-and-reconstructability standard, and they map directly onto what an assurer physically does when they pull a thread.

Criterion One: Provenance Shown

For any number the tool produced, the output must display the named, dated source of every input. Activity data must trace to a specific supplier document; the emission factor must trace to a named database, with its version and date. Not "an emission factor was applied" but "factor X from database Y, version Z, dated D." If the provenance is implied, inferred, or living only in the tool's internal state, the criterion is not met. An assurer cannot accept a source they cannot see.

Criterion Two: Number Reconstructed

Someone who was not present when the tool ran must be able to take the output and rebuild the number from raw evidence, arriving at the same figure. This is the test of reconstructability, and it is stricter than provenance: it is not enough to name the inputs, the arithmetic from inputs to result must be reproducible by an outsider. If the only person who can explain how activity data plus factor became the reported figure is the analyst who ran the tool, the number fails, because assurance assumes that analyst is unavailable.

Criterion Three: Audit Trail Exported

The tool must be able to export the full evidence package in a standalone, readable form that does not depend on a live login to the vendor's platform. The assurance file is a durable artifact; it has to survive a change of tool, a change of staff, and the passage of years. A trail that exists only as a screen inside the product, or only for as long as the subscription is active, is not an exported audit trail. The test is blunt: can you hand the assurer a file, offline, that contains the numbers and their complete lineage?

Criterion Four: Estimation Boundaries Labeled

Every estimated value must be visibly labeled as an estimate, with its method, and separated from measured primary data. Where the tool filled a gap, the output must say so, name the estimation approach (for example spend-based averaging), and tag the value secondary. The failure this criterion catches is the most dangerous in the entire program: an estimate presented as measured activity data. If measured and estimated values are indistinguishable in the output, the tool is laundering estimation into apparent measurement, and the criterion fails regardless of how accurate the estimate happens to be.

A number the tool cannot show the source of, cannot let an outsider rebuild, cannot export offline, and cannot mark as estimated is not a proof of concept. It is a demo wearing a number as a costume.

A Scored Go/No-Go Rubric

Translate the four criteria into a rubric the pilot team scores together, in the open, against the mirror dataset. Score each criterion pass or fail on the hard cases specifically, not on the clean ones, because the clean cases were never the question. The rubric below turns the gate from a judgment call into a defensible record of what the tool did.

CriterionPassFailWeight
Provenance shownEvery produced number displays named, dated source for activity data and factor (version included)Any number lacks a traceable, named source for an inputGating
Number reconstructedAn outsider rebuilt at least one sampled number to the same figure from raw evidenceOnly the operator could explain the number, or the rebuild divergedGating
Audit trail exportedFull lineage exported offline in a standalone, readable formTrail exists only inside the tool or only while logged inGating
Estimation labeledEvery estimate marked secondary with method, separated from primaryAny estimate presented as primary or measured dataGating
Gap behaviorTool flags no-data cases rather than fabricating a valueTool produces a confident number where the answer should be a flagGating
Speed and costMeaningful time or cost saving over the manual baselineNo material efficiency gainSecondary

The weighting is the point. The five assurance criteria are gating: a fail on any one of them is a fail overall, no matter how strong the rest. Speed and cost are secondary: they can inform the vendor conversation and the business case, but they can never rescue a tool that cannot produce a defensible number. A tool that is fast, cheap, and untraceable scores a no-go. A tool that is slower and pricier but clears all five gating criteria scores a go. This ordering is deliberately the opposite of the demo instinct, and enforcing it in writing is how the gate resists the friendly gravity that pulls every pilot toward yes.

Failure Modes That Sink a PoC

Even teams that intend to run an honest PoC fall into a small number of recurring traps. Naming them makes them easier to refuse.

The Curated-Dataset Trap

The most common failure is testing on the clean slice because it is the easiest to assemble and the most flattering to the tool. It produces a confident yes that predicts nothing about the messy population, and it is doubly dangerous because it feels rigorous: fifty real supplier files were processed. The defense is the mirror dataset and its pre-labeled key. If the slice contains no gaps, the PoC cannot report how the tool handles gaps, and that silence will be read as success.

The Moving-Goalpost Trap

When the gate is defined after the tool has run, success gets quietly redefined to fit whatever the tool did well. The factor sources it could not show become "a configuration detail for later"; the silent backfill becomes "something we will govern in production." The defense is writing the gate and the rubric down before the PoC starts and having the assurer, not the pilot team, own the standard. A criterion agreed in advance cannot be softened after the fact without someone having to say so out loud.

The Demo-Dazzle Trap

A tool that parses a scanned PDF in seconds or drafts fluent narrative is genuinely impressive, and impressiveness is not defensibility. The dazzle of a smooth demo pulls attention away from the unglamorous questions, where is the source, can an outsider rebuild this, is this estimate labeled. The defense is the rubric's discipline of scoring the four assurance criteria on the hard cases and treating speed as secondary. Watch what the tool does on the supplier who never replied, not what it does on the clean report.

The Absent-Assurer Trap

Inviting the assurer late, as a courtesy after the decision is effectively made, guarantees that the PoC tested the pilot team's guess at assurer expectations rather than the real thing. The gap between "we think this looks defensible" and "our assurer confirmed this would be defensible" is exactly the gap that reappears as a finding on a live disclosure. The defense is to bring the assurance perspective into the design phase and let their answer to "what would you need to see" become the gate.

A Worked Example: Two PoCs, Same Tool, Different Gate

Watch the same vendor tool run through a demo-style PoC and an assurance-gated PoC, and produce opposite verdicts.

The demo-style PoC. The team feeds the tool fifty clean, machine-readable supplier reports. It parses all fifty, categorizes them, applies factors, and returns a Category 1 total of 18,400 tonnes CO2e in an afternoon. Measured against a hand total, it lands within 8%. The slide reads "92% accurate, ready to scale." No one asked for the source of any factor, no gaps existed in the data, and no assurer saw the output. The verdict is a confident yes, and it is hollow, because nothing about an actual reporting cycle was tested.

The assurance-gated PoC. The same tool gets a deliberately representative slice: forty reports of varying quality, five suppliers who never responded, three scanned PDFs, two reports in non-standard units, and two seeded known-answer cases, one of which has no underlying data and should be flagged, not filled. The gate is set in advance with the assurer: every produced number must be reconstructable to source, with primary and secondary clearly separated and an exportable trail. The results are revealing. The tool parses the clean reports well and even handles two of the three scanned PDFs. But on the five non-responders it silently inserts spend-based averages and presents them as ordinary line items with no secondary label. On the seeded no-data case, it confidently produces a number. And when asked to export the audit trail, it produces a totals summary with no factor sources. Against the gate, this PoC fails: the tool fabricates on gaps, mislabels estimates as primary, and cannot export a defensible trail.

Same tool, same week, opposite conclusions. The demo PoC would have sent this tool to enterprise scale, where its silent backfilling would have laundered estimates into a published, assured Scope 3 number across the whole value chain, and the assurer would have found it. The gated PoC caught all of it on a fifty-line slice, for the cost of a careful design, while there was still no signature and no exposure. The gate did not make the tool worse; it told the truth about the tool earlier, which is the entire purpose of a proof of concept.

Reading a Failed Gate Correctly

A failed gate is not always a rejected vendor. It is a precise diagnosis. The gated PoC above produced an actionable list: the tool needs a configuration that forbids silent backfill, a mandatory secondary label on estimates, a refusal behavior on no-data cases, and a real lineage export. Some of those may be solvable with the vendor before scaling; some may be architectural and fatal. Either way, you now know exactly what to fix or walk away from, in specific terms tied to assurer expectations, rather than discovering it as an assurance finding on a live disclosure. The PoC's value is not the green light; it is the truthful verdict, and a gated PoC that fails has done its job better than a demo PoC that passes.

Key Takeaways

  • A badly designed PoC is a demo you ran yourself: curated data, speed and accuracy metrics, success declared when a number appears, none of which is the question that matters for a regulated discloser.
  • The only success criterion that counts is whether the PoC's output could survive an external assurance engagement, because the number it produces will, at scale, become a published, assured figure.
  • The assurance gate is a single, explicit, pass/fail criterion set before the PoC starts: an assurer must be able to reconstruct a produced number to raw evidence, see primary versus estimated, and find the trail in a readable file.
  • The gate is real only when it is defined in advance, so success cannot be redefined to fit the tool, and tested against a real assurer's expectation rather than the pilot team's guess.
  • The gate tests provenance (named, dated sources for every input) and reconstructability (an outsider can rebuild the number), plus the honest treatment of estimation: a tool that silently backfills gaps and labels them primary fails outright.
  • Design the PoC with a representative data slice including the ugly cases, a real assurer perspective, seeded known-answer cases that include a should-refuse case, and a required assurance artifact as the deliverable, not just the headline number.
  • In the worked example the same tool passes a demo PoC at 92% accuracy and fails the gated PoC by fabricating on gaps, mislabeling estimates, and failing to export a trail, catching at fifty lines what would have been an assurance finding at enterprise scale.
  • A failed gate is a precise diagnosis, not always a rejected vendor: it produces an actionable fix list tied to assurer expectations, while there is still no signature and no exposure.