←
AI for ESG & Sustainability Reporting
Strategic · M4 · lesson 4 of 23 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Avoiding Metrics That Hide Risk
📖
now learning

Avoiding Metrics That Hide Risk

15 min

The reporting AI program at a large manufacturer looked exemplary right up to the week it did not. Every metric on the steering deck was green for eighteen months. The report was generated in record time. The disclosure dashboard showed full coverage, every category populated, no red cells. Leadership held it up as the model other functions should copy. Then the assurer arrived, pulled a thread in the Scope 3 numbers, and the whole thing unravelled: a large share of the "populated" categories turned out to rest on estimates the model had generated to fill gaps, several factors could not be traced to a source, and the primary-data share, which nobody on the steering deck had been tracking, had quietly fallen for a year and a half. The engagement produced a wall of findings. The disclosure was restated. The program that had been the model became the cautionary tale. Nothing had gone wrong with the work that the metrics measured. Everything had gone wrong with what the metrics chose not to measure.

This lesson is about that gap: the vanity metrics that make a reporting program look healthy while risk builds underneath, and the counter-metrics that surface the truth in time to act. Speed is not assurability. A green dashboard is not a defended disclosure. And the most dangerous metric is not a wrong one; it is a flattering one that everybody watches while the number that actually predicts failure goes unmeasured. A strategist's job is to know which metrics reassure and which metrics reveal, and to make sure the revealing ones are on the deck.

What makes a metric a vanity metric

A vanity metric is not a false metric. It is a true metric that measures the wrong thing for the decision at hand, chosen because it moves in a flattering direction and is easy to report. In reporting AI the vanity metrics share a family resemblance: they all measure the ease of producing the disclosure, and none of them measure whether the disclosure can be defended. That is the whole problem in one sentence. In a regulated, externally-assured disclosure, the report is easy to generate and hard to defend, so any metric anchored to generation will look good long before, and often precisely while, defensibility is eroding.

Three vanity metrics do the most damage. The first is speed of report generation: how fast the draft comes together. It feels like the point of the whole program, and it is genuinely valuable, but on its own it rewards skipping checks, because the fastest cycle is the one that verifies the least. The second is dashboard greenness, or completeness by cell: every category populated, no gaps showing, all indicators green. It looks like thoroughness, but a cell is green whether it holds a supplier's measured figure or an average the model invented to close the gap, so greenness can rise precisely as the evidence base weakens. The third is volume of output: pages drafted, datapoints generated, tasks automated. It looks like productivity, but volume of AI-generated text and figures is not the same as volume of defensible text and figures, and a program can generate more while defending less.

There is a reason these three are so seductive to a steering committee specifically. They are the metrics that photograph well. They rise fast, they rise early, and they rise in response to exactly the investment leadership just approved, so they seem to confirm the decision. A leadership audience that has committed budget to a reporting AI program is primed to look for evidence that the commitment was right, and the vanity metrics supply that evidence in abundance, quarter after quarter, right up until the engagement. Worse, they are the metrics a vendor is happiest to instrument for you, because a vendor also wants to show that its tool is working, and its definition of working is your report coming out faster and more complete. So the default dashboard you inherit, the one that ships with the platform, is frequently built entirely from vanity metrics. Left unexamined, the tooling measures its own success and hands you the flattering picture, and the numbers that would contradict it are simply not in the product.

A vanity metric is a true number that measures how easily you produced the disclosure, not whether you can defend it. In a world where the report is easy to generate and hard to defend, watching only the easy-to-move numbers is watching the program build risk in green.

The counter-metrics that surface the truth

For every vanity metric there is a counter-metric that measures what the vanity metric conceals, and the discipline is to put the pair on the deck together so the flattering number can never be read alone. The counter-metrics all measure defensibility rather than ease, which is exactly why they are less comfortable to watch and more important to watch.

Primary-data share, falling

The counter to speed and to dashboard greenness is primary-data share: the proportion of your quantified footprint that rests on primary data, meaning supplier-reported or directly-measured activity data, rather than estimates and averages. When a program gets faster and more complete by letting the model fill gaps, primary-data share falls, because the new speed and the new coverage were bought with estimates. A falling primary-data share is the single most reliable early warning that a green program is hollowing out. It moves before the assurer arrives, which is what makes it a leading indicator, and it is hard to game, which is what makes it trustworthy. If you watch one counter-metric, watch this one, and watch its direction, not just its level: a share that is stable at a modest level is a known posture, but a share that is sliding is a program converting evidence into guesswork.

Assurance findings, rising

The counter to volume and to the general sense that the program is succeeding is the count and severity of assurance findings, read as a trend. Findings are the assurer's verdict on defensibility, and a rising count, especially of high-severity findings, is the clearest signal that the disclosure is becoming harder to defend even as it becomes easier to produce. The catch is that findings are a lagging metric, delivered once a year after the engagement, which is exactly how the manufacturer in the opening got eighteen months down the road before the truth surfaced. That lag is why findings must be paired with the leading counter-metrics: primary-data share and provenance completeness predict the findings months before the assurer confirms them, so a strategist watches the leading indicators to avoid being surprised by the lagging one.

Provenance gaps

The counter to dashboard greenness at the level of the individual number is the provenance gap rate: the proportion of disclosed figures that cannot be traced to a named, dated source. A dashboard cell can be green because it is populated while the number in it has no source anyone can point to, and provenance-gap tracking is what makes that invisible failure visible. Measure it continuously: for every material figure, is there a source that an assurer could follow from the published number back to the evidence? A rising provenance-gap rate means the disclosure is filling with numbers that exist without support, which is the precise condition that fails an engagement. It is the metric that would have caught the manufacturer's untraceable factors long before the assurer did.

Reading the pairs, and why the gap is invisible without them

The reason a vanity metric is dangerous is that, read alone, it is genuinely indistinguishable from success. Fast report generation looks the same whether the speed came from better tooling or from skipped checks. A green dashboard looks the same whether the cells hold measured data or invented estimates. High output looks the same whether the output is defensible or not. The vanity metric contains no information that would let you tell the healthy case from the failing one. Only the counter-metric carries that information, which is why the pairing is not optional.

So a strategist reads in pairs. Speed of generation is read against provenance-gap rate: fast and traceable is a real win, fast and provenance-gaps-rising is the trap. Dashboard greenness is read against primary-data share: fully populated with primary-data share holding is genuine completeness, fully populated with primary-data share falling is coverage bought with estimates. Volume of output is read against findings trend: more output with findings falling is capacity gained, more output with findings rising is risk manufactured at scale. In every pair, the vanity metric tells you the program is working and the counter-metric tells you whether that is true. Remove the counter-metric and you have removed the only number on the deck capable of contradicting the good news, which is exactly why failing programs so often have beautiful dashboards: the contradicting number was never there to be seen.

Notice what this means for how risk actually accumulates. It does not announce itself. There is no quarter where a vanity metric turns red to warn you, because vanity metrics do not measure the thing that is going wrong. Risk builds silently, in the space between the metric that is watched and the metric that is not, and the longer that space goes unmeasured the larger the eventual correction. This is why the discipline is preventive rather than reactive: by the time a problem is visible in the metrics a steering committee habitually watches, it is usually visible because the assurer made it visible, which is the most expensive moment for it to surface. The counter-metrics are not there to explain a failure after the fact. They are there to make the failure visible while it is still cheap to fix, which requires putting them on the deck before there is any sign of trouble, precisely when the program looks most healthy and the temptation to leave well enough alone is strongest.

Worked example: the green program that failed assurance

Return to the manufacturer and put numbers on it, so the mechanism is unmistakable. Every figure is illustrative; the point is the shape. The table below shows what the steering deck tracked, all of it green and improving, alongside what the steering deck did not track, all of it deteriorating.

MetricTypeYear 0Year 1On the deck?
Report generation time (weeks)Vanity159Yes, celebrated
Dashboard completeness (percent of cells populated)Vanity78100Yes, celebrated
Datapoints auto-generatedVanity1,1002,400Yes, celebrated
Primary-data share (percent of quantified)Counter4631No, untracked
Provenance-gap rate (percent of figures untraceable)Counter722No, untracked
Assurance findings (high-severity)Counter211Only at year-end, too late

Read the top three rows and the program is a triumph: generation time cut by forty percent, the dashboard gone from seventy-eight percent to fully populated, datapoints more than doubled. Every one of those numbers is true. Read the bottom three rows and the same program is a disaster in slow motion: primary-data share collapsed from forty-six to thirty-one percent, the provenance-gap rate tripled from seven to twenty-two percent, and high-severity findings, invisible until the engagement, jumped from two to eleven. The two halves describe the same program in the same year. The difference is not the reality; the difference is which metrics the steering deck chose to show.

Now watch the causal chain, because it is the same one every time. The tooling was configured to prioritise completeness and speed, so when supplier data did not arrive, the model filled the cell with an estimate rather than leaving a gap. That single behaviour lit up every vanity metric: generation got faster because gaps stopped blocking the cycle, the dashboard went green because every cell got populated, and datapoint volume soared because the model was generating figures to fill the holes. The same behaviour moved every counter-metric the wrong way: primary-data share fell because the new figures were estimates not evidence, the provenance-gap rate rose because the estimates had no traceable source, and findings climbed because the assurer will not accept a number that exists because the model guessed. One behaviour, opposite readings, and only the counter-metrics could see it. Had the manufacturer put primary-data share and the provenance-gap rate on the deck next to the speed and completeness numbers, the divergence would have been visible in the first quarter of year one, with eighteen months of runway to fix it instead of a restatement to explain.

Fixing the deck before the assurer does

The remedy is structural, not exhortation. You do not fix vanity-metric risk by telling people to be careful; you fix it by making the deck itself incapable of showing a flattering number alone. Three moves do that. First, mandate pairing: no efficiency or completeness metric appears on the steering deck without its counter-metric physically beside it, so speed is never shown without provenance gaps and greenness is never shown without primary-data share. A leadership audience that only ever sees the pairs cannot celebrate a green number it has not seen contradicted or confirmed.

Second, promote the leading counter-metrics to continuous tracking. Primary-data share and provenance-gap rate are computable throughout the cycle, not only at year-end, so track them monthly and put their trend lines, not just their current values, in front of leadership. A trend line makes a slide begin to fall impossible to ignore, which is the whole point, because it is the direction, caught early, that gives you time to act. Third, treat the lagging metric as confirmation, not discovery. Assurance findings should verify what the leading counter-metrics already told you, never surprise you. If the year-end findings are worse than your leading indicators predicted, the failure is not only in the disclosure; it is in a measurement system that let the risk hide, and that measurement system is the strategist's responsibility.

Underneath all of it sits the cardinal rule, which is also the reason the counter-metrics exist. Every figure must trace to evidence, and "the AI estimated it" is not evidence. The vanity metrics measure how efficiently you produced the disclosure. The counter-metrics measure whether the disclosure obeys the rule. A program that watches only the first set is measuring its own speed toward a restatement, in green, and calling it success. A strategist puts the second set on the deck so the program can see the truth while there is still time to change it.

Key Takeaways

  • A vanity metric is not false; it is a true number that measures how easily you produced the disclosure rather than whether you can defend it. In a world where the report is easy to generate and hard to defend, easy-to-move numbers look good precisely while defensibility erodes.
  • The three most dangerous reporting AI vanity metrics are speed of report generation, dashboard greenness or completeness by cell, and volume of output. Each rewards the exact behaviour that fills the disclosure with unsupported numbers.
  • Every vanity metric has a counter-metric that measures what it conceals: primary-data share (falling), assurance findings (rising), and provenance-gap rate. The counter-metrics measure defensibility, which is why they are less comfortable and more important to watch.
  • Primary-data share falling is the single most reliable early warning that a green program is hollowing out, because it moves before the assurer arrives and is hard to game. Watch its direction, not just its level: a sliding share is a program converting evidence into guesswork.
  • Assurance findings are the truth-telling verdict but a lagging one, delivered once a year, which is how a program can look green for eighteen months before failure surfaces. Pair findings with the leading counter-metrics that predict them months earlier.
  • Read metrics in pairs: speed against provenance gaps, greenness against primary-data share, volume against findings trend. A vanity metric read alone is indistinguishable from success, because it contains no information that could contradict the good news.
  • In the worked example, one behaviour, letting the model fill gaps with estimates, lit up every vanity metric green and moved every counter-metric the wrong way in the same year. Only the counter-metrics could see it; without them on the deck, the divergence stayed invisible until the restatement.
  • Fix the deck structurally: mandate that no efficiency metric appears without its counter-metric beside it, track leading counter-metrics continuously as trend lines, and treat year-end findings as confirmation rather than discovery. Every figure must trace to evidence, and a program watching only its own speed is measuring its way to a restatement in green.