←
AI Agent Builders & Citizen Developers
Capable · M1 · lesson 1 of 25 · in progress
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI-Drafted QBR Decks with Verification
📖
now learning

AI-Drafted QBR Decks with Verification

15 min

QBR week is the worst week of the quarter for a CS Ops team. Eight CSMs each need to walk into a customer's executive room with a 12-slide deck that says the right ARR, the right user counts, the right churn signal, the right product adoption story. The CSMs spend Friday night assembling slides instead of preparing for the conversation. The decks ship with copy-paste errors, last-quarter's numbers, and one or two phrases that — let's be honest — nobody verified against Salesforce or Mixpanel. An LLM can draft these decks in under two minutes each. The question that decides whether you ship the workflow or get fired for it is whether you build the verification step that catches the three QBR hallucinations: made-up ARR, fabricated user counts, and inverted churn signals.

Why QBR Decks Are the Killer Use Case

QBRs (Quarterly Business Reviews) are the single highest-stakes customer-facing artifact a CS team produces. They go to executives. They drive renewal conversations. They land in board reports. They influence expansion decisions worth six and seven figures. They are also tedious to assemble — most of the slide content is mechanical aggregation from Salesforce, Gainsight, Mixpanel, Amplitude, or whatever product analytics tool the company uses. The intelligence is in the narrative; the labor is in the data plumbing.

An LLM, given clean inputs, can draft a QBR deck in under two minutes. It pulls ARR from Salesforce. It pulls active user counts from Mixpanel. It pulls usage trend from the product DB. It writes the talking points. It generates the executive summary. The CSM walks in with a 90% complete draft and spends time on the 10% that's actually intelligent — the conversation strategy, the renewal positioning, the expansion ask.

The math on time savings is real. A team of 8 CSMs running ~6 QBRs each per quarter spends roughly 4-6 hours per deck on assembly. That's 32-48 hours of CSM time per QBR cycle, per CSM, per quarter. With an AI-drafted workflow, that drops to 30-60 minutes per deck for the review-and-edit pass. Savings: roughly 25-40 hours per CSM per quarter, or 200-320 hours across an 8-person team.

But you cannot ship this workflow without verification. The cost of one bad QBR — a made-up ARR figure presented to a customer CFO, a "we've grown your usage 40%" slide that is mathematically wrong by 30 percentage points — is six months of trust rebuilding and a renewal at risk. The verification step is the line between an LLM-drafted QBR program and a career-ending mistake.

The Three QBR Hallucination Patterns

Practitioner experience as of May 2026 across multiple deployments converges on three specific failure modes. Memorize these.

Hallucination pattern 1: Made-up ARR

The LLM is given a Salesforce export with ARR data. Somewhere in the pipeline, the export is incomplete — the account doesn't have a current opportunity, or the ARR field is null, or the connector pulled the wrong fiscal year. The LLM, asked to "draft a QBR slide showing the customer's current ARR and growth trajectory," will not say "ARR data is missing." It will produce a confident number that sounds plausible. We've seen one team's draft deck quote $340,000 ARR for an account whose actual ARR was $280,000. Where did $340,000 come from? The model interpolated from the customer's headcount and some industry benchmarks in its training data. It made it up.

The fix is non-negotiable: ARR figures never come from the LLM. The slide template has {{ customer.current_arr }} as a literal field reference resolved by the workflow tool, not as text the model writes. The model writes the narrative around the ARR ("Acme has invested $280,000 in the platform this fiscal year, up from $215,000 last year"), but the numbers $280,000 and $215,000 are injected as template variables. The model never sees the numbers as text it might paraphrase; it sees them as {{ arr_current }} placeholders.

Hallucination pattern 2: Fabricated user counts

Similar shape. "Active users in the last 30 days" should come from Mixpanel or Amplitude. If the Mixpanel query failed silently (timeout, auth issue, customer ID mismatch), the LLM, asked to draft an "adoption metrics" slide, will produce a number. The number is often suspiciously round — "Acme has 142 active users monthly." 142 is not random; the model picked a number that "feels right" for a customer of that size. Actual active users: 38.

The cost of this hallucination is more subtle than the ARR case. The CSM walks into the QBR and says "you have 142 active users." The customer's IT lead, who knows actual usage, frowns. The conversation derails. Trust takes months to rebuild. We've seen this happen.

The fix is the same template-injection discipline, plus a hard rule: if the upstream query returns no data, the workflow halts the deck draft and routes to a "needs manual data review" queue. Never fall through to LLM-generated numbers.

Hallucination pattern 3: Inverted churn signals

The LLM is given a Gainsight Health Score input ("score: 42, trend: declining") and asked to summarize the customer's health in the executive overview slide. The narrative comes out as "The Acme partnership is in a strong position, with positive momentum across all dimensions." Where did "strong position" come from? The model inverted the signal. It read "score: 42, trend: declining" and the surrounding context happened to include words like "growth," "expansion," "investment," and the model defaulted to a positive tone because positive QBRs are statistically more common in the training data.

This is the most insidious hallucination because it doesn't violate a numeric fact — it inverts the qualitative interpretation. The CSM, scanning the draft quickly, might not catch it. The customer reads "strong position" while their own team is internally questioning whether to renew. The dissonance ends the conversation.

The fix is a structured Health summary input plus an explicit polarity check in the prompt: "The customer's health is summarized as {{ health_state }} (one of: healthy, watch, at-risk, critical). Your narrative MUST match this polarity. If health_state is 'at-risk' or 'critical,' the narrative tone is 'we have work to do.' Never describe a critical account as in a strong position." Plus a post-hoc verifier: a Code node that checks whether words like "strong," "excellent," "positive momentum" appear in slides when health_state ∈ {'at-risk', 'critical'} and flags the deck for human review.

The Workflow Architecture

The shape:

  1. Trigger: Scheduled (every Monday at 6am the week of QBR), or manual (CSM clicks "Draft QBR" in a Gainsight CTA).
  2. Pull data: Parallel HTTP nodes hit Salesforce (ARR, owner, contract dates), Gainsight (Health Score, recent Timeline highlights, open CTAs), Mixpanel/Amplitude (usage metrics, feature adoption), Zendesk (ticket counts, escalation history).
  3. Validate inputs: Code node confirms all sources returned data. Missing data → "needs manual review" branch. Stale data (older than 7 days) → flag for review.
  4. Build deck context object: One structured object with all numeric fields as numbers, all categorical fields as enums, all narratives as null (to be filled by LLM).
  5. LLM draft generation: One Anthropic node per slide template, or one big call with the full deck structure. Each call has access to numeric fields as data but must produce only narrative text.
  6. Verification node: Code node runs three checks (number fidelity, polarity, missing fields).
  7. Generate deck: Google Slides API or PowerPoint generation node renders the deck. Template has placeholders for every number; numbers come from step 4, narratives come from step 5.
  8. Review queue: Deck saved to a "draft" folder. CSM gets a Slack notification with the deck link and a verification summary (which inputs were stale, which sections need close review).
  9. Audit log: Postgres row with customer ID, deck URL, input snapshot hash, model used, verification result, timestamp.

The Deck Context Object

This is the most important data structure in the workflow. Get it right and the rest is mechanical.

{
  "customer": {
    "name": "Acme Corp",
    "sf_account_id": "0011A00001abc",
    "industry": "Manufacturing",
    "size_segment": "mid_market"
  },
  "commercial": {
    "current_arr": 280000,
    "prior_arr": 215000,
    "contract_start": "2023-02-01",
    "renewal_date": "2026-08-15",
    "days_to_renewal": 92,
    "expansion_opportunity_open": true,
    "expansion_opportunity_amount": 65000
  },
  "usage": {
    "mau_current": 142,
    "mau_prior_quarter": 118,
    "mau_growth_pct": 20.3,
    "feature_adoption": {
      "analytics_module": "active",
      "integrations_module": "not_adopted",
      "reporting_module": "active"
    },
    "data_freshness": "2026-05-12",
    "data_freshness_status": "fresh"
  },
  "health": {
    "score": 72,
    "state": "healthy",
    "trend": "stable",
    "open_risks": [],
    "recent_wins": ["successful EU rollout", "integrations team hired"]
  },
  "support": {
    "tickets_last_quarter": 8,
    "tickets_prior_quarter": 12,
    "escalations": 0,
    "avg_resolution_hours": 14
  },
  "call_summaries": [/* last 5 Gainsight Timeline entries from previous lesson */]
}

Every number is typed. Every categorical field is an enum. Every "narrative" field is missing (the LLM produces those). The LLM never sees these numbers as text it might paraphrase — it sees them as structured input data.

The System Prompt

You are a Customer Success operations writer drafting talking points for a Quarterly Business Review deck. You work for the SaaS platform; the customer is the audience.

You will receive a structured JSON object with the customer's commercial, usage, health, and support data. Your output is a JSON object with narrative text for each slide. You produce ONLY narrative; you do NOT produce numbers. Numbers come from template injection.

Critical rules:
1. NEVER state a number that is not in the input JSON. If a number is missing or null, write "[needs CSM input]" in your narrative. Do not interpolate from context.
2. The customer's health.state is authoritative. Your tone MUST match the polarity:
  - healthy: confident, partnership-building
  - watch: balanced, identifies areas to improve
  - at_risk: direct, names the problem, proposes a path
  - critical: candid, "we have work to do," focused on recovery plan
3. Reference template variables for numbers: write The customer has invested {{ arr_current }} this fiscal year not the raw number.
4. Each slide narrative is 2-5 sentences. No fluff. No marketing speak. Operations writers, not copywriters.
5. If the input shows missing or stale data (data_freshness_status: "stale" or any commercial/usage field is null), refuse to draft that slide and set the slide narrative to "[DATA QUALITY ISSUE — CSM REVIEW REQUIRED: ...specific issue]".

Slide structure (produce JSON keyed by slide name):
- executive_summary: 4 sentences. State of the partnership.
- commercial_overview: ARR trajectory, renewal status, expansion narrative.
- adoption_story: usage trend, feature adoption highlights and gaps.
- health_assessment: candid health summary tied to health.state.
- support_summary: ticket trends, escalation handling.
- next_quarter_focus: 3 prioritized initiatives for the next 90 days.
- asks: 1-3 specific asks of the customer (executive sponsor meeting, beta participation, case study).

Output: a single JSON object. No prose before or after.

The prompt does five critical things. It forbids numeric paraphrasing. It anchors tone polarity to a structured field. It enforces template-variable references. It demands brevity (2-5 sentences per slide). It defines the data-quality escape hatch.

The Verification Node (The Reviewer Checklist)

This is the difference between a workflow that ships and a workflow that destroys customer trust. The verification step runs three checks programmatically and adds a fourth human-review step the CSM must complete before the deck goes out.

Programmatic check 1: Number fidelity

Extract every number from the LLM's narrative output. For each number, verify it exists in the source data object. Reject if any number in the narrative doesn't match a value in the input JSON (within rounding tolerance — $280,000 and $280K are the same; $280,000 and $284,000 are not).

// extract all numbers from narrative (handles $284K, 142, 20.3%, etc.)
const numberPattern = /\$?(\d+(?:,\d{3})*(?:\.\d+)?)\s?(?:K|k|M|m|%)?/g;
const narrativeNumbers = Array.from(narrative.matchAll(numberPattern), m => normalize(m[0]));
const sourceNumbers = flattenAndNormalize(deckContext);
const orphanNumbers = narrativeNumbers.filter(n => !sourceNumbers.includes(n));
if (orphanNumbers.length) return { __verify: 'fail_number_fidelity', orphans: orphanNumbers };

This catches "made-up ARR" and "fabricated user counts" patterns. Every number in the narrative has to trace back to the source object. If it doesn't, the verifier blocks the deck.

Programmatic check 2: Polarity match

If health.state ∈ {'at_risk', 'critical'}, scan all narratives for positive-tone words: "strong," "excellent," "thriving," "outstanding," "robust," "healthy partnership," "exceptional." Flag any matches for human review. The check isn't a blanket reject — sometimes a critical account has a healthy support relationship — but the flag forces a CSM to verify.

The inverse: if health.state === 'healthy' and the narrative contains words like "at risk," "concerning," "deteriorating," "we have work to do" — flag for review. Polarity mismatches go both directions.

Programmatic check 3: Completeness

Every required slide must have a narrative or an explicit [DATA QUALITY ISSUE...] sentinel. Any slide with an empty narrative or a placeholder like "..." or "TODO" triggers a verification failure. The deck does not render with missing sections.

Human check: The 60-second reviewer checklist

After the programmatic checks pass, the deck is rendered and the CSM gets a Slack notification with a link to the draft and a 60-second checklist:

  1. Does the executive_summary tone match what you'd say in person? (Polarity gut check.)
  2. Are the renewal/expansion numbers correct? (Compare against the Salesforce opportunity tab.)
  3. Are any of the "recent wins" in health.recent_wins actually wins, or did the AI/data pipe surface a false positive?
  4. Did the AI invent any customer-side names, products, or events that aren't in your Gainsight notes?

The checklist is short on purpose. A 20-item checklist gets skimmed. A 4-item checklist gets read. Each item is binary; the CSM clicks "looks right" or "fix this." The "fix this" payload routes back to the CSM for manual edits, and the correction is logged for the weekly feedback review.

Real Deployment Experience

A team of 6 CSMs at a vertical SaaS company shipped this workflow for their Q1 2026 QBR cycle (late March 2026). Roughly 35 QBRs were drafted. Time to first draft per deck: 90 seconds. Time from draft to ship: averaged 35 minutes per deck (CSM edits, verification, customer-side personalization).

Before the workflow: average deck assembly time was 4.5 hours per QBR. After: 35 minutes review + 90 seconds draft. Savings: ~4 hours per deck, ~140 hours across the 35-QBR cycle.

What the verification step caught:

  • 3 made-up ARR values. All three traced to a stale Salesforce sync (the account had an opportunity update that hadn't replicated to the source data warehouse yet). The verification fail routed to manual review, the CSM pulled the fresh number, the deck regenerated.
  • 1 inverted polarity. A critical-health account got a draft that opened with "Strong momentum across the partnership." The polarity check fired. The CSM edited the narrative manually. No customer ever saw the broken version.
  • 2 fabricated feature adoption claims. The model invented a "successful rollout of the AI co-pilot module" in one deck — the customer had not purchased that module. The number-fidelity check passed (no numbers were wrong) but the human reviewer caught it. Lesson: the human checklist needs to include "any invented product names or features?"

The cost saved isn't just labor. The avoided customer-trust damage is the bigger return. One mis-stated ARR in front of a CFO is worth multiple full-time-equivalent quarters of recovery. The verification step is what makes the workflow safe to ship at the executive-customer level.

Why the CSM Must Still Review

There is a temptation, once the verification node passes, to ship the deck directly. Don't.

The verification node catches the things you can check programmatically: number fidelity, polarity, completeness. It does not catch:

  • Tonal mismatch with this specific customer's communication style (Acme prefers direct; the AI wrote diplomatic)
  • Politically sensitive framing (the AI proposed an expansion when the customer's sponsor just announced a hiring freeze)
  • Recent context that isn't in any database (the CSM had a heads-up call last week that the customer's CEO changed)
  • Strategic positioning ("don't mention the integrations gap on slide 4 because we're meeting with their integrations team next month and want to time the conversation")

The AI drafts. The CSM edits. The verification node catches the things the CSM might miss in their 35-minute review. Each layer does what it's best at. None of them ship alone.

The shortest road to losing trust in an AI workflow is shipping the first thing it produces. The shortest road to keeping trust is treating the AI as a fast first draft, the verification as a typo check, and the CSM as the editor. None of those three steps is optional.

What Not to Let the AI Do

A short list of QBR tasks where you should not deploy the LLM in 2026, even with verification:

  1. Don't let the AI write competitive comparisons. The cost of an inaccurate "we beat Competitor X on feature Y" claim is a credibility hit you cannot recover. Competitive content is human-written.
  2. Don't let the AI write the "asks" without CSM input. The AI suggests asks; the CSM picks them. Asking a customer for a case study when they just escalated three tickets is a relationship-killer.
  3. Don't let the AI talk about future roadmap. Product roadmap content goes through product marketing review. The AI has no idea what's been moved, what's been killed, or what's NDA'd.
  4. Don't let the AI quote unverified customer wins. If health.recent_wins contains "successful EU rollout," verify before the deck ships. The AI is just reading the array.
  5. Don't let the AI produce the deck PDF directly to the customer. Render to a draft folder. CSM exports to customer-facing PDF. Manual final step. The 5 seconds of friction is the safety net.

The Feedback Loop from CSM Edits

Every CSM edit becomes a data point. The workflow stores the original LLM draft, the final shipped deck, and the diff. Once a quarter, the CS Ops lead reviews the diffs for patterns. Common patterns observed:

  • "The AI was too verbose; CSMs cut 30%+ of the executive_summary every time." Fix: lower the sentence count from 4 to 3 in the prompt.
  • "The AI keeps suggesting 'executive sponsor meeting' as an ask for accounts under $50K ARR." Fix: add to prompt: "Asks should be proportional to account ARR — case studies and beta participation for SMB, executive sponsor for enterprise."
  • "The AI's 'next quarter focus' is too generic." Fix: add to prompt: "next_quarter_focus items must reference specific feature names or specific customer-stated goals from health.recent_wins or call_summaries; no generic suggestions."

This is the same Friday-review discipline as the call summary workflow in Lesson 1, applied to QBR drafts. Prompt versioning. Holdout testing. Steady improvement.

Key Takeaways

  • QBR decks are the highest-ROI LLM workflow in CS Ops, but only with verification. Time savings: ~4 hours per deck, hundreds of CSM hours per quarter at a 6-8 person team scale.
  • The three QBR hallucination patterns: made-up ARR (interpolated from headcount + training data), fabricated user counts (round numbers that "feel right"), and inverted churn signals (positive tone defaulted when context is mixed).
  • The structural fix is template injection: numbers come from data sources via template variables; the LLM produces only narrative. Numbers never pass through the model as text it might paraphrase.
  • The polarity anchor is structural too: health.state is a typed enum, and the system prompt makes the narrative tone explicitly contingent on it. "If state is at_risk or critical, tone is 'we have work to do.'"
  • The verification node runs three programmatic checks (number fidelity, polarity, completeness) and surfaces a 60-second human checklist for the CSM.
  • Number fidelity check: extract every number from narrative, verify each exists in the source data object, fail on orphans. Catches both made-up ARR and fabricated user counts.
  • Polarity check: if health.state is at_risk/critical, flag positive-tone words; if healthy, flag negative-tone words. Catches inverted churn signals.
  • Completeness check: every slide must have a narrative or an explicit data-quality sentinel. No silent failures.
  • Human review is non-negotiable. The CSM still owns 4 things the verifier cannot check: tonal style match, political sensitivity, missing recent context, strategic positioning.
  • Never let the LLM produce: competitive comparisons, account-asks without input, roadmap content, unverified wins, or the final customer-facing PDF.
  • Real Q1 2026 deployment: 35 QBRs drafted, 4-hour-per-deck savings, ~140 hours saved on the cycle. Verification node caught 3 made-up ARRs, 1 inverted polarity, and 2 fabricated feature claims that human review then escalated.
  • The feedback loop: store original draft and final shipped deck, review the diff each quarter, version the prompt based on observed CSM edit patterns. Steady improvement compounds.