←
AI Agent Builders & Citizen Developers
Capable · M18 · lesson 18 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
The Conversation-Summary-to-Gainsight Workflow
📖
now learning

The Conversation-Summary-to-Gainsight Workflow

15 min

A CSM finishes a 45-minute customer call at 11:14am. By 11:23am — before they have opened their notes app — a properly built workflow has pulled the Gong transcript, scored it against a coaching rubric, written a four-sentence summary to the Gainsight Timeline, posted a Slack ping to the CS lead if a churn signal fired, and updated the customer's Health Score component. The CSM hasn't touched a thing. This lesson is exactly how to build that workflow, end-to-end, with real prompt design, real Gainsight field mappings, and the three failure modes that send operators back to manual note-taking.

Why This Workflow Exists

The math is plain. A CS team running 60 customer calls a week, with each CSM spending 15 minutes after a call writing notes and Gainsight timeline entries, burns 15 hours of CSM time weekly on a task that an LLM does in 18 seconds. At a fully-loaded CSM cost of $90/hour, that's $1,350/week or about $70,000/year of CSM time spent transcribing a call they already had.

The cost isn't just labor. It's signal loss. CSMs run out of time and write three-sentence summaries that miss the churn flag the customer dropped at minute 32. They forget to surface the competitive mention. They don't tag the expansion opportunity because they had three more calls before lunch. The platform fills with thin, lossy notes. Renewal forecasts degrade. The VP of CS looks at QBR prep and sees nothing.

An LLM-driven summary doesn't tire. It applies the same rubric to call 60 of the week as it did to call 1. It catches the phrase "we've been looking at competitors" every time. It flags the customer who said "I'm not sure we'll renew" in the last two minutes. It writes consistent, structured notes that show up in Gainsight, Catalyst, or Vitally as a Timeline entry, a Health Score input, and (when warranted) a Risk flag with a reason.

The Stack as of May 2026

Three vendor categories, picked once per company:

Call transcript source

  • Gong dominates enterprise. Their /v2/calls endpoint exposes transcript, sentiment, topics, trackers, and call participants. The API is well-documented and supports webhook triggers on call completion. Pricing scales with seat count; expect $1,200-1,800/user/year for the AI tier.
  • Chorus (ZoomInfo) ships similar shape — calls API, webhook-on-complete, structured transcript with speaker labels. Slightly thinner integration depth than Gong but better pricing for mid-market.
  • Avoma is the upstart at $19-79/user/month. The /v1/transcriptions endpoint returns clean speaker-tagged transcripts. Smaller AI feature set than Gong but the workflow doesn't care — we're using our own LLM for the summary.

CS platform destination

  • Gainsight is the 800-pound gorilla. The Timeline API (REST) accepts POST /v1/timeline with activity type, subject, body, and arbitrary custom fields. Health Score components update via the Scorecard API. Gainsight's "Calls to Action" (CTAs) trigger from Health Score changes or direct API creation.
  • Catalyst (now part of Totango as of late 2025) offers a cleaner API. POST /v1/customers/{id}/timeline mirrors Gainsight's pattern. Easier for mid-market teams.
  • Vitally is the developer-friendly choice. Their API is well-shaped, their webhooks fire fast, and their custom-field handling doesn't make you cry. Pricing favors smaller portfolios.

Workflow tool

n8n self-hosted or Cloud is the practitioner's default for this pattern because the Gong/Chorus/Avoma APIs require careful authentication, pagination, and transcript parsing that hits Zapier's complexity ceiling. Make is the close second. Zapier works but stops being cost-effective above ~30 calls/day (see the 60,000-task threshold lesson). We will use n8n examples; Make is structurally identical.

Workflow Architecture

The shape:

  1. Trigger: Gong webhook fires on call completion (event type call.completed).
  2. Fetch transcript: HTTP node calls GET /v2/calls/{id}/transcript with the Gong access token.
  3. Pre-process: Code node concatenates speaker-tagged lines, strips filler, truncates if >25k tokens.
  4. Look up customer: Search Gainsight for the account by domain or Gong account ID; pull current Health Score and recent timeline entries.
  5. LLM summarize: Anthropic node with system prompt (the coaching rubric) and user prompt (the transcript + customer context).
  6. Validate output: Code node parses the JSON, validates schema, falls back to a "needs human review" branch if invalid.
  7. Post to Gainsight: HTTP node creates a Timeline entry. If risk flag is "high," create a CTA. If sentiment is "expansion," tag the relevant playbook.
  8. Notify: Slack node pings the CSM and (conditionally) the CS lead.
  9. Audit log: Append a row to a Postgres or Airtable log with the call ID, summary hash, model used, timestamp, and any error.

Nine nodes. About 35 minutes to build the first version. About 4-6 hours total to ship something you'd let near a customer account.

The Coaching Rubric and Prompt Design

This is where most teams get it wrong. They write a prompt like "Summarize the following call transcript in 4 sentences." They get back four sentences that are technically correct and operationally useless. The fix is the rubric.

What the rubric does

The rubric is the structured frame the LLM uses to read the transcript. It defines exactly which dimensions to score and exactly what to extract. A good rubric for a CS call has 7-12 categories. Ours:

  • summary_paragraph (4 sentences max, plain English)
  • sentiment (one of: very_positive, positive, neutral, negative, very_negative)
  • risk_flag (one of: none, watch, at_risk, critical)
  • risk_reason (string, only if risk_flag != none)
  • expansion_signal (boolean) + expansion_reason
  • competitive_mention (array of competitor names mentioned)
  • blockers (array of strings — what is in the customer's way)
  • action_items_csm (array of strings — what the CSM owes the customer)
  • action_items_customer (array of strings — what the customer owes us)
  • topics_discussed (array of tags from a controlled vocabulary)
  • quote_of_concern (string, verbatim — used for QBR slides; empty if none)
  • confidence (one of: high, medium, low — the model's self-rating of summary quality)

The system prompt

You are a Customer Success operations analyst. You read call transcripts between a CSM and a customer and produce a structured summary used by the CS team to update account records and identify churn or expansion signals.

Your output is a single JSON object matching the schema below. No prose before or after.

Apply the rubric strictly:
- sentiment: the OVERALL tone of the customer (not the CSM). Heavy weight on the last 10 minutes — that's where customers signal real intent.
- risk_flag: 'critical' only if the customer used language like 'we may not renew,' 'we're evaluating alternatives seriously,' 'leadership is questioning the investment,' or directly mentioned canceling. 'at_risk' if the customer is frustrated with multiple recent issues, expressed declining engagement, or has unresolved escalations. 'watch' for one-off frustrations without escalation. 'none' if no churn signal.
- expansion_signal: true ONLY if the customer mentioned wanting more seats, additional products, new use cases, or expansion to other teams.
- quote_of_concern: a verbatim quote (15-40 words) that captures the most operationally important thing the customer said. Used in QBR slides. Empty string if nothing rises to that bar.
- confidence: 'low' if the transcript was short (<5 minutes of speaking time), cut off, or contained too much off-topic chatter. 'medium' for normal calls. 'high' for calls with clear signal and abundant context.

If the transcript does not contain meaningful CSM-to-customer conversation (e.g., it's a demo recording, an internal meeting that was miscategorized, or sub-2 minutes), set confidence: 'low' and summary_paragraph to 'Insufficient content for analysis.'

Output schema:
{ "summary_paragraph": "...", "sentiment": "...", "risk_flag": "...", "risk_reason": "...", "expansion_signal": true|false, "expansion_reason": "...", "competitive_mention": [...], "blockers": [...], "action_items_csm": [...], "action_items_customer": [...], "topics_discussed": [...], "quote_of_concern": "...", "confidence": "..." }

That prompt is approximately 480 tokens. It does four important things at once:

  1. Defines the schema so the output is parseable.
  2. Anchors edge cases (what "critical" means, what "expansion" means) so the model doesn't drift.
  3. Lets the model say it doesn't know via the confidence: low escape hatch. This is the single most important trick to reducing hallucinations — a model with no way to admit uncertainty will fabricate.
  4. Defines the bar for quote_of_concern so it's not just any quote — it's the quote operations actually wants.

The user prompt

The user prompt is the transcript plus a small bundle of customer context:

Customer: Acme Corp (ARR: $84,000, on platform since Feb 2023, current Health Score: 72/100)

Recent timeline entries (last 60 days):
- 2026-04-22: QBR completed, expansion conversation tabled
- 2026-04-08: Support ticket #4421 escalated (data export performance)
- 2026-03-30: Renewal conversation initiated, 90 days out

Call participants: Sarah Chen (CSM), Marcus Diaz (VP Operations, Acme), Lin Park (Data Eng, Acme)
Call date: 2026-05-14
Call duration: 47 minutes

Transcript:
[Sarah] Hey Marcus, Lin, thanks for hopping on...
[Marcus] Yeah, no problem. Look, before we get into it...
(continues for ~5,800 tokens)

The customer context is doing serious work. Without it, the model can't tell whether "we're worried about renewal" is from a million-dollar account or a $4k pilot. The recent timeline entries tell the model whether this call is the third escalation in a month or a one-off. The participants tell the model whether the VP of Operations dropping a churn signal is more important than the same words from a data engineer.

Parsing and Validation

The Code node that runs after the LLM call has one job: turn the model's JSON output into a validated object, or branch to "needs human review" if it doesn't.

What can go wrong

  • Malformed JSON. Model wraps the JSON in ```json fences. Or adds prose before/after. Or truncates mid-output because of token limits.
  • Missing required fields. Model returns valid JSON but omits confidence.
  • Invalid enum values. Model returns sentiment: "mixed" when only the five enum values are allowed.
  • Hallucinated quote. Model invents a quote that's not in the transcript. This is the hardest to detect programmatically and we cover the manual check in the next lesson.

The validation code

The validator strips markdown fences, parses the JSON, validates against a schema (we use a simple JSON Schema check), and either returns the parsed object or routes to a "manual review" branch:

// n8n Code node, runs after the LLM output
const raw = $input.item.json.content;
const cleaned = raw.replace(/^```(?:json)?\s*/, '').replace(/```\s*$/, '').trim();
let parsed;
try {
  parsed = JSON.parse(cleaned);
} catch (e) {
  return [{ json: { __status: 'parse_error', error: e.message, raw } }];
}
const requiredFields = ['summary_paragraph', 'sentiment', 'risk_flag', 'confidence'];
const missing = requiredFields.filter(f => !(f in parsed));
if (missing.length) return [{ json: { __status: 'schema_error', missing, parsed } }];
const validSentiments = ['very_positive','positive','neutral','negative','very_negative'];
if (!validSentiments.includes(parsed.sentiment)) {
  return [{ json: { __status: 'enum_error', field: 'sentiment', parsed } }];
}
// quote check: if quote_of_concern is non-empty, verify it appears verbatim in the transcript
const transcript = $('Fetch transcript').item.json.transcript;
if (parsed.quote_of_concern && !transcript.includes(parsed.quote_of_concern)) {
  parsed.__quote_warning = 'verbatim_mismatch';
}
parsed.__status = 'ok';
return [{ json: parsed }];

The verbatim quote check is small but valuable. It catches the most common single hallucination mode in this workflow — the model paraphrasing a real customer concern and presenting it as a verbatim quote. Operators trust verbatim quotes; the model needs to actually deliver them.

Posting to Gainsight (or Catalyst or Vitally)

The Timeline entry write is one HTTP node. The payload shape varies by vendor; the pattern is the same.

Gainsight Timeline POST

POST /v1/timeline
Authorization: Bearer {GAINSIGHT_TOKEN}
Content-Type: application/json

{
  "companyId": "{gainsight_company_id}",
  "activityType": "Call",
  "subject": "Call summary — {{ $json.topics_discussed.join(', ') }}",
  "activityDate": "{{ $('Trigger').item.json.callDate }}",
  "notes": "{{ $json.summary_paragraph }}",
  "customFields": {
    "Sentiment__c": "{{ $json.sentiment }}",
    "Risk_Flag__c": "{{ $json.risk_flag }}",
    "Risk_Reason__c": "{{ $json.risk_reason }}",
    "Expansion_Signal__c": {{ $json.expansion_signal }},
    "Competitive_Mention__c": "{{ $json.competitive_mention.join('; ') }}",
    "AI_Generated__c": true,
    "AI_Model_Version__c": "claude-sonnet-4-5",
    "AI_Confidence__c": "{{ $json.confidence }}"
  }
}

Four important fields beyond the obvious: AI_Generated__c (so humans can filter), AI_Model_Version__c (so you can re-process old entries when you swap models), AI_Confidence__c (low-confidence entries get a CSM review), and a separate Risk_Reason__c field (so you can sort or report on it).

Conditional CTA creation

If risk_flag === 'critical', create a Call-to-Action. Gainsight's POST /v1/cta takes an account, an assignee (the CSM), a type (we use "Risk"), and a description. The description is the LLM's risk_reason with a "review the transcript before any action" line appended. The CTA appears in the CSM's daily worklist with the verbatim quote attached.

Health Score adjustment

Don't let the LLM directly write Health Score numbers. Let it set a categorical signal (positive / neutral / negative) and have a Gainsight Rules Engine rule consume that signal to adjust the Health Score. This separation matters because Gainsight's Health Score logic is a configured business rule, not a model output. Treating it otherwise causes scope creep and audit nightmares.

The Feedback Loop That Makes It Better

An LLM workflow that doesn't get better with feedback is a fire-and-forget mistake. The structure that works:

The CSM correction interface

Each AI-generated Timeline entry has a "Correct" button (a Gainsight custom action that posts to a webhook). When a CSM clicks it, a small modal opens: which fields were wrong, what should they have been, and a free-text "why." That payload lands in a Supabase table with the original transcript, the original output, and the correction.

The weekly review

Every Friday, an operator (typically the CS Ops lead) reviews the week's corrections. Patterns emerge: the model is over-flagging at_risk on calls with one-off frustration. The model is under-flagging expansion_signal when the customer says "we might explore the analytics module next quarter" (the operator considers that a clear signal; the model is too conservative).

Prompt versioning

Updates to the system prompt are versioned. The prompt lives in a Notion doc with a version number. Each version's promotion is gated by running it through a 30-call holdout set and comparing risk/sentiment/expansion outputs to the human-corrected ground truth. The team we work with sees their false-positive rate on risk_flag drop from 18% (v1.0, March 2026) to 6% (v1.4, May 2026) through this loop.

Three Failure Modes and the Fix for Each

Failure mode 1: The miscategorized recording

Gong recorded a 90-minute internal sales rehearsal session and labeled it a customer call because the customer's domain was in the meeting attendees. The transcript has 0% customer voice. The LLM (without an escape hatch) confidently produces a "summary" about how the customer is highly engaged on the analytics roadmap.

Fix: The confidence: low escape hatch in the prompt, plus a pre-LLM filter: if no participant matches the customer's known users (by email domain check), skip the LLM call entirely and route to the manual review queue. We've seen this catch ~3% of miscategorized recordings every month — small but it's the kind of bug that destroys CS team trust the day finance sees a "highly engaged" summary for a churned customer.

Failure mode 2: The truncated transcript

A 2-hour transcript that runs past the model's context window (current Sonnet 4.5: 200k tokens, fine; older models or some providers, less). The naive workflow truncates from the end and loses the last 20 minutes — which is where customers actually drop churn signals.

Fix: If the transcript exceeds ~25k tokens (a tunable threshold), summarize in two passes — split the transcript by speaker turns into thirds, summarize each third separately with the same rubric, then have a final pass synthesize the three sub-summaries. This costs 4 LLM calls instead of 1 but preserves end-of-call signal. The other fix: weight the final 25% of the call more heavily in the prompt ("Pay particular attention to the final 10 minutes of the transcript, where customers signal real intent").

Failure mode 3: The wrong customer mapping

The workflow looks up the Gainsight account by the email domain of the customer participant. The customer is "acme.com" but Gainsight has "acme-corp.com." The lookup fails. The workflow either crashes or — worse — posts the summary to the wrong account.

Fix: Three-step lookup. Try email domain. Try Gong account ID as a stored mapping. Try fuzzy name match with a 90%+ confidence threshold. If all three fail, post the summary to a "needs assignment" queue, not to a wrong account. The cost of an extra workflow path is much lower than the cost of a customer reading another customer's call summary.

Real Numbers from a March 2026 Deployment

A SaaS company with 4 CSMs and ~50 customer calls per week shipped this workflow in mid-March 2026. The before-state:

  • ~12 hours/week of CSM time on post-call notes (3 hours each, 4 CSMs)
  • ~40% of calls had Timeline entries within 24 hours; ~25% got entered within a week; ~35% never got entered
  • Renewal forecasts were based on ad-hoc CSM judgment with no structured churn-signal tracking

After 8 weeks of the workflow:

  • ~1 hour/week of CSM time on post-call review (down from 12)
  • ~98% of calls had Timeline entries within 30 minutes
  • The Risk Flag distribution surfaced 7 at-risk accounts the team hadn't proactively identified; 4 of them got recovery plans in time; 3 churned but the CS lead had advance notice
  • One expansion signal flagged by the workflow (a customer mentioning "we're rolling out to the EU team next quarter") turned into a $42k expansion deal that the CSM had genuinely missed in their post-call notes (verified by reviewing the recording)

The total cost of the LLM portion: about $38/month at 50 calls/week, 4 weeks/month, using Sonnet 4.5 with cache_control on the system prompt. The CSM time saved at $90/hour fully loaded: ~$1,000/week, or about $4,000/month. Net: ~$3,960/month in CSM time freed up.

The point isn't that the LLM is smarter than the CSM. The point is that the LLM never gets tired, never runs out of time, and applies the same rubric to call 50 of the week that it did to call 1. Consistency at scale, not intelligence, is the value.

What to Build This Week

If you're starting from zero on this workflow, the build order:

  1. Day 1 (90 min): Wire the Gong webhook to n8n. Confirm the trigger fires. Pull and print a transcript. Don't even call the LLM yet — verify you can read the data.
  2. Day 1 (90 min): Write v0.1 of the system prompt with the rubric. Run it against 10 historical transcripts in the Claude Console or n8n manually. Read the outputs. Adjust.
  3. Day 2 (2 hours): Add the validation code. Write the Gainsight Timeline POST. Test end-to-end on a single test account.
  4. Day 3 (1 hour): Add the CTA creation on risk_flag = critical. Add the Slack notification. Run for 5 real calls in shadow mode (write to a "shadow" Timeline activity type that doesn't show in the normal feed).
  5. Day 4 (1 hour): Compare shadow outputs against what the CSM actually wrote. Adjust the prompt. Promote.
  6. Weeks 1-4: Run with feedback corrections. Each Friday, review corrections, version the prompt, re-test on the holdout set, ship the new version.

Eight to ten hours of build time. Two-to-three weeks of feedback-loop tuning before the workflow earns full trust. That's the realistic timeline for a CS Ops operator shipping their first conversation-summary-to-Gainsight workflow.

Key Takeaways

  • Call-summary-to-CS-platform workflows save 10-12 hours of CSM time weekly per 50-call cadence and dramatically improve Timeline entry coverage (from ~40% to ~98% within 24 hours in measured deployments).
  • The stack of choice in May 2026: Gong/Chorus/Avoma for transcripts, n8n or Make for orchestration, Gainsight/Catalyst/Vitally for destination. Pick one per category and don't proliferate.
  • The coaching rubric is the prompt — not "summarize the call" but a 7-12 dimension structured output with explicit edge-case anchors and a confidence: low escape hatch.
  • The four hallucination modes to defend against: malformed JSON, missing fields, invalid enums, and the fabricated verbatim quote. Validate all four in a Code node after the LLM call.
  • The verbatim quote check (does quote_of_concern appear in the transcript?) catches the most damaging single hallucination mode in this workflow.
  • Use a three-step customer lookup (domain, Gong account ID, fuzzy name) and route lookup failures to a "needs assignment" queue — never to a wrong account.
  • For transcripts >25k tokens, summarize in thirds and synthesize, or weight the final 25% of the call explicitly in the prompt; customers signal churn in the last 10 minutes.
  • Tag every AI-written Timeline entry with AI_Generated__c, AI_Model_Version__c, and AI_Confidence__c custom fields so CSMs can filter and re-process when models or prompts change.
  • Don't let the LLM write Health Score numbers directly. It writes categorical signals; the Gainsight Rules Engine converts signals to score deltas. Separation of concerns survives audits.
  • The feedback loop matters more than the initial prompt. False-positive rate on risk_flag dropped from 18% to 6% in eight weeks of weekly correction review and prompt versioning at one team we work with.
  • Build order: trigger and transcript fetch first, rubric and dry-run second, validation and write third, conditional CTAs and notifications fourth, shadow mode for a week, then promote. Eight-to-ten hours of build, two-to-three weeks of tuning.