←
AI Readiness & Process Transformation
Capable · M3 · lesson 3 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

AI-Assisted Interview Synthesis at Scale

15 min

The wall of the third-floor conference room is covered in sticky notes, four hundred of them, in five colors whose meaning nobody fully remembers anymore. It is Thursday. The interviews finished last Friday. The consultant who ran them has spent the week clustering, re-clustering, and quietly deciding which twenty stakeholders' words survive into the summary deck, and by now the summary says roughly what the consultant expected it to say before the first interview started. Meanwhile, in the steering committee that will read that deck, a vice president will lean forward at slide six and ask the only question that ever matters about people data: "Who said that?" And the whole assessment will live or die on whether the answer is defensible. This lesson is about doing that week of synthesis in an afternoon with AI, without losing the one property that keeps people data from becoming a weapon: the ability to trace every claim back to a human being who actually said it, and to protect that human being while you do.

The Dimension Every Assessment Skips

You have spent two chapters of this level assessing processes and data. Now comes the dimension that the arithmetic says matters most and that real assessments skip most often: people. Boston Consulting Group's 10-20-70 rule, the recurring arithmetic of this program, holds that successful AI transformation is 10 percent algorithms, 20 percent technology and data, and 70 percent people and process. MIT's autopsy of the 95 percent of stalled pilots found the same thing from the failure side: the killer pattern was adoption without transformation, tools used but nothing changed, and adoption is not a technical variable. It is twenty people deciding, individually and mostly silently, whether the new thing deserves their cooperation.

So why does the people dimension get skipped? Not because assessors think it is unimportant. It gets skipped because the evidence is inconvenient. Process evidence lives in ticket systems and SOPs (standard operating procedures, the documented way work is supposed to happen). Data evidence lives in tables you can profile. People evidence lives in conversations: twenty stakeholders' hopes, fears, workarounds, and grudges, none of it written down anywhere, all of it obtainable only by sitting across from someone for half an hour and listening. Historically, converting twenty such conversations into findings took the week of sticky notes you just watched, and the output arrived pre-filtered by whoever wrote the summary, because a human synthesizer under time pressure keeps what confirms the emerging story and drops what complicates it. The synthesis was slow, expensive, and quietly biased, so most assessments replaced it with a survey, or with the sponsor's opinion of what the team thinks, which is how pilots launch into workforces that were never actually asked.

AI collapses that week into an afternoon. A language model is genuinely excellent at the mechanical middle of qualitative synthesis: extracting distinct statements from transcripts, clustering them into themes, counting who said what. This is one of the highest-leverage uses of AI in the entire assessment toolkit, and this lesson teaches it end to end.

But before the method, one warning that shapes everything after it. People data has a property that process data does not: it gets weaponized. Nobody gets marched into a manager's office because of what a cycle-time histogram revealed. People do get marched into offices because of what they said in an interview. "Who said that?" is never a methodological question in a steering committee; it is a political one, asked by someone deciding whom to blame, lobby, or discount. Which means traceability, the ability to walk every claim back to its source, is not a nicety in interview synthesis. It is survival gear, both for your findings and for the people who trusted you with their honesty.

The deliverable this lesson builds is the Interview Synthesis Pack, and it has four components:

  • Themes with strength counts: each finding scored by how many distinct speakers voiced it.
  • Representative quotes with speaker codes: exact words, attributed to anonymized codes, never to names.
  • Tensions: the places where stakeholders contradict each other, documented instead of averaged away.
  • The traceability index: a locked appendix mapping every claim and quote back to its transcript and line number.

Everything below is the method for producing those four things honestly.

Interview Craft First: You Cannot Synthesize What You Did Not Collect

A tempting mistake, now that AI makes synthesis cheap, is to treat the interviews themselves as raw material of no particular craft. Resist it. The synthesis is a lens; it cannot add signal that the conversation never captured. Garbage transcripts in, confident garbage themes out. So the method starts a step earlier than the AI does, with the interview itself.

The 30-minute readiness interview guide

You do not need ninety minutes and a discussion guide the size of a phone book. You need thirty minutes and five questions, asked in order, with silence allowed after each one:

  1. "What does this process feel like on a bad day?" Not "describe the process": you already mapped it in Chapter 2.1. Bad days are where the truth lives, because bad days are what people remember in sensory detail and what the official process map never shows.
  2. "What would you fix first?" This surfaces priorities, and priorities differ by seat in ways the synthesis will need. A clerk's first fix and a director's first fix are usually different problems wearing the same process name.
  3. "What have you already tried?" The most respectful question in the set, and the most diagnostic. It surfaces the workaround economy, the unofficial fixes people built when the official channel failed them, and it tells you whether the organization has already burned trust on previous improvement attempts.
  4. "What would make you trust an AI-assisted version of this?" Note the phrasing: not "do you trust AI," which invites a culture-war answer, but "what would make you trust," which invites a specification. People answer this question with startling precision: "show me why it flagged the invoice," "let me override it without a ticket," "do not let it touch bank details."
  5. "What would make you route around it?" The mirror question, and the one nobody asks. Every stakeholder can describe the exact conditions under which they would quietly stop using the new system, and those conditions are your adoption risk register, dictated to you for free.

Consent, recording, and the attribution promise

Before the first question, two sentences of housekeeping that determine the quality of everything after. First, consent to record: you will be transcribing these conversations for AI-assisted analysis, and people have a right to know that. Second, and this is the one assessors skip, the attribution promise, stated up front and kept absolutely: "Everything you say is anonymized by default. In any output, you are a speaker code, not a name. If we ever want to use a quote in a way that could identify you, we clear it with you first."

Say it in those words, at the start, every time. Not because it is polite, though it is, but because of a hard mechanical fact: the synthesis is only as honest as people felt safe being. An interviewee who suspects their words will surface in a deck with their fingerprints on them gives you the corporate answer, and twenty corporate answers synthesize into a beautiful theme map of an organization that does not exist. The attribution promise is not a courtesy layered on top of the method. It is the method's data-quality control, applied at the moment of collection.

Coding With AI: One Transcript at a Time

Now the AI earns its keep. Qualitative researchers call this step "coding": breaking raw conversation into discrete, labeled items. Done by hand it is the most tedious work in the assessment; done with AI it is fast, and done with AI carelessly it is fast and wrong. The discipline has two stages, and the order matters.

Stage one: the extraction prompt, one transcript at a time

Feed transcripts to the model one at a time, never as a merged blob. Per transcript, the prompt is:

"From this transcript, list every distinct concern, hope, workaround, and factual claim the speaker expresses. For each item: tag it with the speaker code, quote the speaker's exact words, and note the line number where the words appear. Do not paraphrase inside quotation marks. Do not merge similar items. If something is implied but not stated, label it INFERRED."

Every clause of that prompt is a scar from a real failure. "Exact words" and "do not paraphrase inside quotation marks" exist because the model's default behavior is to tidy human speech into something smoother than what was said, and a tidied quote is a small fabrication wearing quotation marks. "Line number" exists because it makes the traceability index automatic instead of a reconstruction project. "One transcript at a time" exists because a model given ten transcripts at once starts blending speakers, and blended speakers are the disease this whole lesson vaccinates against.

The output is a coded item list per interview: typically 25 to 45 items for a thirty-minute conversation, each row carrying speaker code, type (concern, hope, workaround, claim), exact quote, and line number.

Stage two: the cross-transcript merge

Only when every transcript is individually coded do you merge. Paste all the coded item lists together (items, not raw transcripts) and prompt:

"Cluster these coded items into themes. For each theme: name it, list the item IDs it contains, and report its strength as a count of distinct speakers, not a count of mentions."

That last instruction is the single most important quality rule in the merge, so sit with it. One person saying something five times is one person. A model counting mentions will report the most talkative interviewee's pet grievance as your dominant theme; a model counting distinct speakers reports how widely a view is actually held. "Strength: 11 of 14 speakers" is evidence. "Strength: 23 mentions" is a measure of who talks a lot. When a theme's breadth gets challenged in a meeting, only the first number survives the challenge.

The merge output is your theme table: nine or so themes, each with a name, a strength count, two or three representative quotes with speaker codes, and the list of coded items behind it. Two-thirds of the Interview Synthesis Pack, produced in an afternoon.

Tension Hunting and the Traceability Index

The highest-value move: hunt tensions explicitly

Here is the move that separates a competent synthesis from a valuable one. After the merge, ask one more question: "Where do these interviewees contradict each other? List every point on which two or more speakers make incompatible claims, with the opposing quotes and speaker codes side by side."

You have to ask explicitly, because no language model volunteers contradiction. The model's deepest instinct is coherence: given twenty conflicting voices, it will average them into a plausible middle unless instructed otherwise. You met this failure mode in Chapter 2.1 as the Output Skeptic's pattern four, smoothing, the averaged contradiction, and interview synthesis is smoothing's natural habitat, because human organizations are made of contradictions. Managers say the process works; the people doing it describe a thriving workaround economy. Operations wants the AI to be fast; compliance wants it to be controllable. The clerk wants fewer exceptions; the analyst's expertise, and quiet pride, is exception handling.

Tensions are not noise in your findings. Tensions ARE the finding, because tensions are where pilots die later. Every one of those contradictions is a fault line that a deployment will eventually stand on: the pilot that satisfies operations' speed requirement violates compliance's control requirement, and nobody discovers the collision until month four, in production, expensively. A synthesis that surfaces four tensions in week one has just bought the design team four early warnings at interview prices instead of incident prices. And run the logic in reverse for your quality control: a synthesis of fifteen or more stakeholders that shows zero tensions is not describing a harmonious organization; it is exhibiting a smoothing failure. Real organizations disagree. If your themes do not, the model averaged something away, and you go hunting for it.

The traceability index: survival gear

The fourth component of the pack is the one that saves you in the meeting. The traceability index is a simple table, one row per claim and per quote used anywhere in your outputs: claim or quote, speaker code, transcript file, line number. The extraction prompt's line-number discipline means this table mostly builds itself.

Then the access rule, which is where the survival gear becomes survival: the index lives in the locked appendix; the deck shows only speaker codes. The mapping from codes to names exists in exactly one place, access-controlled, held by you. Because the "who said that?" moment is coming. A theme will sting someone powerful, and the political question will land on you in front of the room. With the pack behind you, the answer is:

"Three people, across two different teams, and the attribution stays anonymized per the interview promise."

Study that sentence, because it is engineered. "Three people" answers the evidentiary challenge: this is not one malcontent. "Across two teams" kills the next move, which is discounting the finding as one department's grudge. "Per the interview promise" reframes the refusal to name names from evasion into integrity: you are not hiding your evidence, you are keeping a commitment the room heard you were going to make. That sentence survives politics. "Um, I'd have to check my notes" does not, and neither do the twenty people who trusted you.

In people data, a claim you cannot trace is a claim you cannot defend, and a source you cannot protect is a source you will never get honesty from again.

Verification: Quotes Are Load-Bearing Evidence

By now the verification habit from Chapter 2.1 should be reflex, but interview synthesis has its own signature fabrication mode, so the checks are specific.

The misattribution spot-check

The dominant failure here is not the invented quote from nowhere; it is subtler. It is the paraphrase hardened into quotation marks: the model captures the gist of what AP-03 said, restates it more crisply than AP-03 ever spoke, and wraps it in quotes with AP-03's code on it. Or it attributes speaker A's sentiment to speaker B, because the two discussed the same topic and the model's attention slipped between transcripts. Both are fabrications, and both are undetectable by reading the synthesis alone, because the fake quote is always more quotable than the real one.

The check: for every transcript, sample three coded items and walk them back to the raw text. Open the transcript, go to the cited line number, and compare word for word. The quote matches exactly, or the item fails. For 14 transcripts that is 42 comparisons, perhaps 40 minutes of work, and it gives you a measured error rate for the extraction instead of a hope. And one rule with no sampling discount: any quote that will appear in the deck gets verified, 100 percent, no exceptions. Quotes are load-bearing evidence in people assessment, the qualitative equivalent of financial figures, and they get the same treatment this program gives every AI-touched figure: verified before anyone else sees them.

The omission check: the single-voice signal

Strength counts create their own blind spot: a synthesis organized around "how many people said it" will structurally bury the thing only one person said. So run the closing prompt: "List every concern that appears in only one interview."

This is the single-voice signal, and it has saved more assessments than any other five-minute check in this chapter, because organizational knowledge is unevenly distributed on purpose. The one person who mentioned the compliance restriction was the only person in your sample whose job required knowing about it. A strength count of 1 of 14 does not mean unimportant; it sometimes means the only qualified witness spoke. Review every single-voice item and ask: is this idiosyncratic, or is this the one specialist in the room? The second kind goes into the findings with its strength honestly labeled and its significance explained.

The failure story: composite voices

Now the cautionary tale, illustrative but drawn from a pattern you will be offered by every eager model and every deadline. An assessor, call him Marcus, finishes a 16-interview synthesis and finds the quote section untidy: real humans speak in fragments, and the fragments look weak on slides. So he accepts the model's helpful offer to merge related quotes across speakers into composite "representative voices": one polished paragraph per theme, distilling what "the frontline" thinks, smooth as a press release. The deck looks magnificent.

In the readout, a director goes quiet at slide nine. Then: "That quote. The first half is my sentence, from my interview. I never said the second half." He is right: the second half belongs to a clerk two levels down, welded on because the model judged the sentiments related. And now watch the collapse propagate, because it does not stop at one slide. If one quote is a composite, every quote in the deck is suspect. The director says so, out loud. The room recalibrates every finding as possibly manufactured. Sixteen honest interviews, weeks of goodwill, and the assessor's credibility, all spent in ninety seconds, and the next assessor to request interviews at this company will be told exactly why people decline.

Composite quotes are fabrication wearing a media-training smile. The tidiness that makes them attractive is the tell: real evidence is ragged. If a theme needs a cleaner voice than any single human provided, write your own summary sentence and label it as yours, above real quotes attributed to real speaker codes. Never let the machine put words in a mouth. That is not a style rule. It is the difference between evidence and fiction.

The Worked Example: Fourteen Interviews, One Afternoon

Here is the full method run once, on the invoice-exception process this level has been assessing since Chapter 2.1. All numbers are hypothetical and illustrative; the shape is what you are copying.

The assessor books 14 interviews across the seats that touch the process: six AP (accounts payable) clerks, two team leads, two Treasury analysts, two IT staff, one compliance officer, one vendor manager. Thirty minutes each, five questions, recording consented, attribution promise stated up front all 14 times. Coding runs one transcript at a time through the extraction prompt at roughly 25 minutes per transcript AI-assisted, including the human skim of each item list: about six hours of coding total, spread across the week as transcripts arrive. The merge, tension hunt, and checks fill one afternoon.

The merge produces nine themes. The top three tell the story:

ThemeStrengthRepresentative quote (speaker code)
Fear of blame when the AI mislabels an exception11 of 14 speakers"If it codes it wrong and I pass it, that's my name on the error." (AP-04)
Frustration with the queue wait9 of 14"An exception sits two days before a human even sees it." (TL-01)
Quiet pride in exception-handling craft6 of 14"Anyone can process a clean invoice. The weird ones are where I earn my job." (AP-02)

Look at that third theme, because it is the change-management goldmine nobody expected. No survey would have found it; no sponsor predicted it. Six of fourteen people are telling you, unprompted, that exception handling is not a chore the AI relieves them of but a craft the AI threatens to take. A rollout that pitches "the AI handles the weird ones" walks straight into that pride; a rollout that pitches "the AI clears the trivial ones so you get the genuinely weird ones faster" recruits it. Same tool, opposite adoption curve, and the difference was audible only in the interviews.

The tension hunt returns four tensions, including the classic pair: both team leads describe the escalation process as "working as designed" while five of six clerks describe the workaround economy that exists because it does not; and Treasury wants exception turnaround compressed while compliance wants every exception to gain a review step. Four fault lines, mapped before a single design decision is made.

The omission check flags two single-voice items. One is idiosyncratic. The other is the compliance officer, alone in the sample, stating that vendor bank-detail changes carry a regulatory handling restriction: the same restriction the Chapter 2.3 privacy audit flagged in the data. Read what just happened: the interviews and the data audit, two independent instruments, corroborated each other. That is not luck. That is the assessment working as a system, and it is why single-voice items get reviewed instead of buried: 1 of 14 was, once again, the only person whose job was to know.

The misattribution spot-check samples 42 items and catches two paraphrases hardened into quotes; both corrected against the raw text, and every deck-bound quote verified at 100 percent. Total synthesis effort: one afternoon plus the coding hours, against the classic week of sticky notes, with a traceability index the sticky-note wall never had. The pack goes into the assessment file with the theme table, four tensions, two single-voice flags, and a locked index of every claim.

What to Do Monday Morning

This lesson becomes a capability the first time you run one real interview through the full chain. Here is the sequence.

  1. Draft your interview guide from the five questions: bad day, first fix, already tried, what would build trust, what would make you route around it. Adapt the wording to your process, keep the order, keep it to thirty minutes.
  2. Book the first four interviews, choosing seats that see the process differently (a doer, a lead, a downstream consumer, a specialist), and open every one with the attribution promise stated in full: anonymized by default, speaker codes in all outputs, quotes cleared before any identifying use.
  3. Code one transcript with the extraction prompt: every distinct concern, hope, workaround, and claim; speaker code, exact words, line number; INFERRED labeled as such. One transcript at a time, always.
  4. Run the misattribution spot-check on that first coded list: sample three items, walk each back to the transcript line, compare word for word. Your first measured extraction error rate is worth more than any general opinion about AI accuracy.
  5. Start the traceability index file now, with the first transcript's rows, and set its access control before there is anything sensitive in it. The index that exists from day one costs nothing; the index reconstructed the night before the readout costs a weekend and your confidence.

Key Takeaways

  • Treat people readiness as the 70 percent it is: BCG's 10-20-70 and MIT's adoption-without-transformation finding both locate pilot failure in people and process, yet people readiness is the dimension assessments most often skip because its evidence lives in conversations.
  • Collect before you synthesize: run the 30-minute, five-question readiness interview (bad day, first fix, already tried, trust conditions, route-around conditions) with consent to record and the attribution promise stated up front, because the synthesis is only as honest as people felt safe being.
  • Code one transcript at a time with the extraction prompt: exact words, speaker code, line number, no paraphrase inside quotation marks, INFERRED labeled; merged-blob coding is where speaker blending begins.
  • Count speakers, not mentions: report every theme's strength as distinct voices (11 of 14), because one person saying it five times is one person, and only speaker counts survive challenge in the room.
  • Hunt tensions explicitly and treat a zero-tension synthesis of fifteen-plus stakeholders as a smoothing failure (the Output Skeptic's pattern four), because tensions are the fault lines where pilots die later.
  • Build the traceability index (claim, speaker code, transcript, line) into a locked appendix, show only codes in the deck, and answer "who said that?" with breadth plus the promise: "three people across two teams, attribution anonymized per the interview promise."
  • Verify quotes like figures: spot-check three coded items per transcript against raw text, verify 100 percent of deck-bound quotes, and never accept composite "representative voices," which are fabrication wearing a media-training smile.
  • Run the omission check for single-voice concerns, because the one person who mentioned the compliance restriction is often the only qualified witness, and carry the pack forward: themes feed the resistance heat map next lesson, tensions feed the people scorecard after that.