Hands-On Practice: Review AI Summaries, Identify Errors, Revise
Renata runs technical recruiting for a 180-person health-tech company, and she had grown comfortable letting AI summarize her interview notes. It saved her real time: a panel of four interviewers would leave her with maybe 900 words of raw observations per candidate, and the model would compress that into a tidy paragraph she could drop into the hiring channel. Then one Tuesday a hiring manager approved a candidate partly on the strength of a summary line that read "led a team of five engineers." Renata went back to the original notes to find the detail, and it was not there. The interviewer had written "worked on a five-person team." Nobody had lied. The model had quietly promoted a team member into a team lead, and a hiring decision had absorbed the inflation without anyone noticing.
Why Review Is the Real Skill
Reading an AI summary is easy. Reviewing it, finding what it got wrong, understanding where it went wrong, and fixing it before it influences a decision is the actual skill, and it is what separates AI that supports your judgment from AI that quietly misleads you. This session is hands-on practice in exactly that. You will work through realistic summaries against the notes they came from, identify the problems, and revise them. By the end you should have a checklist you can run on every summary before it touches a hiring conversation, and enough pattern recognition to know which errors your own tooling tends to produce.
The reason review is necessary at all comes down to what a summary is. An AI summary is a compression, and every compression loses information and occasionally invents it. The model is optimizing for a clean, confident, readable paragraph, not for fidelity to your notes. When the source material is thin, it pads. When the source is ambiguous, it resolves the ambiguity in whatever direction reads most smoothly, which is usually the more flattering direction. None of this is malicious and none of it is rare. It is the normal behavior of a system built to produce plausible text, and it means a summary is a draft to be verified, never a finding to be trusted on sight.
Renata reframed her job accordingly. She is not a reader of summaries; she is an editor of them. Her standard is simple and strict: nothing goes into a hiring conversation until she has checked it against the notes it claims to summarize. That one habit, applied to every candidate rather than to the ones that happen to look suspicious, is what makes the time savings safe to keep. The habit only works if it is unconditional, because the summaries that need review the most are precisely the ones that read as if they do not.
The Seven-Point Review Checklist
Before a summary informs any decision, Renata runs it against seven checks. They are fast once they are habitual, and together they catch nearly everything that goes wrong. Work them in order the first several times; after that they collapse into a single pass.
Fact accuracy. Every explicit fact in the summary must trace to the notes: years of experience, company names, project names, specific technologies. If the summary states a number, that number has to appear in the source. This is the check people assume they are already doing and usually are not, because a specific number reads as evidence of care rather than as a claim requiring verification. Treat every concrete detail as something to look up, not something to absorb.
Distinction between explicit and inferred. The summary should mark what the candidate actually said apart from what the model concluded from it. "Stated experience with Python" is explicit and traceable. "Likely comfortable across the stack" is an inference, and it should read like one on the page. The danger is not that models infer; sensible inference is useful. The danger is inference that arrives wearing the clothing of a stated fact, because a hiring manager reading it downstream has no way to tell the two apart and will weigh them the same.
Context preservation. A flagged concern without its context is not information; it is a smear. "Four roles in six years" means one thing in isolation and something entirely different once you know what drove each move. Ask of every red flag in the summary whether the explanation travelled with it. If the summary flags without explaining, the reviewer downstream will supply their own explanation, and the one they invent will usually be less generous than the truth in the notes.
Evidence behind claims. Major assessments need supporting examples. "Strong problem solver" with nothing attached is decoration, and a hiring manager cannot do anything with it except agree or disagree on instinct. "Strong problem solver: walked through how he isolated a race condition under load" is a finding, because it names the observation the judgment rests on and lets a reader disagree with the interpretation while still keeping the fact.
Balance. The summary should reflect the weight of the evidence across the whole interview rather than overweighting one memorable moment or one interviewer's tone. Panels produce uneven material: one interviewer writes three paragraphs, another writes two lines. A model summarizing that will naturally give the verbose interviewer more influence over the verdict, which is a weighting decision nobody made deliberately.
Tone accuracy. Uncertainty in the notes must survive into the summary. "I am not sure about their depth with container orchestration" is a statement about the interviewer's confidence, not about the candidate's skill, and it must not become a flat "lacks container orchestration experience." Models tend to launder hedged observations into confident verdicts because confident prose reads better, and that laundering is one of the most consequential distortions in the whole category.
Missing information. The most expensive errors are omissions, because they leave no trace on the page. Ask what the notes contained that the summary dropped, and what you wish were there that is not: a soft-skill observation, a stated concern, a compensation flag, technical depth on a specific area, an observation about how the candidate worked with the rest of the panel. This is the only check that requires reading the source in full rather than scanning it for confirmation.
Worked Review One: Hallucination and Inflation
Here is a candidate Renata processed this week. The role is Senior Backend Engineer, and the panel left her roughly 850 words of notes. The model produced this summary:
"Daniel has eight years of backend engineering experience with deep expertise in distributed systems. He architected the company's real-time claims processing pipeline and led a team of five engineers. A strong, clear communicator who explains technical decisions well. No concerns."
Clean, confident, and ready to paste. That is exactly the kind of summary that gets approved without a second look, so Renata slows down and checks each claim against the source notes, which read:
"Backend engineer for several years, mostly on monolithic services earlier in career. At current employer, worked on the claims processing system, one of a five-person team. Explained his debugging approach clearly when asked. One interviewer noted he seemed hesitant on the system-design question about scaling writes, said he would want to read up before committing to an approach. Salary expectation came in above the band we discussed."
Now the gaps are visible. "Eight years" is a fabrication; the notes say "several years." "Deep expertise in distributed systems" appears nowhere in the source. "Architected the pipeline" overstates "worked on the claims processing system, one of a five-person team," and "led a team of five engineers" is the same promotion error that burned her before: membership on a five-person team became leadership of one. "No concerns" erased two real ones, a hesitation on a scaling question and a salary expectation above the band. The only line that survives intact is the note about clear communication, and even that should carry its evidence.
It is worth naming the failure types precisely, because naming them is what makes them findable next time. The summary contains a hallucinated detail (eight years, distributed-systems expertise), an inflated inference (architected, led the team), and two omitted concerns (the scaling hesitation and the compensation gap). The hallucinations would have made Daniel look stronger than the evidence supports. The omissions would have hidden the two things the hiring manager most needed to weigh. How you catch this is unglamorous and reliable: check every specific claim against the notes. If the summary says "eight years" and the notes say "several years," that is a hallucination, full stop, regardless of how plausible eight years might be for someone at that career stage.
The Revision: Before and After
Renata does not throw the summary out. Discarding it wastes the work and tempts her to write the summary from scratch under time pressure, which produces its own errors. She revises instead, keeping what is true and restoring what was lost. The revised version reads:
"Daniel has several years of backend engineering experience, earlier work mostly on monolithic services. At his current employer he contributed to the claims processing system as one of a five-person team; the notes do not indicate he led the team or owned the architecture. Communication: strong, explained his debugging approach clearly when prompted. Open concerns to weigh: one interviewer found him hesitant on a system-design question about scaling writes, and he said he would want to read up before committing to an approach. His salary expectation came in above the band we discussed. No demographic inferences drawn from the notes."
The revised summary is less flattering and far more useful. It tells the hiring manager exactly what the panel observed, separates what Daniel said from what the model guessed, and surfaces the two concerns that should shape the next conversation rather than burying them under a confident "no concerns." Notice also what the revision does with the absence of evidence: it says the notes do not indicate leadership, rather than asserting that Daniel was not a leader. That distinction matters, because the failure was in the summary, not in the candidate. The point of review is not to make candidates look worse. It is to make the record accurate enough to decide on.
Renata also closes the loop on her prompt. The recurring failure, turning team membership into team leadership and dropping stated concerns, is a pattern rather than an incident, so she adds two lines to her summarization prompt: "Do not assign leadership or ownership the notes do not explicitly state. List every concern or hesitation noted by any interviewer; never conclude 'no concerns' unless the notes contain none." The next three summaries she runs hold up. The error was fixable at the source once she had caught it at the output, which is the general shape of this work: catch by hand, then fix upstream so you stop catching the same thing by hand.
Worked Review Two: Lost Context Behind a Red Flag
The second summary in Renata's queue contained a single line under concerns: "Red flag: candidate changed jobs 4 times in 6 years." Nothing in that sentence is factually wrong, which is what makes it a useful example. Every fabrication check passes. The problem is entirely one of what was left out.
The original notes read: "Changed jobs 4 times in 6 years. First move: company acquired. Second move: wanted to learn new tech stack. Third move: spouse relocated. Current role: 3 years, no plans to leave." With the context restored, the picture inverts. Three of the four moves had external or deliberate causes, an acquisition, an intentional growth decision, and a partner's relocation, and the candidate has been stable in the current role for three years with no plans to leave. The raw count of four moves in six years is technically accurate and practically misleading, and a hiring manager scanning a shortlist will read "red flag" and stop there.
How to catch this one: review every red flag in the summary and ask whether it carries an explanation. If a concern is merely flagged, go back to the notes and ask what accounts for it. The revision is straightforward once you have looked. "Changed jobs 4 times in 6 years. Context: acquisition, intentional growth-seeking, spouse relocation. Currently in a stable role for 3 years with no plans to leave." The concern is still on the page, because it is legitimate to note a pattern of movement, and a reviewer who wants to probe it in a later conversation still can. What changed is that the reader now has what they need to weigh it rather than react to it.
This category deserves particular care in recruiting because context loss is not evenly distributed across candidates. Career interruptions, relocations, and non-linear paths are more common for some groups than others, so a summarization habit that strips context from movement will systematically disadvantage the candidates whose careers were shaped by circumstance rather than by choice. Restoring context is an accuracy practice and a fairness practice at the same time.
Worked Review Three: Rich Observations Flattened Into Ratings
The third summary read: "Communication: Adequate. Team collaboration: Adequate. Leadership: Not assessed." It is compact, it is honest in tone, and it destroys most of what the panel actually saw.
The notes behind it read: "In technical round, very clearly explained his debugging approach and asked clarifying questions from the system designer. In team round, facilitated discussion when two engineers disagreed, acknowledged both sides, helped reach consensus. Quiet overall, but thoughtful." The summary reduced "clearly explained" and "facilitated discussion between two disagreeing engineers" to the word "adequate," twice. That is not a compression; it is an erasure. Soft skills are the material most often lost this way, because rich behavioral observation does not survive being converted into a rating scale, and a model asked to produce a tidy structured summary will reach for the scale.
Catching it requires a direct comparison: put the summary's soft-skills section beside the corresponding notes and ask whether detailed observations have collapsed into vague ratings. The revision restores the evidence under each heading. "Communication: Strong. Clearly explained debugging approach. Asked clarifying questions of the system designer. Team collaboration: Strong. Facilitated discussion between two disagreeing engineers, acknowledged both sides, helped the group reach consensus. Leadership: Not directly assessed." The ratings moved from "adequate" to "strong" not because Renata decided to be generous but because the evidence in the notes supports the stronger reading and the original summary was not reflecting it.
One caveat travels with this scenario. "Quiet overall, but thoughtful" is an observation about style, and it is exactly the kind of note that gets converted downstream into a communication verdict. Quiet is not the same as a weak communicator, and the notes here show the opposite: a candidate who explained clearly and facilitated a disagreement. Preserve the style observation if it is useful, but never let it override the behavioral evidence sitting next to it.
A Repeatable Review and Revision Process
Renata runs the same six-step pass on every summary, and it takes her about four minutes once it is routine. First, scan for obvious inflation: any specific fact or achievement that looks larger than you remember from the room, or that carries a precision the notes are unlikely to have supported. Second, check every major claim against the notes, line by line, with particular attention to years of experience, titles, ownership language, and claimed achievements. Third, assess context: does every flagged concern carry its explanation, and does every strength carry its evidence, or are you looking at bare flags and bare ratings?
Fourth, hunt for omissions by reading the notes for anything the summary failed to mention, since this is the step that catches the invisible errors and the only one that cannot be done by staring at the summary alone. Fifth, fix what you found, either by editing the summary directly or by handing the model specific feedback and asking for a revision; direct editing is faster for one-off problems, while a prompt revision is the right move once the same problem has appeared more than twice. Sixth, reread the result cold and ask two questions: does this feel accurate, and does it represent the candidate fairly? Four minutes per candidate, weighed against the cost of one hire made on inflated or incomplete information, is not a close call.
Anti-Patterns That Defeat the Review
Three habits quietly undo all of this. The first is accepting summaries without verification: the summary says "twelve years of experience" and nobody ever checks it against the resume or the notes, so a hallucinated number becomes a hiring rationale and then becomes an assumption everyone downstream shares. It happens because verification feels like distrust of a tool that has been right often enough to earn the benefit of the doubt. The defense is to spot-check every specific number and claimed achievement against the source, every time, especially the ones that make the candidate look strong.
The second is using a summary that is eighty percent there because finishing the review feels like more work than it is worth. The missing twenty percent is almost never the easy part; it is usually the soft-skill nuance or the stated concern that turns out to matter. A team that evaluates technical fit from a summary with no soft-skills content discovers the communication problem after the hire rather than before it. The defense is to revise or expand the weak section before relying on it, which usually means one targeted follow-up request rather than a full rewrite.
The third is trusting the summary's tone over your own recollection. "Quiet in the panel" becomes "poor communicator," and a candidate is mischaracterized on personality rather than on ability. This one is dangerous because the summary's verdict is stated confidently and your memory is not, so the document wins by default. The defense is to compare every tone assessment against both the actual notes and your recollection of the room, and to demote any confident verdict the evidence does not support back to the observation it was built from.
Practice
Work these in order on real summaries from your own pipeline where you can. Each exercise isolates one error type before the final one puts them together.
- Spot hallucinations. Take an AI summary, real or from the examples above. Identify three specific claims in it. Check each one against the original notes. Record any hallucination or inflation you find, and note which of the three you would have accepted without checking.
- Add missing context. Take a summary that flags a concern, whether that is job movement, a communication style, or a gap. Compare it to the original notes. Write down what context is missing, then revise the flag so it carries its explanation with it.
- Verify soft skills. Take the soft-skills section of a summary and put it beside the corresponding notes. Are the detailed observations preserved, or have they been reduced to ratings? Revise the section so every rating is followed by the specific behavior that supports it.
- Run a full review and revision. Take one complete AI summary and work the whole seven-point checklist: accuracy, explicit versus inferred, context, evidence, balance, tone, and missing information. Document every issue you find, then produce a revised summary you would be willing to attach your name to.
- Build your own checklist. Based on the issues you actually caught in the exercises above, write a personal review checklist. Which errors does your tooling produce most often on your kind of interviews? Order your checklist so the most frequent errors get checked first.
- Fix one error at the prompt. Take the single most repeated error from your log and write an explicit constraint that would prevent it, in the style of "never assign ownership the notes do not state." Add it to your summarization prompt and re-run three summaries to see whether it held.
Reflection
- In your experience with AI summaries so far, which error type do you meet most often: hallucinations, lost context, or flattened soft-skill assessments?
- How long does it actually take you to review one summary against its notes, and is that time investment worth the quality gain in your context?
- What is one error pattern you could prevent entirely by changing your summarization prompt rather than catching by hand?
- If you re-reviewed your recent hiring decisions with this checklist, what might you discover about the information those decisions were made on?
- Which of your hiring managers reads the summary rather than the notes, and what does that mean for how carefully the summary needs to be edited?
- Where in your process would a hallucinated detail be caught by someone else, and where would it travel all the way to an offer unchallenged?
Glossary
- Hallucination. Information generated by the model that was not present in the source material, such as a specific number of years the notes never stated. Always verify specific facts in a summary against the original notes.
- Inflated inference. A conclusion the model drew that the source supports only weakly or not at all, typically in the flattering direction, such as membership on a team becoming leadership of it.
- Context loss. When a summary flags a concern without preserving the circumstances that give the concern its meaning, leaving a technically accurate statement that misleads the reader.
- Omission. Information present in the notes that the summary silently dropped. The hardest error class to catch, because it leaves no visible trace in the output.
- Soft skills. Interpersonal qualities such as communication, collaboration, and leadership. Frequently underrepresented in AI summaries because rich behavioral observation does not survive conversion into a rating.
- Spot-checking. Verifying key facts in a summary against the original source. The essential defense against hallucination, and the step most often skipped when a summary reads well.
- Revision. Editing a summary, or supplying the model with specific feedback and asking it to produce a corrected version, based on issues identified during review.
Related Lessons
This practice session sits between the lessons that explain the errors and the lessons that build systems around them.
- AI Summarization: Capabilities and Typical Errors is the natural precursor, since it explains why compression produces exactly the three failure types you practiced catching here.
- Common AI Errors in Recruiting: Hallucinations, Misinterpretations, Omissions supplies the fuller taxonomy behind the checklist, including error types that appear outside summarization.
- Verification Techniques: Spot-Checking Facts, Sources, and Candidates extends step two of the review process into a general method for checking claims against sources.
- Formatting and Organizing Summaries for Decision-Making covers the structure question this lesson only touches, including how to lay out a summary so evidence stays attached to each judgment.
- Flagging Red Flags and Concerns Without Bias goes deeper on the context-loss scenario, and on why stripping context from a concern lands unevenly across candidates.
- Feedback Loops: How to Report AI Errors and Improve System Performance is where the errors you catch here become logged patterns and prompt fixes rather than one-off saves.
Closing
The difference between AI-supported recruiting and AI-misled recruiting is one step, and that step is verification. When you review every summary against the notes it came from, you catch hallucinations, restore lost context, and rescue the observations that a rating scale flattened, all before any of it influences a decision. Renata's near miss was not caused by a bad tool. It was caused by a good tool producing a confident sentence that nobody checked, which is the ordinary way this fails.
This week, use AI to summarize a set of interview notes and review each summary with the checklist. Document the issues you find rather than just fixing them, because the log is what tells you which errors are patterns and which were flukes. Over a few weeks you will learn how your particular tooling behaves on your particular kind of interviews, and that knowledge is what turns review from a chore into a fast, targeted pass. Verification is where AI-assisted recruiting becomes trustworthy recruiting, and once you have a review process you can rely on, you can use the tool with far more confidence than you could without one.
Key Takeaways
- A summary is a draft, not a finding. Compression loses real information and invents plausible detail, almost always in the flattering direction. Nothing reaches a hiring decision until it has been checked against the notes it claims to summarize.
- Run the seven-point checklist every time. Fact accuracy, explicit versus inferred, context preservation, evidence behind claims, balance, tone accuracy, and missing information catch nearly everything that goes wrong.
- Learn the three failure types by name. Hallucinated detail (a number or skill not in the notes), inflated inference (membership promoted to leadership), and omitted concern (a stated hesitation erased by a confident "no concerns").
- Always spot-check specific facts. Years of experience, titles, project ownership, and claimed achievements are where hallucination concentrates, because precise detail reads as evidence of care rather than as a claim to verify.
- Red flags must carry their context. "Changed jobs four times in six years" is accurate and misleading; the same line with the acquisition, the growth move, the relocation, and three stable years in the current role is usable information.
- Soft-skill assessments need evidence, not ratings. "Clearly explained his debugging approach, facilitated a disagreement between two engineers" beats "adequate," and a quiet style is not a communication weakness.
- Omissions are the dangerous errors. A fabricated fact leaves a trace you can catch; a dropped concern leaves the page looking clean. Read the notes for what the summary failed to mention, not only for what it got wrong.
- Do not accept a summary that is eighty percent there. The missing portion is usually the soft-skill nuance or the stated concern that matters most. Revise or expand the weak section before you rely on it.
- Revise, do not discard. Keep what is true, restore what was lost, and separate what the candidate said from what the model guessed. State that the notes do not support a claim rather than asserting the opposite.
- Fix recurring errors at the prompt. When the same failure appears across candidates it is a pattern. Add an explicit constraint to the summarization prompt and re-test, so you stop catching the same error by hand.
- Document what you catch. Over time the log reveals how your tooling summarizes your specific interview context, and that pattern knowledge makes each review faster and more targeted.
- Four minutes of review protects every decision downstream. A short, repeatable verification pass is cheap against the cost of a single hire made on inflated or incomplete information.
Frequently Asked Questions
Does reviewing every summary cancel out the time AI saved me? No, and the arithmetic is worth doing explicitly rather than assuming. Renata's panels produce around 900 words of raw notes per candidate; reading and synthesizing that by hand is a substantially longer job than reading a generated paragraph and checking its claims against the source, which takes her about four minutes once the pass is routine. What review costs you is the fantasy that the summary is finished when it appears. What it buys you is a document you can actually put in front of a hiring manager. If your review is taking far longer than a few minutes per candidate, that is usually a signal that your prompt is under-constrained rather than that review is too expensive.
What if I no longer have the original notes to check against? Then you do not have a summary you can verify, and you should treat it accordingly. Use it as a memory aid for your own recollection, not as evidence in a decision, and do not forward it as though it were a record of what the panel observed. The practical fix is upstream: make raw notes a retained artifact rather than a disposable input, so every summary has a source it can be checked against later. This also matters if a hiring decision is ever questioned, since a generated summary with no underlying notes is a weak thing to have to explain.
Should I edit the summary myself or ask the model to revise it? Both work, and the choice depends on whether you are fixing an incident or a pattern. Editing directly is faster for a one-off problem and gives you exact control over the wording, which matters when the fix involves careful phrasing such as noting that the notes do not indicate leadership. Asking for a revision with specific feedback is better when several sections need rework, and it has a useful side effect: the feedback you write is usually a draft of the prompt constraint you should add permanently. Once the same correction has come up more than twice, stop revising outputs and change the prompt.
The summary is harsher than I remember the candidate being. Is that also an error? Yes, and it is the tone-accuracy check doing its job. Errors run in both directions, though the flattering direction is more common. Go back to the notes and find the specific observation the harsh verdict was built from. If the notes say "quiet overall, but thoughtful" and the summary says "weak communicator," the summary has converted a style observation into an ability judgment, and that is a mischaracterization you should correct. Demote the verdict back to the observation and let the reader weigh it. The same applies when the summary is kinder than the room was.
How do I keep the review itself from introducing bias? Anchor every edit to the notes rather than to your impression. The revision standard is that each claim in the final summary can be traced to a specific line in the source, and that anything the source does not support is either removed or explicitly marked as not indicated. Do not add demographic inferences, and do not let a candidate's style substitute for evidence about their ability. Context loss deserves particular attention here, because career interruptions and relocations are not evenly distributed across candidates, so a habit of stripping context from movement will disadvantage some groups more than others. Restoring context is a fairness practice as much as an accuracy one.
What do I do when the summary and my memory of the interview disagree? Treat the notes as the tiebreaker, not the summary and not your memory. Memory is reconstructive and shifts toward whatever you have since read, which is precisely why an inflated summary is so effective at rewriting your recollection of a candidate. If the notes are silent on the disputed point, the honest resolution is to record that the point is not established rather than to pick a side, and to raise it as an open question for the next conversation with the candidate. An open question is a legitimate output of review. A confident verdict built on nothing is not.
Skill.re