←
AI for Recruiters
Capable · M12 · lesson 12 of 27 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Formatting and Organizing Summaries for Decision-Making

15 min

Devon runs talent acquisition at a 250-person health-tech company, and last quarter a clinical data engineer search nearly fell apart in the final stretch. Five finalists, a four-person hiring panel, and a debrief meeting that ran 70 minutes because nobody could hold all five candidates in their head at once. The interviewers had each written thoughtful notes, and Devon had asked an AI assistant to summarize them. The summaries were accurate. They were also useless for the meeting: each one was a different shape, three paragraphs of prose for one candidate and a wall of bullets for another, with the panel flipping between documents trying to remember whether candidate three was the one who was strong on pipelines or the one who was strong on stakeholder communication. The content was fine. The format was killing the decision. Devon spent the next two weeks fixing the format, not the notes, and the next debrief for a different role took 22 minutes.

Format Is the Decision Interface

A great summary that is hard to scan is worth less than a mediocre summary that is easy to scan, because the summary is not the product. The decision is the product. The summary is the interface a hiring manager uses to reach that decision, and like any interface, its job is to surface the right information in the right order with the least possible friction. When Devon's panel was comparing five finalists, the question was never "is this a good write-up." The question was "who do we advance, and can we defend that choice." Everything about format should serve that second question.

This matters more as volume rises. Reading one candidate summary is forgiving; format barely registers. Comparing five, or screening fifteen, is where structure earns its keep. A panel member should be able to scan any single summary in about two minutes and walk away knowing three things: what this person is strong at, what the concerns are, and how they stack up against the others. If your format does not deliver those three things quickly, the panel will fall back on memory and gut, which is exactly where inconsistency and bias creep in.

Three Structures and When to Use Each

There is no single right format. There are three workhorses, and the skill is matching the structure to the decision in front of you.

The scorecard is the default for evaluating one candidate against the role. It rates a fixed set of dimensions, each with a short evidence note. Use it after every interview, for every candidate, because its discipline pays off later: when all candidates share the same dimensions, the scorecards become comparable almost for free.

The narrative is a short structured prose summary, three or four tight sections, used when nuance matters more than speed. It fits a single high-stakes hire, an executive search, or a candidate whose story does not reduce cleanly to ratings. Narrative is not an excuse for rambling; it is prose with headings, ordered by importance, kept under roughly 300 words.

The comparison table is for the moment of choice between finalists. One row per candidate, one column per dimension, scannable in a single glance. It does not replace the individual scorecards; it sits on top of them as the summary-of-summaries the panel looks at together. Devon's 22-minute debrief was built around a single comparison table projected on the screen, with the individual scorecards available for anyone who wanted to drill into the evidence.

A practical rule: scorecard per candidate during the pipeline, narrative only when the role genuinely demands it, comparison table at the finalist stage. Most teams overuse narrative because it feels thorough. It usually just feels long.

Order By Decision-Relevance

Within any summary, sequence is a design choice, not an afterthought. The principle is progressive disclosure: put the information a decision-maker needs first, first. A reader scanning quickly should get the headline in the opening 20 seconds and the supporting detail only if they choose to keep reading.

For a candidate scorecard, Devon settled on this order. One, a one-line profile: role, level, and immediate read. Two, the overall fit call: strong, possible, weak, or unable to assess, stated up front rather than buried at the bottom, with a single line of reasoning attached. Three, key strengths, three to five bullets, each with evidence. Four, areas of concern, two to four bullets, each with why it matters. Five, the dimension ratings. Six, the fit signals for this specific environment. Seven, suggested next steps. Putting the fit call near the top feels backwards to people used to building toward a conclusion, but a hiring manager who is scanning ten summaries wants the verdict before the reasoning, then dips into the reasoning for the ones that matter.

The same logic governs the comparison table: the leftmost columns should be the dimensions that most drive the decision for this role, and the rightmost column should be the overall call. The eye reads left to right and lands on the verdict last, after the evidence that supports it.

If the summary lives in a digital document rather than on paper, let visual encoding do some of the scanning for you. A consistent color convention, green for strengths, red for concerns, yellow for anything the interview left unclear, lets a panel member take in the shape of a candidate before reading a word of it. Keep the convention identical across every summary for the role, since the value comes entirely from its predictability, and never let color carry information that the text does not also state plainly.

A Worked Scorecard for a Clinical Data Engineer

Here is the template Devon built, filled in for one finalist. The role's decision dimensions were fixed in advance with the hiring manager: technical depth, healthcare data experience, communication, collaboration, and learning agility. Ratings use a four-point scale (Strong, Good, Moderate, Limited) and every rating carries one line of evidence.

Profile. Amara, senior data engineer, 6 years experience. Interviewed for the Clinical Data Engineer role.

Overall fit: Strong. Deep pipeline expertise plus direct healthcare data exposure; minor gap in real-time systems is coachable.

Strengths.

  • Designed a batch ingestion pipeline processing 40 million records nightly at her current employer (stated explicitly in the technical round).
  • Has worked directly with HL7 and FHIR healthcare data formats; named specific de-identification challenges she has solved.
  • Explained a complex schema migration in plain language to the non-technical interviewer without prompting.

Concerns.

  • All pipeline experience is batch; our roadmap includes streaming, which she has not built (explicit). Matters because year-one work is 30 percent real-time.
  • Has only worked on teams of 10-plus; we are a team of 4. Unclear how she handles broad ownership (inferred from how she described her past role).

Dimension ratings.

  • Technical depth: Strong. Walked through partitioning and idempotency tradeoffs unprompted.
  • Healthcare data experience: Strong. Direct HL7/FHIR work, named real de-identification problems.
  • Communication: Strong. Tailored explanations to each interviewer's background.
  • Collaboration: Good. Described positive cross-functional work; no conflict examples offered.
  • Learning agility: Good. Self-taught FHIR on a prior project; eager about streaming gaps.

Fit signals for our environment. This section answers a narrower question than the dimension ratings do: what specifically suggests this person would thrive here, on this team, rather than being good at the job in general? Each line gets a yes, a no, or an honest unclear, and unclear is a legitimate answer that tells the panel where to probe next.

  • Comfortable with ambiguity: Yes. Taught herself FHIR mid-project when nobody else on the team knew it.
  • Drives own priorities: Unclear. Every example came from a team of ten or more where priorities were set above her.
  • Supports others' growth: Unclear. Limited evidence from the interview either way.

Next steps. Reference check on independent ownership at small-team scale; short technical follow-up on streaming fundamentals.

That fills one page and scans in under two minutes. Now multiply by five finalists with the identical structure, and the comparison table writes itself:

  • Amara: Technical Strong, Healthcare Strong, Comms Strong, Collab Good, Learning Good, Overall Strong.
  • Bode: Technical Strong, Healthcare Moderate, Comms Good, Collab Good, Learning Strong, Overall Possible.
  • Chen: Technical Good, Healthcare Strong, Comms Strong, Collab Strong, Learning Good, Overall Strong.
  • Dilip: Technical Strong, Healthcare Limited, Comms Moderate, Collab Good, Learning Good, Overall Possible.
  • Esme: Technical Moderate, Healthcare Strong, Comms Good, Collab Strong, Learning Strong, Overall Possible.

The panel can now see in one glance that Amara and Chen are the strong calls, that the choice between them turns on technical depth versus collaboration, and that nobody is guessing. That is the difference between a 70-minute debrief and a 22-minute one.

Consistency, Comparability, and Fair Evaluation

The single most important formatting rule is also the one most often broken: evaluate every candidate for a role on the same dimensions, in the same template. The moment Amara is rated on "technical depth, healthcare experience, communication" and Bode is rated on "experience, culture, energy," the comparison is broken. You cannot honestly say one is stronger than the other, because you measured different things. Inconsistent dimensions do not just slow the decision; they corrupt it.

This is not only an efficiency point. It is a fairness and compliance point. Structured, consistent evaluation against job-related dimensions is a well-established way to reduce the influence of unstructured impression, and unstructured impression is exactly where bias operates. When every candidate is scored on the same job-relevant criteria with evidence attached, the evaluation is more defensible if a hiring decision is ever questioned, and disparities are easier to detect because you are comparing like with like. Recruiters working in jurisdictions with rules on automated employment decision tools, such as New York City's Local Law 144, should also remember that consistent, documented, evidence-backed criteria are the foundation any bias audit relies on. The format is not a cosmetic choice; it is part of the audit trail. To keep the dimensions genuinely consistent, lock them with the hiring manager before the first interview, not after, and resist the temptation to add a one-off dimension for a single candidate.

Length Discipline and Surfacing Evidence

Two habits separate a decision-ready summary from a data dump. The first is length discipline. A scorecard that runs two pages is a scorecard nobody reads to the end. Cap strengths and concerns at three to five bullets each, hold dimension notes to one sentence, and keep the whole thing to a page. When you direct an AI assistant to produce the summary, the constraint has to be explicit, because models tend to pad when source notes are thin. Tell Claude or ChatGPT the exact structure, the bullet limits, and the word ceiling, and the output arrives decision-ready instead of needing a trim.

The second habit is surfacing evidence. "Strong communicator" is an opinion. "Strong communicator: tailored her schema explanation to the non-technical interviewer without prompting" is an observation someone else can evaluate. Every rating should carry the evidence behind it, and the summary should distinguish what was explicitly stated from what is being inferred. "Said she has shipped FHIR pipelines" is explicit; "seems comfortable with ambiguity" is inferred from how she described uncertain situations. Marking that line protects the panel from treating a hunch as a fact, and it protects the candidate from being downgraded on something nobody actually observed. When you build the prompt, instruct the model to cite a specific example for every claim and to flag inferences as inferences. A summary built that way is not just faster to act on; it is more honest about what it knows.

Building and Testing Your Template

Turning all of this into a reusable template is a short, concrete process. Start by fixing the decision dimensions with the hiring manager for the specific role; five is usually right, three is thin, seven is too many to scan. Write the section headings in decision-relevance order. Write the summary prompt that tells the AI exactly that structure, those bullet limits, the evidence requirement, and the instruction to use identical dimensions across candidates so the outputs stay comparable. Then test it on three real interviews before you trust it: does it give the panel what they need, does it scan in two minutes, does it support a side-by-side comparison. Refine what is missing or bloated, and only then make it the team standard. Devon keeps three templates, one each for engineering, product, and clinical roles, because the right dimensions differ by job family even though the format discipline is identical across all three.

Three Anti-Patterns to Avoid

Narrative summaries as decision documents. The long, flowing write-up that opens "Amara interviewed very well, she has a lot of pipeline experience, and in the architecture round she talked about how she thinks about data quality" reads pleasantly and costs the panel five minutes per candidate. Decision-relevant details are scattered through the prose, no two summaries put them in the same place, and comparison becomes an act of memory. Reserve narrative for the rare role that genuinely needs it, and give everything else bullets, tables, and consistent sections.

Inconsistent evaluation dimensions. When one candidate is assessed on strategic thinking, execution, and communication, and the next on experience, communication, and team skills, you have produced two documents that cannot be compared. You will still feel like you are comparing them, which is the dangerous part; you are actually comparing two different measurements and calling the difference a preference. Fix the dimensions per role and apply them to everyone.

Vague assessments. "Great communicator," with nothing behind it, fails twice. Nobody else on the panel can evaluate whether your bar matches theirs, and you will not remember in three weeks what made you write it. The repair is always the same: attach the observation. "Great communicator: explained the system architecture clearly, asked clarifying questions before designing, and engaged all three interviewers" gives the panel something to agree or disagree with.

Practice Exercises

  • Build your template. For a role you are hiring now, write down the key decision dimensions, the section headings in decision-relevance order, the information that matters most to the panel, and the parts that must be scannable in seconds.
  • Summarize with it. Take a real set of interview notes and produce a decision-ready summary using your template. Ask whether it is scannable and whether it gives your team what they need to move.
  • Create a comparison table. Summarize three candidates for the same role, then build the side-by-side table. Can you see clear differences, or do the candidates blur together because the dimensions are too vague?
  • Test and iterate. Share a summary with your hiring team and ask them directly: can you scan this in two minutes, do you have everything you need to decide, and what is missing? Their answers, not your instincts, tell you what to change.
  • Build a template library. Create two or three templates for different role families, such as engineering, product, and sales, and test each with the teams that will use it. Keep the one that gets the fastest, most confident decisions.

The Vocabulary of Decision-Ready Summaries

  • Structured format. Information organized with bullets, tables, and consistent headings so that scanning and comparison are easy.
  • Progressive disclosure. Presenting information in order of importance, with the most critical material first and the supporting detail after.
  • Decision dimensions. The specific factors you evaluate every candidate on for a role, held constant across the whole slate.
  • Comparative table. A side-by-side evaluation of candidates on identical dimensions, built for the moment of choice.
  • Evidence-based assessment. An evaluation supported by specific examples or quotes from the interview, which is more reliable than an interpretive one.

What Good Formatting Buys You

Summary formatting is the difference between a hiring team that can decide and a hiring team that is drowning in information. Get it right and the team moves faster and decides better, because the effort goes into weighing candidates rather than into reconstructing what each write-up was trying to say.

So take whatever format your team uses today and rebuild it once, this week, into the decision-ready structure: profile, fit call, strengths, concerns, dimensions, fit signals, next steps. Run it on a live slate and watch how much faster your panel gets through the evaluation. That single reformatting pass usually buys back more time than any other change you can make to how you use AI on interview notes.

Reflection

  • How do you summarize candidates today? Is the result scannable, and could your team decide from it without opening the full interview notes?
  • What information do you most often wish were in a summary and find is not? What is that telling you about your current process?
  • If you could see every candidate for a role on a single page, what would you most need to know about each one?
  • How would your hiring improve if every summary were structured the same way and easy to compare?

Key Takeaways

  • Format is the decision interface, not decoration. The summary's job is to get a hiring manager to a fast, defensible decision; structure that serves scanning and comparison beats polished prose that does not.
  • Match the structure to the decision. Scorecard per candidate during the pipeline, narrative only when nuance truly demands it, comparison table at the finalist stage as the summary-of-summaries the panel decides from.
  • Order by decision-relevance. Lead with the profile and the overall fit call, then strengths, concerns, dimensions, fit signals, and next steps, so a quick scan surfaces the verdict in the first 20 seconds.
  • Consistency enables fair comparison. Evaluate every candidate on the same job-related dimensions in the same template; it is what makes comparison honest, bias harder, and decisions defensible under rules like NYC Local Law 144.
  • Keep it to a page and cite evidence. Cap bullets, hold dimension notes to a sentence, attach an observation to every rating, and mark inferences as inferences rather than facts.
  • Make the AI do the structure. Specify the exact format, bullet limits, word ceiling, evidence requirement, and shared dimensions in the prompt, so tools like Claude or ChatGPT return decision-ready summaries instead of prose you have to reshape.
  • Test before you standardize. Run the template on three real interviews, confirm it scans in two minutes and supports comparison, then keep one template per role family.