Hands-On Project: Conduct a Structured Interview with AI-Assisted Prep
Theo recruits for a 150-person logistics firm in Columbus, and he used to run interviews the way most teams do: four interviewers, four conversations, four wildly different sets of questions. When the panel met to decide, the loudest voice usually won. The VP would say "great culture fit" and everyone nodded. After a senior hire washed out in four months, Theo looked back at the file and realized the panel had never agreed on what they were actually evaluating. They had run four interviews and made one decision based on whoever spoke first in the debrief. This project rebuilds that broken process into a structured interview: a consistent question set, a shared scorecard, independent scoring before discussion, and AI used to do the drafting work without ever making the decision.
What You Are Building
This is a hands-on project, so it ends in artifacts rather than in agreement. By the time you finish you will hold four things for one real open role: a question set mapped to named competencies, an anchored scorecard that says in writing what each point on the scale means, independent scores submitted by every interviewer before anyone discusses the candidate, and a debrief record showing how the panel got from those scores to a decision. Produce those four for a single requisition and you have run a structured interview, with most of the next one already built.
The order is doing work that is easy to lose. Questions come from the job profile, not from the interviewers' preferences. The scorecard is written before anyone meets the candidate, not reconstructed afterwards to justify a feeling. Scores are committed before the debrief, not produced in it. Reorder those steps and the structure evaporates while the paperwork remains, which is why teams conclude structured interviewing does not work when what actually happened is that it did not get run.
Why Structure Beats Gut Feel
Decades of hiring research point in one direction: structured interviews predict job performance far better than unstructured ones. The reason is mechanical, not mysterious. When every candidate is asked the same job-relevant questions, evaluated against the same rubric, and scored before the panel talks, you remove the two biggest sources of noise: questions that drift from candidate to candidate, and a debrief where early opinions anchor everyone else. Structure is also what makes a process defensible. Consistent, job-related questions and documented scoring are exactly what EEOC guidance expects when you have to show that a hiring decision was based on the role, not on a protected characteristic.
The two sources of noise are worth separating, because teams usually fix one and assume they have fixed both. Question drift is the visible problem: when four interviewers improvise, the candidate is effectively assessed on four different jobs, and the panel then compares scores that were never measuring the same thing. Anchoring in the debrief is the invisible problem and the more expensive one, because even a panel that asked identical questions can converge on the first confident opinion voiced in the room, at which point the careful question design bought nothing. Theo's old process failed on both counts at once.
Theo's project has four parts: build a question set and scorecard from the job profile, prepare interviewers, score independently, then calibrate in a debrief. AI accelerates the first part. Humans own every part after it. He builds it on a live requisition, a Senior Operations Analyst with a real hiring manager and a real deadline, because a structured interview designed on a hypothetical role never gets stress-tested by the awkward parts of a real one.
Step One: Draft the Question Set from the Job Profile
Theo started with the job profile for the open role, a Senior Operations Analyst, and fed it to an AI assistant with a tight prompt: extract the four core competencies the role actually requires, then draft two behavioral questions and one situational question per competency, each with a short explanation of what a strong answer demonstrates. In ninety seconds he had a draft of twelve questions where building them from scratch would have taken an afternoon.
The prompt is the part of this step to copy, because of how much structure it encodes. It names the source document, so the model works from the role rather than from a general idea of what analysts do. It fixes the number of competencies, forcing prioritization instead of a sprawling list nobody can score. It specifies the mix, two behavioral and one situational, so you get both what a candidate has done and how they reason about a situation they have not yet faced. And it demands a statement of what a strong answer demonstrates, which is what makes verification possible: without it you are reading questions and guessing at intent. A vague prompt produces a plausible list you cannot audit.
The draft was a starting point, not a finished product, and the human-verification step is where the real work happened. Theo read every question against two tests. First, relevance: did the question measure something the job genuinely needs, or did the AI pad the list with generic filler. He cut three questions that tested skills the role did not use. Second, bias: did any question assume a particular background, penalize a non-traditional career path, or stray toward protected territory like family status, age, or origin. The AI had drafted one question asking how the candidate balanced "demanding family commitments" with deadlines; Theo struck it immediately, because it invited exactly the kind of inference that has no place in a hiring decision. The lesson is the pattern for all AI-assisted prep: the model drafts at speed, the human verifies for relevance and fairness before anything reaches a candidate.
That family-commitments question deserves a second look because of how ordinary it is. It does not read as hostile; it reads like an interviewer trying to be understanding, which is precisely why it survives in processes where nobody reviews questions in advance. The problem is not intent. It is that the question invites the candidate to disclose family status, and once disclosed, that information sits in the room influencing a decision it has no business influencing. Written questions reviewed before the loop catch this. An improvised question is reviewed by nobody and leaves no trace that it was asked.
Step Two: Build the Scorecard and the Rubric
A consistent question set is only half the structure. The other half is a shared scorecard that tells four different interviewers what a 1 and a 4 actually look like. Theo had the AI draft a four-point anchored rubric for each competency, then edited it for clarity. On a 1-to-4 scale, a 1 meant the candidate showed no relevant evidence or a clear gap, a 2 meant partial evidence with notable concerns, a 3 meant solid, role-ready evidence, and a 4 meant strong evidence that exceeded the bar with specific examples. The anchors matter because "4" means nothing until everyone agrees on what earns it.
Look at what those anchors are built from. Every one is a statement about evidence, not about the candidate as a person: absence of evidence, partial evidence with concerns, evidence sufficient for the role, evidence beyond the bar with specific examples. An interviewer scoring against them has to point at something the candidate said, which is a very different discipline from scoring against an impression. Anchors phrased as personal qualities, "confident," "impressive," "a self-starter," restore exactly the subjectivity the scorecard was built to remove, and they do it while looking rigorous, because they still produce a number.
He then split the four competencies across the four interviewers so each owned the questions they were best positioned to assess, with one shared competency all four scored to create a calibration point. Every interviewer got the same scorecard, the same anchors, and the same instruction: write your scores and your evidence notes before the debrief, and bring them to the meeting unchanged. The scorecard is the artifact that turns four opinions into four comparable measurements.
The shared competency is the cleverest part of the design and the part most often skipped. If each interviewer scores only their own competency, you get four measurements of four different things and no way to tell whether your interviewers are calibrated to each other at all. One competency scored by everyone gives you that reading. Tight clustering means the anchors are being applied consistently and the individually owned scores can be trusted further. A scatter means you have learned something about your interviewers before learning anything about the candidate, on this loop rather than after a bad hire.
Step Three: Prepare the Interviewers
Materials do not run themselves. Before the loop, Theo gives each interviewer their assigned competency, the questions attached to it, the full anchored rubric, and one instruction that carries most of the weight: score and write evidence notes before the debrief, and bring them unchanged. He states it as a rule rather than a suggestion because the pressure to soften a score arrives later, in a room, from a colleague who outranks you, and a rule set in advance is what an interviewer can point to then. Evidence notes are the half of that instruction most likely to be dropped under time pressure and the half that makes disagreement resolvable, since when two interviewers differ the only way forward is comparing what each of them heard. Theo asks for short, specific notes tied to what the candidate said rather than summaries of impressions, which quietly improves the interviewing itself: someone who must produce evidence listens for it and asks the follow-up that produces it.
The Worked Example: Independent Scoring Before the Debrief
Here is the part that does the heavy lifting against groupthink. After the loop, all four interviewers scored the shared competency, problem-solving under operational pressure, independently, submitting before anyone discussed the candidate. Their raw scores came in at 4, 2, 3, and 2. That spread is not a problem to hide. It is information. A two-point gap between the highest and lowest score signals that the interviewers saw genuinely different things, and that is precisely the disagreement a structured process is designed to surface rather than bury.
Contrast this with Theo's old process. In an unstructured debrief, the interviewer who spoke first and gave a confident "4" would have anchored the room, the two who scored a 2 would have softened to "well, maybe I missed something," and the panel would have converged on a high score within five minutes, never examining why two experienced interviewers had real doubts. The pre-committed independent scores made that impossible. With the spread on the table, the debrief became an evidence review: the interviewer who scored a 4 pointed to a specific example of the candidate diagnosing a routing failure under deadline, while the two who scored a 2 noted the candidate had deflected a direct follow-up about a decision that went wrong. After comparing evidence, not impressions, the panel calibrated to a shared 3, with one interviewer holding at 2 and documenting why. The number moved because the evidence justified it, not because the loudest voice prevailed. The dissent was preserved in the record rather than erased.
Notice that the two pieces of evidence are not in conflict. The candidate genuinely did diagnose a routing failure under deadline, and genuinely did deflect a question about a decision that went wrong. Both are true, and an unstructured debrief would have forced the panel to pick one story. The structured version let them hold both, which is a more accurate picture of a person than either score alone, and it was available only because two interviewers wrote down what they heard before anyone told them what the answer was.
The numeric point is worth stating plainly. Before calibration, the score spread was two full points on a four-point scale, wide enough that the decision could have gone either way depending on who talked first. After an evidence-based debrief, the spread narrowed to one point and the reasoning was explicit. Independent scoring did not manufacture false agreement; it produced honest, documented agreement and kept the disagreement that mattered visible.
The preserved dissent is the detail to steal. A panel that records only its final consensus throws away its most useful signal, because the interviewer who held at 2 has left behind a specific, dated, evidence-backed reservation. If this hire struggles in the same area a year from now, that note is the fastest route to understanding what the process saw and chose to accept. If the hire thrives, it is evidence the panel weighed a real concern rather than missing it. Clean unanimous numbers usually mean someone stopped arguing, not that everyone agreed.
Step Four: Run the Calibrated Debrief
Theo facilitated the debrief with three rules. Scores are submitted before the meeting and cannot be changed silently; any change is announced with the evidence that prompted it. The session reviews evidence competency by competency, not candidate by candidate, so the panel argues about what was observed rather than about overall vibes. And the recruiter watches for the tells of groupthink: rapid consensus, deference to seniority, and the word "fit" used as a substitute for a specific, role-related reason. AI can help here too, summarizing the four interviewers' written notes into a single comparison view so the panel sees agreements and conflicts at a glance, but the summary is a draft Theo verifies against the raw notes before the panel relies on it. The model organizes the evidence; it never casts a vote.
Panels resist the competency-by-competency rule, because the instinct is to go around the room collecting overall verdicts, which feels efficient and is the fastest possible route to anchoring. Working one dimension at a time bounds the first thing anyone says, and it stops a strong showing in one area from silently inflating scores in unrelated ones, which is the halo effect operating live in a meeting.
The AI summary needs one guardrail spelled out. It is genuinely useful, because reading four sets of notes in a live meeting is slow and the panel will skim. But a summary is a compression, and compressions drop the outlier detail one interviewer noticed and nobody else did, which in a hiring debrief is often the most important sentence in the file. Theo verifies the comparison view against the raw notes before the panel works from it, and keeps those notes open during the discussion so any interviewer can pull their own evidence back into the room.
Step Five: Document, and Use the Record
The structured process produces a natural paper trail, and Theo treats it as a compliance asset, not paperwork. For each candidate, the file holds the question set, the anchored rubric, every interviewer's independent scores with evidence notes, and the calibrated outcome with any preserved dissent. That record is what lets the firm show, if ever challenged, that the decision rested on consistent job-related criteria applied the same way to every candidate. It also lets Theo audit fairness over time: if scores trend lower for one group on a particular competency, the documentation tells him whether the question, the rubric anchors, or an individual interviewer is the source, so he can fix the specific cause instead of guessing.
That diagnostic property is the payoff most teams never collect, and it exists only because the record is granular. A file holding one overall verdict per candidate can tell you a disparity exists and nothing about where it came from. Per-competency scores from named interviewers with evidence notes can distinguish three problems that look identical in aggregate: a question that disadvantages a group regardless of who asks it, an anchor ambiguous enough that interviewers read it differently, and one interviewer scoring differently from colleagues on the same competency. Three findings, three different fixes, and without the granularity you either do nothing or change everything.
One boundary is worth drawing clearly. In this project, AI prepares questions, drafts the rubric, and summarizes notes. It does not score candidates and it does not rank them. That keeps the work on the right side of the line: the interview decision is made by humans on documented evidence. If the firm ever introduced AI that actually scored candidates, that would trigger additional obligations, including bias-audit and notice requirements under laws like NYC Local Law 144 where applicable, which is a good reason to keep the scoring firmly in human hands for now.
Anti-Patterns
Shipping the AI's question list without the relevance and bias pass. This is taking the twelve drafted questions into the loop because they read well. It happens because the draft arrives fast and fluent, and padding does not look like padding until you check each question against a named competency. What goes wrong is that generic questions eat interview time owed to the competencies the role actually needs, and a question like the family-commitments one Theo struck invites disclosure of protected information that then sits in the room affecting a decision. The counter is the two-test read on every item before anything reaches a candidate, deleting rather than rewriting whatever fails it.
A scale with no written anchors. This is handing four interviewers a 1-to-4 scorecard and assuming the numbers mean the same thing to all of them. It happens because the scale itself looks like the structure, and writing anchors feels like restating the obvious. What goes wrong is that you now hold four incomparable measurements dressed as data, and the panel spends the debrief discovering that one interviewer's 3 is another's 4 instead of discussing the candidate. The counter is anchors phrased in evidence, as Theo's are, never anchors phrased as personal qualities like confident or impressive.
Scoring in the room. This is opening the debrief by asking everyone for their number. It happens because it feels efficient and because chasing pre-submissions is one more errand before the meeting. What goes wrong is the whole mechanism: the first or most senior voice anchors the room, quieter interviewers adjust toward it without noticing, and the panel records agreement it never had. Theo's panel scored 4, 2, 3, 2 in private, a spread that would have collapsed into a confident 4 within five minutes had the numbers been spoken aloud in sequence. The counter is submission before the meeting, with any later change announced alongside the evidence prompting it.
Treating score spread as a problem to reconcile away. This is seeing a two-point gap and running the debrief as a negotiation toward a number everyone can live with. It happens because disagreement feels like a process failure and a unanimous panel is easier to report upward. What goes wrong is that the panel lands on a compromise nobody's evidence supports, and the reservation behind the low score, in Theo's case a candidate deflecting a direct follow-up about a decision that went wrong, disappears unexamined. The counter is to treat the spread as the finding, move numbers only where evidence justifies it, and keep any surviving dissent in the record with its reasoning.
Letting "fit" carry an unexamined reservation. This is accepting culture fit as a scoring reason, the phrase that carried Theo's old panel to the hire that washed out. It happens because the word sounds like consensus and unpacking it feels confrontational in a meeting. What goes wrong is that fit has no anchors and cannot be evidenced, so it becomes a container for anything an interviewer cannot or will not state directly, including reactions that are not job-related at all. The counter is to treat the word as a prompt rather than an answer: ask which competency is being described and what the candidate said that produced the reaction, then record whatever comes back, including silence.
Letting the model score, rank, or stand in for the notes. This is asking the AI which candidate is strongest, or working the debrief from its summary without checking it. It happens because the model will answer, the answer sounds measured, and reading four sets of raw notes live is genuinely slow. What goes wrong is twofold: a summary compresses away the outlier observation only one interviewer made, and a model that scores candidates changes the legal character of the process, triggering obligations such as the bias-audit and notice requirements under laws like NYC Local Law 144 where applicable. The counter is Theo's boundary, stated in the project documents rather than assumed, with the raw notes open throughout the debrief.
Build Checklist
Work these in order against one real open requisition. Each step produces a piece of the finished artifact set.
- Pull the job profile and extract competencies. Prompt the model to name the core competencies the role requires, then draft two behavioral and one situational question per competency, each with a statement of what a strong answer demonstrates. Fix the competency count in the prompt so you get prioritization, not a list nobody can score.
- Run the two-test read on every question. Name the competency each question measures and delete any you cannot map. Then check whether the question assumes a background, penalizes a non-traditional career path, or invites disclosure of family status, age, or origin. Expect to cut; Theo cut three of twelve.
- Write anchors in the language of evidence. Draft a four-point anchor per competency, then edit each one until it describes what the candidate said or showed rather than what kind of person they are. An anchor an impression could satisfy is not finished.
- Assign competencies and pick your shared one. Give each interviewer the competency they are best positioned to assess, and name one that everybody scores. That shared score is your calibration reading on the panel itself.
- Brief the panel with the rule, not the suggestion. Distribute the questions, the full rubric, and the instruction to submit scores with evidence notes before the debrief and bring them unchanged, stated as a rule so an interviewer under pressure has something to point to.
- Collect scores before anyone talks. Take submissions in writing before the meeting opens, and read the spread on the shared competency first. It tells you whether you are about to discuss a candidate or discover a calibration problem.
- Run the debrief competency by competency. Review evidence one dimension at a time, announce any score change with the evidence behind it, and treat "fit" as a request for a specific role-related reason. Record the outcome and every surviving dissent with its reasoning.
- File the full set and use it later. Store the questions, the rubric, all independent scores with notes, and the calibrated outcome. Then read across candidates one competency at a time to check whether scores trend differently by group, and whether the source is the question, an anchor, or an individual interviewer.
Reflection
- In your last panel debrief, who spoke first, and how close was the decision to what they said?
- If a rejected candidate asked why, which interviewer could point to something the candidate actually said?
- Where does "fit" appear in your interview records, and what is it standing in for?
- Do your interviewers score any competency in common, and if not, how would you know whether they are calibrated?
- When your panel last disagreed by two points, what happened to the low score, and is its reasoning still in the file?
- Which of your questions could you not map to a competency, and why is it still being asked?
Glossary
- Structured interview. An interview where every candidate for a role gets the same job-relevant questions against the same rubric, with scores committed before the panel discusses them.
- Competency. A capability the role genuinely requires, extracted from the job profile, which questions measure and scores are recorded against.
- Behavioral and situational questions. The first asks what a candidate has actually done, producing evidence of demonstrated capability; the second asks how they would reason about a situation not yet faced, producing evidence of judgment where experience is thin.
- Anchored rubric. A written definition of what each point on the scale means, phrased in evidence rather than personal qualities, so a 4 means the same thing to every interviewer.
- Shared competency. The one competency every interviewer scores, creating a calibration point that reads the panel's consistency rather than the candidate.
- Independent scoring. Submitting scores and evidence notes before the debrief and bringing them unchanged, which stops the first or most senior voice setting the room's answer.
- Evidence note. A short, specific record of what the candidate actually said, attached to a score, and the only thing that makes a disagreement resolvable.
- Score spread. The gap between highest and lowest independent score on a competency, treated as information about what interviewers saw rather than a problem to reconcile away.
- Anchoring. The effect where an early stated opinion sets the range everyone else adjusts within, which pre-committed scores prevent.
- Preserved dissent. A minority score kept in the record with its reasoning rather than erased into consensus, which is what makes the file useful when performance is reviewed later.
- Calibrated debrief. A session reviewing evidence competency by competency, requiring any score change to be announced with the evidence prompting it, and treating "fit" as a request for a role-related reason.
- Audit trail. The per-candidate file holding questions, rubric, all independent scores with notes, and the calibrated outcome with dissent, supporting both a fairness review and a defense of the decision.
Related Lessons
- Interview Preparation: Candidate Research & Structured Question Generation is the technique this project applies, covering how to get a usable question set out of a job profile.
- Structured Evaluation: Avoiding Halo Effects and Confirming Bias develops the mechanism behind independent scoring and the competency-by-competency debrief.
- Debrief Support: Synthesizing Panel Feedback goes deeper on the AI summary step, including how to stop a compression dropping the observation only one interviewer made.
- What Fair Hiring Looks Like: Structured Processes and Consistency is the fairness case for everything here, and the material to read before explaining to a skeptical hiring manager why the process changed.
- Avoiding Bias in Prompts: Language, Examples, and Assumptions covers the upstream half of the family-commitments problem, since a prompt carrying assumptions produces questions carrying the same ones.
- Documenting Decisions: Clear Records for Legal and Fairness Review is the record layer this project generates, and what turns a preserved dissent into evidence.
Closing
Theo's old process was not careless. Four experienced people spent real time with a candidate and reached a decision they believed in. What it lacked was any mechanism for the four of them to discover that they had been evaluating different things, and any record that would have let someone reconstruct the reasoning after the hire washed out four months later. The failure was not bad judgment. It was judgment with nowhere to be recorded and nothing to be compared against.
What replaces it is not complicated. Questions drawn from the job profile and read twice, once for relevance and once for bias. A four-point scale whose anchors describe evidence. One competency everyone scores so you can tell whether the panel is calibrated. Numbers committed before anyone speaks. A debrief that moves one competency at a time and demands evidence for every change. A file that keeps the dissent. Run it once on a live requisition and you own the materials; run it again and most of the preparation is already done. The AI makes the first hour cheap. Everything that makes the decision trustworthy happens after that, in human hands, on the record.
Key Takeaways
- Structure is what makes interviews predictive and defensible. Consistent job-relevant questions, a shared rubric, and documented scores predict performance better and are exactly what EEOC guidance expects when you must show a decision was role-based, not biased.
- AI drafts; humans verify. Let the model extract competencies and draft questions and anchors from the job profile in seconds, then verify every item for relevance and for bias before it reaches a candidate. Theo cut three of twelve, including a family-commitments question the AI had drafted.
- Anchor the scorecard so a 4 means the same thing to everyone. Written anchors per competency turn four opinions into four comparable measurements, and every anchor should describe evidence rather than a personal quality.
- Have everyone score one shared competency. That overlapping measurement reads whether the panel itself is calibrated, on every loop instead of after a hire goes wrong.
- Score independently before the debrief. Pre-committed scores stop the first or most senior voice anchoring the room. Theo's panel scored 4, 2, 3, 2, a two-point spread an unstructured debrief would have collapsed into a fast, false "4."
- Treat score spread as information, not a problem. An evidence-based debrief narrowed it from two points to one and preserved the lone dissent with its reasoning. The number moved because the evidence justified it, not because someone talked first.
- Run the debrief on evidence, competency by competency. Announce any score change with the evidence behind it, and treat "fit" as a demand for a specific, role-related reason rather than as an answer.
- Keep scoring human and document everything. AI prepares and summarizes but does not score or rank, which keeps the decision human and the audit trail clean. If AI ever scored candidates, obligations like NYC Local Law 144 would apply.
Frequently Asked Questions
Does structuring the interview mean interviewers cannot follow up on interesting answers? No, and a structure that prevented follow-up would be worse than no structure. What is fixed is the question set: every candidate for the role gets the same job-relevant questions mapped to the same competencies. How an interviewer probes an answer is where their skill lives, and the evidence-note requirement encourages probing, because someone who must record what the candidate said will ask the follow-up that produces it. In Theo's loop, the reservation behind the two low scores came from exactly such a follow-up.
Our panel almost always agrees. Is independent scoring worth the overhead? Frequent unanimity is the reason to check, not the reason to skip it. A panel that discusses first and scores after will agree at a rate that says more about the discussion than about the candidates, because the first confident opinion sets the range everyone else adjusts within. Scoring independently costs a few minutes and tells you which kind of agreement you have. Tight clustering on the shared competency means your panel really is calibrated. A scatter like Theo's 4, 2, 3, 2 means your prior unanimity was a debrief artifact, and you have just recovered information that was being destroyed on every loop.
What if the panel cannot agree even after reviewing the evidence? Then you record the disagreement rather than resolving it artificially. Theo's panel calibrated to a 3 with one interviewer holding at 2 and documenting why, which is more useful than a manufactured consensus at either number. The written dissent gives the decision an honest picture, strong demonstrated capability alongside an unresolved question about how the candidate handles their own mistakes, and gives whoever reviews this hire's performance later a specific, dated reservation to check against.
Can we let the AI score candidates if a human reviews the scores afterwards? That is a different process with different obligations, and this project deliberately does not go there. In Theo's design the model drafts questions and anchors and organizes written notes into a comparison view he verifies against the raw notes; it does not score or rank anyone. A tool that scores candidates changes the legal character of the process, triggering additional requirements including the bias-audit and notice obligations under laws like NYC Local Law 144 where applicable. Keeping scoring in human hands costs little, since the expensive part of preparation was always the drafting, which is exactly the part AI does well.
Skill.re