Interview Preparation: Candidate Research & Structured Question Generation
Marcus runs talent acquisition at a 600-person healthcare technology company, and his five recruiters fill roughly 90 roles a year. The bottleneck is not sourcing. It is the twenty minutes before every interview when a hiring manager opens a resume for the first time, skims it on the way to the conference room, and improvises questions on the spot. This lesson follows Marcus as he uses AI to research candidates without crossing legal or ethical lines, generate a structured question bank tied to the competencies that actually matter, and produce a one-page briefing pack so every interviewer walks in knowing exactly what they own.
The Twenty-Minute Problem Before Every Interview
Marcus watched the failure play out on a Senior Backend Engineer loop last quarter: four interviewers, four different sets of questions, two of them overlapping almost entirely on "tell me about a hard bug," and nobody assigned to probe the one thing the role actually hinged on, which was experience scaling a system past a few thousand concurrent users. The candidate was strong and got an offer, but the debrief was a mess because no two interviewers had assessed the same thing. Four people had spent an hour each and produced four impressions that could not be added together.
The obvious diagnosis is the wrong one. The interviewers were not lazy or unskilled. They were handed a resume with no shared plan and twenty minutes to build one, and under those conditions everyone reaches for the question they personally find revealing. That is why the overlap happened: "tell me about a hard bug" is a good question, which is exactly why two people chose it independently. And it is why the gap happened: nobody had been told they owned scalability. Improvisation does not distribute coverage; it concentrates it wherever the interviewers' instincts agree.
AI does not fix the deeper problem, which is that the team had no shared structure. What it does is make structure cheap enough that busy people will actually use it. Most teams already know structured interviewing is better and still do not do it, and the reason is almost never disagreement about the evidence; it is that building a competency-mapped question bank and rubric for every open role is hours of work per requisition and nobody has the hours. Change the cost and the behaviour changes with it.
Why Structured Interviews Are Worth the Prep
The reason to invest in this at all is that structured interviews are one of the better-evidenced practices in hiring. A structured interview asks every candidate for the same role the same core questions, scored against the same defined criteria, in the same general sequence. It consistently predicts on-the-job performance better than the unstructured "let's just chat and see how it goes" approach, and it is far easier to defend if a hiring decision is ever challenged.
This matters under United States employment law. The EEOC enforces Title VII, which prohibits employment discrimination based on protected characteristics, and a consistent, job-related, documented process is exactly what you want if you ever have to show that two candidates were evaluated on the same basis. Improvised interviews, where each candidate gets different questions, are the opposite of that. So when Marcus uses AI to build question banks, he is not just saving time. He is converting an ad-hoc process into a consistent, documented one, which is both fairer and more defensible.
The discipline to hold onto: AI generates the structure and the draft questions, but a human confirms that every question is job-related and that no question strays into protected territory. AI will happily write a friendly icebreaker that asks about family or where someone is "originally from." You delete those. The tool accelerates; it does not get a vote on what is appropriate to ask. Notice that such questions do not read as hostile; they read as an interviewer building rapport, which is precisely why they survive where nobody reviews questions in advance, and why the review has to be a scheduled step rather than a thing a careful person happens to do.
Responsible Candidate Research: What to Look Up and What to Leave Alone
Before Marcus builds questions, he wants the interview to feel prepared rather than generic. Candidate research is where good intentions go wrong fastest, so he works from a simple rule: research the professional record relevant to the role, and stop there. The legitimate inputs are the candidate's resume, the portfolio or work samples they submitted, a public professional profile if they have one, and published work like talks, articles, or open-source contributions they chose to make public. The off-limits inputs are personal social media, anything that reveals a protected characteristic, and anything the candidate did not put forward as part of their professional presence.
| Input | Status | Why |
|---|---|---|
| Resume and submitted portfolio or work samples | Legitimate | The candidate put these forward for this role specifically |
| Public professional profile | Legitimate | Part of the professional presence the candidate chose to publish |
| Published talks, articles, open-source contributions | Legitimate | Work the candidate chose to make public, relevant to the role |
| Personal social media | Off-limits | Not offered as professional presence, and a reliable way to import bias |
| Anything revealing a protected characteristic | Off-limits | Cannot lawfully or fairly inform the decision, and cannot be unseen once seen |
| Anything the candidate did not put forward professionally | Off-limits | Outside what the hiring decision requires |
If any candidate is based in the EU or the United Kingdom, GDPR applies to how their personal data is collected and used, including the principle of data minimization, which means you should only process what is necessary and relevant for the hiring decision. Pulling a candidate's personal Instagram into your prep is not necessary and not relevant. It is also a reliable way to import bias you cannot later remove from your judgment.
That last sentence is the part recruiters underrate, because it holds even for people acting in complete good faith. Information about someone's family, health, politics, or beliefs does not sit inertly in your head waiting to be consciously discounted. Once you have seen it, you cannot demonstrate that it played no part, and neither can you fully know. The boundary protects the quality of your own judgment as much as the candidate's right to be assessed on their work. The organizing question is not "can I find this" but "did the candidate offer this for this decision," and the answer to the first is almost always yes.
The Research Synthesis Prompt
Here is the research synthesis prompt Marcus uses, pasting in only the resume and submitted portfolio notes:
"You are helping me prepare for an interview. Below is a candidate's resume and the portfolio notes they submitted for a Senior Backend Engineer role. Summarize: (1) the two or three accomplishments most relevant to building and scaling backend systems, (2) any specific technologies or architectures they claim deep experience with, (3) gaps or ambiguities in the resume worth clarifying in the interview. Do not speculate about anything not stated in these documents. Do not infer or comment on the candidate's age, gender, nationality, or any personal characteristic. If something is unclear, say so rather than guessing."
Read what each part of that prompt is buying. Pasting only the permitted documents enforces the research boundary at the point of use rather than leaving it to memory; the model cannot summarize material it was never given. Fixing the output at two or three accomplishments forces prioritization instead of a restatement of the resume. Asking for gaps and ambiguities changes the interview rather than just the prep, because it converts what the documents do not say into a specific thing to ask about.
The instruction not to infer or comment on age, gender, nationality, or any personal characteristic is doing work that the input boundary alone does not do. A resume you were entitled to read still carries signals a model will cheerfully generalize from: graduation dates, a name, a location, the language of a previous employer. A summary that mentions any of them has moved a protected characteristic from a document nobody was scrutinizing into a briefing document interviewers will read closely, which is a worse position than not having asked. The instruction keeps the summary about the work.
That final instruction matters. It tells the model to flag uncertainty instead of inventing a plausible-sounding detail, which is the single most common way AI research goes wrong: a confident summary of an accomplishment the candidate never actually claimed. The consequence in an interview room is specific and awkward. An interviewer opens with a question premised on a project the candidate did not do, and the candidate now has to correct the interviewer while under evaluation, which costs them composure and costs the company credibility. Marcus checks the summary against the source documents before he walks in, and the check is quick precisely because the summary is short and traceable.
Start From Competencies, Not Questions
The mistake most teams make is starting from questions. Marcus starts from competencies, because a question bank is only as good as the criteria underneath it. For the Senior Backend Engineer role, he and the hiring manager agreed on four competencies that actually predict success: systems design and scalability, debugging and incident response, collaboration across teams, and technical communication. Everything flows from those four.
The agreement with the hiring manager cannot be delegated to a model, and it surfaces disagreement early rather than in the debrief. Naming four competencies is a forcing function: it requires the manager to say what the role actually hinges on and to accept that anything left off will not be systematically assessed. On the loop that went wrong, nobody had performed that exercise, which is why scalability had no owner.
Building the Question Bank
He gives AI the competencies and asks for behavioral questions in a specific format. Behavioral questions ask about real past situations rather than hypotheticals, on the logic that what someone has actually done predicts future behavior better than what they say they would do. His prompt:
"Generate a structured interview question bank for a Senior Backend Engineer. Use these four competencies: systems design and scalability, debugging and incident response, cross-team collaboration, technical communication. For each competency, write: (1) one primary behavioral question asking about a specific past situation, (2) two follow-up probes to get past a rehearsed answer to specifics, (3) a three-level rubric describing what a weak, solid, and strong answer looks like. Keep questions job-related and free of anything touching protected characteristics."
The follow-up probes are the requirement most likely to be dropped and the one that determines whether the bank is usable. A single behavioral question invites a rehearsed answer, and rehearsed answers are smooth, complete, and almost entirely uninformative, because a candidate who has told a story ten times has sanded off every detail that would let you evaluate it. Two probes per question build the escape route into the artifact. Generating the rubric alongside keeps the question and the standard for judging it attached, which stops the bank degrading into a list of questions people ask and then score by feel.
A worked example of the systems-design entry the model produced, after Marcus edited it:
Primary question. "Tell me about a backend system you designed or significantly re-architected to handle growth. What was the load before and after, and what were the hardest tradeoffs you made?"
Follow-up probes. "What did you measure to know the change worked?" and "What would you do differently if you rebuilt it today?"
Rubric. A weak answer describes the system only in vague terms and cannot name concrete numbers or tradeoffs. A solid answer names real load figures, a specific bottleneck, and the change that addressed it. A strong answer does all of that and also articulates the tradeoff that was deliberately accepted, such as added operational complexity in exchange for horizontal scalability, and what they monitored to catch the downside.
Notice how the primary question is built. It asks for a specific past system rather than a philosophy of scaling, then asks two things a person who did the work would know and a person who watched it happen would not: the load before and after, and the hardest tradeoff. The two probes attack from opposite sides, one testing whether the candidate closed the loop or simply shipped and assumed, the other testing whether they have reflected. The three rubric levels are separated by evidence, not polish: vagueness, then concrete specifics, then specifics plus an owned tradeoff and the monitoring that went with it.
Editing the Rubric, Because Bias Hides There
Marcus reviews every rubric line, because the rubric is where bias hides. A line like "communicates with confidence" rewards a personality style, not a job skill, so he rewrites it to "explains a technical tradeoff clearly enough that a non-specialist could follow it." Same competency, now measuring the job requirement rather than charisma.
The pattern generalizes to every rubric you will ever edit. "Communicates with confidence" cannot be evidenced: two interviewers scoring the same answer against it will produce different numbers and neither can say why, because the line describes a reaction in the interviewer rather than a behaviour in the candidate. It also carries known biases, since confidence is read differently depending on who displays it and norms about assertiveness vary widely, so the line rewards a background rather than a capability. The rewrite fixes both at once, because it can be checked against what was actually said and it happens to be what the job requires.
The test Marcus applies to any rubric line is whether an interviewer scoring against it would have to point at something the candidate said. If not, the line measures an impression, and the number it produces will look like data while carrying none. This is also why the review has to be a human pass rather than another prompt: the model drafted that phrasing because it is everywhere in the material it learned from. AI drafts the rubric; a human confirms it is job-related.
Dividing the Loop So Interviewers Do Not Overlap
Four competencies and four interviewers suggests an obvious move: assign each interviewer a competency to own. This is the fix for the overlap problem Marcus saw last quarter, where two interviewers both asked about debugging and nobody owned scalability. With a shared bank and clear ownership, the loop covers all four dimensions once each, at depth, instead of covering one dimension four times shallowly.
He asks AI to turn the bank into a coverage plan: "Given these four competencies and a four-person interview loop, assign one primary competency to each interviewer, list the secondary competency they can probe if time allows, and flag any competency that needs two interviewers because it is high-risk for this role." For this role the model flagged systems design as worth a second look, so Marcus assigned it as a secondary to the interviewer with the most architecture experience. Every competency now has a clear owner and the most important one has a backup, decided before anyone walks into a room.
The high-risk flag is the part of that prompt worth keeping when you adapt it. Competencies are not equally consequential, and for the one that determines whether the hire works at all, a single interviewer's read is a thin basis. Deliberate double coverage there gives two independent readings where it counts without doubling the cost of the loop. Marcus still made the assignment himself, matching the second read to the interviewer with the most architecture experience, which is a judgment about his own colleagues the model has no basis for.
The One-Page Briefing Pack
The deliverable that changes interviewer behavior is a single page per interviewer, generated in seconds once the bank and the assignments exist. Marcus prompts: "Create a one-page interview briefing for the interviewer assigned to systems design and scalability. Include: the candidate's two most relevant accomplishments, their assigned competency and why it matters for this role, their primary question and two follow-ups, the three-level scoring rubric, and a reminder to score against the rubric immediately after the interview rather than relying on memory."
That final reminder is deliberate. The biggest threat to a structured interview is not bad questions; it is interviewers who score from memory hours later, when impressions have blurred into a halo effect. A briefing that ends with "score now, in the rubric, before you discuss with anyone" protects the independence of each evaluation, which is what makes the debrief meaningful. When four people score independently against the same rubric and then compare, disagreement is signal. When they score from memory after chatting in the hallway, disagreement is just noise.
The two halves of that instruction are separate protections. "Score now" guards against decay, since specific evidence fades within hours into a general sense of how it went. "Before you discuss with anyone" guards against contamination, since a hallway conversation is enough to move a number, and once it has moved the second score is an echo rather than independent evidence.
What the Prep Cycle Actually Costs
Marcus's whole prep cycle for a loop now runs about 25 minutes: research synthesis, competency confirmation with the hiring manager, question-bank generation, coverage assignment, and four briefing packs. The old way was twenty improvised minutes per interviewer with no shared structure to show for it. That comparison is the argument to make to a skeptical hiring manager, because it is not a request for extra time. It is a reallocation of time already being spent, from four people improvising in parallel to one person preparing once, with an artifact the next loop for that role largely inherits.
Anti-Patterns
Researching everything you can find. This is opening a candidate's personal social media alongside their resume on the theory that more context makes a better-prepared interviewer. It happens because the material is one search away, curiosity feels like diligence, and nobody is asked to justify what they looked at. What goes wrong is that you acquire information about family, health, beliefs, or origin that cannot lawfully or fairly inform the decision and cannot be unseen, and for EU or UK candidates you have processed personal data beyond what the hiring decision requires, against GDPR's data minimization principle. The counter is Marcus's boundary applied as a rule: resume, submitted portfolio, public professional profile, published work, and nothing the candidate did not put forward professionally.
Taking the AI's summary into the room unchecked. This is reading the research synthesis on the way to the interview and treating it as the candidate's record. It happens because the summary is fluent, short, and arrives at exactly the moment you have no time. What goes wrong is that the most common AI research failure is a confident description of an accomplishment the candidate never claimed, and it surfaces live, in front of the candidate, who must correct an interviewer while being evaluated. The counter is Marcus's pair of controls: instruct the model to flag uncertainty rather than guess, then trace every claim back to a line in the resume or portfolio before the loop.
Starting from questions instead of competencies. This is asking a model for good interview questions for a role and building the loop from what comes back. It happens because questions are the visible artifact and a list of them feels like progress, while competencies feel like an abstraction to get past. What goes wrong is that you cannot score answers consistently without criteria underneath them, so four interviewers score against four private standards and nothing guarantees the thing the role hinges on gets assessed. The counter is to fix three to five competencies with the hiring manager first and generate everything else from that agreement.
Shipping rubric lines that measure personality. This is accepting "communicates with confidence," "shows strong ownership," or "impressive presence" because they sound like standards. It happens because that phrasing is pervasive in interview material and therefore in what a model drafts, and because such a line still produces a number, which makes the scorecard look rigorous. What goes wrong is that it describes a reaction in the interviewer rather than a behaviour in the candidate, so two people scoring the same answer disagree and neither can say why, and it rewards norms about assertiveness rather than the job. The counter is Marcus's test: if scoring would not require pointing at something the candidate said, rewrite the line until it would.
Leaving coverage to chance. This is running four interviewers off a shared question bank without assigning anyone a competency to own. It happens because the bank feels like the structure and the assignment step looks like administrative overhead. What goes wrong is exactly what Marcus saw: interviewers converge on the questions they personally find revealing, a good question gets asked twice, and the competency the role hinges on gets asked by nobody. The counter is a written coverage plan naming a primary competency per interviewer, a secondary to probe if time allows, and deliberate double coverage on the highest-risk competency.
Scoring later, and socially. This is filling in the scorecard that evening, or after comparing notes with a colleague in the hallway. It happens because scoring feels like paperwork that can wait and because talking about a candidate straight afterwards is the most natural thing in the world. What goes wrong is two separate failures: specific evidence decays into a general impression within hours, so the score reflects a feeling, and any conversation before scoring makes the second score an echo rather than independent evidence. The counter is the briefing's closing instruction, with the rubric in the interviewer's hand so scoring immediately is possible.
Practice Prompts
Work these against one real open requisition, in order. Each produces a piece of the artifact set the next loop for that role reuses.
- Agree the competencies before you write anything. Sit with the hiring manager and name three to five competencies that actually predict success in the role. Make the trade explicit: anything not on the list will not be systematically assessed. Note which single competency the role most hinges on.
- Draw your research boundary in writing. List what you will read for this candidate: resume, submitted portfolio or work samples, public professional profile, published talks, articles, or open-source contributions. List what you will not: personal social media, anything revealing a protected characteristic, anything the candidate did not put forward professionally. For any EU or UK candidate, confirm what you are processing is necessary and relevant to the hiring decision.
- Run the research synthesis with the guardrails in the prompt. Paste only the permitted documents and ask for the two or three most role-relevant accomplishments, the technologies they claim deep experience with, and the gaps worth clarifying. Instruct the model not to speculate beyond the documents, not to infer or comment on age, gender, nationality, or any personal characteristic, and to say when something is unclear rather than guessing.
- Verify the summary against the source. Trace each claimed accomplishment back to a line in the resume or portfolio before it reaches a briefing pack. Delete anything you cannot find.
- Generate the bank in the full format. For each competency, ask for one primary behavioral question about a specific past situation, two probes designed to get past a rehearsed answer, and a three-level rubric. Require questions to be job-related and free of anything touching protected characteristics.
- Edit every rubric line to the evidence test. For each line, ask whether scoring against it would require pointing at something the candidate said. Rewrite anything that describes a personality trait into the job requirement it was standing in for, the way "communicates with confidence" becomes "explains a technical tradeoff clearly enough that a non-specialist could follow it."
- Build the coverage plan. Assign one primary competency per interviewer, name a secondary they may probe if time allows, and give the highest-risk competency a deliberate second reader chosen by you, based on who is best placed to judge it.
- Produce one page per interviewer. Include the two most relevant accomplishments, the assigned competency and why it matters, the primary question and both probes, the rubric, and the instruction to score immediately and before discussing the candidate with anyone.
Reflection
- On your last loop, which competency did nobody own, and when did you find out?
- What did you read about your last candidate that they did not put forward for the role, and what did it change?
- If a rejected candidate asked what they were evaluated on, what document would you open?
- Which line in your current scorecard could two interviewers score differently without either being able to explain why?
- How long after your last interview did you write your score, and had you spoken to anyone about the candidate first?
- Where in your process does someone check that an AI-generated summary matches the source documents?
Glossary
- Competency. A capability the role genuinely requires, agreed with the hiring manager before any question is written, and the criterion every question and rubric line traces back to. The evidence test is what keeps a rubric line tied to one: would scoring against it require pointing at something the candidate actually said? If not, the line measures an impression.
- Data minimization. The GDPR principle that you should only process personal data necessary and relevant for the purpose, which for EU and UK candidates makes the research boundary a legal requirement rather than only good practice.
- Research boundary. The rule that prep draws on the resume, submitted portfolio, public professional profile, and published work, and never on personal social media, anything revealing a protected characteristic, or anything the candidate did not put forward professionally.
Related Lessons
- Hands-On Project: Conduct a Structured Interview with AI-Assisted Prep is where this technique gets run end to end on a live requisition, through independent scoring and a calibrated debrief.
- Research Synthesis: Building Candidate Context from Multiple Sources goes deeper on the synthesis step, including how to combine permitted sources without losing traceability to the original document.
- Privacy Boundaries: Data Sharing, Tool Selection, and Compliance covers where candidate documents are allowed to go once you paste them into a tool, which is the half of the research boundary this lesson treats as given.
- Avoiding Bias in Prompts: Language, Examples, and Assumptions is the upstream half of the rubric problem, since a prompt carrying assumptions produces rubric lines carrying the same ones.
- Structured Evaluation: Avoiding Halo Effects and Confirming Bias develops why scoring immediately and independently is what protects the debrief.
- Debrief Support: Synthesizing Panel Feedback picks up after the loop, once four independent scores against a shared rubric exist to be compared.
Closing
The Senior Backend Engineer loop that went wrong did not go wrong because anyone was careless. Four capable people each did their honest best with a resume and twenty minutes, and the result was two interviewers asking the same good question and nobody assessing the thing the role hinged on. That outcome was produced by the conditions, not the people, and it reproduces itself on every loop run under the same conditions no matter how experienced the interviewers are.
What Marcus changed is not complicated. Competencies agreed with the hiring manager before any question exists. Research held inside a boundary the candidate defined by what they chose to put forward, with the model instructed never to speculate or infer. A bank of primary questions, probes, and three-level rubrics generated in minutes and then edited line by line so every standard describes evidence rather than personality. A coverage plan with a named owner per competency and a second reader where it matters. One page per interviewer, ending with the instruction to score now, in the rubric, before talking to anyone. The AI made the structure cheap. Everything that makes it fair, the boundary, the verification, the rubric edits, stayed exactly where it belongs.
Key Takeaways
- Start from competencies, not questions. A question bank is only as good as the criteria beneath it. Agree with the hiring manager on three to five competencies that actually predict success in the role, then have AI generate questions and rubrics for each. Questions without defined criteria produce inconsistent scoring.
- Structure is what makes interviews fair and defensible. Asking every candidate the same job-related questions scored against the same rubric predicts performance better than improvised conversation and aligns with the consistency the EEOC looks for under Title VII.
- Research the professional record and stop there. Resumes, submitted portfolios, public professional profiles, and published work are fair game. Personal social media and anything revealing a protected characteristic are not. For EU and UK candidates, GDPR's data minimization principle makes that boundary a legal requirement, not just good practice.
- Put the no-inference instruction in the prompt. Tell the model not to infer or comment on age, gender, nationality, or any personal characteristic, because a document you were entitled to read still carries signals a summary would otherwise surface into a briefing.
- Make AI flag uncertainty instead of inventing detail. The most common research failure is a confident summary of an accomplishment the candidate never claimed, and it surfaces live, in front of the candidate. Instruct the model to say when something is unclear, and verify against the source before you walk in.
- Edit the rubric, because bias hides there. Lines like "communicates with confidence" reward personality, not job skill, and cannot be evidenced. Rewrite every line until scoring against it would require pointing at something the candidate said. AI drafts the rubric; a human confirms it is job-related.
- Assign one competency per interviewer and double the risky one. Clear ownership stops the loop probing one dimension four times and missing three others, and a second reader on the highest-risk competency gives two independent readings where it counts.
- Score immediately, and before you talk to anyone. Scoring late records a feeling rather than evidence, and scoring after a hallway conversation records someone else's. Independent scoring against a shared rubric is what makes debrief disagreement signal instead of noise.
Frequently Asked Questions
Does a shared question bank stop interviewers following up on interesting answers? No, and a structure that prevented follow-up would be worse than none. What is fixed is the core question set: every candidate for the role gets the same job-relevant questions mapped to the same competencies. How an interviewer probes an answer is where their skill lives, which is why the bank ships two probes with every primary question rather than leaving probing to instinct, and an interviewer who wants to go further from there should.
The candidate has a public personal account. Can I look at it if it is public? Public is not the test; offered is. The boundary is what the candidate put forward as part of their professional presence for this role, which is why a submitted portfolio and a published conference talk are in scope and a personal account is not, regardless of its privacy setting. You cannot unsee what you find, so information about family, health, or beliefs sits in your judgment whether or not you intended to weigh it, and neither you nor anyone else can later demonstrate that it played no part. For EU and UK candidates, GDPR's data minimization principle turns the same boundary into a legal requirement.
Can I skip the rubric edit if the AI's rubric already looks reasonable? A rubric that looks reasonable is exactly the failure case. Phrasing like "communicates with confidence" is pervasive in interview material, so it is what a model drafts by default, and it reads as a legitimate standard while measuring the interviewer's reaction rather than the candidate's behaviour. It still produces numbers, which makes the scorecard look rigorous while rewarding assertiveness norms that vary by background. Run the evidence test on every line, and note that asking the model to check its own draft is not a substitute, since it applied that phrasing in the first place.
Skill.re