Prompting for Resume Screening, Sourcing, and Research
Marcus runs technical recruiting at a 220-person B2B analytics company, and on a normal week he is the only recruiter covering four open engineering roles. When a data engineer requisition opened, 140 applications landed in nine days. He had screened resumes with AI before, but his prompts were lazy: "Summarize this resume." The summaries were fine and useless. They told him what the candidate said about themselves, not whether the person could do the job he was actually hiring for. This lesson is the framework Marcus rebuilt his screening around, the one that turned a generic summarizer into a structured screening assistant that saved him roughly eight minutes per resume without making a single hiring decision on his behalf.
Why Generic Prompts Fail
Resume screening is one of the highest-volume tasks in recruiting. You are looking at dozens of applications per role, and you need to quickly separate "not a fit" from "might be interesting" from "strong contender." It is exactly where prompting pays off most, and exactly where prompting badly wastes the most time. The reason "summarize this resume" produces weak output is that the model has no definition of good. A resume summary written without a target role optimizes for nothing. It will faithfully restate the candidate's self-description, lead with whatever the candidate led with, and stay silent on the gaps that matter to you because it does not know what matters to you. Marcus was getting eloquent paragraphs that left him exactly as uncertain as before he ran them.
The fix is to stop asking the model to summarize and start asking it to evaluate against a specific, stated standard. Instead of a generic summary, you ask it to extract specific qualifications, flag concerns without bias, surface cultural fit signals, and highlight what is actually relevant to this hire. A screening prompt that works has five layers, each of which closes one of the gaps a generic prompt leaves open: role specification, qualification extraction, red flag identification, cultural fit signals, and decision support. The layers are sequential because each one depends on the one before it, and once you have built two or three of them for real roles you have the beginnings of a prompt library you will reuse every week.
Layer One: Role Specification
Be brutally specific about what you are hiring for, because the job title is never enough. Consider a senior product manager search. The weak prompt is "Screen this resume for a product manager role." The strong version reads: "Screen this resume for a senior product manager role at a B2B SaaS company selling to healthcare organizations. This person will lead the roadmap for our compliance and integration products, working with a team of 2 engineers and supporting 15 mid-market customers. Key success metrics: ability to manage ambiguity, translate customer needs into product decisions, and drive execution with a small team." That second prompt tells the model exactly what good looks like. It is not about product management skills in the abstract; it is about the skills relevant to one specific situation.
Marcus applies the same discipline to his data engineer role. His opening reads: "Screen this resume for a data engineer at a mid-market analytics company. The team is three engineers. This person owns ETL pipelines and the warehouse layer end to end, works in Snowflake, Python, Airflow, and AWS, and must be able to take a project from ambiguous request to production without daily direction." That paragraph names the stack, the team size, and the autonomy expectation, which are the three things that will actually differentiate his candidates. Everything downstream gets sharper because the standard is now concrete rather than assumed.
Layer Two: Qualification Extraction
Tell the model which qualifications matter most and how to present them. Do not leave this to chance. For the product manager role, the extraction instruction lists five items: years of B2B SaaS experience; the specific industries or customer segments the candidate has worked with; the team size they have led or influenced; experience with product analytics or data-driven decision-making; and evidence of cross-functional collaboration with engineering and design. Then the instruction that most recruiters skip: for each qualification, note whether it is explicitly stated or inferred. This is far more useful than a generic summary, because you get structured information you can compare across candidates rather than ten different narratives to re-read.
The stated-versus-inferred flag is the single highest-value line in the layer. When the model marks a claim as inferred rather than stated, Marcus knows exactly which qualifications to verify in a screen instead of taking the model's reconstruction at face value. It converts a silent risk, meaning the model quietly filling a gap with a plausible guess, into a visible to-do item. Structured extraction also makes candidates genuinely comparable. When every resume comes back in the same shape, Marcus can scan ten of them in the time a generic summary takes to read one, because he is comparing the same fields on every candidate.
Layer Three: Red Flags Without Bias
Red flags are real, but they are usually about role fit rather than about the person, and the line between the two is exactly where bias enters. A model might flag "worked at 5 companies in 6 years." That pattern has at least three readings: opportunity-seeking, which is positive; poor fit repeated, which is negative; or external circumstances such as layoffs and relocations, which is neutral. Your prompt has to help the model distinguish. The instruction that does this asks four things of every concern: what is it; what might cause it, with multiple explanations acceptable; would this actually be a concern for this specific role; and what would you want to ask the candidate to understand it better.
The fourth question is what turns a red flag from a filter into a conversation starter, and it is the difference between screening and sorting. Instead of silently dropping a candidate with a two-year gap, Marcus gets a suggested screen question: ask what they worked on during that period and whether it was caregiving, study, contract work, or job search. The gap stops being a verdict and becomes something to understand. This framing also keeps Marcus on the right side of EEOC disparate-impact concerns, because gaps and tenure patterns correlate with caregiving and health status, both of which touch protected characteristics under Title VII and the ADA. A prompt that treats a tenure pattern as automatically disqualifying has encoded an assumption you would never write down as a policy.
Layer Four: Cultural Fit Signals
Culture fit is real, and it is also the single most common vector for bias in screening, so good prompts help you see signals without making unfair assumptions. The instruction asks whether anything in the resume suggests comfort with ambiguity and change; willingness to wear multiple hats; communication skills and the ability to work cross-functionally; and evidence of a learning and growth mindset. For each signal, the model must point to specific examples from the resume. Marcus adapts the same four to his own role, asking for evidence of taking initiative and ownership, learning new tools quickly, and clear written communication, each with a cited line.
Then comes the constraint that does the protective work, and it is the most important sentence in the entire prompt: do not make assumptions about personal characteristics, background, or identity, and do not infer cultural fit from educational choices, geographic location, or demographic signals. If a signal is not supported by a concrete example, say so rather than guessing. That constraint is what prevents the model from using implicit bias to "assess" culture fit. The citation requirement is what makes the output honest. A model asked whether a candidate "seems like a culture fit" will happily generate a vibe. A model required to point at the exact resume line that demonstrates ownership either finds one or admits it cannot, and both outcomes are useful.
Layer Five: Decision Support, Not Decisions
End the prompt with a clear decision framework. Do not ask the model to decide; ask it to prepare you to decide. Marcus requests a one-line summary of the candidate's fit drawn from a fixed set of four: strong fit, meaning clear qualifications and no red flags; possible fit, meaning the candidate meets most criteria with some concerns to explore; weak fit, meaning missing key qualifications or significant concerns; and unable to assess, meaning insufficient information. The fixed vocabulary matters as much as the content, because four consistent labels can be counted, sorted, and later audited for selection-rate disparities, while four different phrasings of "pretty good" cannot.
Notice what the question is and is not. It is "what is this candidate's likely fit," never "should we interview this person." The second question asks the model to make a human decision that Marcus owns and is accountable for. The first organizes the evidence so that Marcus can make that decision in seconds. It is also a less biased output, because a fit assessment against stated criteria can be checked against those criteria, whereas a yes-or-no recommendation hides its reasoning inside a single word.
Building the Complete Screening Prompt
The five layers assemble into one reusable prompt. Marcus's finished version for the data engineer role opens with the role specification: he is screening for a data engineer at a mid-market analytics platform, the team is three people, and the role requires infrastructure knowledge covering data warehousing, ETL pipelines, and SQL; hands-on engineering in Python or similar; and the ability to take ownership of projects independently. The analysis section asks for years of data engineering or related experience; the specific tools and technologies used, covering databases, pipeline tools, and cloud platforms; evidence of building or maintaining production systems rather than prototypes; and the candidate's demonstrated ability to work independently versus within large teams.
The red flag section names three role-specific concerns to look for: long gaps or very frequent job changes; a mismatch between job title and actual responsibilities; and a technology stack significantly different from the one in use, which he names explicitly as Snowflake, Python, Airflow, and AWS. For each red flag, the model must suggest questions to ask. The cultural fit section asks for evidence of taking initiative and ownership, the ability to learn new tools quickly, and communication skills, each tied to specific examples.
The prompt closes with four constraints that carry most of its fairness weight. Do not make assumptions about the candidate's background, identity, or demographics. Do not penalize educational choices, geographic location, or career paths that are simply different from a typical trajectory. If information is missing, note it explicitly rather than inferring. And format the response as a structured list with sections for Qualifications, Red Flags, Cultural Fit Signals, and Overall Fit Assessment. That last constraint looks like formatting housekeeping and is actually what makes 140 outputs comparable. What comes back is organized information that supports a decision instead of a paragraph that describes a person.
A Worked Example: 140 Resumes, One Afternoon
Here is what the five layers produced on Marcus's data engineer role. He assembled the single reusable prompt with a placeholder for the resume text, ran it across all 140 applications, and got back 140 structured assessments. The numbers below illustrate how the math works on a screen like his; they are not a benchmark. Generic summaries had been taking him about ten minutes per resume to read and still leave him uncertain, so 140 resumes was effectively impossible to do well in a week. The structured output came back in a fixed format he could scan in roughly ninety seconds each: 140 resumes in about three and a half hours of focused review instead of a theoretical twenty-three.
Of the 140, the assessment marked 31 strong fit, 44 possible fit, 58 weak fit, and 7 unable to assess. Marcus did not trust those labels blindly. He read every strong and possible fit himself, which is 75 resumes, and spot-checked 15 of the weak-fit and unable-to-assess group to confirm the model was not dropping qualified people for the wrong reasons. It had mislabeled two: a strong candidate marked weak because their pipeline work was described in a portfolio link the model could not see, and one possible-fit who was actually weak on the autonomy requirement. Catching those two is precisely why a human reads the borderline cases, and it is also why the "unable to assess" label exists rather than defaulting a thin resume to weak.
One fairness check Marcus ran afterward: he confirmed the strong-fit rate did not diverge sharply across the demographic groups he could estimate, using the four-fifths rule as a rough screen. If one group's selection rate had fallen below 80 percent of the highest group's rate, that would have been a signal to audit the prompt and the underlying applicant pool, not to proceed. This is the same disparate-impact logic the EEOC applies to any selection procedure, and it applies to AI-assisted screening exactly as it applies to a human reviewer. The prompt is a selection procedure. Treat it like one.
Sourcing and Research Prompts
Sourcing is detective work. You find a candidate online and you want to understand quickly what they have actually done, what they are interested in, and whether they are likely to engage. The same prompting discipline applies, with an extra emphasis on what the model may and may not use. When Marcus finds a candidate on GitHub whose profile suggests data engineering work, he does not ask "is this person good." He asks the model to research and summarize the public repositories and their purpose; the technologies the person works with most frequently; evidence of open source contributions or mentoring; any indication of their current role or company if it is public; and geographic location if it is public. Two hard constraints close the prompt: use only publicly available information, and do not speculate about private details. He caps the answer at 200 words.
That prompt orients him in under a minute on the questions that decide whether to reach out at all. Does this person really do data engineering, or is the profile misleading? Are they active in open source? Do they seem interested in learning new tools? Those answers inform whether and how he makes contact. Note the tension the location field creates and resolve it deliberately: knowing a publicly stated location is logistical information, while inferring anything about fit or motivation from that location is the bias the layer four constraint exists to prevent. The prompt may collect a fact; the screening standard still may not use it as a proxy.
Comparative Prompts: Screening Multiple Candidates
Comparative prompts are the other high-leverage move, because AI can surface patterns across resumes that are hard to hold in your head one at a time. When Marcus has three finalists, he attaches all three and asks the model to compare them on years of relevant experience; depth of technical skills, extracting the specific tools and frameworks for each; evidence of leading projects or teams; and communication skills and clarity in describing impact. He asks for a comparison table. At the end he asks three closing questions: which candidate seems most ready for a senior role, which would benefit most from mentorship, and which has the most diverse skill set.
He is not asking the model to rank people, and the distinction is worth guarding. He is asking it to lay the evidence side by side so that his own ranking is grounded in the same dimensions for every candidate, which is itself a consistency safeguard against the order effects and halo bias that creep into sequential resume review. Comparative prompts also help you calibrate. You see not just who is qualified, but who is qualified relative to the others, and quite often you discover that the criterion you thought was decisive is one all three candidates meet, which tells you your real differentiator is somewhere else entirely.
Anti-Patterns
Screening without context. This is asking the model to screen a resume without explaining what the role actually requires, and its purest form is the two-word prompt: "Screen this resume." It fails because the model has no idea what good looks like, so you get a generic summary that may highlight irrelevant skills and miss what is actually important. Avoid it by always beginning your screening prompt with a detailed description of the role, the team, and what success looks like in the first 90 days. If you cannot write that description, the problem is not the prompt; it is that the role is not yet defined well enough to screen for.
Asking the model to make the hiring decision. This is ending your screening prompt with "should we interview this person" or, worse, "based on all the above, do you think we should interview this candidate, yes or no." It fails because you are asking a system to make a human decision. It should support your decision-making, not replace it; you own the hiring decision. Avoid it by asking the model to organize information instead: what is this candidate's fit assessment, or what questions would you recommend asking. Then you decide whether to interview, and the reasoning behind that decision stays visible and yours.
Biased red flag identification. This is red flag language that is code for discrimination, and it is the most dangerous of the three because it usually sounds reasonable to the person writing it. An example: "Are there any warning signs that this person might not be committed? Look for signs they might be looking for work-life balance or that they might leave for personal reasons." It fails because "looking for work-life balance" and "might leave for personal reasons" encode assumptions about parenthood, caregiving, health, and other protected characteristics. This is discriminatory. Avoid it by focusing on role-specific concerns: gaps in relevant skills, misalignment with the tech stack, or short tenures in similar roles that suggest the role was not a fit. Always require specificity and role-relevance.
Practice
- Build a role-specific screening prompt. Write one for a role you are currently hiring for, including all five layers: role specification, qualification extraction, red flag identification that is role-specific rather than discriminatory, cultural fit signals, and decision support.
- Run a comparative analysis. Take two contrasting resumes, one strong fit and one weak fit for a role you know well. Write a prompt that compares them and identifies what makes one stronger than the other.
- Write a sourcing research prompt. You have found a candidate on LinkedIn. Write a prompt that tells the model to research and summarize their actual experience based on publicly available information, the technologies they are most skilled with, and any indication they are open to new opportunities. Decide first what you would want to know before reaching out.
- Identify bias in a screening prompt. Review this one: "Screen this resume for a marketing manager role. Look for red flags like job hopping, gaps in employment, or anything that suggests they might not be serious about a long-term commitment. Any sign they might want flexible hours is a concern." What is biased about it? Rewrite it to be fair and role-specific.
- Test and iterate. Write a screening prompt, use it on a real resume, review the output, and identify what would make it more useful for your decision-making. What questions did the output fail to answer?
Reflection
- What do you currently spend the most time on when screening resumes? What information are you looking for that you never find, because you are not specifically asking for it?
- What red flags have you historically used in screening? Are they role-specific and job-focused, or do they encode assumptions about identity or non-traditional career paths? How might you reframe them?
- Think about a recent hire who worked out really well. What did their resume show you that made you want to interview them, and what signals were you looking for?
- When comparing candidates, what framework do you use to decide who is stronger? How would that framework look written out as a structured prompt?
Glossary
- Red flag. A pattern or piece of information in a candidate's background that might indicate they are not a fit for the specific role. Role-specific red flags focus on actual job fit, not discrimination.
- Sourcing. The process of finding candidates through research, networking, or outreach. Effective sourcing involves understanding what candidates are actually looking for and what skills they have.
- Qualification extraction. Identifying and listing the specific skills, experience, and achievements from a candidate's materials. AI can organize qualifications in a structured way that supports comparison.
- Cultural fit signal. Evidence from a candidate's background suggesting they would thrive in your specific team and company culture. Should be based on observable behaviors and outcomes, not assumptions about identity or background.
- Comparative analysis. Reviewing multiple resumes side by side to understand relative fit and to calibrate what you are actually looking for in candidates.
Related Lessons
- Prompt Anatomy: Structure, Context, and Constraints is the general form of the five-layer pattern used here, and worth reading first if the layer structure feels arbitrary.
- Iterating with AI: Following Up, Clarifying, and Refining is the natural next step: what to do with a screening output that is close but not right, and how to refine a prompt across attempts.
- Avoiding Bias in Prompts: Language, Examples, and Assumptions goes deeper on the constraint language in layers three and four, including the phrasings that quietly encode assumptions.
- Flagging Red Flags and Concerns Without Bias expands layer three into a full treatment of role-relevant concerns and how to convert each into a question.
- Extracting Key Qualifications and Fit Indicators covers layer two in more depth, including how to structure extraction for roles where the qualifications are hard to state.
- Hands-On Practice: Build Your Prompt Library is where the prompts you write here get organized into something reusable across requisitions.
Closing
Resume screening is one of the highest-leverage places to use AI in recruiting. Every hour you save on screening is an hour you can spend on relationship-building, interviewing, and closing great candidates. But the leverage only works if your prompts actually extract what you need to know, which means strong prompts turn resume review from a chore into structured intelligence gathering rather than into a faster version of the same uncertainty.
Take a role you are currently hiring for and build a screening prompt using the five-layer framework. Test it on three to five resumes. Notice what information you are actually using to make screening decisions and what information the prompt is providing, and iterate on the gap between the two. Over time you will develop screening prompts that feel like they were written specifically for your needs, because they were. Marcus's data engineer prompt is now the base template for every engineering requisition he runs, and he rewrites only the role specification layer each time.
Key Takeaways
- Evaluate against a stated standard, do not summarize. Screening prompts need context about the specific role, the team, and what success looks like. Never ask a model to screen without explaining what good means in your situation.
- Build the prompt in five layers. Role specification, qualification extraction, red flag identification, cultural fit signals, and decision support, in that order, because each layer depends on the one before it. The result is organized information that supports faster, better decisions.
- Mark inferred versus stated. When the model labels which qualifications it reconstructed rather than read directly, you know exactly what to verify instead of trusting the reconstruction.
- Turn red flags into questions. Red flags should be role-specific and based on actual job-fit concerns, never coded discrimination. Require the model to explain each concern, offer multiple explanations, and suggest a follow-up question rather than treating it as a filter. This is also what keeps you clear of EEOC disparate-impact and ADA concerns.
- Require citations for fit signals, and constrain what may be used. A model forced to point at a specific resume line either finds evidence or admits it cannot. Tell it explicitly what not to do: do not infer motivation from demographics, do not assume anything about identity, and do not penalize educational choices, geographic location, or non-traditional career paths.
- Ask for a fit assessment, never the decision. "What is this candidate's fit assessment" prepares you to decide. "Should we hire this person" outsources a judgment you own and are accountable for.
- Comparative prompts help you calibrate. Screening several candidates together shows you relative fit and often reveals that your real differentiator is not the criterion you assumed.
- Read the borderline cases and run a fairness check. Spot-check the labels the model assigns, read every strong and possible fit yourself, and use the four-fifths rule to catch selection-rate disparities before they compound.
Frequently Asked Questions
Is it a problem that the model never sees the whole applicant pool at once? For screening, no, and it is arguably an advantage: assessing each resume against a fixed, stated standard is more consistent than assessing it against whoever happened to apply the same week. Comparative prompts are where you deliberately reintroduce the pool view, and you use them at the finalist stage where relative judgment is the actual question. Keep the two uses separate, because a per-candidate standard that quietly drifts toward "better than the last one I read" is how order effects get back in.
What do I do with the "unable to assess" bucket? Read those resumes yourself, and treat the size of the bucket as feedback on your prompt rather than on the candidates. Marcus had seven, one of which turned out to be a strong candidate whose work lived behind a portfolio link the model could not open. A resume that is thin on the dimensions you asked about is not the same as a candidate who is weak on them, and collapsing the two is exactly the kind of quiet exclusion that produces disparate outcomes without anyone deciding to exclude anyone.
Can I let the AI auto-reject the weak-fit group to save time? No. The whole design of layer five is to produce an assessment you act on, not a decision the tool executes. You own the hiring decision, and the labels are only as good as the prompt that produced them, which is why Marcus spot-checked the weak-fit group and found a mislabel. Beyond the accuracy problem, an automated rejection removes the human review that both good practice and, in a growing number of jurisdictions, the law expect from an automated selection procedure.
How do I know my screening prompt is not itself biased? Read it back looking for any criterion that is not job-related, since that is where bias hides in plain sight, and pay particular attention to red flag language and to anything that mentions background, location, education, or career trajectory. Then check the outputs rather than trusting the reading: apply the four-fifths rule to the selection rates your labels produce, and if one group's rate falls below 80 percent of the highest group's, audit the prompt and the applicant pool before you continue screening.
Skill.re