Avoiding Bias in Prompts: Language, Examples, and Assumptions
Leah is a technical recruiter at a 90-person fintech startup, filling about eight engineering roles a quarter with AI doing her first-pass resume summaries. She thought of herself as fair-minded until she ran an experiment. She fed the model her standard screening prompt, the one that asked it to flag "a rockstar engineer, passionate, aggressive in problem-solving, willing to work long hours." Then she ran the same fifteen resumes through a neutral rewrite. The two passes surfaced noticeably different candidates. Her prompt, not the resumes, had been doing the sorting. Leah's realization is the heart of this lesson: when you ask an AI to look for something, you are not asking a neutral question. You are asking through a lens shaped by your own assumptions, and the model faithfully amplifies whatever bias that lens contains.
The Four Sources of Bias in Prompts
Bias enters prompts from four places. It is worth naming them, because once you can name a source you can audit for it. The four are gender-coded language, demographic assumptions, cultural shortcuts, and loaded language. Most biased prompts contain more than one, and Leah's opening example contained all four.
None of this is usually intentional, which is exactly why it survives review. Consider three phrases that appear in almost every recruiter's vocabulary. "Look for someone who is a good cultural fit" very often encodes "someone like us." "We need a real team player" very often encodes gender expectations. "Passionate about tech" very often encodes assumptions about a candidate's background and education, about who had the free time and the family context to treat technology as a hobby before it was a job. Each phrase feels neutral to the person typing it. Each one hands the model a target it will then hit with great precision. When you pass a biased prompt to an AI, the AI does not soften it. It amplifies it, consistently, across every candidate in the batch.
Source One: Gender-Coded Language
Certain words carry gender associations whether or not you intend them. Male-coded terms include "aggressive," "competitive," "dominant," "take charge," and the startup vocabulary of "rockstar," "ninja," and "guy," along with sports metaphors like "knock it out of the park." Female-coded terms include "nurturing," "supportive," "collaborative," "team player," "people-focused," and "detail-oriented." Neither set is wrong, and both describe real and valuable qualities. The problem is asymmetry. If Leah's prompts only ever use male-coded language to describe the ideal engineer, the model learns to surface assertive styles and quietly downweight everything else, encoding a preference she never consciously chose. The fix is gender-neutral phrasing: "excellent engineer" instead of "rockstar," "works effectively across teams" instead of "team player," "results-focused" instead of "aggressive sales approach."
The lists run longer than the obvious offenders, and the less obvious entries are the ones that survive a careful edit. On the male-coded side, add "ambitious" and "warrior" to the vocabulary above, and add the whole family of sports metaphors, not just the famous ones: "go the distance" belongs on the list alongside "knock it out of the park." On the female-coded side, add "helpful," add "communication skills" as a standalone requirement, and add any reference to caregiving or relationship-building as a proxy for how a person works. The mechanism is worth stating plainly, because it explains why this matters more with AI than with a human reader. When you write male-coded language into a prompt, the model becomes more likely to flag candidates with assertive styles as strong. When you write female-coded language, it emphasizes relational skills instead. Both qualities matter in real jobs. The distortion comes from using only one register to describe your ideal candidate, because then the model optimizes for that register alone.
The effect compounds when the AI generates text, not just when it screens. When Leah asked the model to draft a job description from her bullet points, it picked up on the tone of her prompt and ran with it. A prompt seeded with "rockstar" and "crush it" produced a posting full of competitive, high-octane phrasing, and the kind of language research on job ads associates with smaller applicant pools among women. The same bullet points fed through a prompt that said "write in a clear, welcoming, professional tone" produced a posting that read as open to a wider range of people without losing any technical specificity. The lesson is that the model mirrors your register. Watch for clusters, too: one male-coded word is easy to miss, but "aggressive, driven, dominant, relentless" stacked together creates a strong signal. When you find one coded word, scan the whole prompt for its friends.
Source Two: Demographic Assumptions
Prompts often smuggle in assumptions about age, nationality, education, or life stage. "Recent graduate" assumes age and signals a preference for early-career talent, which can run afoul of age-discrimination protections. "Native speaker" assumes nationality and excludes fluent bilingual candidates who speak with an accent. "Someone from a top university" assumes education and correlates with socioeconomic status. A stray gendered pronoun, "early in his career," signals a male assumption to the model. Each of these pre-filters for a group before the AI has even read the candidate. Focus instead on what drives performance: "early-career engineer comfortable learning on the job," "fluent English speaker," "evidence of strong technical fundamentals," which can come from a bootcamp, self-teaching, or a state school as easily as an Ivy.
The legal stakes here are real, not abstract. In the United States, the Age Discrimination in Employment Act protects workers who are 40 and older, so a prompt that filters for "recent grad" or "digital native" can produce a candidate pool that exposes the employer to an age-discrimination claim. The EEOC treats hiring tools, including AI-assisted ones, as subject to the same disparate-impact scrutiny as any other selection procedure. The four-fifths rule offers a rough screen for that impact: if the selection rate for one group falls below four-fifths of the rate for the highest-selected group, that is a signal worth investigating. Leah cannot run a formal adverse-impact study on every prompt, but she can do the cheap version: ask whether a phrase describes the job or describes a person, and strip anything in the second category. "Comfortable with on-call rotation" describes the job. "Young and high-energy" describes a person, and a protected characteristic at that. The first survives; the second gets cut.
Source Three: Cultural Shortcuts
Culture fit is real, but recruiters describe it with shortcuts that exclude people who differ from them. "We're a young, energetic startup" often means "we're all under 35." "We grab drinks after work" assumes alcohol and after-hours availability, disadvantaging caregivers and people with different lifestyles. "We're all passionate about the mission" assumes one particular expression of commitment. When you code culture as drinks or as visible passion, you screen out people who engage differently. Describe the substance instead: "a team that moves quickly and enjoys rapid iteration," "strong communication and comfort with direct feedback." Let candidates match the substance, not the style.
Source Four: Loaded Language
Some recruiting language carries built-in assumptions. "Willing to work long hours" often means expecting unpaid overwork and disadvantages caregivers. "Self-starter" can mean "someone who needs no structure," penalizing people who thrive with clear direction. "Strong communicator" sometimes means "extroverted." "Executive presence" correlates with height, gender presentation, and confidence in ways that are simply biased. Replace loaded terms with specific, observable behaviors: "able to set priorities independently," "able to explain technical concepts clearly and ask clarifying questions." Specificity is the antidote.
"Culture fit" itself belongs on this list as much as on the previous one, because used without specificity it simply means "like us." That is the whole content of the phrase. If you cannot say what the fit consists of, in terms a candidate could read and assess themselves against, the phrase is not describing your culture. It is describing your resemblance to the people already there.
Worked Example: Rewriting Leah's Prompt
Here is Leah's original engineering prompt: "We're looking for a rockstar engineer who's passionate about building world-class software. We need someone aggressive in problem-solving and willing to work long hours to ship great products." Four problems: "rockstar" and "passionate" are vague and male-coded, "aggressive" is male-coded, "willing to work long hours" assumes unpaid overwork, and "world-class" is empty. Her rewrite: "We're looking for an engineer with strong problem-solving skills who can debug complex systems, work effectively with a team, and ship features that solve real customer problems. You should be comfortable with rapid iteration and learning new tools. We work in focused sprints and value quality over hours."
When Leah re-ran her fifteen-resume test with the rewritten prompt, the candidate set shifted. Across her last quarter of eight hires, she began applying a simple discipline: every screening prompt gets the gender-flip test before use. The measurable change was in her interview slate, which went from 2 of 10 women in the prior quarter to 4 of 10 once the prompts stopped quietly coding for assertiveness. These are her own pipeline figures, illustrative of the effect rather than a published statistic, but the mechanism is the lesson: the prompt was the filter, and fixing the prompt widened the funnel without lowering the bar.
Second Worked Example: A Culture-Fit Prompt
Screening prompts are not the only place bias hides. Leah also used a prompt to score how well candidates would "fit the team," and that one was worse than her engineering prompt because it dressed up similarity as a virtue. The original read: "Rate how well this candidate would fit our culture. We're a young, scrappy team that grabs drinks after work, moves fast, and is obsessed with the mission. Look for someone who'll vibe with us." Almost every clause is a problem. "Young" is an age proxy. "Grabs drinks" assumes alcohol and after-hours availability and disadvantages caregivers. "Vibe with us" is an explicit instruction to reward people who resemble the existing team, which is the literal definition of affinity bias automated at scale. "Obsessed with the mission" is the vagueness problem again: it sounds like a standard but sets none.
Her rewrite reframed fit as alignment with how the team actually works rather than who the team happens to be: "Assess this candidate against how our team operates. We ship in short iterations and adjust based on customer feedback, so look for comfort with changing priorities and shipping incrementally. We give and receive direct feedback, so look for evidence of clear communication and openness to critique. We care about the product domain, so look for genuine interest in the problem space, expressed in any form. Do not infer fit from age, hobbies, or social style." The rewritten prompt evaluates behaviors a candidate can demonstrate and a recruiter can defend, and the closing instruction tells the model what not to use, which matters because models will otherwise reach for whatever correlates with the vague target.
The same repair works on the candidate-facing version of that text, the paragraph in a job posting where most teams write "we need someone who gets our vibe." Describe the actual dynamic instead: we move quickly, we give each other direct feedback, we celebrate wins together, and we support each other when things get hard. Then state the expectation about commitment honestly rather than in code: we expect you to be engaged and committed to the mission, and that shows up as the quality of your work rather than the quantity of your hours. That last sentence does real work, because it replaces the unspoken "long hours" expectation with a standard a caregiver, a person with a disability, or anyone with a life outside the office can meet on equal terms.
Building a Bias-Checked Prompt Library
One good rewrite helps for one search. The leverage comes from turning rewrites into a reusable library so the team is not re-deriving fairness from scratch each time a role opens. Leah keeps a shared document of approved prompt blocks, each one already run through her two tests, organized by the task it does: resume screening, job-description drafting, outreach messages, and fit assessment. A new recruiter starts from an approved block instead of a blank box, which is where most biased phrasing creeps in.
A practical library has a few moving parts:
- Approved blocks. Vetted, reusable prompt text for each recurring task, written in neutral, behavior-based language. These are the default starting point.
- A banned-words list. A short, living list of terms that have failed review: "rockstar," "ninja," "native speaker," "recent grad," "young," "culture fit," "executive presence," "long hours." When one shows up in a new prompt, it is a prompt to rewrite, not necessarily a violation, but always a flag.
- Before-and-after pairs. A handful of worked rewrites like the two in this lesson, so the reasoning is visible and teachable rather than a rule handed down.
- A change log. A note of what was changed and why, so the library improves as the team learns and so an audit later can show the work.
The library also makes the whole effort auditable. If a candidate or a regulator ever asks how the company uses AI in hiring, "here are our vetted prompts, our banned terms, and our review notes" is a far stronger answer than "we trusted whoever was typing." Leah reviews the library once a quarter, retiring blocks that stopped matching how roles are actually written and adding new banned words as she catches them. The point is not to freeze a perfect set of words. It is to make the fair version the easy default and the biased version something you have to go out of your way to type.
Two Tests That Catch Most Bias
Two quick checks catch the majority of biased prompts. The gender-flip test: mentally swap the gender of the person you are describing. If the prompt suddenly reads strangely, it was gender-coded. The specificity test: can you state the same skill in neutral, observable terms? If not, the language is probably loaded. Leah runs both before any prompt enters her library.
Three Anti-Patterns
Assuming gender in prompts. "He should be someone who..." signals a gender preference even when unintentional and excludes people who do not identify that way. Use "they" or no pronoun. Conflating education with ability. "Only candidates from top universities" pre-filters for privilege, since prestige correlates with socioeconomic status, race, and nationality; define the skills and look for evidence of them from any source, including bootcamps, self-teaching, public universities, and apprenticeships. Vague culture-fit language. "Someone who fits our culture, which is about passion and commitment" filters for similarity, because people from different backgrounds express commitment in different ways; replace it with specific behaviors candidates can evaluate themselves against, like "values rapid feedback," "comfortable with ambiguity," and "takes ownership."
The Vocabulary, in One Place
Four terms and one test carry most of the weight in this lesson, and it helps to be able to name them when you are reviewing a colleague's prompt rather than your own.
- Gender-coded language is words and phrases that carry implicit masculine or feminine associations even though they are being used to describe a neutral job attribute. "Aggressive" is male-coded; "nurturing" is female-coded.
- Demographic assumptions are assumptions encoded in a prompt about a candidate's age, nationality, education, or life stage. They pre-filter for particular groups before any evaluation happens.
- Cultural shortcuts are vague phrases used to describe culture fit that actually encode an expectation of similarity. "Good vibe" and "our kind of person" are the pure forms.
- Loaded language is words carrying implicit assumptions about the preferred worker: "self-starter," "willing to work long hours," "executive presence."
- The specificity test is the check that ties them together: can you describe the same behavior or skill in neutral, specific, observable language? If you cannot, the language you have is probably biased.
Practice
Reading about coded language does very little. Catching it in your own writing does almost everything, so work through these in order, on prompts you actually use.
- Audit your language. Take three recruiting prompts you have written or used recently. Review each one against all four sources: gender-coded language, demographic assumptions, cultural shortcuts, and loaded language. Then identify which source shows up most often in your own writing, because that is the habit worth breaking first.
- Rewrite for neutrality. Choose the one prompt with the most obvious bias. Rewrite it to remove gendered language, demographic assumptions, and cultural shortcuts, while keeping it specific and honest about what you actually need. Vagueness is not neutrality.
- Run the gender-flip test. Take a recruiting prompt and mentally flip the gender of the person you are describing. Does it still sound right? If it reads differently, the prompt was gender-coded, and it needs a rewrite before it touches a candidate.
- Build a bias-free role description. Pick a role you are hiring for now and write a prompt describing the role, the team, and the culture without a single gender-coded word, demographic assumption, or shortcut. Be specific about what genuinely matters for the work.
- Test both versions on the AI. Write a prompt using coded or loaded language and run it on real content. Note the output. Then rewrite the prompt to be neutral and run it on the same content. Compare the two results side by side and name exactly what changed. This is the exercise that converts belief into evidence.
Reflection
Sit with these four questions before you move on. They are more useful written down than answered in your head.
- Where do you think your own prompts are most likely to carry bias: in how you describe culture, in your word choices, or in your assumptions about the ideal candidate?
- When you read other people's job descriptions and recruiting prompts, what language do you notice that might encode bias, and how would you rewrite it?
- Think about your last few hires who worked out really well. Did they actually match the culture-fit description you use in your prompts, or did they bring something different that also worked?
- How might writing more specific, less biased prompts improve your hiring outright, by surfacing people you would otherwise never have seen?
Putting It to Work This Week
Fair hiring depends on fair prompts. If you encode bias into a prompt, that bias flows through everything downstream of it: the outputs the AI hands you, how you read candidates, who you interview, who you hire. The reverse is equally true and more encouraging. When you are deliberate about writing fair, specific, neutral prompts, the fairness compounds through the entire process, because every stage after the prompt inherits it.
So make it concrete. This week, audit three prompts you actively use. Identify the bias in each. Rewrite one with the bias removed, test both versions on the same content, and notice the difference in what comes back. Then update your prompt library one block at a time, so the fairer, more specific version becomes the one your team reaches for by default.
Related Lessons
Fair prompts are one link in a chain, and the links on either side matter.
- Prompt Anatomy: Structure, Context, and Constraints covers the mechanics of a well-built prompt. Bias removal is a great deal easier when the prompt already has clear structure and explicit constraints to hang the fairness instructions on.
- Fairness Checks: Identifying Gender, Age, Disability, and Other Bias Signals is the natural next step, applying the same fairness lens to the output side: reviewing what the AI gives back for bias before it influences your decisions.
- Hands-On Practice: Build Your Prompt Library turns the library described above into a structured build, which is where the one-off rewrites in this lesson become a durable team asset.
- How AI Can Perpetuate or Amplify Bias explains the underlying mechanism, why a model amplifies rather than neutralizes the bias in your instructions, and why your wording carries so much weight.
- Gender, Age, Disability, Race, and Socioeconomic Bias in Recruiting gives the fuller picture of the bias categories this lesson touches, including the protected characteristics that carry legal consequence when a prompt filters on them.
Frequently Asked Questions
If I remove all the personality from my prompts, won't I just get generic candidates? No. Stripping coded language does not mean stripping standards. You are trading vague signals like "rockstar" for sharp ones like "can debug a production incident under time pressure." Specific, demanding, observable criteria raise the bar; they do not lower it. What you lose is the accidental filter for assertiveness or sameness, not your ability to ask for excellence.
Is it actually illegal to write a biased prompt? The prompt itself is not the violation; the hiring outcome is what gets scrutinized. But a prompt that filters on age, sex, national origin, or another protected characteristic can produce a discriminatory result, and the EEOC treats AI-assisted selection tools the same as any other. Frameworks like the four-fifths rule for adverse impact and the ADEA's protection of workers 40 and older apply regardless of whether a human or a model did the sorting. Writing fair prompts is how you stay on the right side of an outcome you would otherwise have to defend.
The AI gave me a great candidate list. Why second-guess the prompt? Because a list that looks great can still be quietly narrow. A biased prompt produces confident, plausible results; that is exactly why it is dangerous. The candidates it surfaces are often genuinely good, which makes it easy to conclude the prompt worked. Run the gender-flip test and check whether a different but equally valid framing would have surfaced a different set. If it would, the prompt was choosing, not just retrieving.
How do I check a prompt fast when I'm in a hurry? Run the two-question version: does any phrase describe a person rather than the job, and would the prompt read strangely if you flipped the candidate's gender? If both come back clean, you have caught most of the common problems. For anything that goes into repeated use, promote it to your vetted library so the check only has to happen once.
Key Takeaways
- Bias comes from four sources. Gender-coded language, demographic assumptions, cultural shortcuts, and loaded language; audit every prompt for all four.
- Gender-coded words steer AI output. Asymmetric use of male- or female-coded terms tilts which candidates surface; choose language that codes as neither.
- Drop demographic proxies. Screen on demonstrated skills, not university pedigree, age, nationality, or life stage.
- Describe culture by behavior. "We value direct feedback" beats "we're a close-knit team" because it lets candidates self-assess against substance.
- Replace loaded terms with observable behaviors. "Self-starter" and "executive presence" hide bias; "sets priorities independently" does the work without it.
- Run the gender-flip and specificity tests. If a prompt reads differently when you flip genders, or relies on vague descriptors, it needs a rewrite before it touches a candidate.
Skill.re