←
AI for Recruiters
Proficient · M18 · lesson 18 of 32 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Hands-On Project: Design a Communication Campaign with AI Support

15 min

Nadia sources engineers for a 90-person climate-tech startup in Seattle, and her outreach used to be a single cold message blasted to fifty passive candidates a week. The reply rate hovered around 9 percent, most of those were polite declines, and she had no idea which messages worked because she sent each one slightly differently and tracked nothing. When her hiring manager asked her to fill three backend roles in a quarter, she knew one-shot outreach would not get there. This project rebuilds her outreach into a designed campaign: a multi-touch sequence, AI-personalized templates, an A/B test on the part that matters most, and a human review on every message before it leaves her outbox.

What You Are Building

This project ends in a working campaign, not a plan for one. The deliverable is five artifacts you can hand to a colleague and have them run without asking you a question: a sequence specification saying how many touches go out on which days and what job each one does, a template set with the variable slots marked, a written review checklist the human applies before any message sends, one A/B test defined on a single variable with a stopping rule, and a measurement sheet naming the small set of numbers you will actually look at. Build those for one role and you have a campaign. Build them once and the second role reuses most of it.

Notice that four of the five artifacts are constraints rather than content. The templates are the only part that produces words. The sequence spec constrains when, the review checklist constrains what may leave the outbox, the test definition constrains what you are allowed to conclude, and the measurement sheet constrains what counts as working. That ratio is deliberate. AI makes the words cheap, and once words are cheap the thing that determines whether a campaign helps or damages you is the set of rules governing what happens to them. Nadia's old process had unlimited words and no rules, which is precisely why she could not say what worked.

Step One: Design a Campaign, Not a Message

The first shift is mental. A passive candidate who is happily employed rarely replies to a first message, no matter how good it is. Reply rates climb when outreach is a planned sequence of touches spaced over time, each adding a reason to engage rather than simply repeating the ask. Nadia designed a four-touch sequence over fourteen days for her backend role, and the discipline that makes it a sequence rather than four copies of the same nudge is that every touch has a distinct job.

TouchDayJob it doesThe ask
One0Short personalized opener naming a specific reason she reached outA reply
Two4Follow-up adding one concrete detail about the team and the problem they are solvingA reply
Three9Brief value-add: something genuinely relevant to the candidate's own workNone
Four14Polite close that leaves the door open and makes clear this is the last messageNone

The cadence is deliberate. Four touches over fourteen days is persistent without tipping into harassment, and the day-fourteen close is as important as the opener, because a clean exit protects the candidate experience and the employer brand for everyone who does not reply this quarter.

Touch three is the one most people cut, and it is the one carrying the design. It asks for nothing. Its only job is to send something the candidate would find useful whether or not they ever talk to Nadia, which is what changes the character of the sequence from pursuit into a series of contacts worth receiving. If you cannot think of anything genuinely relevant to send a particular candidate at touch three, that is a useful signal about how well you actually understand their work, and the honest response is to improve the research rather than to fill the slot with another nudge dressed as a resource.

Step Two: Build Templates That Do Not Read Like Templates

Nadia built each of the four touches as a template with bracketed variables for the elements that change per candidate: the specific project or repository that prompted the outreach, the technology overlap with her open role, and a genuine, non-generic hook. AI does the personalization work at scale. For each candidate, she feeds the model the public profile details she is permitted to use and her approved template, and it drafts a version that weaves the specific hook into the fixed structure. What took her six minutes per message by hand now takes about ninety seconds.

The structure of the template is what keeps the model useful and bounded. The fixed text carries everything that must be true for every candidate: what the company does, what the role is, the opt-out, the professional register. The bracketed slots carry only what is genuinely candidate-specific. That split means the model is not writing a message from scratch each time and inventing new claims about the company along the way; it is filling defined openings inside text a human already approved. It also makes review fast, because the reviewer knows exactly which parts could have changed and can read those first.

The danger of AI personalization is the uncanny middle: messages that are personalized enough to feel automated but not enough to feel human, the "I see you worked at [COMPANY] and that's so impressive" register that every candidate now recognizes and ignores. What makes that register fail is not that it is generic but that it is a compliment with no content behind it, which reads as a merge field wearing a smile. The fix is specificity that could only apply to this person: not that they worked somewhere, but what they built there and why it connects to the problem Nadia is hiring against. A hook a hundred candidates could receive is not a hook.

Step Three: Put a Human Review Between Draft and Send

Nadia's guardrail is the rule that anchors this entire project: a human reviews every AI-drafted message before it sends. Not a sample. Every one. The review is fast, fifteen to twenty seconds per message, and it checks three things: is the personalization specific and true, does the tone sound like Nadia rather than a bot, and is there anything off, a hallucinated detail, an awkward phrase, an assumption the profile does not support. The AI drafts at volume; the human guarantees that nothing embarrassing or false reaches a real person.

Sampling is the tempting compromise and the one that does not work here, because of how the errors are distributed. A model that hallucinates a project the candidate never worked on does not do it at a steady rate you can catch by checking one message in ten; it does it on the candidates whose public information was thinnest, which is a specific subset rather than a random one. Reviewing a sample tells you the average message is fine and leaves the worst message in the batch unread. Since the cost of the review is fifteen to twenty seconds and the cost of one confidently wrong message sent under the company's name is a candidate telling their network about it, the arithmetic is not close.

Write the three checks down rather than holding them in your head, because a checklist that exists on paper survives a busy Thursday and an unwritten habit does not. Accuracy first, since a false detail is the only error that is unrecoverable once sent. Tone second, since a message that does not sound like the person whose name is on it undermines every message after it. Then the catch-all, anything off, which is where the reviewer's judgment lives and where most real saves happen, because the problems that matter are usually not on a list.

The Worked Example: A/B Testing the Subject Line

Nadia wanted to know what actually moved replies, so she ran a clean A/B test on the highest-leverage variable: the subject line of touch one. The subject line decides whether the message is opened at all, so it is the right place to test first. She split her week-one batch of one hundred passive backend candidates into two groups of fifty, identical in every way except the subject line.

Version A was a conventional role-forward subject: "Senior Backend Engineer role at a climate-tech startup." Version B was specific and curiosity-driven, referencing the candidate's own work: "Your work on distributed queues, and a problem we're chasing." Everything downstream, the body, the cadence, the touches, stayed identical, so any difference in replies could be attributed to the subject line alone. Version A returned 9 replies from 50 sends, an 18 percent reply rate. Version B returned roughly 14 replies from 50, a 27 percent reply rate. The specific, candidate-referencing subject line won by nine percentage points, which on her quarterly volume is the difference between scraping by and comfortably filling the funnel.

The two subject lines are worth reading side by side, because the difference is not that one is cleverer. Version A describes what Nadia wants. Version B describes something the candidate already cares about and then hints that there is a reason it is being mentioned. The candidate reading version B in a crowded inbox has to open it to find out what the connection is, and that is the entire mechanism. It also means version B cannot be written generically: it names distributed queues because this candidate works on distributed queues, so the subject line is itself a personalized field rather than a fixed string, which is what makes the AI drafting step and the test compatible rather than in tension.

Two disciplines made the test trustworthy. First, she changed only one variable; if she had also changed the body, she would not know which change drove the lift. Second, she let the test run to a reasonable sample before declaring a winner rather than reacting to the first three replies. Once Version B won, it became her default subject template, and she queued the next test on touch-two timing. The campaign improves because each cycle isolates one variable, measures it, and folds the winner back into the standard.

Be careful about what the result licenses you to say. The comparison the test supports is version B against version A, on this role, with this batch, at this moment, because those are the only two things held identical. It does not license a claim about the reply rate her old undesigned blast produced, which came from different candidates in a different context with nothing controlled. Treating an A/B result as evidence about anything outside the test is the fastest way to accumulate confident beliefs that are not true, and the discipline that makes the next test worth running is refusing to overclaim from this one.

Outreach at scale touches privacy law, and Nadia treats compliance as part of the design rather than an afterthought. Under GDPR for candidates in the EU and CCPA for candidates in California, the data she uses to personalize is personal data, and candidates have rights over it. Three rules govern her campaign. Every message includes a clear, working opt-out, and an opt-out is honored immediately and permanently across the whole sequence, so a candidate who declines at touch one never receives touches two through four. She uses only data she has a legitimate basis to use and stores it in the applicant tracking system with a defined retention period, not in a personal spreadsheet that lives forever. And she is ready to honor access or deletion requests, which means knowing where every candidate's data sits.

The opt-out rule is not just legal hygiene; it is brand protection. A candidate who asks to be removed and then gets three more automated touches will remember the company, and not kindly. Building the suppression logic into the sequence from the start, so a reply or an opt-out halts the remaining touches, is what lets AI-assisted volume coexist with respect for the individual.

The word to hold onto in that rule is automatic. An opt-out honored by Nadia remembering to remove someone from a list is an opt-out that fails on the week she is busiest, which is exactly the week the sequence is largest. Suppression has to be a property of the system: the reply or the opt-out lands, the remaining touches for that candidate stop, and nobody has to do anything. The same applies to the storage rule. A personal spreadsheet is not a policy violation because spreadsheets are bad; it is a violation because nobody can answer a deletion request against a file they do not know exists, and knowing where every candidate's data sits is a precondition for honoring the rights the law grants.

Step Six: Measure What the Campaign Is Doing

Nadia tracks a small set of numbers that tell her whether the campaign is healthy: open rate by subject-line variant, reply rate by touch, positive-reply rate as distinct from total replies, and opt-out rate. The opt-out rate is her early warning system. If it climbs, the messaging has drifted from personalized to spammy, and she pulls back regardless of what the reply rate says. She also watches the gap between total replies and positive replies, because a campaign that generates many "no thanks" responses is loud but not effective. The point of measurement is not to prove the campaign works; it is to learn which touch, which subject line, and which hook actually earns engagement, so the next iteration is better than this one.

Two of those four numbers exist specifically to stop the other two from misleading her, which is why the set has to be read together. Reply rate on its own rewards anything that provokes a response, including irritation. Positive-reply rate separates the responses that could become a hire from the ones that are a candidate closing a loop. Opt-out rate is the count of people the campaign actively cost her, and it is the one metric that can veto a good-looking result: a rising opt-out rate alongside a rising reply rate means the campaign is getting louder rather than better, and Nadia pulls back on the strength of the opt-out number alone. Reply rate by touch, meanwhile, tells her which part of the sequence is doing the work and therefore where the next test belongs.

One discipline keeps the whole system honest: the human review never gets automated away as volume grows. It is tempting, once the templates are good and the A/B winner is set, to let the AI send directly. Nadia refuses, because the review is the only thing standing between scaled personalization and a hallucinated detail going out under the company's name. The campaign scales the drafting. It does not scale away the judgment.

Anti-Patterns

Sending four copies of the same nudge and calling it a sequence. This is spacing the original cold message across fourteen days with the wording changed. It happens because building four touches with distinct jobs takes real thought, and touch three, the one that asks for nothing, looks like the easiest to convert into another ask. What goes wrong is that repetition without new information reads as pressure, so the candidate who ignored touch one has a stronger reason to ignore touch four, and the opt-out rate climbs while the reply rate does not. The counter is Nadia's structure, where every touch has a stated job: open, add a concrete detail about the team and the problem, give something useful with no ask, then close cleanly.

Personalizing into the uncanny middle. This is the "I see you worked at [COMPANY] and that's so impressive" register, personalized enough to feel automated and not enough to feel human. It happens because the model will happily produce a compliment from any profile, and a compliment reads like personalization while costing nothing to generate. What goes wrong is that candidates recognize the pattern immediately and it lands worse than an honest generic message, because it signals that the sender wanted the appearance of research without doing any. The counter is a hook that could only apply to this person, naming what they built and why it connects to the role, and treating an empty touch-three slot as a research problem rather than a slot to fill.

Reviewing a sample instead of every message. This is spot-checking one message in ten once the templates look reliable. It happens because sampling is the standard quality-control move and because the review feels redundant when the last forty messages were fine. What goes wrong is that model errors are not evenly distributed: hallucinated details cluster on the candidates whose public information was thinnest, so the sample tells you the average message is fine while leaving the worst one in the batch unread. The counter is the fifteen-to-twenty-second review on every draft, against a written checklist, which is cheap next to one confidently false message going out under the company's name.

Testing two things at once, or calling the winner early. This is changing the subject line and the body together, or declaring version B the winner after three replies. It happens because there are always several improvements you want to make and because an early lead looks like a result. What goes wrong is that a two-variable test cannot attribute the lift to either change, so the winner becomes your default for reasons you cannot identify, and a call made on a handful of replies is a call made on noise that will not repeat. The counter is Nadia's pair of disciplines: change exactly one variable with everything downstream held identical, and let the test reach a reasonable sample before folding the winner into the standard.

Honoring opt-outs by hand. This is treating removal as something the recruiter remembers to do rather than something the sequence does. It happens because at small volume manual removal works, and the failure only appears once the campaign is large enough to matter. What goes wrong is that a candidate who asked to be removed receives touches two through four during the busiest week of the quarter, which is a legal exposure under GDPR and CCPA and, just as damaging, a story that candidate tells. The counter is suppression logic built in from the start, so a reply or an opt-out automatically halts the remaining touches, with candidate data held in the ATS under a defined retention period rather than a personal spreadsheet nobody can search when a deletion request arrives.

Reading reply rate as the scoreboard. This is judging the campaign on total replies and reporting the number upward. It happens because reply rate is the easiest metric to collect and the one that sounds like success. What goes wrong is that reply rate rewards anything provoking a response, including irritation, so a campaign drifting toward spam can show a rising reply rate right up until the damage is visible elsewhere. The counter is Nadia's set read together: positive-reply rate separated from total replies, reply rate broken out by touch so you know which step is working, and opt-out rate treated as an early warning with power to veto a good-looking result.

Build Checklist

Work these in order against one live role. Each step produces one of the five artifacts.

  • Write the sequence specification. Fix the number of touches, the day each goes out, and the job each one does in a sentence. Include at least one touch that asks for nothing, and a final touch that closes cleanly and says so.
  • Build the templates with the variables marked. Write the fixed text carrying everything true for every candidate, including the opt-out, then bracket only the genuinely candidate-specific slots: the project or repository, the technology overlap, the hook.
  • Draft one message with AI and rewrite the hook by hand. Feed the model your approved template and the profile details you are permitted to use, then check whether the hook it produced could be sent to anyone else. If it could, it is not a hook.
  • Write the review checklist down. Accuracy, tone, and anything off, in that order, in a form a colleague could apply. Time yourself on ten messages so you know what the review actually costs before you decide whether you can afford it.
  • Define one A/B test with a stopping rule. Name the single variable, confirm everything downstream is held identical, and write in advance how many sends you will wait for before reading the result. Start with the subject line of touch one.
  • Wire the suppression logic before the first send. Confirm that a reply or an opt-out automatically halts the remaining touches, and that candidate data sits in the ATS with a retention period so you can answer an access or deletion request.
  • Set up the measurement sheet. Open rate by variant, reply rate by touch, positive-reply rate separate from total replies, and opt-out rate. Decide now what opt-out number would make you pull back regardless of replies.

Reflection

  • How many touches does your current outreach have, and can you state a distinct job for each one?
  • Read your last outreach message. Which sentence could only have been sent to that person?
  • What proportion of your AI-drafted messages does a human actually read before they send?
  • If a candidate replied "please remove me" today, what would have to happen for touches two through four to stop?
  • What is the last thing you changed about your outreach, and how did you decide it was an improvement?
  • Do you know your positive-reply rate as distinct from your total reply rate, and your opt-out rate at all?

Glossary

  • Touch. One message in a sequence, defined by the day it sends and the job it does, as distinct from a repeat of the previous ask.
  • Cadence. The spacing of touches over time, set to be persistent without tipping into harassment; Nadia's is four touches over fourteen days.
  • Value-add touch. A message that asks for nothing and sends something genuinely relevant to the candidate's own work, which changes the sequence from pursuit into contacts worth receiving.
  • Closing touch. The final message, which leaves the door open and states plainly that it is the last one, protecting the candidate experience for everyone who does not reply.
  • Bracketed variable. A marked slot in a template holding only genuinely candidate-specific content, which bounds what the model may change and makes the review fast.
  • Hook. The specific, true reason this candidate is being contacted. A hook a hundred candidates could receive is not a hook.
  • Uncanny middle. Personalization sufficient to feel automated but not to feel human, the "I see you worked at [COMPANY]" register, which lands worse than an honest generic message.
  • Human review gate. The fifteen-to-twenty-second check on every AI-drafted message before send, covering accuracy, tone, and anything off, applied to every message rather than a sample.
  • A/B test. A comparison of two versions differing in exactly one variable, with everything downstream held identical so the difference in outcome can be attributed to that variable.
  • Single-variable discipline. Changing one thing per test, which is what makes the result attributable and the winner worth adopting as the new default.
  • Stopping rule. The sample size decided in advance at which a test will be read, which prevents a winner being called on the first few replies.
  • Suppression logic. The automatic halting of remaining touches when a candidate replies or opts out, built into the sequence rather than depending on the recruiter remembering.
  • Legitimate basis. The requirement that personalization data be information you are permitted to use, which under GDPR and CCPA makes candidate data personal data with rights attached.
  • Retention period. The defined lifespan of stored candidate data, held in the ATS rather than a personal spreadsheet so access and deletion requests can actually be honored.
  • Positive-reply rate. Replies that could become a hire, tracked separately from total replies because a campaign generating many declines is loud but not effective.
  • Opt-out rate. The early warning number, which signals drift from personalized toward spammy and can veto a good-looking reply rate on its own.

Closing

Nadia's old outreach was not lazy. She wrote every message herself and sent fifty a week, which is more effort than most sourcing gets. What it lacked was structure: no second touch, no way to tell one message from another after the fact, and no number that would have told her whether a change helped. Effort without structure produces exactly her result, a reply rate she cannot explain and cannot improve.

The campaign that replaces it is five short documents and one rule. Four touches with distinct jobs over fourteen days. Templates whose fixed text is approved and whose brackets are genuinely specific. A human reading every draft for fifteen seconds before it sends. One variable tested at a time against a stopping rule written in advance. Suppression and retention wired in rather than remembered. And a measurement sheet where the opt-out rate can overrule a good-looking reply rate. The AI makes personalization affordable at fifty candidates a week. Everything else on that list is what makes it something a candidate is glad to receive.

Key Takeaways

  • Design a sequence, not a single message. Nadia's four touches over fourteen days each do a distinct job: open, add a concrete detail about the team and the problem, give something useful with no ask, then close cleanly so everyone who did not reply is left on good terms.
  • Use bracketed templates so the model fills slots rather than writing from scratch. Fixed approved text carries what must be true for every candidate; the brackets carry the project, the technology overlap, and the hook. Drafting drops from six minutes per message to about ninety seconds.
  • Avoid the uncanny middle. A compliment with no content behind it reads as a merge field wearing a smile. Specificity that could only apply to this person is what separates a hook from the "I see you worked at [COMPANY]" register.
  • Review every AI-drafted message, not a sample. Errors cluster on the candidates with the thinnest public information, so sampling leaves the worst message unread. Fifteen to twenty seconds against a written checklist of accuracy, tone, and anything off.
  • Test the highest-leverage variable first, one at a time. Nadia tested only the touch-one subject line across two groups of fifty. A specific, candidate-referencing line drew a 27 percent reply rate against 18 percent for the conventional one, a nine-point lift, and the result is attributable only because everything downstream was identical.
  • Do not overclaim from a test. The comparison holds for version B against version A on this role and this batch. Reading it as evidence about anything outside the controlled comparison is how confident false beliefs accumulate.
  • Build consent and opt-out into the sequence, automatically. Under GDPR and CCPA, candidate data is personal data. Every message carries a working opt-out, suppression halts the remaining touches without anyone remembering, and data lives in the ATS with a retention period rather than a forever spreadsheet.
  • Measure to learn, and never automate away the review. Reply rate by touch, positive-reply rate, and opt-out rate read together, with opt-out able to veto a good-looking result. Scale the drafting, not the judgment that keeps a false detail from going out under the company's name.

Frequently Asked Questions

Is a fifteen-second review on every message realistic at volume? It is the constraint the campaign is designed around, which is why the templates are structured the way they are. The reviewer is not reading a fresh message; they are reading approved fixed text with a small number of bracketed slots, and they know exactly which parts could have changed. Time yourself on ten messages before deciding you cannot afford it, and weigh the result against the alternative, since one confidently false detail sent under the company's name costs more than the review does across a whole batch. If the honest answer is still that you cannot review everything, that is a signal to send fewer messages rather than to review fewer.

Does a four-touch sequence risk annoying candidates? The cadence is set to make that unlikely and the metrics are set to catch it if it happens. Four touches over fourteen days is persistent without tipping into harassment, touch three asks for nothing, and touch four closes cleanly and says it is the last message. Beyond design, the opt-out rate is the check: if it climbs, the messaging has drifted from personalized toward spammy, and Nadia pulls back regardless of what the reply rate says. The version worth worrying about is not four touches but four repetitions of the same ask, which is a different thing wearing the same shape.

What should I test after the subject line? Whatever the reply-rate-by-touch numbers point at, which is why that metric is broken out by touch rather than reported as a single figure. Nadia's next test was touch-two timing, because the subject line had been settled and the second touch was where her sequence looked weakest. The rule that matters is not which variable comes next but that you test one at a time with everything else held identical and a stopping rule written in advance, then fold the winner into the standard before starting the following test. A campaign improves by accumulating attributable results, and a test you cannot attribute contributes nothing to that.

Can we let the AI send directly once the templates are proven? That is the specific temptation Nadia refuses, and the reason is that "proven" describes the template rather than the message. The fixed text has been reviewed and is stable. What has not been reviewed is the personalization the model produced for this particular candidate from this particular profile, which is exactly where a hallucinated project or an assumption the profile does not support enters. The campaign scales the drafting from six minutes to ninety seconds and leaves the judgment in place, and the fifteen-to-twenty-second review is the entire price of that arrangement.