Hands-On Project: Design a Workflow for a Specific Recruiting Challenge
Elena is a recruiting manager at a regional logistics company, and every autumn the same wave arrives. The distribution center needs 60 warehouse associates hired in four weeks to handle peak shipping volume. Last year the requisition drew 540 applications, and Elena and one coordinator drowned in them. Resumes piled up, qualified candidates went cold while they waited a week for a first call, and the team ended up rejecting hundreds of people with a form email nobody had actually read. They hit the number, barely, but the process was exhausting and Elena could not have defended a single rejection decision if anyone had asked. This project is about fixing that. You will follow Elena as she takes one concrete challenge, maps it honestly, redesigns it step by step with verification checkpoints, and puts guardrails around the parts where AI touches a hiring decision, and you will produce the same thing she did: a design document you could take back to your organization and implement.
How This Project Works
This is a hands-on project rather than a lecture. You will work through six stages in order, on a real recruiting challenge from your own organization, and each stage produces a written artifact that the next stage consumes. Select a challenge, map the current state, identify improvement opportunities, design the improved workflow, define measurement, and plan the implementation. The templates here scaffold your thinking so you do not start from a blank page, and by the end you will hold a design document containing all six pieces.
Work it individually or with a small group; the group version is better, because the argument about which step is really the bottleneck is where most of the learning happens. Write your answers down at each stage rather than holding them in your head. A workflow that exists only in someone's head cannot be reviewed, taught, or defended, which is the same reason the final stage of the project is documentation.
Stage 1: Choose One Concrete Challenge
A good workflow project starts with a challenge that has three properties. It causes pain, meaning it genuinely frustrates you or your team by wasting time, producing poor quality, or creating fairness concerns. It is specific, which is the difference between "our recruiting is slow" and "our engineering screening is too slow because we get 200 applications for five roles." And it is actionable, meaning it sits within your influence to change rather than being determined by factors outside your control.
Four examples show the right level of resolution. A campus recruiting team has high no-show rates for first-round interviews: candidates schedule, then do not appear, wasting hiring manager time and frustrating everyone. A company has an offer acceptance rate of 60 percent, losing candidates to competitors after extending offers, with some accepting and then backing out days before their start date. A team notices that referral candidates move through the process in one week while non-referral candidates take four, creating inconsistency and possibly biasing hiring toward people with existing networks. A hiring team finds its first interview round takes too long because each of five interviewers is assessing the same skills, wasting interview time and producing inconsistent feedback.
Elena's challenge qualifies on all three properties. It is painful because the volume crushes a two-person team. It is specific because it is not "hiring is slow" but "we must screen 540 applications and hire 60 warehouse associates in four weeks." And it is actionable because the screening process is hers to redesign, even though the headcount target and the calendar are fixed. Write your own challenge down in one or two sentences before going further. Elena's reads: reduce the time it takes to fairly screen a high-volume seasonal applicant pool so the team can fill 60 roles in four weeks without sacrificing fairness or candidate experience. Every design decision that follows is judged against that sentence.
Stage 2: Map the Current State
Before redesigning anything, document what actually happens today. Create a diagram of your current process for this challenge and capture six things. The starting point: does a candidate apply, do you reach out, does a referral come in? The steps in order, including every decision point where candidates might advance, be rejected, or be held. The bottlenecks: where do candidates wait and where does work pile up? The time each step takes, counting both process time, meaning the work itself, and wait time, meaning the interval during which nothing happens to the candidate. Who is involved at each step: recruiter, hiring manager, interview panel, someone else? And where the handoff risk lives, meaning the points where information gets lost or misunderstood as it passes between people.
Elena's map ran like this. An application lands in the applicant tracking system. The coordinator opens it, reads the resume, and decides whether the candidate meets the basic bar: legal work eligibility, ability to lift a stated weight, availability for the required shifts, and a reasonable commute. Borderline cases wait in a pile for Elena. Cleared candidates get a phone screen, then an in-person interview, then an offer.
The map exposed the bottleneck immediately. Manual resume review was eating roughly six hours of screening time per requisition batch when applications surged, and because it was sequential, candidates at the back of the queue waited days. Wait time, not work time, was losing the best applicants to faster employers, and that distinction only became visible because the map recorded both. The rejection step was the other weak point: under time pressure the team sent rejections fast and without consistency, which is exactly the condition that produces unfair and indefensible decisions. Elena documented all of it as her baseline, because she could not prove improvement later without numbers now.
Stage 3: Identify the Improvement Opportunities
With the map in hand, interrogate it with five questions. Where is cycle time longest, and where do candidates wait the most? Where are decisions unclear or inconsistent? Where could process redesign, with no AI involved at all, eliminate a bottleneck outright? Where would AI add genuine value by supporting human judgment, reducing volume on a scarce resource, or improving consistency? And where is human judgment irreplaceable, meaning which touchpoints must be protected?
The third question deserves more weight than it usually gets. Plenty of recruiting bottlenecks are process problems wearing a technology costume, and a scheduling delay caused by sequential handoffs is fixed by running steps in parallel, not by adding a model. Look for the process fix first, because it is cheaper, faster to implement, and does not introduce a new source of error.
For each opportunity you identify, write three things: what you want to improve, whether it is a process problem or a decision problem, and whether AI could help or whether it requires process redesign. That classification is the whole point of the stage. Elena's answers were direct. Cycle time is longest at the initial screen. Decisions are least consistent at rejection. AI adds genuine value by extracting and structuring the four basic qualifications from each application so a human is not reading 540 resumes line by line. And human judgment is irreplaceable for any decision to reject a candidate and for any borderline call.
That last answer defines the guardrail that shapes her entire redesign. AI in this workflow does extraction and flagging, never rejection. It reads an application and reports, with the supporting text, whether each of the four qualifications appears to be met, missing, or unclear. It does not decide. It does not rank candidates against each other on anything resembling a demographic proxy. A human makes every advance-or-reject call. Elena wrote this down as a non-negotiable so that nobody under deadline pressure could quietly let the AI auto-reject the unclear pile.
Stage 4: Design the Improved Workflow
Now draw the new workflow in the same diagram format as the current state, so the two can be compared directly. A good design shows six things: parallel paths where work no longer needs to happen in sequence; human touchpoints where judgment is essential, annotated with what information the human needs and what criteria they are assessing; automation or AI placed where it supports decisions without replacing judgment; measurement points where you will gather data on efficiency, quality, fairness, and experience; simplified handoffs and eliminated steps; and the information flows showing what data moves where.
One constraint governs the whole exercise. Unless the challenge is severe, your new workflow should not be dramatically different from the current state. Modest improvements are far more likely to succeed than radical redesigns, because they can be absorbed by a team that is also doing its job. Elena's redesign adds one AI step, two verification checkpoints, and a scheduling change. That is all.
Step 1: AI extraction. As applications arrive, an AI step reads each one against a fixed prompt and returns a structured summary: work eligibility, lifting requirement, shift availability, and commute, each marked met, missing, or unclear, and each accompanied by the exact application text supporting the call. The prompt forbids inferring anything about gender, age, national origin, or any other protected characteristic, and instructs the model to mark anything ambiguous as unclear rather than guessing.
Verification checkpoint 1. Before trusting the extraction at scale, Elena hand-checks the AI's structured output against the source application for the first 30 candidates. The extraction has to agree with a human read at least 90 percent of the time on the four fields before the team relies on it. This is the test that turns AI output from a hopeful shortcut into a trusted input.
Step 2: Human triage. The coordinator works from the structured summaries rather than raw resumes. Candidates clearly meeting all four qualifications advance to phone screen. Candidates clearly missing a hard requirement are reviewed by a human before any rejection. The unclear pile goes to Elena. Because the AI surfaced the supporting text, triage is fast and every decision has a documented basis.
Step 3: Parallel scheduling. Phone screens are booked through self-scheduling links so candidates do not wait for the coordinator to call, running advancement and scheduling in parallel rather than in sequence. This attacks wait time directly, and it is the pure process fix in the design: no AI involved.
Verification checkpoint 2: human review before rejection. No candidate is rejected on the AI's say-so. A human reviews every rejection, and for the unclear pile the default is to advance to a quick human look rather than to drop. This checkpoint is both a fairness control and a quality control.
Step 4: Consistent rejection communication. AI drafts the rejection note from an approved template, a human approves it, and it goes out promptly and uniformly. Consistency here is a fairness asset, not merely a courtesy, because it removes the variation that time pressure had been introducing.
The Before and After
The redesign's value shows up in time and in defensibility. Before, manual screening ran about six hours per requisition batch and candidates waited days in a sequential queue. After, with AI handling extraction and humans working from structured summaries, screening time per requisition fell to roughly two hours, and self-scheduling cut the wait between application and first contact from days to under 48 hours for cleared candidates. Across a 540-application surge, that is a large reclaim of a scarce team's time, redirected to the human judgment the workflow deliberately protects.
Just as important, every advance and every rejection now carries a documented basis pulled from the application itself, which is exactly what Elena could not produce the year before. The workflow is not dramatically more complex than the old one, and that restraint is deliberate. Modest, defensible improvements survive contact with a real peak season in a way that a wholesale reinvention, adopted by nobody once the pressure arrives, does not.
Guardrails, Adverse Impact, and Privacy
Because AI now touches the screen, Elena added explicit fairness guardrails beyond the human-review checkpoints. The extraction prompt is barred from using or inferring demographic proxies, and the four qualifications it checks are all genuine, job-related requirements. After the first wave of screening she runs an adverse-impact check on the AI-assisted screen using the four-fifths rule: she compares the rate at which candidates advance past the initial screen across the demographic groups she is able to track, and if any group's advance rate falls below four-fifths of the highest group's rate, she treats it as a flag to investigate the prompt and the criteria, not as an acceptable cost of speed.
She also keeps the workflow inside privacy rules. Candidate data fed to the AI step is limited to what the screen actually needs, handled under the company's GDPR and CCPA obligations, and not retained beyond the requisition. And she documents the whole design, because a workflow that lives only in her head cannot be audited, trained on, or defended. If the company operates anywhere that regulates automated employment decision tools, such as under New York City's Local Law 144, the documentation and the human-decision design are what keep the workflow on the right side of the line.
Stage 5: Define Measurement Before Launch
Define your metrics before the new workflow goes live, so you can judge the design honestly rather than relying on the feeling that things went more smoothly. Take one metric from each of four families. One efficiency metric covering speed, resource utilization, or cost. One quality metric covering retention, performance, or promotion. One fairness metric such as conversion rate by demographic group or time-in-process by demographic group. And one experience metric such as net promoter score, satisfaction, or willingness to recommend.
For each metric, record four things: the current baseline drawn from your current-state map, the target you expect after the change, how you will measure it, and how frequently. A metric without a baseline cannot demonstrate improvement, and a metric without a cadence quietly becomes an annual exercise. Elena's set reads: for efficiency, screening time per requisition, targeting a drop from six hours to two. For quality, the offer-acceptance and 30-day retention rate of seasonal hires, to confirm that faster screening did not lower the bar. For fairness, advance rates past the initial screen by demographic group, watched against the four-fifths threshold. For experience, time from application to first contact, plus a short candidate satisfaction pulse.
Notice that the fairness and quality metrics exist specifically to catch the design succeeding at the wrong thing. It is entirely possible to halve screening time while quietly raising the rejection rate for one group or lowering the calibre of hires, and a measurement set built only around speed would report that as a triumph.
Stage 6: Plan the Implementation
A design that never survives contact with the team is not a result. Seven principles govern the move from document to practice, and they double as a checklist for reviewing anyone else's workflow design. Identify the specific challenge precisely, not "improve sourcing" but "reduce the time it takes to find senior machine learning engineers." Map the current state honestly, including what is working, not only what is broken. Identify where AI can help and be specific about which step. Design the human touchpoints deliberately, naming what judgment is required at each. Plan monitoring so you will know whether it is working. Plan iteration, designing for a first version rather than perfection and building in feedback loops. And document: write the workflow down, train the team on it, and maintain the documentation as it changes.
The iteration principle is the one most often skipped. A first version that goes live and gets corrected on real evidence beats a perfect design that is still being refined when peak season starts. Elena's extraction prompt was revised twice in the first week on the strength of what checkpoint 1 revealed, which is the system working rather than the system failing. Good workflows are simple, measurable, and continuously improved on the basis of results.
Anti-Patterns
Over-automation. This is designing to automate everything technically possible rather than placing automation strategically where it adds value without replacing judgment. It is seductive because each individual automation looks like a saving, and the compounding effect is a workflow where nobody can explain why a particular candidate was rejected. The defense is Elena's explicit written boundary: AI extracts and flags, humans decide, and any proposal to extend AI into a decision has to clear the same bar the original design did.
Ignoring the stakeholder perspective. This is designing a workflow that works beautifully from your seat but that hiring managers, screeners, or candidates dislike or simply will not use. It happens because the designer is the person with the clearest view of the problem and the least exposure to the daily friction of the solution. The defense is to walk the design through with each group before launch and to ask what would make them abandon it under pressure, then fix that.
Underestimating change management. This is producing a genuinely good workflow and failing to consider whether your team will actually adopt it. A new process competes with an old one that people already know, and under deadline the old one wins by default. The defense is to plan adoption as part of the design: who trains whom, what the first week looks like, and who is accountable for the new steps actually happening.
Misaligned incentives. This is designing a workflow that achieves your stated objective, usually speed, at the cost of something else that matters, usually quality or fairness. It is the failure mode that measurement is designed to catch, which is why the metric set spans four families rather than optimizing the one you set out to improve. If your only metric is the one your design targets, you have built a system that cannot tell you what it broke.
Practice Prompts
- State the challenge. Select a specific recruiting challenge from your organization and write it in one sentence. Explain what makes it a problem and how it affects hiring, then test it against the three properties: painful, specific, actionable.
- Map the current workflow. Document the process for that challenge, including cycle time split into process time and wait time, every decision point, every handoff, and the bottlenecks. This is your baseline, so capture the numbers.
- Identify three opportunities. For each, specify whether it is a process problem or a decision problem, and what solution you would pursue. Be honest about which ones need no AI at all.
- Design the improved workflow. Show parallel paths, human touchpoints with the information and criteria each requires, where automation sits, verification checkpoints, and simplified handoffs. Keep the change modest enough to adopt.
- Define the measurement approach. Name one metric each for efficiency, quality, fairness, and experience, with a baseline, a target, a method, and a frequency for each.
Reflection
- If you actually implement the workflow you have designed, what is the one thing most likely to go wrong, and how would you mitigate that risk?
- Who is the most important stakeholder to convince about your new workflow, and what specifically would convince them?
- What is one thing about your challenge that you did not anticipate when you started this design?
- If you had to pick the single change most likely to improve your situation, what would it be, and is it in your design?
- Which step in your design would be the first to be abandoned under deadline pressure, and what would keep it in place?
Glossary
- Design document. A comprehensive description of a workflow covering current state, improvements, roles and responsibilities, measurement, and implementation approach. It is the deliverable of this project.
- Stakeholder. Anyone affected by or involved in the workflow, including recruiters, hiring managers, interview teams, and candidates.
- Change management. The work of helping people adopt new ways of working, including communication, training, and change leadership.
- Process time and wait time. Process time is the work itself; wait time is the interval during which nothing happens to the candidate. Mapping both separately is what reveals which one is actually costing you candidates.
- Verification checkpoint. A defined point where AI output is checked against a human read or an approval gate before it is relied upon.
- Four-fifths rule. The rule of thumb that a group advancing at less than four-fifths of the highest group's rate is a flag for adverse impact warranting investigation.
Related Lessons
This project brings together the whole workflow design chapter, and several lessons feed it directly.
- Workflow Mapping: Understanding Current Flows supplies the mapping technique that stage 2 depends on, including how to capture wait time separately from process time.
- Process Redesign: Where AI Can Add Value Without Displacing Human Judgment develops the distinction at the heart of stage 3, between a process problem and a decision problem.
- Human Touchpoints: Strategic Moments for Human Review covers how to choose and equip the points where human judgment is protected.
- Metrics and Monitoring: Tracking Efficiency, Quality, and Fairness extends the four-family metric set in stage 5 into an ongoing monitoring system.
- Piloting and Iteration: Testing Workflows and Gathering Feedback is the natural next step once your design exists and needs a first real run.
- Hands-On Project: Document Your Recruiting Process and AI Use takes the documentation principle in stage 6 and turns it into its own deliverable.
Closing
Workflow design is a learnable skill, and it improves with repetition rather than with study. The more designs you work through, the better you get at seeing where the real bottleneck is, at telling a process problem from a decision problem, and at placing AI where it earns its position instead of where it is easiest to install. Elena's redesign was not clever. It was one AI step, two checkpoints, and a scheduling change, chosen because they addressed what her map had shown her.
What made it work was the discipline around it: a challenge stated in one sentence, a baseline captured before anything changed, an explicit written rule that AI never rejects a candidate, a hand-checked accuracy gate before the extraction was trusted, an adverse-impact check on the results, and four metrics defined before launch. Take this project back to your organization and iterate on it with your team, and treat the first version as something to correct rather than something to defend.
Key Takeaways
- Start with a real, specific, actionable challenge. A painful problem stated in one sentence keeps every design decision honest. "Screen 540 applications and hire 60 associates in four weeks" beats "make hiring faster."
- Map the current state completely before redesigning. Capture steps, decision points, who is involved, handoff risk, and both process time and wait time. You cannot prove improvement without a baseline, and the map is what reveals that wait time was losing the best candidates.
- Look for process fixes before AI. Simplify handoffs, create parallel paths, and eliminate unnecessary steps first. Self-scheduling attacked wait time with no model involved at all.
- Be intentional about where AI is used. Support decision-making, reduce volume on a scarce resource, or improve consistency. Use AI for extraction and flagging, never for rejection, and write that boundary down so nobody crosses it under deadline pressure.
- Protect human touchpoints and equip them. Identify where judgment is essential and make sure those points have clear criteria and the information they need, including the supporting text behind every AI call.
- Build a verification checkpoint into every place AI touches the flow. Hand-check extraction accuracy against a human read before trusting it at scale, and require human review before any rejection.
- Measure the before and after with real numbers. Screening time per requisition from six hours to two, first contact from days to under 48 hours. Reclaimed time gets redirected to the judgment the workflow protects.
- Define measurement upfront across four families. One metric each for efficiency, quality, fairness, and experience, each with a baseline, a target, a method, and a cadence, so you find out whether the design achieved its objective or broke something else.
- Put fairness and privacy guardrails on any AI-assisted screen. Forbid demographic proxies in the prompt, run a four-fifths adverse-impact check on advance rates, limit and properly handle candidate data under GDPR and CCPA, and document the design so it can be audited and defended, including where automated employment decision tools are regulated.
- Prefer modest, defensible improvements over radical redesigns. One AI step, two checkpoints, and a scheduling change survive a real peak season far better than a wholesale reinvention nobody adopts under pressure.
Frequently Asked Questions
How do I know whether my challenge is specific enough to design against? Test it by asking whether you could tell, three months from now, that it had improved. "Our recruiting is slow" fails that test because nothing about it names a measurable quantity. "We must screen 540 applications and hire 60 associates in four weeks" passes, because every element of it is countable and every design decision can be checked against it. If your statement does not contain a number, a stage, or a role, it is a theme rather than a challenge, and you will end up designing a workflow that improves something you never defined.
What if my bottleneck turns out to be something AI cannot help with? That is a successful outcome of stage 3, not a failure of it. A large share of recruiting bottlenecks are sequencing and handoff problems, and the fix is running steps in parallel, removing an approval, or letting candidates schedule themselves. Those changes are cheaper, faster to implement, and introduce no new source of error. The classification exercise, process problem or decision problem, exists precisely so you do not spend a quarter deploying a model against a problem that a calendar link would have solved.
How do I set the accuracy bar for a verification checkpoint? Set it before you look at any results, and set it against what an error costs at that step. Elena required agreement with a human read on at least 90 percent of fields across the first 30 candidates before relying on the extraction, which is defensible for a step whose output a human still reviews. A step where AI output flows onward with less human scrutiny warrants a tighter bar, and any step where an error would directly reject a candidate should not be automated at all. What matters most is that the gate exists, that the number is written down in advance, and that failing it means the step does not go live.
Our applicant tracking system will not let us change the workflow. Is this project still useful? Yes, and it may be more useful. The design document is what turns a vague complaint about the system into a specific, evidenced request: here is the step, here is the wait time it creates, here is what the redesigned flow requires, and here is the measured improvement we expect. Much of the design will also work outside the system entirely, since self-scheduling, a written rejection review step, and a stratified quality check are process disciplines rather than features. Design the workflow you need, then negotiate the constraint with evidence rather than assumption.
How much should I involve hiring managers in the design? Early and specifically, because they determine whether the workflow survives. Show them the current-state map first, since the wait times are usually more persuasive than any argument you could make, and ask them directly what would cause them to abandon the new steps in a busy week. Involving them in the design of the criteria rather than presenting a finished rubric buys ownership you cannot get any other way. Then show them the measurement once it exists, because a practice people helped design and can see working is the one that is still running next season.
Do I need to disclose to candidates that AI is used in the screen? Treat it as required rather than optional, and check the rules that apply where you hire, since jurisdictions that regulate automated employment decision tools attach specific notice obligations to their use. Beyond compliance, disclosure is part of the defensibility Elena was after: a workflow where AI extracts qualifications and a human makes every advance-or-reject call is straightforward to explain, and explaining it plainly is far easier than explaining after the fact why you did not. Keep the disclosure specific about what the AI does and does not do, and keep the documentation that shows the description is accurate.
Skill.re