Piloting and Iteration: Testing Workflows and Gathering Feedback
Lena Vasquez is the recruiting operations lead at a 900-person logistics company, where a team of eight recruiters screens heavy inbound volume for warehouse and driver roles. Leadership has asked her to roll out an AI-assisted resume-screening and outreach workflow across the whole team. Lena's instinct is the right one: do not roll it out, pilot it first. This lesson follows her as she designs and runs a structured six-week pilot, decides what success means before she starts, gathers feedback from both numbers and people, iterates once, and reaches a defensible go, no-go, or scale decision. The pilot is not a delay before the real work. It is the experiment that determines whether the real work is worth doing.
Why Pilot Instead of Rolling Out
The temptation Lena resists is to flip the new workflow on for all eight recruiters at once. A full rollout treats an untested change as a finished one, and if the workflow has a flaw, that flaw now lives in every search simultaneously with no clean way to tell whether outcomes got better or worse. A pilot is both risk management and learning. It contains any damage to a small group, and it surfaces the edge cases that look invisible on paper: the referral that does not fit the new flow, the candidate type the AI screen mishandles, the handoff that confuses people, the moment where a recruiter quietly works around a step because nobody explained it.
A pilot also buys something a rollout cannot: credibility. When Lena can later show her team and leadership that the new workflow measurably beat the old one with real recruiters and real candidates, adoption stops being an argument. People resist change they were told to accept and embrace change they watched work. Framing matters here as much as design. Lena tells everyone involved, plainly, that this is a pilot and that the point is to learn: "This is new, we are figuring it out, let us gather feedback and improve it together." That framing lowers defensiveness, invites people to report problems rather than hide them, and makes it psychologically possible for the pilot to fail, which is the only condition under which its success means anything. The objection she expects, and gets, is that hiring is urgent and a pilot is a delay nobody can afford. The honest answer is that a pilot is risk prevention rather than overhead: a few weeks spent catching a bias problem or a quality regression is far cheaper than discovering the same problem after it has been running across every open req, and the second kind of discovery costs more than time.
Where the Pilot Sits: Design, Pilot, Audit, Iterate, Scale
Before designing anything it helps to see where a pilot sits in the larger arc, because it is one phase of five and most failures come from skipping one of the other four. Phase 0 is design: map the current workflow, design the AI integration, and define the success criteria. Phase 1 is the pilot itself, commonly run with one or two people, 50 to 100 candidates and two to four weeks, long enough to learn what works and what breaks. Phase 2 is audit, and it is the phase teams skip: stop, analyze the pilot data, check quality, fairness and efficiency deliberately, and name the issues and their fixes rather than carrying an impression forward. Phase 3 is iterate, making the changes the audit identified, retraining the team, and adjusting criteria, gates or tools. Phase 4 is scale: roll out to the full team, keep monitoring weekly and monthly, and stay ready to pause if something emerges at volume that the pilot never surfaced.
Those sizes are a starting shape rather than a rule, and Lena's pilot runs longer and wider because her volume and her exposure both warrant it. The principle underneath the phases is what transfers: a short pilot that catches a problem is worth more than a long rollout that hides one, so pilot early, audit properly, and iterate fast rather than perfectly.
Designing the Pilot: Scope, Cohort, and Control
Lena designs the pilot to be small enough to manage and large enough to mean something, which is the central tension of pilot design. She scopes it to two high-volume role families, warehouse associate and driver, because those carry enough weekly volume to expose variance. Scope needs an explicit boundary in both directions: what is in, and equally what is out, because a pilot that quietly expands to cover an extra req or an extra team stops being comparable to anything. She sets the duration at six weeks, long enough to move past the awkward first week and capture a real rhythm, short enough to decide without stalling the broader rollout. Longer pilots gather more data and delay every decision that depends on them; shorter ones move fast and may never encounter the situations that break the workflow.
Her cohort is six of the eight recruiters running the new AI-assisted workflow, deliberately mixed: two early enthusiasts, two steady middle-of-the-road recruiters, and two skeptics, because skeptics find the problems enthusiasts talk themselves past. Everyone who touches the workflow belongs in the cohort, not just the screeners, so hiring managers and interviewers are briefed and included. The remaining two recruiters keep working the old manual workflow on comparable roles as a control group, so Lena can compare test against control rather than against a fuzzy memory of how things used to feel. She holds the rest of the environment as steady as she can, asking that job descriptions, hiring managers and req volume stay consistent during the six weeks, and she writes down any change she cannot prevent so she can account for it later.
| Design decision | Too small | Too large | Lena's choice |
|---|---|---|---|
| Scope | One req; nothing generalizes | Every role family; no clean comparison | Two high-volume role families |
| Duration | Ends inside the learning curve | Delays every dependent decision | Six weeks |
| Cohort | One recruiter; personality, not process | The whole team; no control left | Six of eight, mixed by attitude |
| Control | None; comparison against memory | Not a real risk here | Two recruiters on the old workflow |
Defining Success Before You Start
Lena decides what success looks like before a single candidate flows through the new workflow, because metrics defined after the fact tend to flatter whatever happened. She picks two quantitative measures and one qualitative one. The primary metric is time-to-screen, the hours from application to a screening decision, where the old baseline averages about 22 hours. The secondary metric is screen-to-interview quality, measured as the share of AI-screened candidates the hiring manager rates as genuinely interview-worthy, baselined at about 55 percent. The qualitative measure is recruiter satisfaction with the workflow, captured on a simple 1-to-5 scale. Each has a baseline attached, because a success criterion without a baseline is an opinion.
Crucially, Lena writes a go/no-go rule in advance, so the decision is not relitigated under pressure at the end. Her rule: the pilot scales if time-to-screen drops at least 30 percent without screen-to-interview quality falling below the 55 percent baseline, and if recruiter satisfaction averages 3.5 or higher. If speed improves but quality drops below baseline, it is a no-go pending fixes. If nothing moves, it is a no-go. Naming the threshold up front is what keeps the pilot honest, because it commits her to accepting an answer she might not want. Vaguer criteria are worth writing down too, provided they are made specific: "no major breakdowns" means nothing until you define what counts as one, and "the team says it feels faster" is a real criterion only if you decide in advance how you will ask.
Two further criteria sit outside that rule as gates rather than as targets, because neither is something a strong speed number can buy back. The first is fairness: advancement rates broken out by demographic group, with no disparate impact detected and an impact ratio at or above 0.80 for every group. A pilot is the cheapest moment in the entire lifecycle to detect bias, because the population affected is still small and nothing has been distributed across every open req. It is also where the compliance clock starts, since where a rule such as NYC Local Law 144 applies, a bias audit is owed before an automated employment decision tool is deployed rather than after it has been running, and EEOC guidance expects documented validation rather than an assurance that someone looked. The second gate is candidate experience: satisfaction should not drop and candidate feedback should be neutral or better. Lena writes both as pass/fail conditions rather than as numbers to optimize, so a strong efficiency result cannot be traded against either.
Running the Pilot: Practical Mechanics
Once the pilot starts, the work shifts from design to observation, and the single most valuable habit is documenting as you go. Lena keeps a running log of what actually happens: the moments where the process breaks down, where people get confused, where a handoff fails, where someone has to ask a question the instructions should have answered. Written at the time, in a sentence, these entries become the richest feedback she has, because by week six nobody will remember the week-two friction that got worked around and never fixed. She also documents the variance the workflow encounters. Not every candidate arrives the same way; some are referred, some sourced, some apply directly. Recording how the workflow handled each path tells her whether it accommodated referrals gracefully and whether sourced candidates moved faster, which is exactly the knowledge she needs to refine it later. What she logs is organized rather than freeform, because a pilot generates far more observation than anyone can sort afterwards, and five dimensions cover almost everything worth capturing.
| Dimension | What to track | Why it matters |
|---|---|---|
| Efficiency | Time per candidate at each stage, recruiter hours, how long each decision takes | Are we actually faster, or do the new gates add overhead? |
| Quality | Advancement rates, the quality of the AI's own output such as parsing accuracy, hiring manager feedback on screened candidates | Did speed cost quality, and are the AI's decisions any good? |
| Fairness | Advancement rates by demographic group, and whether any group moves through faster or slower | Early detection of bias, while the affected population is still small |
| Process adherence | Did recruiters follow the workflow, skip any gates, or complain about specific steps? | The design may need changing to fit how real teams work |
| Edge cases | How unusual candidates fared: career changers, international applicants, nontraditional backgrounds | Workflows usually fail on the unusual cases, so catch them here |
The second habit is supporting the team rather than watching them struggle. A new workflow feels awkward for the first week regardless of how good it is, and people will need clarification. Lena makes sure someone is explicitly on the hook to field questions, because problems nobody reports do not get solved, and a recruiter who quietly invents a workaround has both hidden a defect and corrupted the measurement. The third is measuring as the pilot runs rather than only at the end. She calculates her three metrics weekly, which lets her distinguish a genuine workflow problem from a first-week learning curve that will resolve on its own. Early numbers are noisy and she treats them as direction rather than verdict, but a metric moving the wrong way in week two is worth understanding in week two.
Gathering Feedback: Several Methods, Used Together
Lena gathers feedback continuously rather than waiting for the end, and she uses several methods because each reveals something the others miss. She runs short check-ins at the one-week and three-week marks, asking each pilot recruiter what is working, what is not, and what they need, which lets her course-correct mid-pilot instead of collecting complaints too late to act on them. Around that spine, five other methods do distinct jobs.
- Observation. Lena or a designated observer watches a few live screening sessions. This is where the real friction hides, because you see where recruiters hesitate, where they need to ask for help, and where they improvise around a step. People rarely report the workarounds they have stopped noticing.
- Individual conversations. One-on-one, people say things they will not say in a group. The useful questions are comparative and open: how did this feel compared to the old way, what surprised you, what would you change.
- Group debrief. Bring the participants together for an hour, take notes, and let the discussion run. Group dynamics surface ideas that individuals would not raise alone, because one person naming a frustration gives three others permission to agree.
- Structured survey. A short instrument rating specific aspects of the workflow, how clear the instructions were on a 1-to-5 scale, how long a given step took, one thing that could be better, gives you quantitative feedback you can compare across people and across weeks.
- Candidate pulse. Send a brief survey to candidates who went through the pilot workflow: how clear was the process, how long did you wait between stages, would you recommend this company. This is the only method that reports the side of the experience your recruiters cannot see.
- Hiring manager feedback. Ask the managers receiving the screened candidates one comparative question: is the quality of what we are sending you better or worse than usual? They see the workflow's output without seeing its mechanics, which makes their read the closest thing you have to an independent check on screening quality.
Used together these give a complete picture, and the division of labor between them is what matters. The numbers tell Lena whether outcomes moved. The conversations tell her why. A pilot that produces only metrics will show her that quality dipped without telling her that the cause was a wording problem in the screening prompt, and a pilot that produces only anecdotes will leave her arguing about whose impression is right.
Iterating Without Over-Iterating
At the three-week mark Lena synthesizes the feedback and resists the urge to change ten things. She will always receive more feedback than she can act on, and the first job is triage: some of it points at genuine problems, some at preferences, some comes from people living inside the workflow daily and some from stakeholders with far less context. A workable discipline is to synthesize everything and identify the top three things to improve, then act on fewer than three. Early data shows time-to-screen has dropped sharply, but skeptics report the AI screen is down-ranking strong candidates who describe forklift certification in nonstandard wording, dragging quality. That is a real problem pointing at the screening prompt, so it earns an iteration. The general rule behind her timing is to iterate once, early enough that the remaining weeks can measure the effect: around the two-week mark on a four-week pilot, the three-week mark on Lena's six-week one. Test one or two small changes, see whether they help, and build on what works rather than waiting until you have accumulated ten improvements and can no longer tell which one mattered.
Lena adjusts the prompt to recognize the certification variants and communicates the change plainly: "You flagged that the screen missed certified candidates, here is the fix, let us see if it helps." That sentence does more work than it looks. Telling people what changed and why shows that feedback drives decisions, which is what makes the next round of feedback honest. Then she stops iterating. Changing the workflow every week would prevent her recruiters from building any muscle memory and would muddy the measurement, since she would never know which change drove which result. Never iterating is the opposite failure, ignoring what your team is teaching you. One substantive change at the midpoint, measured before and after, is the discipline that sits between them.
The measurement half of that discipline is easy to skip and easy to regret. After the change Lena watches screen-to-interview quality specifically, to confirm the fix moved the metric rather than merely making the team feel heard. Changes that please people without improving outcomes are the most persuasive kind of false progress, and some changes improve one dimension while quietly degrading another, buying speed at the cost of quality. Measurement is the only thing that tells the difference, which is why the metric panel has to exist before the pilot starts rather than being assembled once the results are in.
Testing Variations: A/B Tests Inside a Pilot
Once the workflow is running smoothly, a pilot can do more than answer yes or no; it can start optimizing, and the instrument for that is a small A/B test comparing two variants of a single decision. The discipline is the pilot's own, at smaller scale: change exactly one thing, hold everything else constant, and decide in advance which outcome you are reading. Three tests show the shape of what is worth asking in recruiting.
- Screening criteria weight. One group screens using the AI score plus their own soft judgment; the other screens on the AI score alone. Which group advances the better candidates, judged downstream by the hiring manager?
- Quality gate strictness. One group documents a reason for every screening decision; the other documents rejections only. Which ends up with the better audit trail, and does the lighter version produce more mistakes?
- Communication style. One group sends templated rejection emails, the other personalized ones. Which draws better candidate feedback, and is the difference worth what personalization costs?
Keep these small: roughly 25 candidates per variant, run over one to two weeks, with the outcome measured and the learning written down where the next person will find it. Small is the point rather than a compromise, because a test large enough to feel definitive is large enough to be a second pilot. Read the sample size honestly in the other direction too. A difference between two 25-candidate arms is a direction worth exploring, not a finding to institutionalize, and the accurate way to report it is that the test favored one variant, not that it proved one.
Red Flags and the Rollback Plan
Every pilot needs a way to stop, decided before it starts, because the moment you need a rollback plan is the moment nobody has the composure to design one. Lena writes hers into the pilot document alongside the go/no-go rule, as red flags that trigger a pause rather than a debate.
- Quality drops more than 10 percent. If the candidates being advanced are meaningfully weaker than baseline, pause and investigate, reverting to manual screening until the cause is found and fixed.
- The impact ratio falls below 0.80 for any group. If any group is being systematically disadvantaged, pause immediately rather than at the next scheduled review, audit the criteria, roll back, and fix the bias before the question of scaling is even raised. The same 0.80 line that a pilot must clear to be called a success is the line at which a running pilot stops, and the further below it the reading sits, the less the pause is a judgment call.
- The team cannot sustain it. If the pilot recruiters have come to dread the workflow because the gates feel excessive or the tool fights them, redesign before scaling. A process nobody wants to use will be worked around at volume, whatever the metrics said in week six.
The plan itself has to be more than an intention. Write down the steps to revert to the manual process, make sure the team is trained on that fallback so reverting is not itself a scramble, and name the person who owns the decision to pull the trigger. When you do roll back, document what went wrong and why, because that record is the input to the next iteration and it is what stops the same failed fix being attempted twice.
The Go, No-Go, or Scale Decision
Before the decision comes the audit, and Lena works a fixed checklist so that reading the result is a procedure rather than an argument.
- Pull all the pilot data, time-to-screen, quality and the fairness ratios, and compare it against the baseline rather than against the first week.
- Review the recruiter feedback in full: what worked, what did not, what they would change.
- Review the candidate feedback for themes rather than for individual complaints.
- Audit a sample of 20 screening decisions, documenting the rationale for each, and look for bias or quality problems the aggregate metrics would not show.
- Check the edge cases specifically: how did career changers, international applicants and nontraditional backgrounds fare?
- Debrief as a team, share the findings, and get buy-in for whatever comes next.
- Update the workflow documentation so it records what changed and why.
At six weeks Lena reads the results against the rule she wrote at the start. Time-to-screen fell from about 22 hours to about 13, a drop of roughly 40 percent, clearing the 30 percent threshold. Screen-to-interview quality dipped to 51 percent in the first three weeks, then recovered to about 58 percent after the prompt fix, landing above the 55 percent baseline. Recruiter satisfaction averaged 3.8, above the 3.5 bar, with the skeptics moving from doubtful to cautiously positive once their certification concern was addressed. The control group's metrics held flat, which tells Lena the gains came from the workflow and not from some company-wide change in the same period.
Because every condition of her predefined rule is met, the decision is go, and specifically scale. She rolls the workflow out with the iterated prompt baked in, the control recruiters fold in, and she keeps measuring the same three metrics during rollout to catch any drift at larger scale. The rollout is staged rather than simultaneous, and the first stage deliberately goes to recruiters who were not in the pilot cohort, because people who did not help build the workflow are the ones who will find the instructions that only ever made sense to their authors. Two go first and get watched closely, then four, then everyone, with the metrics checked at each step and a standing agreement to pause and debug rather than push through if something moves. On that pattern a full rollout for a team of five or more typically takes four to six weeks. Had quality stayed below baseline, the answer would have been no-go pending another iteration, and Lena would have said so plainly. That willingness is exactly what made the eventual go decision trustworthy. A pilot run by someone who was never going to accept a negative result is not an experiment; it is a rehearsal with extra steps, and everyone involved can tell the difference.
Anti-Patterns
The pilot that is too small to learn from. A company pilots a new screening workflow with five hires over two weeks. Five candidates do not carry enough variance to test whether the workflow handles different candidate types, and two weeks is not enough to see whether it holds up under normal load. The pilot proves that five specific people can follow a process and teaches nothing about whether it works more broadly. It happens because teams want to test quickly before committing, so they shrink the scope until it is affordable. What goes wrong is that you scale a workflow validated only under artificial conditions, and it breaks in production. The fix is to pilot with enough volume and enough time to encounter real variance; four weeks and thirty to fifty hires is a far better size than two weeks and five, and if you cannot allocate that today, the allocation is still worth making rather than running a test that cannot answer the question.
The pilot where everything else changes too. You pilot a new workflow with one team while the old workflow continues elsewhere, and midway through, leadership rewrites the job description for the pilot roles. Or the hiring manager changes. Or a new sourcer joins. Several things move at once and you cannot attribute the outcome to any of them. It happens because organizations are living systems and change constantly; a truly isolated pilot is difficult to arrange. What goes wrong is that you finish the pilot without clean learning, unable to say what worked. The fix is to minimize other changes deliberately, asking for consistency in team, job descriptions and volume for the duration, and to document every change you could not prevent so you can factor it into the analysis rather than discover it afterward.
Dismissing negative feedback as resistance. The pilot indicates that the new workflow is slower and that people dislike it. But leadership has already committed to the change, so the feedback gets recategorized as resistance and the rollout proceeds. The team never fully adopts the workflow and it underperforms permanently. It happens because once leadership commits publicly, admitting the change might not work becomes psychologically and politically expensive, and momentum plus face-saving does the rest. What goes wrong is that you deploy a change your team does not believe in, that does not work in practice, and that teaches everyone to be cynical about the next change too. The fix is to run pilots with genuine openness to the answer being "this does not work," and to say so out loud when it is. Your team will trust you more for listening than for being consistent.
Practice Prompts
- Design a pilot on paper. For a workflow change you are actually considering, define the scope (which roles, which time period), the duration, the participants including at least one skeptic, and the success criteria. Write the go/no-go rule as a sentence with numbers in it.
- Build the feedback plan. Decide what you will observe and when, who you will interview and at which marks, what survey you will run, what you will ask candidates, and which metrics you will track weekly. Assign each method to the question it is meant to answer.
- Rehearse the uncomfortable result. Imagine you have just finished the pilot and your team reports that the new workflow takes longer than the old one, but you believe in the change. Write out how you would think it through: what questions would you ask, and specifically what evidence would change your mind?
- Draft the debrief. Design a one-hour end-of-pilot session with an agenda and discussion questions built to surface the feedback you actually need rather than the feedback that is comfortable to give.
- Test the riskiest assumption first. For the change you are considering, name the single assumption most likely to be wrong. Design a two-week mini-experiment that would test just that assumption, and say what result would falsify it.
Reflection
- Think about a major process change in your organization. How was it communicated? Did you feel consulted, and did the team embrace it or work around it?
- If you were to pilot the workflow change you are considering, what are you most nervous about testing? What does that nervousness tell you about where to look first?
- Who in your organization is most likely to give you honest feedback about whether a change is working, and how will you actually get that person's input?
- When should a pilot end and rollout begin? What decision rule would you be willing to write down before you start?
- What would have to be true for you to abandon a change you had publicly advocated for? Answer that now, while nothing is at stake.
Glossary
- Pilot. A small-scale, time-boxed test of a new workflow before full rollout, run to reduce risk and generate learning rather than to demonstrate a conclusion already reached.
- Control group. A comparable group that continues working the existing workflow during the pilot, so results can be compared against a concurrent reference rather than against memory.
- Go/no-go rule. A decision threshold written before the pilot starts, specifying the metric movements required to scale, so the decision cannot be relitigated once results arrive.
- Iteration. A deliberate change to the workflow made in response to feedback, measured before and after so its effect can be distinguished from the team's feelings about it.
- Scope creep. Expanding a pilot beyond its defined boundaries mid-flight, which destroys comparability and makes clean learning impossible.
- Structured feedback. Feedback gathered through consistent methods, surveys, scheduled check-ins, defined interview questions, rather than ad hoc conversation, so responses can be compared across people and over time.
- Variance. The range of different situations and candidate types that flow through a workflow. A pilot needs enough volume and duration to encounter it, because workflows usually fail on the unusual cases.
- A/B test. A small comparison of two variants of one decision, run inside a working pilot with everything else held constant, to answer a narrow optimization question cheaply.
- Rollback plan. The documented steps for reverting to the previous process, written before the pilot starts and paired with red flags that trigger it, with a named owner of the decision.
- Edge case. An unusual candidate or path the workflow was not designed around, such as a career changer or an international applicant. Edge cases are where workflows break, so a pilot has to log how each one was handled.
Related Lessons
- Metrics and Monitoring: Tracking Efficiency, Quality, and Fairness supplies the measurement panel a pilot is read against, including how to baseline before you change anything and how to separate a metric you can attribute from one you cannot.
- Hands-On Project: Design a Quality Audit Plan builds the standing audit process that takes over once a pilot ends and the workflow becomes business as usual.
- Change Leadership: Driving Adoption While Managing Resistance goes deeper on the adoption side of the rollout that a successful pilot earns you.
- Workflow Mapping: Understanding Current Flows is the prerequisite step, because you cannot pilot a redesigned workflow until you have documented the one you are replacing.
- Hands-On Project: Design a Workflow for a Specific Recruiting Challenge is where the design that this lesson tests gets built in the first place.
Closing
Pilots bridge the gap between design and implementation. On paper a workflow looks logical; in practice it is messier, because real people have questions, edge cases arrive, and timing rarely cooperates. A pilot lets you work through all of that at small scale, cheaply, before the flaws are distributed across every open req. The organizations that do this well do not deploy workflows and hope. They deploy deliberately, gather feedback through more than one channel, iterate once with intent, measure whether the iteration worked, and scale only when a rule they wrote in advance says they have earned it. Treat the pilot as the experiment it is, and the rollout stops being an argument.
Key Takeaways
- Pilot before you roll out. A pilot contains risk to a small group, surfaces edge cases that are invisible on paper, and builds the credibility that makes a later rollout stick. People embrace change they watched work.
- Scope for variance and keep a control. Choose enough volume and time to expose real variation, and keep a comparable group on the old workflow so you compare test against control rather than against memory.
- Mix the cohort, especially the skeptics. Include enthusiasts, steady performers and doubters. Skeptics find the problems enthusiasts talk past, and winning them over is the strongest signal of real adoption.
- Define success and a go/no-go rule before you start. Name the metrics, the baselines and the threshold in advance. A pre-committed rule keeps the decision honest when it is made under pressure.
- Document as it runs, and support the team. Keep a running log of breakdowns, confusion and how the workflow handled referrals, sourced and direct candidates. Make sure someone answers questions, because problems nobody reports never get fixed.
- Gather both numbers and voices, continuously. Track metrics weekly and pair them with observation, one-on-ones, a group debrief, a structured survey and a candidate pulse. Numbers show whether outcomes moved; conversations explain why.
- Iterate once, deliberately, and measure the change. Triage feedback into problems and preferences, make one substantive midpoint fix, communicate it plainly, hold everything else steady, and confirm before and after that the metric moved rather than just the mood.
- Make fairness a gate, not a target. Track advancement rates by demographic group throughout the pilot and require an impact ratio at or above 0.80 for every group before scaling. A pilot is the cheapest moment to detect bias, and where rules such as NYC Local Law 144 apply, the bias audit is owed before deployment rather than after.
- Write the rollback plan before you need it. Name the red flags that stop the pilot, a quality drop beyond 10 percent, an impact ratio below 0.80 for any group, a team that cannot sustain the workflow, then document the steps back to the manual process, train the team on them, and name who decides.
- Be willing to say no. The credibility of a go decision depends on having genuinely been open to a no-go. Read the results against the rule you wrote, and accept the answer it gives.
Frequently Asked Questions
How long should a pilot run? Long enough to move past the first-week learning curve and encounter real variance, short enough that the decision does not stall everything behind it. Lena chose six weeks for continuous, high-volume roles. For seasonal or cyclical hiring, one full hiring cycle is usually the better unit than a fixed number of weeks, because a calendar window that misses the cycle tests nothing. The failure mode at both ends is real: too short and you learn only that a few people can follow instructions, too long and you have delayed a decision you could have made a month ago.
What if we cannot spare a control group? Then say so and adjust your confidence accordingly. A control is what lets you distinguish your workflow change from everything else happening in the same six weeks, and without one you are comparing against a baseline that may itself have drifted. If a concurrent control is genuinely impossible, tighten the alternatives: hold the environment as constant as you can, document every change you could not prevent, and lean harder on stage-level metrics that the change plausibly affects rather than on end-to-end numbers that everything affects.
The team likes the new workflow but the metrics have not moved. What now? That combination is worth taking seriously rather than resolving in either direction quickly. Satisfaction is a real outcome, since a workflow people hate will not survive contact with a busy quarter, but it is not the outcome you set out to buy. Check first whether the metric had time to respond and whether your sample is large enough to show a change of the size you expected. If both are fine, the honest reading is that the workflow is more pleasant and not yet more effective, which is a no-go against a rule that required movement, and a reason to look at whether you pointed the change at the actual bottleneck.
How do we handle feedback that contradicts itself? Sort it by proximity and by kind. Feedback from people running the workflow daily carries more weight on how it functions than feedback from stakeholders with less context, though the latter may see consequences the operators cannot. Then separate problems from preferences: "the screen misses candidates who phrase a certification differently" is a defect with a fix, while "I preferred the old layout" is a taste that may resolve as people acclimate. Act on the top few defects, log the preferences, and tell people which category you put their input in so they know they were heard.
Do we need another pilot after we iterate? Usually not, if the change was contained and you measured its effect within the pilot window, which is what Lena did by fixing the prompt at week three and reading quality over the remaining weeks. A second pilot is warranted when the iteration is large enough to be a different workflow rather than an adjustment to the same one, because at that point your earlier weeks are measuring something you no longer intend to deploy. If you are unsure, ask whether a person reading your pilot report could tell which version produced which numbers.
Can we run several pilots at once? Run them sequentially wherever they share people, because the same recruiter cannot meaningfully run two new workflows at the same time and neither result would be readable afterwards. Run them in parallel only where the teams are genuinely separate. Sequential is the safer default even when parallel is available, since the first pilot almost always teaches you something that improves the design of the second, and running them together throws that away for a few weeks of calendar time.
The pilot shows the workflow is making biased decisions. What now? Roll back rather than iterate in place, and do it before the next batch of candidates goes through. Then find where the bias is actually coming from, which means examining the data, the criteria, the tool and recruiter judgment rather than assuming the tool is at fault. Fix it, retrain the tool if that is the source, and run another pilot to confirm the ratio moved rather than trusting that the fix was sound. What you do not do is scale a workflow you already know is producing biased decisions on the theory that you will fix it during the rollout.
Skill.re