Building a Pilot Timeline and Success Criteria
Ekundayo owns a coffee roasting operation in Richmond, Virginia, a roastery with a small retail counter and a wholesale business supplying about thirty local cafes. When he decided to try an AI tool for managing his wholesale client communications, he did what most small business owners do. He signed up, poked around for a few weeks, and then said to himself around week eight, "I think it's helping, but I'm not sure." He could not say whether the tool had reduced the time he spent on follow-up emails. He had not measured it before he started, and he had never decided in advance what success would look like. Eight weeks in, he had nothing concrete to show and no defensible reason either to keep the subscription or to cancel it.
The Hour That Prevents This
The fix is simple and takes about an hour before you start any AI pilot. You write down a timeline and you define what winning looks like. That is the whole intervention. It is not a project plan, not a business case, and not a document anyone outside your business will ever read. It is one page that settles three things in advance: when the pilot starts and ends, which single number you are watching, and what you will do at the end depending on where that number lands. The framework below is built for a ten-person business rather than a Fortune 500.
Settling those things early matters because the end of a pilot is the worst possible moment to decide what the pilot was for. By week eight you are attached to the tool, or tired of it, or quietly embarrassed about the money. Anything you write at that point is shaped by the decision you already want to make. A criterion written before you had any feelings about the tool is the only version that survives contact with your own preferences later.
Phase One: Setup and Baseline
Weeks 1 and 2 are for configuring the tool, training whoever will use it, and establishing your baseline measurement. You are not yet measuring results from the AI. You are measuring the before state, which is the only thing in this whole exercise you cannot go back and collect later. Once the tool is in the workflow, the old numbers are gone for good, and every claim you make afterwards becomes an argument rather than a comparison.
During setup, document how the process works right now, without AI. How long does it take in a typical week? How many errors or problems come up? What does a good output look like next to a bad one, in enough detail that somebody else could tell them apart? Write the answers somewhere you will still be able to find them at the end of week eight, because that is when they matter.
Setup is also when the people who will use the tool learn it, and that part gets rushed far more often than the configuration does. If three people will be sending these emails once the tool is in production, all three need to have used it on real work before the active pilot starts. Otherwise week three measures their learning curve rather than the tool, and you will read the flat result as a failure of the software. Two weeks is generous for configuration and thin for training, so spend the difference on the training side.
Ekundayo's baseline would have been simple. How many hours per week do I spend writing and managing wholesale client follow-up emails? He could have timed it for a single week. Fifteen minutes of observation, recorded in a note on his phone, would have produced the one number his entire pilot later needed and did not have. Instead he had nothing to compare against, which is why week eight gave him a feeling rather than a finding.
Phase Two: The Active Pilot
Weeks 3 through 8 are where you use the tool under real conditions with real stakes, on live client work rather than examples you picked because they suited the tool. Six weeks is usually the right pilot window for a small business AI tool. It is long enough for a genuine pattern to emerge out of ordinary week-to-week variation, and short enough that you are not spending months on something that is not working. Much shorter and you are mostly measuring novelty.
Track your primary metric weekly during this phase, not monthly. Weekly numbers catch problems while you still have room to fix them. If the tool is helping, you will often see an early signal by week three or four, once the initial learning curve flattens out. Treat that signal as a hint rather than a verdict, and treat its absence as a reason to look at how the tool is actually being used before you conclude anything. If the numbers are still flat by week four, you have time to adjust and still decide by week six, instead of drifting toward week twelve with nothing to show.
Whoever will use the tool in real operation should be inside the pilot, not watching it. A pilot that only the owner ran tells you very little about what happens when three people use it differently, at different speeds, with different tolerance for a draft that comes back slightly wrong. If including them is genuinely impractical, run it yourself, but treat adoption as an untested question rather than a solved one and expect the rollout to surface problems the pilot could not have caught.
Track what goes wrong as well. Keep a failure log: every workaround, every error, every time somebody skips the tool and does the job the old way. A line each is enough, with a note on what the task was. These entries tell you whether the problem is fixable, meaning the setup or the prompt needs adjusting, or fundamental, meaning the tool does not fit the workflow. Two very different pilots can produce the same disappointing number at week eight, and only the log tells you which one you are actually holding.
Phase Three: The Decision Point
The end of week 8 is your go/no-go moment. You compare the metric against your baseline, you read back through the failure log, and you make a deliberate decision: expand the tool into more of the work, modify the approach and run a second short pilot, or cancel the subscription. Deliberate is the operative word. The call gets made on a date you set in advance, out loud, whether or not you feel ready to make it.
Modify deserves a word, because it is the option people misuse. Modifying means you have a specific, named reason to believe a different setup would clear the bar, drawn from the failure log rather than from optimism, and you run a second short pilot with a fresh criterion to test it. It is not a way of extending the first pilot under a new name. If you cannot say in one sentence what you are changing and why the log points at it, the honest option is one of the other two.
What you do not do is extend the pilot indefinitely because you have not decided yet. That is how a pilot drifts past its own window and arrives at week twelve without an answer, still costing money every month it runs. An inconclusive result after eight weeks is itself a result. It usually means the tool is not delivering enough value to justify its cost and the complexity it adds to your workflow. You do not need to watch something fail actively before you are allowed to stop paying for it.
Defining Success Criteria
Success criteria are specific, measurable, and decided before the pilot starts. They answer one question: under what conditions would I call this a success? A weak criterion is "the tool saves me time." Nobody can argue with it, which is exactly what is wrong with it. Did it save thirty seconds or three hours? Both satisfy the sentence, so the sentence cannot settle anything at week eight, which is the only moment it was written for.
A strong criterion for Ekundayo would read like this: by week eight, time spent on wholesale client follow-up emails is down from a baseline of 4.5 hours per week to 2.5 hours per week or less. It is dull, it is checkable in a minute, and two people reading it would reach the same verdict. Strong success criteria share four properties:
- One primary metric. Not five metrics, one. The single most important number for this use case, chosen while you are calm.
- A specific target. Not "improvement" but a stated figure, either a percentage or a move from one number to another.
- A deadline. By week eight, not "eventually." A criterion without a date never comes due.
- A baseline. You cannot measure improvement without knowing where you started, which is why phase one exists.
Choosing the one primary metric is harder than it sounds, because most tasks improve in several directions at once and all of them feel important. The test is which number you would quote if somebody asked why you kept the tool. For an email workflow that is usually time, for a scheduling tool it might be errors, for a drafting tool it might be how many drafts go out without a rewrite. Pick the one that would change your mind, and let the others sit in the secondary list where they belong.
You are allowed to pick a modest target. There is a temptation to write something ambitious because it sounds serious, and then to quietly discount it later when the tool lands short. A pilot that clears a modest bar honestly teaches you more than one that misses an impressive bar you invented for an audience that does not exist.
Secondary Criteria Worth Tracking
Your primary metric tells you whether the tool works. Secondary criteria tell you whether it is worth keeping even when the primary metric improves. A tool can look good on the headline number while half the team quietly avoids it, or while it generates rework that lands somewhere the primary metric cannot see. Four secondary criteria are worth carrying through the pilot, and none of them needs more than a line in the same note.
| Secondary criterion | What it tells you | Where the number comes from |
|---|---|---|
| Adoption rate | What percentage of eligible uses actually went through the tool | Count of tasks done with the tool against the total eligible tasks that week |
| Error or correction rate | How often the output needed significant revision before it could be used | Your failure log, marking the drafts that needed a major rewrite |
| Staff experience | Whether the tool makes the job easier or harder for the people doing it | Asking them directly, in those words |
| Cost per unit of value | Whether the price is proportionate to the work produced | Monthly subscription divided by units of work the AI completed |
Each of these can overturn a good headline result. If you or your team consistently skipped the tool, low adoption is telling you something the time saving is hiding. A high correction rate means the tool is creating work rather than saving it, because somebody is now writing the draft twice. And when you ask the people using it whether their job got easier or harder, their answer matters on its own terms, not only as a proxy for the number.
A Sample Pilot Document
You do not need a formal document. A one-page note is fine, written in whatever you already use. Here is what Ekundayo should have written before he started, and what took him about an hour to reconstruct badly afterwards from memory:
Tool being tested: AI email assistant
Use case: Writing and managing wholesale client follow-up emails
Pilot window: Weeks 1 and 2 setup; weeks 3 through 8 active pilot; week 8 decision
Baseline: Currently spending 4.5 hours per week on this task, measured during setup week
Primary success criterion: Time on this task reduced to 2.5 hours per week or less by week 8
Secondary criteria: Fewer than 20% of AI drafts require major rewrites; all three team members using the tool at least 80% of the time
Decision rule: If the primary criterion is met, expand to other client communication types. If it is not met, cancel the subscription at the end of week 8.
That is the whole document. One page, written in advance, followed at week eight regardless of how comfortable or uncomfortable the tool feels by that point. The decision rule is the part people leave out, and it is the part that does the work. Writing down what you will do if the pilot fails, at a moment when you have not yet spent two months on it, is the closest thing a small business has to a defence against its own sunk costs.
Anti-Patterns
Starting the pilot before recording the baseline. This is the single most common failure and the only one that cannot be repaired. If the tool went live on Monday and you start measuring on Wednesday, the before number no longer exists anywhere except in your impressions of it. Everything afterwards is a debate about how things used to feel.
Writing a criterion you cannot lose. "The tool saves me time" and "the team finds it useful" are not criteria, they are moods. If there is no result that would count as a failure, you have not defined success, you have arranged to declare it. Read your criterion back and ask what number would force you to cancel. If you cannot answer, rewrite it.
Moving the goalposts at week eight. The tool lands short of the target, and suddenly there is a case that the real benefit was always quality rather than time. Sometimes that case is genuine, in which case run a second pilot with quality as the primary metric and a fresh baseline. Do not retrofit it onto the pilot that just missed.
Extending the pilot because the result is unclear. Unclear is a result. Four more weeks of ambiguity costs four more subscription payments and produces the same shrug, because the reason for the ambiguity is usually the design of the pilot rather than its length.
Judging the pilot on the primary metric alone. The headline number can improve while adoption sits at a third of eligible tasks and the correction rate quietly doubles somebody's workload. Read the secondary criteria before you expand anything.
Treating an early signal as the verdict. A promising week three usually reflects your own close attention and the easy cases you happened to run first. It is a reason to keep going to week eight, not a reason to stop early and roll the tool out to everyone.
Practice Prompts
Use these with an AI assistant while you are designing the pilot, not while you are trying to interpret one that has already finished.
- Baseline design: "I am about to pilot an AI tool for [task] in my [type of business]. Tell me exactly what I should measure this week, before I start, and how to capture it in under fifteen minutes."
- Criterion writing: "Here is my draft success criterion: [paste it]. Tell me whether it has one primary metric, a specific target, a deadline, and a baseline. Rewrite it so it has all four."
- Stress test: "Here is my pilot plan and my success criterion. Describe three ways this pilot could produce a misleading result, and what I would need to be tracking to notice each one."
- Decision rehearsal: "Here is my baseline, my target, and my week eight number. Argue the case for expanding, the case for modifying, and the case for cancelling, and tell me which parts of my evidence are too thin to support any of them."
Reflection
- Think of an AI tool you are paying for now. Could you state, in one sentence and one number, what it was supposed to improve and whether it did?
- What is the before number for the task you are most tempted to hand to AI next, and how long would it actually take you to measure it this week?
- If your current pilot fails, what would you have to admit, and to whom? That answer usually explains why the criterion was written vaguely.
- What date is in your calendar for the go/no-go decision? If there is no date, what has been serving as one?
Glossary
Baseline: The measurement of how the task performs before AI is introduced, recorded during setup and impossible to reconstruct once the tool is live.
Primary success criterion: The single metric, target, deadline, and baseline combination that decides whether the pilot succeeded.
Secondary criterion: A supporting measure such as adoption, correction rate, staff experience, or cost per unit of value, used to check whether a good headline result is real.
Failure log: A running list of workarounds, errors, and skipped uses during the active pilot, which separates fixable setup problems from a fundamental mismatch.
Decision rule: The sentence written before launch stating what you will do if the criterion is met and what you will do if it is not.
Go/no-go point: The fixed date at the end of the pilot when the decision is made, chosen in advance so it cannot be postponed by degrees.
Related Lessons
- Selecting the Right Pilot Project comes before this one: it covers how to choose a task that is worth putting through a timeline at all.
- Setting Baseline Metrics Before AI Adoption goes deeper on phase one, which is the step this lesson depends on most and the one most often skipped.
- Launch, Monitor, and Optimize Your Pilot covers the weekly running of the active pilot window in more detail.
- From Pilot to Production: Scaling What Works picks up at the go decision, when a successful pilot has to survive more people and less of your attention.
- When AI Isn't Working: Recognizing Negative ROI is the companion to the no-go decision, for tools that are costing more than they return.
Closing
Ekundayo's problem was never the tool. It might well have been saving him time, and he may have cancelled something that worked. What he lost was the ability to know either way, and he lost it in the first week, before the tool had done anything at all. The repair costs about an hour and fits on one page: a baseline measured before launch, one criterion with a number and a date, a weekly line in a note, and a decision rule written while you still have no feelings about the answer. Everything else in a pilot is optional.
Key Takeaways
- Measure your baseline before the pilot starts. You cannot know whether an AI tool improved something unless you documented how the work went before you added it, and that number cannot be recovered afterwards.
- Run a six-week active pilot inside an eight-week window. Two weeks of setup and baseline, six weeks of real use, and a decision at the end of week eight.
- Define one primary success criterion with a specific number, a deadline, and a baseline. "Saves me time" is not a criterion. "Down from 4.5 hours per week to 2.5 by week eight" is.
- Track the primary metric weekly, not monthly. Weekly tracking catches problems while there is still time to adjust before the decision point.
- Keep a failure log alongside the metric. Workarounds, errors, and skipped uses are what tell you whether a weak result is a setup problem or a bad fit.
- Write your go/no-go rule before you start. Deciding in advance what "not working" looks like is what stops the goalposts moving at week eight.
- Check the secondary criteria before expanding. Adoption, correction rate, staff experience, and cost per unit of value can each overturn a good headline number.
- An inconclusive result after eight weeks is a result. Insufficient evidence of value is a reason to stop; you do not need to see active failure first.
Frequently Asked Questions
What if I cannot measure the task cleanly? Most small business tasks resist clean measurement, and rough numbers are still enormously better than none. Time one representative week with a note on your phone and accept that it is approximate. The point of a baseline is not precision, it is having something specific to compare against so that week eight produces a comparison rather than an impression. An approximate before number and an approximate after number, measured the same way, will answer the question.
Can I run two pilots at once? Only if they touch different tasks and different people. Two AI tools aimed at the same workflow in the same eight weeks produce a combined result you cannot separate, and you will end up keeping both or dropping both for the wrong reasons. If they genuinely do not overlap, running them in parallel is fine, as long as each has its own baseline, its own criterion, and its own decision date.
The tool clearly helps, but my primary metric barely moved. Now what? That is a real and common outcome, and it usually means you picked the wrong primary metric rather than the wrong tool. Do not quietly swap the criterion at week eight. Note what you actually observed, then design a short second pilot with the metric that reflects it, and a fresh baseline. The discipline you are protecting is that criteria get set before results, not after.
What if the pilot succeeds but the price rises at full volume? Check the pricing tiers during setup rather than at week eight, and carry cost per unit of value as a secondary criterion. A tool that clears the time target while costing far more at production volume has not really passed, because the criterion was about value and price is half of that calculation.
Skill.re