Selecting the Right Pilot Project
Blessing runs an independent bookstore in Asheville, North Carolina: twelve hundred square feet, one full-time employee, and herself. When she decided to try AI last October, she wanted to start big. Her plan was an AI-powered personalized recommendation engine that would email every customer a custom reading list based on their purchase history. She spent six weeks researching vendors, talking to a developer, and worrying about her customer database. Then she burned out and tried nothing at all. Meanwhile her competitor two miles away started using an AI tool to write Instagram captions and saved four hours a week in the first month. The competitor's pilot was small, boring, and immediately useful. Blessing's was ambitious, complex, and never launched.
Why Your First Pilot Determines Everything
Your first AI pilot sets the tone for everything that follows. Get it right and your team believes the effort is worth it, which makes the second project easier to start and the third one easier still. Get it wrong, or never launch it at all, and skepticism hardens into a settled opinion that AI is a distraction for businesses larger than yours. That opinion is expensive to shift later, and it is usually formed by one bad experience rather than by any considered judgment.
So the goal of a first pilot is not to solve your biggest problem. It is to produce a small, visible win that proves AI works in your specific business, on your actual work, with your actual constraints. Think of it like learning to swim. You do not start in the deep end. You start in the shallow end, get comfortable with the mechanics, and build toward the deep end from there. A pilot is the shallow end, and treating it as anything grander is how six weeks of vendor research turns into nothing.
The Four Criteria for a Good Pilot
A good first AI pilot scores well on four dimensions. Rate each candidate project from 1 to 5 on each one, and be honest about the low scores rather than generous, because a candidate that scores badly here will usually fail for exactly the reason the low score identified.
The four work together rather than independently, and the pattern to watch for is a candidate that scores at the top on one dimension and at the bottom on another. Customer pricing is highly repetitive and completely unsuitable as a first pilot, because a mistake there costs money and trust. A report you produce a few times a year might be perfectly safe to get wrong and still a poor choice, because you will have forgotten the whole exercise before it comes around again. A strong first pilot is unremarkable across all four rather than outstanding on any one.
1. Repetition
The task happens often. Daily or weekly is ideal. A task you do once a month is not worth automating first, because you will not accumulate enough repetitions to see a benefit while you still care about the experiment, and you will have forgotten how the workflow goes by the time it comes around again. Look for the things you do the same way, over and over, every single week.
2. Low stakes if wrong
If the AI makes a mistake, the consequences should be minor and easily fixed. A misworded social media draft? You edit it before posting. An incorrectly generated invoice? You catch it before it goes out. Compare that with an AI making errors in customer pricing or employee scheduling, where mistakes cost money and trust and are often discovered by the customer rather than by you. Start with tasks where a bad output is annoying rather than damaging.
3. Clear before and after
You need to be able to measure the difference between doing the task manually and doing it with AI help. Time saved per week. Number of drafts produced per hour. Fewer customer complaints. If you cannot measure it, you cannot learn from it, and you cannot make the case to yourself or to your team that the pilot worked. Vague benefit is how pilots end in a shrug and a subscription nobody cancels.
4. No new tools required
Your first pilot should use an AI tool you already have access to, or one with a free tier. Adding a new subscription, a new login, and a new learning curve on top of an already unfamiliar task multiplies the number of ways the pilot can stall. ChatGPT, Claude, or a similar conversational AI tool you can open today is enough for a first pilot. The point is to learn the process, not to evaluate vendors, and vendor evaluation is a project of its own that will happily consume six weeks.
Scoring Your Pilot Candidates
Start by listing five tasks you do every week that feel tedious or time-consuming. Write your own five first, then check them against the menu below, which is there to prompt your thinking rather than to be the answer. Common candidates for small businesses include:
- Writing replies to common customer emails, such as questions about hours, pricing and availability
- Drafting social media posts
- Summarizing meeting notes or phone calls
- Creating first drafts of job postings
- Writing product or service descriptions
- Generating follow-up emails after appointments
Then score each one from 1 to 5 on the four criteria. The highest-scoring task is your pilot. If two tie, take the one you personally do most often, because you will notice the difference faster and you will not have to interpret somebody else's account of how it went.
| Criterion | What a high score looks like | What a low score looks like |
|---|---|---|
| Repetition | The task recurs daily or weekly, in much the same shape each time | The task comes around monthly or less, or looks different every time |
| Low stakes if wrong | A bad output is caught and fixed before it leaves the business | An error reaches a customer, a price, or a schedule before anyone sees it |
| Clear before and after | You can state a number today and the same number in four weeks | The benefit would only ever be describable as "it feels easier" |
| No new tools required | Runs on a tool you already have open, or one with a free tier | Needs a new subscription, a new login, and a vendor decision first |
What Blessing's Pilot Looked Like
When Blessing ran this exercise honestly, the winner was clear, and it was nothing like her original plan. It was replying to the same ten customer email questions she got every week: store hours, event tickets, gift wrapping, and special orders. She spent about forty-five minutes a week on them. Repetitive, low stakes, easy to time, and possible with a tool she already had open. It scored well on all four criteria precisely because it was unglamorous.
She started a pilot where she pasted each incoming email into Claude and asked it to draft a friendly reply in her store's voice. Week one saved her twenty-two minutes. Week two saved thirty-one, as her prompt got sharper and she stopped rewriting drafts that were already fine. By week four she was reviewing and sending AI drafts in under a minute each. Nothing about that result would impress anyone at a conference, which is part of why it happened at all.
It is worth comparing the two projects directly, because they belonged to the same owner in the same month. The recommendation engine scored badly on every criterion: it ran rarely, it touched customer data, it had no measurable before state, and it needed a developer and a vendor decision before anything could be tried. The email replies scored well on all four. Blessing's instinct about which project mattered more to her business may well have been right. Her instinct about which one to do first was what cost her six weeks.
Set Your Success Bar Before You Start
Before you launch, write down exactly what success looks like. Make it specific and time-bound, and write it while you still have no emotional stake in the answer. A bad success criterion sounds like "the AI helps with emails." A good one sounds like this: after four weeks, I spend less than fifteen minutes per week on routine customer email replies, and the quality is at least as good as my manual replies.
Note that the good version carries a quality condition alongside the time condition. Without it, a pilot can pass on speed while quietly degrading the thing customers actually receive, and time saved on worse work is not a saving. Decide in advance what "at least as good" means to you, even if the standard is only that you would have been happy to send it yourself.
A four-week timeline works well for a first pilot. It is long enough to smooth out the learning curve, which is real and which will make week one look worse than the tool deserves, and short enough that you stay motivated and remember why you started. Set a calendar reminder for week four so the evaluation actually happens, and evaluate honestly when it does.
What to Document During the Pilot
Keep a simple log. A notes app on your phone is enough, and anything more elaborate tends to become a second project. Four things are worth capturing, and none of them takes more than a moment at the end of the task:
- Time spent on the task each week, before and after AI
- Quality issues, meaning outputs that needed significant editing before you could use them
- Prompts that worked, saved with the exact wording, whenever the AI gives you a result you would send as is
- Surprises, meaning things the AI did that you did not expect, good or bad
That log becomes the evidence you use at week four to decide whether to expand the pilot, change your approach, or move on to a different use case. It also does something less obvious: the saved prompts are the beginning of the process document you will need if anyone else ever runs this task, and the surprises are usually where you learn what the tool is genuinely good and bad at.
The best pilot is the one you actually finish, not the one that would impress a tech journalist.
What Makes a Pilot Fail
Scope creep. You start with email replies, and then decide to also try AI for social media, product descriptions and the newsletter, all inside the same four weeks. Now it is impossible to say what worked and what did not, the log has three shapes, and the week four evaluation collapses into a general impression. One pilot, one task.
Vague prompts. If you type "write me an email reply" and get a mediocre result, that is usually an incomplete instruction rather than a limit of the tool. Pilots fail when owners give up after the first bad output instead of spending ten minutes improving the prompt. Prompting is a skill you build during the pilot, and week one output is not a fair test of it. That said, some tasks genuinely sit outside what the tool does well, and your quality log is what tells the two apart: a specific instruction that still produces unusable output several times over is telling you something about the task, not about your wording.
No measurement. If you do not track time before and after, you cannot evaluate the pilot and will fall back on gut feel, which is unreliable in both directions. It flatters a tool you enjoyed using and it undersells one that saved real time without feeling dramatic. Track the numbers, even roughly.
After the Pilot: What Comes Next
If your four-week pilot succeeds, meaning time saved, quality maintained and the team comfortable, you have two options: embed the process into your normal workflow, or expand to a second pilot task. Do not try to do both at once. Embed first. That means writing down the prompt you use, making the steps repeatable, and making sure anyone else on your team can run it the way you do rather than approximately.
Embedding is more concrete than it sounds. It means the prompt lives somewhere other than your chat history, the steps are written in the order you actually do them, and the person who is not you can produce an acceptable output on their first attempt without asking you a question. Until that is true, the pilot has produced a personal habit rather than a business process, and personal habits disappear the week you are away.
Blessing eventually built her recommendation email system, but only after six months of smaller pilots had made her comfortable with AI tools. By then she knew how to write a prompt, how to evaluate an AI output, and which vendor she trusted enough to work with. The deep-end project worked because she had already learned to swim, and the six weeks she originally spent researching it had taught her almost nothing by comparison.
Anti-Patterns
Starting with the most impressive project. The recommendation engine was the most interesting idea Blessing had and the worst possible first pilot: infrequent, technically complex, dependent on customer data she was rightly nervous about, and impossible to evaluate quickly. Ambition is not the problem. Ambition first is.
Researching instead of piloting. Six weeks of vendor comparison feels like progress and produces no evidence about your own business. If a decision is taking more than an afternoon before anything has been tried, the pilot has been replaced by a procurement exercise.
Picking a task nobody can measure. "Being more creative" and "improving our brand voice" are not pilot candidates, because there is no week four number to look at. Choose something countable first, even if it is not the most valuable thing on your list.
Running three pilots at once because they all seem small. Individually each is manageable. Together they produce one confusing month, no clean comparison, and a strong chance that all three quietly stop.
Judging the tool on week one output. Week one measures your prompting, your setup, and your unfamiliarity with the whole thing. That is exactly why the pilot runs four weeks rather than one.
Skipping the quality half of the success bar. A pilot that halves your time while producing replies you would not have sent yourself has not succeeded, it has shifted the cost somewhere you are not looking, usually onto the customer.
Practice Prompts
Use these while you are choosing and setting up a pilot, before you have committed to a candidate.
- Candidate generation: "I run a [type of business] with [number] staff. Here are the tasks I do every week: [list them]. Which of these are repetitive, low stakes if the output is wrong, and easy to measure? Rank them and say why."
- Stakes check: "Here is the task I am considering for an AI pilot: [describe it]. Describe the worst realistic outcome if the AI gets it wrong and nobody notices before it leaves my business."
- Success bar: "Help me turn this into a specific, time-bound success criterion with both a time condition and a quality condition: [describe your goal]."
- Prompt improvement: "Here is a prompt I used and the output it produced: [paste both]. The output was not usable because [reason]. Rewrite the prompt and explain what you changed."
Reflection
- Which AI idea have you been researching rather than testing, and what is the smallest version of it you could run this week?
- List the tasks you repeated most often last week. Which of them would you not mind an AI getting slightly wrong?
- What number would tell you your pilot worked, and could you measure it today without building anything?
- Think of a tool you abandoned after one disappointing output. Was the tool wrong, or was the instruction incomplete, and do you have any record that would settle it?
Glossary
Pilot: A short, deliberately limited test of an AI tool on one recurring task, run to produce evidence rather than to solve the biggest problem in the business.
Success bar: The specific, time-bound statement of what would count as the pilot working, written before launch and covering quality as well as speed.
Pilot log: The running record of time spent, quality issues, prompts that worked and surprises, kept during the pilot and used as evidence at the evaluation point.
Scope creep: The addition of further tasks or tools during a pilot, which makes the result unattributable and is the most common reason first pilots teach nothing.
Free tier: A usable no-cost level of an AI product, which lets a first pilot proceed without a subscription decision or a vendor evaluation.
Embedding: Turning a successful pilot into normal practice by writing down the prompt and the steps so that anyone on the team can repeat it consistently.
Related Lessons
- Planning Your First AI Pilot Project frames the wider planning work this selection step sits inside.
- Building a Pilot Timeline and Success Criteria is the next step: phases, baselines, and the go or no-go decision date.
- Setting Baseline Metrics Before AI Adoption covers how to capture the before number that your success bar depends on.
- From Pilot to Production: Scaling What Works picks up after a successful pilot, when the process has to survive more people and less of your attention.
- Common Prompting Mistakes and How to Fix Them addresses the vague-prompt failure that ends more first pilots than any tool limitation does.
Closing
The gap between Blessing and her competitor was not budget, technical skill, or nerve. It was the size of the first step. One of them picked a task so small it was almost embarrassing to describe and had a result inside a month. The other picked something worth describing and had nothing at all. Pick the boring task, score it honestly on the four criteria, write down what winning looks like before you start, and keep a log you would be willing to show someone. The ambitious project stays on the list. It just goes second.
Key Takeaways
- A pilot's goal is a visible win, not a moonshot. A small, credible success builds the confidence and the skill to tackle bigger problems later.
- Score candidates on four criteria: repetition, low stakes if wrong, clear before and after, and no new tools required. Rate each from 1 to 5 and respect the low scores.
- Common winning pilots include drafting routine email replies, writing social media posts, and summarizing meeting notes: weekly tasks with clear quality standards.
- Use a tool you already have. A new subscription, a new login and a vendor decision on top of an unfamiliar task is how a first pilot stalls before it starts.
- Write your success bar before you start, with a time condition and a quality condition, and set a calendar reminder for the evaluation date.
- Keep a simple log of time spent, quality issues, prompts that worked and surprises. It is your evidence at week four and the seed of your process document.
- One pilot, one task. Scope creep is the most common reason first pilots fail to produce any clear learning.
- Embed before expanding. Document the process so anyone on your team can repeat it before you move to a second use case.
Frequently Asked Questions
What if the task I most want to automate is high stakes? Keep it on the list and do not make it your first pilot. High-stakes tasks need you to already know how the tool behaves on ordinary work, what its failure modes look like, and how to check its output quickly. Run a low-stakes pilot first to build that judgment, then approach the high-stakes task with a review step designed in from the start rather than added after something goes wrong.
How do I measure a task I have never timed? Time it for one week before you start, which is the whole of the setup work. Note the minutes on your phone each time you do the task, and accept that the number will be rough. A rough before number measured the same way as your rough after number will answer the question. A precise number you never collected will not.
My employee is skeptical. Should I include them in the pilot? Usually yes, provided the task is genuinely low stakes. A skeptic who watches from outside stays a skeptic and can tell themselves any story about the result. A skeptic who runs the tool for four weeks either changes their mind or produces the most useful failure log you will get, because they will notice every rough edge an enthusiast would have worked around without mentioning.
What if the pilot half works? That is the most common outcome and it is informative rather than disappointing. Look at the log: if the failures cluster in one type of request, narrow the scope of the workflow to the part that worked and embed that. A tool that reliably handles most of your recurring email types is a real result, as long as you have documented which ones it does not handle and what happens to those.
Do I need to tell customers that AI drafted a reply? That depends on your industry, your obligations and your own standards, and it is worth deciding deliberately rather than by default. What the pilot changes is only who writes the first draft; you remain the sender and you remain responsible for what goes out. If a reply would embarrass you if its origin were known, that is a signal about the reply rather than about disclosure policy.
Skill.re