Innovation Metrics: Measuring What Matters
Andile, who was born in Johannesburg, now runs a two location auto tinting shop in Charlotte. He spent $800 last year on three different AI experiments: a chatbot for lead capture, an AI scheduling tool, and a ChatGPT subscription for writing service descriptions. He could not tell you whether any of them made a difference. He measured nothing, so he learned nothing, and this year, with no data to guide him, he is facing exactly the same guess about where to put his next $800. The money is not the painful part. The painful part is that a year has passed and he knows no more about what works in his business than he did at the start of it.
The Problem With Measuring Innovation
Most business owners track revenue, costs, and customer counts. These are the right metrics for running the business, and there is nothing wrong with any of them. They are simply the wrong metrics for evaluating innovation experiments, and using them for that purpose is why so many owners end up in Andile's position, certain that they tried things and unable to say what happened. The mistake is understandable, because these are the numbers already on the screen and the ones the accountant asks about.
Revenue does not move because you tried a new chatbot. It might move because the chatbot captured 15 more leads per month, which converted at your normal rate, which added revenue over six months. The chatbot is three steps upstream of revenue, and every step adds delay and noise from things that have nothing to do with the experiment. If you only measure revenue, you will never know whether the chatbot worked, because by the time revenue moves, a season has turned, a competitor has opened, and you changed two other things.
Innovation metrics live closer to the experiment. They measure what the experiment was designed to do, not what it eventually might trickle down to. That is a lower bar than proving a return, and deliberately so. The goal in a ninety day window is not to prove that a tool made you money. It is to learn enough to make the next decision better than a guess.
Measure the cause, not just the effect. Revenue is an effect. Lead conversion rate is a cause.
The Three Types of Innovation Metrics
Type 1: Activity metrics, or did we actually try something?
These confirm that the experiment happened. They are not very interesting on their own, and no owner has ever been excited by a count of tools tried, but they establish that you have something to evaluate at all. They also catch a failure mode that no amount of sophisticated measurement will surface, which is the experiment that was decided on, paid for, and never actually run. Track them at the quarter level and keep the list short:
- Number of AI experiments launched this quarter
- Number of team members who used a new AI tool at least once per week
- Time spent testing new tools, in hours per month
If your activity metrics are zero, meaning you keep planning to try things and never do, that is important information rather than an embarrassment. It means the bottleneck is starting, not measuring, and no scorecard will fix it. Activity metrics also catch the quieter version of the same problem: a tool that was purchased, configured, and then used by nobody. An experiment nobody ran has no result to interpret, and treating it as a failed experiment teaches you the wrong lesson about the tool.
Type 2: Output metrics, or did the experiment produce what it was supposed to?
These measure the immediate product of the experiment, meaning what the AI tool actually generated or changed. They are usually available within days rather than months, which makes them the metric you watch while the experiment is still running and the one that tells you whether to adjust the setup. An output metric should be countable without judgment: someone should be able to produce the number from a report or a log without deciding what qualifies.
- Number of leads captured by the AI chatbot per month
- Average time to schedule an appointment, before and after the AI scheduling tool
- Number of service descriptions rewritten using AI, compared with those left alone
For Andile, the chatbot output metric would have been straightforward: how many website visitors started a conversation with the chatbot, and how many of those conversations ended with a contact form submitted or a phone number collected. Neither number requires special software, and both were available to him the entire time. The reason he did not have them is that nobody wrote down, before launch, what the chatbot was supposed to produce.
Type 3: Outcome metrics, or did it move a business number?
These connect the experiment's output to a real business result. They take longer to show up, and they are the ones that tell you whether the experiment was worth doing rather than merely whether it functioned. Each outcome metric should already exist in your business before the experiment starts, because a measure invented for the experiment has no history behind it, and a number with no before value cannot tell you that anything changed.
- Leads captured by the chatbot that converted to paying jobs, as a conversion rate
- Average appointment no show rate, before and after AI scheduling reminders
- Customer satisfaction score for jobs where AI written descriptions were used
The relationship between output metrics and outcome metrics is your most valuable learning, and it is the reason you measure both rather than picking one. If the chatbot captured 30 leads per month, which is a good output, but only 1 converted to a booking, which is a terrible outcome, the chatbot is not the problem. Your follow up process is. You would not have known that from either number alone, and you would certainly not have known it from revenue.
A Simple Innovation Scorecard
You do not need complex software for any of this. A single shared spreadsheet, with one row per experiment, is enough, and the discipline lives in the columns rather than in the tool. What makes it work is that every experiment ends up on the same page, in the same shape, so that a year later you are reading a comparable set of rows rather than scattered notes in different places. These are the columns worth having:
- Experiment name
- What it was designed to do, in one sentence
- Start date and end date
- Monthly cost
- Output metric, what it produced, with before and after numbers
- Outcome metric, what changed in the business, with before and after numbers
- Verdict: keep, modify, or kill
The two before columns are the ones people skip, and skipping them is fatal. A before number has to be recorded before the experiment starts, because afterwards it is unrecoverable and you will be reduced to comparing a real measurement against a memory. If the honest answer is that you do not know your current no show rate, that is your first week of work, and it is worth doing even if the experiment never happens.
If Andile had filled out this scorecard for each of his three experiments, he would know which one to keep, which to adjust, and which to cancel. Instead he renewed all three, because he felt like they were probably helping. That sentence is what an unmeasured year sounds like from the inside. It is not carelessness; it is the only conclusion available to someone with no data, and it costs him the same money again this year.
Choosing What to Measure Before You Launch
The three types are only useful if you pick your specific metrics before the tool goes live. The way to do that is to finish this sentence about the experiment: this tool is supposed to produce more of, or less of, one particular thing. Whatever fills that blank is your output metric. Then ask what that thing is supposed to change about the business, and that answer is your outcome metric. If either question is hard to answer, you have found a problem with the experiment rather than with the measurement.
Be suspicious of any metric that requires an interpretation to produce. If someone has to judge whether a conversation counted, the number will drift as the person doing the counting gets more or less generous. Prefer measures your existing systems already record, which for most small businesses means the appointment book, the point of sale, the inbox, and the website analytics you have never opened. Those sources are unglamorous, but they have history behind them, and history is what turns a measurement into a comparison.
Finally, decide who reads the number and when. A metric with no scheduled reader is not a measurement, it is a stored file. Put the month two output check and the month three verdict in a calendar, with a name attached, at the same moment you set the experiment up. It takes a moment then and it is the difference between a cycle that ends in a decision and one that ends in a renewal notice.
The 90-Day Cycle
Innovation works best in 90 day cycles. That is long enough to collect meaningful data and short enough to pivot before you waste money on something that is not working, and it maps neatly onto three months of ordinary business rhythm rather than requiring a special calendar. The point of the cycle is not the length. It is that the ending is fixed in advance, so the decision arrives on a date rather than whenever you happen to notice the charge. Structure each cycle this way:
- Month 1: launch the experiment, and set up tracking from day one rather than as an afterthought.
- Month 2: check the output metrics. Is the tool producing what it was supposed to produce? If not, adjust the setup, not the measurement.
- Month 3: check the outcome metrics. Did the output produce a business result? Make the keep, modify, or kill call.
The instruction to adjust the setup rather than the measurement is the one that gets violated most often. When month two shows a disappointing output, the tempting move is to redefine what you were counting until the number looks acceptable. That converts an experiment into a justification. If the chatbot is not capturing conversations, change where it appears on the site or what it says first, and leave the definition of a captured conversation exactly where you set it.
Most business owners treat AI experiments as open ended subscriptions, which is why they accumulate. A 90 day cycle forces a decision, and the deadline is doing real work: it is much easier to cancel a tool at a date you set in advance than to cancel one that has quietly become part of the furniture. If you cannot make a verdict after 90 days, you do not have enough data, which means you did not set up the right metrics at the start. That is a finding about your measurement, not a reason to extend.
What Counts as a Successful Experiment?
Not every experiment needs to produce a financial win to be valuable. An experiment that clearly fails, in a way that lets you see why it failed, is more valuable than an experiment where you cannot tell either way, because the first one narrows the field for next time and the second one leaves you exactly where you started.
A useful reframing: any experiment that produces a clear verdict is a success. "This chatbot did not capture enough leads to justify what it costs" is a success. "We are not sure whether the scheduling tool helped" is a failure, and it is a failure of measurement rather than a failure of the tool. The distinction matters because the two call for completely different responses. One tells you something about chatbots in your business. The other tells you to fix how you set experiments up.
Set your decision threshold before you start. Something like "we will keep this tool if it generates at least 10 booked appointments per month that we can trace back to it" turns the end of the cycle into arithmetic rather than a debate with yourself. Then measure. Then decide. A threshold written in advance is the only reliable protection against the retrospective reasoning that renewed all three of Andile's subscriptions.
Anti-Patterns
Measuring revenue and calling it evaluation. Revenue is several steps downstream of anything an AI tool does, and it moves for reasons that have nothing to do with your experiment. By the time it responds, the season has changed and you have altered other things. Use output and outcome metrics that sit close enough to the experiment to be attributable.
Setting up tracking after launch. The before number cannot be reconstructed once the experiment is running. If you launch in week one and start measuring in week three, you have permanently lost the comparison, and everything that follows is an argument rather than a result.
Moving the threshold once you can see the numbers. A verdict criterion invented at the end of the cycle will always be met, because you will pick one that the data satisfies. Write it down before launch, in the scorecard row, where you will have to look at it again.
Running several experiments at once in the same part of the business. Launch a chatbot and a scheduling tool in the same month and the no show rate becomes unattributable. Stagger the experiments, or accept in advance that you will only be able to read the combined result.
Renewing because it feels helpful. Andile's three subscriptions all survived on a feeling. A feeling is a legitimate reason to run an experiment and never a legitimate reason to conclude one. If the scorecard row is empty at day 90, the honest verdict is not keep; it is that you failed to measure, and the tool has not earned another year on the strength of that.
Practice Prompts
Use these with an AI assistant while you are setting up an experiment, not after it has finished.
- Metric design: "I am about to try [tool] in my [type of business] to solve [problem]. Suggest one activity metric, one output metric, and one outcome metric, and for each one tell me exactly where the number would come from."
- Before numbers: "Here is what I am about to change. List the measurements I need to record before I start, and tell me which of them become impossible to recover once the experiment is running."
- Threshold setting: "Here is what this tool costs per month and here is what it is supposed to produce. Help me write a single keep or kill threshold I could evaluate at day 90 without arguing with myself."
- Verdict review: "Here is my scorecard row with before and after numbers for output and outcome. Argue both sides of the keep, modify, or kill decision, and tell me which numbers are too weak to support either case."
Reflection
- List every AI tool you are currently paying for. For how many of them could you state what it was supposed to produce and whether it did?
- Which subscription have you renewed without evaluating it, and what would you need to have recorded three months ago to evaluate it now?
- What is your current no show rate, lead conversion rate, or equivalent baseline? If you do not know it, what is stopping you from measuring it this week?
- Think of an experiment you would call a failure. Do you know why it failed, or only that it did not obviously work?
Glossary
Activity metric: A count confirming that an experiment actually happened, such as experiments launched, people using a tool weekly, or hours spent testing.
Output metric: A measure of what the experiment directly produced, such as conversations captured or time to schedule an appointment, available within days of launch.
Outcome metric: A measure of the business result the output was supposed to cause, such as a conversion rate or a no show rate, which takes longer to appear.
Before number: The baseline measurement recorded prior to launch, without which the after number cannot be interpreted and cannot be reconstructed later.
Verdict: The keep, modify, or kill decision made at the end of a cycle, and the actual deliverable of an innovation experiment.
Decision threshold: The result you commit to in advance as the condition for keeping a tool, written before launch so that it cannot be adjusted to fit the data.
Related Lessons
- Setting Baseline Metrics Before AI Adoption covers the before numbers in detail, which is the step this lesson depends on most.
- Defining Leading and Lagging AI Metrics develops the distinction between measures that sit close to the cause and measures that sit close to the effect.
- When AI Isn't Working: Recognizing Negative ROI is the companion to the kill verdict, for tools that are actively costing more than they return.
- Rapid Prototyping and Experimentation Frameworks covers how to design the experiment that this scorecard then evaluates.
- Building Your AI Impact Dashboard is where these scorecard rows go once you are running enough experiments to need a view across them.
Closing
Andile's problem was never the $800. Plenty of businesses spend that on an experiment and learn something worth several times as much. His problem was that the money bought three tools and no knowledge, so the second year starts exactly where the first one did. The fix is unglamorous and takes an afternoon: one shared sheet, one row per experiment, the before numbers recorded, the threshold written down, and a date ninety days out when someone has to say keep, modify, or kill out loud.
Key Takeaways
- Revenue is too far downstream to evaluate an AI experiment. Use output and outcome metrics, which sit close enough to the change to be attributable to it.
- Three types of metric: activity, meaning did we try it; output, meaning did it produce what it should; and outcome, meaning did it move a business number.
- The relationship between output and outcome is the real learning. Strong output with weak outcome points at a process downstream of the tool, not at the tool.
- A single shared spreadsheet is enough. Complexity in the tracking tool does not improve the quality of the decision; recording the before numbers does.
- Use 90 day cycles: launch in month one, check outputs in month two, and make the keep, modify, or kill call in month three.
- Set your verdict threshold before you start, which is the only reliable defense against retrospective rationalization.
- A clear no is as valuable as a clear yes. What matters is a definitive verdict rather than an open ended subscription you never evaluate.
- If your activity metrics are zero, the bottleneck is starting, not measuring. Fix that first, because there is nothing to evaluate until something launches.
Frequently Asked Questions
What if I cannot attribute the outcome cleanly? Attribution gets harder the further downstream you go, which is exactly why the output metric exists. If you cannot cleanly connect bookings to the chatbot, you can still count how many conversations it captured and how many ended with a phone number, then look at that alongside your booking numbers. Partial attribution and an honest note about its limits beats no measurement.
Can I run more than one experiment at a time? Yes, provided they touch different parts of the business. Two tools affecting the same outcome metric in the same quarter produce a combined result you cannot separate. If you must overlap, decide in advance that you are evaluating the pair rather than each one, and write that in the scorecard row.
What if the tool is free? The subscription is rarely the main cost. Setup time, team attention, the data you move into it, and the switching cost when you leave are all real, and a free tool that produces nothing still occupies the slot where a paid one that works could have been. Run the same cycle, with the cost column reflecting time rather than money.
90 days feels short for something seasonal. Should I extend it? Extend the cycle if your business genuinely turns on a longer rhythm, but set the longer window in advance rather than discovering at day 90 that you need more time. The problem with extension is almost never the calendar; it is that it usually arrives as a way of postponing a verdict the data has already supported.
Who should own the scorecard? One named person, usually whoever approves the spend. Shared ownership of a measurement document reliably produces a document nobody updates. The person who signs the renewal is the person who should have to look at the empty cells before signing it again.
Skill.re