Workflow Analysis: Finding AI Opportunities
Ramona Thibodeau, a process improvement analyst at the Nevada Department of Motor Vehicles, spent three weeks building an AI proposal she was proud of. Her idea: automate the customer complaint routing system. It was the first process that jumped out when her division director said, "Find us some AI use cases." Hundreds of complaints per week, lots of categories, lots of routing decisions. It seemed obvious. She was two days from presenting when a colleague asked one question: "What percentage of complaints require a supervisor to read the full case file before routing?" Ramona checked. It was 68 percent. The pitch collapsed. She had almost recommended automating a process that was mostly judgment calls.
That near-miss reshaped how Ramona worked. She stopped trusting first impressions and built a method instead. The method matters because not every task is suitable for AI. Some are better served by plain process improvement, some by simple rule-based automation, and some by leaving the humans alone. The work is telling those apart before you spend the money, and doing it in a way you can defend afterwards.
The Building Inspector Method
A building inspector does not walk up to a house and eyeball it. The inspector goes room by room, item by item, with a standard checklist. Plumbing fittings. Electrical panel. Load-bearing wall clearances. Each item either passes or it does not, and the inspector's personal opinion of the house is irrelevant to the report.
Workflow analysis for AI works the same way. You do not brainstorm or hold a whiteboard session asking "what could AI help with?" That produces a list of wishes ranked by whoever spoke most confidently. Instead you go process by process, step by step, with a documented checklist: systematic, repeatable, and defensible in a budget hearing or an Inspector General review. Systematic mapping would have surfaced that 68 percent figure on day one, which is the entire argument for doing it before the enthusiasm rather than after.
Mapping a Workflow Step by Step
Start with a single process, end to end. For each step, document six things.
- Who performs this step? Name the role, not the person. "Eligibility Technician II," not "Maria."
- How long does it take? Measure in minutes. Observe it. Average over at least 30 instances.
- What inputs does this step require? Paper forms? Database records? A phone call?
- What decision or output does this step produce? An approval, a routing decision, a letter?
- What can go wrong? List the three most common failure modes: missing documents, system timeouts, ambiguous answers. Then ask the sharper version: what would actually happen if this step were skipped or done wrong, and who would notice?
- How often is a judgment call required that the written rules do not cover? Estimate the percentage of cases requiring discretion beyond the written procedure. This is the most important question on the list.
The visual tool for this is a swim-lane diagram, a process map where each row is one role and arrows show handoffs between roles. Handoffs are where delays accumulate and where AI can sometimes replace a manual step without disrupting the rest of the process. Flowcharts, timelines, and written descriptions all work; what matters is that the artifact is explicit enough that someone who does not do the job can read it and ask a hard question. Ramona built swim-lane diagrams for five DMV processes in two weeks, and the maps immediately showed her which steps were isolated enough to automate.
Six Characteristics That Make a Task AI-Suitable
Once you have mapped a workflow, evaluate each step against six characteristics: the inspector's checklist, where each item passes, partially passes, or fails.
Repetitive
The same type of task is performed dozens or hundreds of times per day, so a model has patterns to learn and the savings recur. A step performed three times per week is rarely worth automating, because the implementation cost will never pay off no matter how well the model works.
Data-Heavy
The step processes large amounts of structured data: pulling fields from forms, cross-checking records across databases, comparing values against thresholds. This is the kind of work where volume favors the machine. Tasks that rest on unstructured narrative text carry higher risk, because the thing being judged is harder to define and harder to test.
Rule-Based with Exceptions
Most instances follow clear written rules while a minority, under 20 percent, require human judgment. AI can handle the general cases and flag the edge cases, but only if the flagging works: exceptions must route to a human reviewer without breaking the process. If the exception rate exceeds 30 percent, the task is not AI-suitable yet, and the right project is codifying the rules rather than buying a model.
High Volume
Volume does two things at once. It multiplies small efficiency gains into meaningful ones, and it supplies enough historical examples to train and test a model, at minimum 1,000 labeled examples and ideally 10,000 or more. A labeled example is a completed past case where you know both the inputs and the correct outcome.
Well-Defined Success Criteria
You can measure whether the AI got it right. "Approved" or "denied" is measurable. "Handled appropriately" is not. Write a concrete test before you build anything: here are 200 cases, here are the correct answers, here is the score. Anything less makes accountability impossible, and an accountability gap in a government process is not a technical debt you can carry quietly.
Trainable from Existing Data
Historical records of past decisions exist, are accessible, and are usable. This is where agencies hit a wall. Records may sit in a legacy system built in 2003, in scanned PDFs on a shared drive, or in paper files in a warehouse. "Exists" is not enough. Data must be digitized, labeled, and legally cleared for AI training, which you confirm with your privacy officer and against your records retention schedule before you promise anyone a timeline.
Counter-Signs: When to Stop Reading the Checklist
Five warning signs cut the other way, and any one of them should slow a candidate down regardless of how well it scores elsewhere. A task that requires real-time adaptation, where the context changes constantly and the right answer this month is not the right answer next month, gives a model nothing stable to learn. A task that requires nuanced human judgment about subtle social, ethical, or contextual factors invites the model to oversimplify in ways that are hard to see in aggregate metrics.
A task that involves physically interacting with the world is outside the scope of the software you are buying, however much the demo suggests otherwise. A task that requires understanding rare events is a data problem that no amount of modeling fixes: if you have ten examples of the thing you need to detect, the model has ten examples too. And a task where an error would cause serious harm demands accuracy so high it may not be reachable, which means the question is not whether the model is good but whether any achievable model is good enough.
What Good and Bad Fits Look Like
Two DMV processes Ramona mapped scored well on all six characteristics. Address verification during license renewal runs at approximately 900 transactions per day, draws on two structured databases, follows a clear rules hierarchy, and has 14 years of historical records behind it. Renewal eligibility checks, verifying no outstanding suspensions, no unpaid fines above the threshold, and valid insurance on file, follow the same pattern. Both are strong candidates, and neither is glamorous, which is usually a good sign.
Two processes she initially liked failed quickly. A child welfare risk assessment tool failed on three characteristics at once: exceptions exceeded 40 percent of cases, success criteria were contested among caseworkers who did not agree on what a correct assessment looked like, and errors, a missed abuse indicator or an unwarranted family separation, were irreversible. Complex zoning variance decisions failed similarly, because they require contextual judgment that no structured dataset captures. Note that both failures were visible from the map. Neither required a pilot to discover.
The AI Opportunity Matrix
After scoring candidate tasks, plot them on a two-by-two grid: the AI Opportunity Matrix. In the version used here the horizontal axis is feasibility, meaning data readiness, rule clarity, and exception rate, and the vertical axis is impact, meaning staff time saved, money saved, and citizen burden reduced. Different templates assign the axes differently, so label yours explicitly on the slide; the quadrant reasoning is what carries the meaning, not the orientation.
The high-impact, high-feasibility quadrant is where you start, and starting anywhere else is how first AI budgets get spent badly. High impact with low feasibility is your investment backlog: tasks that matter but need data cleanup or rule clarification first. Do not pursue them yet, and do not pretend the feasibility problem will resolve itself during procurement. High feasibility with low impact is where agencies often spend their first AI budget on easy wins that disappoint everyone at the review. Low on both is not a project at all.
Ramona's matrix put address verification and renewal eligibility in the high-impact, high-feasibility corner. Complaint routing, the idea she had almost pitched, landed in the backlog quadrant: high potential impact, but feasibility was low because of the 68 percent exception rate. She placed it there with a note that routing rules need to be codified before AI can touch it, which turned a killed proposal into a defined piece of future work.
Evaluating a Single Candidate
For each candidate that survives the matrix, work through a fixed set of questions. What exactly is the task, stated specifically enough that two people would scope it the same way? What is the volume, weekly or monthly? What is the current cost in staff time, in hours per week? What are the potential savings, remembering that partial automation is the normal outcome: if AI handles half the task, what time is actually freed? What is the implementation cost to build, train, and deploy? What is the payback period? What is the risk if the system fails, expressed as what happens to a person? And what is the human oversight cost, which is the question most proposals forget entirely.
Payback is simple arithmetic and worth doing explicitly rather than by feel. If a system costs $100k and saves $50k annually, the payback is two years. Whether two years is acceptable depends on your funding cycle and on how confident you are in the savings estimate, and both of those belong in the write-up next to the number. A payback figure with no stated implementation cost behind it is not a finding; it is a hope, and reviewers who have seen a few of these will treat it accordingly.
Worked Example: Permit Applications
A city planning department receives 500 permit applications per month. The process runs in five steps: intake, checking application completeness, at 2 hours per application; initial review, determining whether the application meets basic requirements, at 1.5 hours; specialist routing, sending the file to the right specialist by permit type, at 0.5 hours; detailed review by the specialist, which varies widely from 10 to 40 hours; and the final decision by city staff, at 1 hour.
The analysis follows the checklist. Steps one through three are repetitive and rule-based, and are the plausible candidates. Step four demands significant human judgment and is much less suitable. Step five is a decision by an accountable official and is not a candidate at all. Ranking within the first three follows the hours: completeness checking is the largest, at 2 hours per application, and has clear rules about what makes an application complete, so it is high feasibility and high impact and goes first. Basic requirement checking at 1.5 hours is next. Specialist assignment at 0.5 hours is medium on both counts, because the assignment rules have to be learned, and is worth pursuing only if the first two succeed.
Compute your own totals from the per-application hours and your own volume rather than importing a headline number. The totals quoted in the underlying course materials for these three steps do not reconcile with an annual period at 500 applications a month; they reconcile as monthly figures. That is exactly the kind of error that survives into a budget submission unchallenged, because a large hours-saved number is the part everyone wants to believe. Multiply it yourself, state the period, and show the inputs so the next reader can check you.
Worked Example: Compliance Monitoring
An environmental agency monitors whether facilities comply with regulations. Data arrives from self-reports and inspections through manual entry, staff review that data weekly to identify non-compliance, and staff then write custom non-compliance notices. Three steps, three different verdicts. Improving data entry accuracy with AI is possible, but the value sits downstream. Flagging non-compliance patterns automatically would save analysis time and is the real opportunity. Drafting notices is feasible with humans reviewing every draft before it goes out, since a notice is a legal communication to a regulated party.
The current process consumes 20 hours per week of staff analysis time, and implementation cost is low to medium, essentially connections to the data sources plus the analysis rules. Be careful with the savings estimate here, because the tempting version conflates two different quantities. A system that flags 80 percent of non-compliance cases does not remove 80 percent of the analysis time. Someone still reviews the flags, someone still handles the 20 percent the system missed, and the missed cases are found by the same manual scan the system was supposed to replace. Estimate the time saved by timing the reduced task, not by borrowing the detection rate.
Red-Flag Patterns to Avoid
Three patterns consistently produce failed government AI projects. First, high-stakes individual decisions with limited data: if the output determines whether someone receives benefits or faces enforcement action and you cannot validate the model rigorously, do not automate it. Second, tasks requiring cultural or contextual judgment, such as a caseworker reading a family situation or a planner weighing neighborhood character, where the operative knowledge is not in any database. Third, tasks where errors are irreversible: a benefits denial that triggers homelessness, an enforcement flag that damages employment. These require human decision authority, with AI in a support role only.
Documenting the Business Case
Each candidate use case needs a number before it goes to leadership. Calculate time savings in staff-hours per week, then multiply by the fully loaded cost per FTE, meaning salary, benefits, and overhead, typically 1.35 times base salary for Nevada state employees. A renewal eligibility check saving 4 minutes per transaction, at 600 transactions per day, recovers 40 staff-hours daily. At $42 per hour fully loaded for a DMV Technician I, that is $1,680 per day, roughly $420,000 per year in recaptured capacity.
Two cautions about that figure. Recaptured capacity is not the same as budget savings unless positions actually change, and a reviewer will ask which one you mean, so say so in the write-up. And every step of the calculation should be visible on the page: minutes saved, transactions per day, loaded hourly rate. Add error rate reduction and citizen wait time improvement where you can measure them and leave them out where you cannot. These numbers anchor your budget justification and give procurement a benchmark to write into the vendor contract, which is a second reason to keep them honest.
The Two-Week Workflow Sprint
Ramona ran her analysis as a structured two-week sprint with three to five staff, mixing process owners, frontline workers, and one IT representative. Week one was structured interviews plus direct observation of at least 30 live transactions per process, because what people describe and what people do diverge in predictable directions. Week two was swim-lane diagram review, six-characteristic scoring, and matrix placement.
The output is a ranked list of five candidate use cases, each with a matrix score, a data readiness assessment, and a time-savings calculation. The ranking goes to the division director with a one-page summary per candidate, enough to brief the budget office and start a procurement conversation. Documented evidence, room by room, just as the inspector's checklist produces. The value of the format is that a skeptical reader can retrace every judgment, which is what makes the recommendation survive contact with a budget hearing.
Anti-Patterns
- The obvious use case. The process that jumps out when someone asks for AI ideas is the one nobody has measured. Ramona's complaint routing looked ideal from every angle except the one that mattered. Map before you pitch.
- Brainstorming instead of inspecting. A whiteboard session produces a list ranked by enthusiasm and seniority. A checklist produces a list ranked by evidence, and only one of those survives an Inspector General review.
- Skipping the exception-rate question. It is the slowest question to answer, because someone has to sample real cases, and it is the one that most often kills a proposal. That is precisely why it gets deferred, and deferring it moves the discovery from the mapping week to somewhere deep in the build.
- Borrowing the detection rate as the savings rate. "The model catches 80 percent, so we save 80 percent of the time" ignores review time and the manual scan still needed for the remainder. Time the reduced task instead.
- A payback figure with no cost behind it. Any payback period rests on an implementation cost; if that number is missing from the proposal, the payback is decoration. Show both or neither.
- Counting recaptured hours as budget savings. Hours freed are real, but they become money only if positions or contracts change. Say which one you are claiming before someone else asks.
- Trusting "the data exists." Records in a 2003 legacy system, in scanned PDFs, or in a warehouse are not training data until they are digitized, labeled, and cleared for this use by your privacy officer against the retention schedule.
- Starting in the easy-wins quadrant. High feasibility with low impact produces a working system nobody cares about, and it spends the credibility you needed for the project that mattered.
Practice Prompts
- Pick one workflow in your agency and map it: what are the steps, who performs each, how long does each take, what inputs does it need, what output does it produce, and what happens if it is skipped or done wrong? Document it as a swim-lane diagram.
- For that workflow, score each step against the six characteristics and against the five counter-signs. Which steps pass? Which fail, and on which specific item?
- Estimate the exception rate for the step you think is the best candidate, by sampling real cases rather than asking someone's impression. Compare your estimate with what the process owner told you.
- Build the two-by-two matrix, label your axes explicitly, and plot every step from your map. Which land in the start-here quadrant, and which belong in the investment backlog with a note about what would move them?
- For your highest-priority candidate, answer the full evaluation set: volume, current staff hours, realistic savings, implementation cost, payback period, risk if it fails, and human oversight cost.
- Compare two candidates from your agency and write the paragraph explaining which you would pursue first and why, in language a budget office will read. Then have a colleague who did not do the mapping try to poke a hole in it.
Reflection
Think about a process you would have nominated for AI a month ago, before reading any of this. Now ask the question Ramona's colleague asked: what percentage of cases require someone to read the full file before deciding? If you do not know, that is the finding. Then ask a second question that is easier to avoid: if this process were automated and it went wrong, who would be harmed, how would they find out, and how would they get it undone? Candidates that survive both questions are worth a sprint. Candidates that fail either one have just saved you three weeks.
Glossary
- Swim-lane diagram: A process map in which each row is a role and arrows show handoffs between roles, making delays and isolated steps visible.
- Exception rate: The share of cases requiring human discretion beyond the written procedure. The single most predictive number in this analysis.
- Labeled example: A completed past case where both the inputs and the correct outcome are known, and therefore usable for training or testing.
- AI Opportunity Matrix: A two-by-two grid plotting candidate tasks by impact and feasibility, used to rank candidates and to justify the ranking.
- Feasibility: In this matrix, the combination of data readiness, rule clarity, and exception rate. It is about whether the thing can be built and validated, not whether it would be useful.
- Fully loaded cost per FTE: Salary plus benefits plus overhead for a full-time equivalent position, the multiplier used to convert staff-hours into dollars.
- Payback period: Implementation cost divided by annual savings. Meaningless without an explicit implementation cost.
- Counter-sign: A characteristic that argues against AI suitability regardless of how well a task scores on the positive checklist.
Related Lessons
- Building an AI Use Case: From Idea to Business Case is the next step for whatever survives your matrix.
- Requirements Gathering for AI turns a scored candidate into something a vendor can build against.
- Risk Classification: Safety-Impacting vs. Rights-Impacting is where the red-flag patterns become formal obligations.
- Data Quality and AI Performance explains what "trainable from existing data" really demands.
- How AI Projects Differ from Traditional IT covers why the usual project intake process misses these questions.
- Building the AI Business Case develops the financial justification beyond the single time-savings figure.
- Testing and Validating AI Systems is how you meet the well-defined success criteria you wrote down.
Closing
The building inspector does not skip the electrical panel because the house looks fine from the outside. You do not skip the exception rate question because the process seems simple. Every element of this method exists to replace an impression with a measurement, and the payoff is not just better project selection. It is that when someone asks in a hearing why you chose this process and not that one, the answer is a document rather than a recollection.
Ramona's proposal died two days before she was due to present it, and that was the good outcome. The bad outcome is the same proposal surviving the presentation, winning funding, and failing eighteen months later in front of an audience that includes the people it was supposed to serve. A method that kills your favorite idea early is doing exactly what you built it for.
Key Takeaways
- Intuition-based AI opportunity identification fails. The obvious use case often collapses when you map it step by step and measure the exception rate.
- Swim-lane diagrams surface the truth. Document every step, every role, every handoff, and every failure mode before evaluating any task for AI suitability, including what happens if a step is skipped.
- The six-characteristic checklist is the core screen. Repetitive, data-heavy, rule-based with exceptions, high volume, well-defined success criteria, and trainable from existing data; a task should pass most of these before advancing.
- Five counter-signs override a good score. Real-time adaptation, nuanced human judgment, physical interaction, rare events, and serious harm from error each argue against automation on their own.
- Exception rate is the most important single number. Tasks where more than 30 percent of cases require human discretion are not ready, and the fix is codifying rules rather than buying a model.
- The AI Opportunity Matrix prevents misallocated budgets. Plot candidates on impact against feasibility, label your axes, start in the high-impact high-feasibility corner, and park high-impact low-feasibility work in a named backlog.
- High-stakes, irreversible decisions require a human in the loop. Benefits denials, enforcement actions, and child welfare determinations carry equity and legal obligations that preclude full automation.
- Evaluate candidates on the full set of questions. Volume, current cost, realistic savings, implementation cost, payback, failure risk, and human oversight cost; a payback figure without an implementation cost is not evidence.
- Do your own arithmetic and show it. Staff-hours recovered times fully loaded FTE cost, with the period stated and the inputs visible, survives a budget hearing, an IG review, and a FOIA request. Borrowed totals do not.
- A two-week sprint produces a defensible ranked list. Three to five staff, structured interviews, direct observation of live transactions, and documented scoring; the output is evidence, not opinion.
Frequently Asked Questions
What if leadership has already picked the use case?
Map it anyway, and map one or two alternatives alongside it. The analysis takes two weeks and produces a document, which changes the conversation from a disagreement about opinions into a disagreement about evidence. If the chosen process scores well, you have strengthened the proposal at low cost. If it scores badly, you have the exception rate, the data readiness assessment, and a named alternative ready, which is a far better position than objecting on instinct.
How do I estimate the exception rate without a long study?
Sample real cases and count. Pull a set of recently completed cases, and for each one determine whether the written procedure alone determined the outcome or whether someone exercised discretion. The number will be uncomfortable and it will differ from what the process owner estimates, which is itself useful information. Sampling is not a research project; it is an afternoon, and it is the cheapest way to avoid Ramona's three wasted weeks.
Is a task with a high exception rate permanently unsuitable?
No, and framing it that way loses the opportunity. A high exception rate usually means the rules were never written down, not that the work is inherently unruly. Codifying those rules is a genuine improvement project with value whether or not AI ever follows it, and it moves the candidate along the feasibility axis. Park it in the backlog with a specific note about what would have to change, exactly as Ramona did with complaint routing.
Our data is in paper files. Is that a stopper?
It is a cost, not necessarily a stopper, but it has to be costed before the AI project is approved rather than discovered afterwards. Digitization, labeling, and clearance for this specific use are three separate pieces of work with three separate owners, and the third one is not yours to decide: your privacy officer and your records retention schedule govern whether the data may be used for AI training at all. Get that answer early, because a negative answer ends the project regardless of how good the digitization plan is.
What is the smallest useful version of this method?
One process, six questions per step, and an honest exception rate. You can do that in a few days without a sprint team, and it will separate serious candidates from wishes. The full two-week sprint earns its cost when you need a ranked list across several processes and a document that will be read by a budget office, an Inspector General, or a procurement officer, each of whom will want to see how you got there.
Skill.re