Designing Multi-Step AI Pipelines
Lieselotte runs a seven-person boutique PR firm in Chicago. Her team handles local restaurant and retail clients: press releases, media lists, coverage reports. Every new client kicked off the same tedious sequence. Someone would gather the client's background info, someone else would research relevant media contacts, a third person would draft a pitch email, and then the account manager would compile everything into a briefing doc. The handoffs were messy and things fell through. The sequence took two to three days when it should have taken four hours. She had tried using AI for individual steps, asking ChatGPT to draft pitches, for example, but each step still required someone to manually pass output to the next person. Then she heard the phrase "AI pipeline" from another small agency owner and started asking what that actually meant in practice.
What a Pipeline Actually Is
An AI pipeline is a sequence of AI operations where the output of each step becomes the input of the next, with some processing or transformation happening between steps. Think of a factory assembly line. Raw material goes in one end, each station does one specific job, and the finished product comes out the other end. No station has to know what every other station does. Each just receives input, does its job, and passes output forward. That independence is what makes a pipeline debuggable later, when one station starts producing something the next one cannot use.
For Lieselotte's client onboarding, a pipeline might look like this:
- Client completes an intake form (raw input)
- AI summarizes the intake into a one-paragraph client brief (Step 1 output)
- AI uses the brief to generate a target media list query (Step 2 output)
- A database lookup retrieves matching journalists and outlets (Step 3 output)
- AI drafts a personalized pitch email for each journalist using the brief and their beat (Step 4 output)
- AI compiles everything into a formatted new-client briefing document (Final output)
Each step is simple on its own. The power comes from connecting them in sequence so they run without manual handoffs, which is also why the interesting design work is at the joins rather than inside any single step.
Trace one piece of information through that list and the design becomes visible. The client brief produced in the second step is read again in the fifth, where the pitch is drafted from the brief and the journalist's beat together. That single output is therefore load-bearing twice, and anything vague in it degrades two later steps rather than one. Steps that feed several downstream consumers are the ones worth defining most carefully, and they are usually the summarizing steps near the front.
Building Your First Pipeline Without Code
You do not need a developer to build a basic AI pipeline. Tools like Zapier, Make, and n8n can connect steps without writing code. What they cannot do is decide what your steps are, which is the part that determines whether the finished pipeline is reliable. Here is how to think through the build, in four stages.
Step 1: Map the current manual sequence
Draw it out, even on paper. What happens first? What information does each step need? What does each step produce? What triggers the next step? If you cannot describe the current manual process clearly, you cannot automate it. Most small business owners discover that their process has unclear handoffs, places where nobody is sure who is responsible for passing information forward. Fix the process before you automate it, or you will automate the confusion and then spend a weekend debugging a tool for a problem that was never in the tool.
The question about triggers is the one people skip, and it is where manual processes hide their dependence on a person noticing something. In Lieselotte's old sequence, the media research began when somebody remembered that the background gathering was finished. That is not a trigger, it is a habit, and habits do not survive translation into a pipeline. Every step needs an event that starts it: a form submitted, a previous step completed, a record written. Where you cannot name one, you have found the handoff that was quietly failing already.
Step 2: Identify what can be AI-automated and what cannot
Not every step in a pipeline needs to be AI. Some steps are lookups that retrieve from a database. Some are formatting: take this content, put it in a template. Some are routing: if the client is in the food industry, use template A; if retail, use template B. AI is the right tool for steps that require generating text, summarizing content, or making judgment-like decisions about language and tone. It is the wrong tool for steps that just retrieve or route, and those are better handled by simple automation logic that behaves identically every time.
The distinction is practical rather than philosophical. A routing rule that picks a template by industry either matches or it does not, and you can read the rule when something goes wrong. The same choice handed to a generative step is a decision you cannot inspect afterward, made slightly differently each time it runs. In Lieselotte's sequence, retrieving matching journalists and outlets is a lookup for exactly this reason: the correct answer already exists in a database, so nothing is gained by asking a model to produce it.
Step 3: Define inputs and outputs precisely
For every AI step in the pipeline, write down exactly what text or data goes in, what the AI is asked to do with it, and what format you need the output in. This specificity is what makes AI in a pipeline reliable. A vague instruction produces variable output; a precise instruction produces consistent output that the next step can rely on. The next step is not a person who can interpret an unusual answer, so anything ambiguous becomes a defect further down the line.
For example, rather than "summarize the client intake form," define it as: "From the intake form responses, write a 100-word client summary in this format: [Client name] is a [type of business] in [city], serving [target customer]. They are seeking PR support for [specific goal]. Key differentiators: [three bullet points]. Tone: [professional/casual/technical]." That definition produces output the pitch-drafting step can use without human interpretation, because every field the later step needs is guaranteed to be present and in a known place.
Step 4: Build and test one step at a time
Do not build the full pipeline and then test it. Build Step 1, test it with five real examples from your business, and confirm the output is consistently usable. Then build Step 2 and test the two-step sequence. Add steps one at a time, testing at each addition. This way, when something breaks, you know exactly which step broke it. Building all of it first and testing at the end leaves you with a failure somewhere in a chain and no way to localize it except by taking the chain apart.
Use real examples from your own business for those tests, not invented ones. Real intake forms come with blank fields, clients who described their goal in one word, and the occasional answer typed into the wrong box, and those are precisely the cases that break a step whose instructions assumed everything would be filled in. Five of them will tell you more about the step definition than any number of clean, imaginary examples, because the clean cases were always going to work.
How Lieselotte Built Hers
She used Zapier connected to Claude via the Anthropic API, which Zapier supports through its AI Actions feature. The intake form was Google Forms. The briefing doc output was a Google Doc created automatically from a template. None of those pieces was exotic, and none of them required anyone on her team to write code.
The two ends of that setup are worth noticing. A form is a good pipeline entry point because it produces the same fields in the same order every time, which is what makes the summarizing step downstream reliable. A document built from a template is a good exit point for the same reason in reverse: the shape of the output is fixed, so the final step fills known slots rather than inventing a layout. Constraining both ends leaves the AI steps a narrow, well-specified job in the middle.
Total build time was about eight hours over two weekends, including testing. Time per new client onboarding after the pipeline is about forty-five minutes of human attention: reviewing AI outputs, making minor edits, sending the pitch emails. That is down from two to three days of fragmented work. Her team now handles roughly 40% more new client onboardings per quarter with the same staff.
The pipeline did not replace anyone. It removed the coordination overhead that had been the actual bottleneck, which is worth stating plainly to a team before you build anything, because the version of this story people imagine when they hear "automation" is a different one.
What Makes Pipelines Break
The most common failure in AI pipelines is poorly defined step boundaries. When the output of one step is vague, the next step cannot use it reliably. The fix is always at the prompt level: be more specific about the format and length of the output, not just the content. Owners tend to reach for a different tool at this point, when the tool was never the problem. The symptom is easy to recognize once you know it: the pipeline works on the examples you built it with and produces something odd on a later real case, because that case exercised a part of the instruction you never wrote down.
The second most common failure is over-automation, meaning trying to remove all human judgment from the process. A pipeline that emails clients automatically without any human review will eventually send something that should not be sent. Keep humans in the loop at the steps that matter: final approval of client-facing content, anything involving money, anything irreversible. Those three categories are worth writing down before you build, because they are much harder to reinstate after a pipeline has been running smoothly for a month.
Lieselotte's numbers show what a review point costs and what it buys. Her remaining human time per client goes on reviewing AI outputs, making minor edits, and sending the pitch emails, and the send stays with a person because a pitch to a journalist cannot be recalled. That is the trade the pipeline is actually making: it removes the coordination between steps, which was the bottleneck, while leaving the judgment at the end, which never was.
Anti-Patterns
- Automating a process nobody can describe. If the handoffs are unclear, the pipeline inherits the confusion and hides it behind a tool. Map the manual sequence on paper first, including what triggers each step.
- Using AI for retrieval and routing steps. A lookup and an if-then rule behave the same way every time and can be read when they go wrong. Reserve AI for generation, summarization, and judgment about tone.
- Writing loose instructions for a step whose output feeds a machine. "Summarize the intake form" is fine when a colleague reads the result, and a defect when the next step expects specific fields in a specific format.
- Building the whole chain before testing any of it. When a finished pipeline produces a bad briefing document, an untested chain gives you no way to localize the fault. Build one step, test it against five real examples, then add the next.
- Removing the human from client-facing sends. A pipeline that emails clients with no review will eventually send something it should not. Client-facing content, anything involving money, and anything irreversible stay with a person.
- Selling the pipeline to your team as a headcount story. Lieselotte's removed coordination overhead, not people. If the team believes otherwise, you get quiet resistance at exactly the review steps you need taken seriously.
Practice Prompts
Use these while designing your own pipeline, replacing the bracketed parts with your details.
- "Here is a process my team runs manually: [describe each step, who does it, and what they hand to the next person]. Break it into discrete pipeline steps, and for each one state what goes in, what comes out, and what triggers the next step."
- "For each step in this pipeline, tell me whether it is a generation step, a summarization step, a judgment step, a lookup, a formatting step, or a routing step: [list steps]. Flag any step I assigned to AI that would be more reliable as plain automation logic."
- "Rewrite this vague instruction into a precise step definition specifying the input, the task, the output format, and the output length: [paste your current instruction]."
- "From the intake form responses, write a 100-word client summary in this format: [Client name] is a [type of business] in [city], serving [target customer]. They are seeking PR support for [specific goal]. Key differentiators: [three bullet points]. Tone: [professional/casual/technical]."
- "Here are the steps in my pipeline: [list steps]. Which are client-facing, involve money, or are irreversible? For each, propose where a human review point should sit and what the reviewer should check."
Reflection
Pick the process in your business with the most handoffs. Can you describe, without hesitating, what triggers each step and what each produces? Every hesitation marks a place where an automated version would fail quietly, and where fixing the process comes before automating it.
Look at the steps you would hand to AI. How many are actually lookups, formatting, or routing dressed up as something cleverer? Moving those to plain automation logic makes the chain steadier and its failures readable.
Which steps in your process are irreversible, involve money, or reach a client directly? Decide now where the human review sits, and how you would explain the pipeline to the people whose coordination work it removes.
Glossary
- Pipeline. A sequence of operations where the output of each step becomes the input of the next, with processing or transformation between steps.
- Step boundary. Where one step ends and the next begins, expressed as what goes in, what comes out, and in what format. Vague boundaries are the most common cause of failure.
- Trigger. The event that starts a step, such as a completed intake form. Every step in an automated sequence needs one.
- Routing step. A step that sends work down one branch or another by rule, such as choosing a template by industry. Better handled by automation logic than by AI.
- Lookup step. A step that retrieves existing records, such as matching journalists and outlets from a database. Deterministic, and not a job for AI.
- Human in the loop. A deliberate review point inside an otherwise automated sequence, placed at client-facing, financial, or irreversible steps.
- Over-automation. Removing human judgment from steps that require it, usually discovered the first time the pipeline sends something it should not.
Related Lessons
- Multi-Step AI Workflow Architecture
- Integration Platforms: Zapier, Make, IFTTT for AI
- Error Handling and Monitoring AI Workflows
- Testing and Validating AI Workflows Before Launch
- Zapier AI Actions: Step-by-Step Guide
Closing
Lieselotte's onboarding did not get faster because AI writes better pitches than her team does. It got faster because each step's output was defined tightly enough that the next step could use it without a person in between. Eight hours of building over two weekends turned two to three days of fragmented coordination into about forty-five minutes of human attention per client. Map the sequence, decide which steps genuinely need AI, define every input and output precisely, add one step at a time, and keep a person on what you cannot take back.
Key Takeaways
- A pipeline connects AI operations in sequence, with each step's output feeding the next step's input. The value is in removing manual handoffs between steps.
- Map the current manual process before automating anything. If the handoffs are unclear now, automating them will automate the confusion.
- Not every pipeline step needs AI. Use AI for generation and judgment steps; use simple automation for retrieval and routing steps.
- Define inputs and outputs precisely for every AI step. Specify the format and length of AI output so the next step can use it reliably.
- Build and test one step at a time. Add steps sequentially and test at each addition so you know exactly where problems originate.
- Keep humans in the loop for irreversible or client-facing final steps. Pipelines should remove coordination overhead, not remove judgment from things that require it.
- No-code automation platforms with AI steps can build most small business pipelines. Start there before investing in custom development.
Frequently Asked Questions
Do I need a developer to build an AI pipeline? No. No-code automation platforms such as Zapier, Make, and n8n connect steps without code. Lieselotte's ran on Zapier with Claude through the Anthropic API, a Google Form for intake, and a Google Doc built from a template.
How long does a first pipeline take to build? Lieselotte's took about eight hours over two weekends including testing, covering intake, brief, media query, lookup, pitch drafting, and the final briefing document.
Which steps should not be AI? Lookups and routing. Retrieving matching journalists from a database and choosing a template by industry are deterministic jobs that plain automation does identically every time.
My pipeline produces inconsistent results. What do I fix first? The step definitions, not the tool. Vague output from one step is unusable input for the next, so specify format and length rather than only content. That is the most common failure in AI pipelines.
How much should I test before adding the next step? Test each step with five real examples from your business, confirm the output is consistently usable, then build the next step and test the sequence. Testing only at the end leaves you guessing which step broke.
Where do humans stay in the loop? Final approval of client-facing content, anything involving money, and anything irreversible. A pipeline that emails clients unreviewed will eventually send something it should not have.
Skill.re