Troubleshooting AI Integrations: Common Issues
Zainab runs a four-person interior design studio in Miami. She spent a Saturday building a Zapier workflow that would take client inquiry forms, send them to ChatGPT for a project scope summary, and email her that summary before she got on the consultation call. She turned it on Sunday night. Monday morning, three consultation calls happened and she had zero summaries. She spent Tuesday troubleshooting. By Wednesday she had fixed it, but if she had known the small set of failure points that account for most AI integration breakages, she would have found it in twenty minutes on Monday morning instead of spending her Tuesday on it.
How to Read Error Messages
The first skill in AI integration troubleshooting is reading the error message without panicking. Most errors in tools like Zapier and Make, and in direct API integrations, fall into a small number of categories, and the message itself tells you which category you are in if you know how to read it. The words may look technical, but they are doing a simple job: naming which part of the chain refused to cooperate. Almost nothing here requires you to understand code.
Most platforms show you errors in a run history or execution log. In Zapier it is the Task History tab. In Make it is the scenario history panel. Start there before you start guessing. The log tells you three things that no amount of theorising will: which step failed, what data that step actually received, and what the service said back. Guessing skips all three, and the most common troubleshooting mistake is rebuilding a workflow that was never broken in the place you rebuilt.
The Six Most Common Failure Types
Failure 1: Authentication error (401 Unauthorized)
What it means: the integration tried to connect to a service, usually an AI API or a connected app, and was rejected because the credentials it presented were wrong or expired.
Why it happens: API keys expire or get rotated. An OpenAI API key that worked when you set up the integration three months ago may since have been replaced. The same applies to OAuth tokens, the behind-the-scenes credential that lets Zapier reach into your Gmail on your behalf, which can expire if the integration has been idle for a while.
Fix: go to the settings of the specific step that failed and look for a "Reconnect Account" or "Test Connection" button. Reconnect, re-enter the API key or re-authorize the OAuth connection, and test the step in place rather than assuming. Then check the AI provider's own dashboard to confirm the key is still active and has not hit a usage limit.
Zainab's outage was in this family. Her OpenAI API key had a spending limit she had set months earlier and forgotten about. The key itself was still valid, but the account had reached its monthly cap, so every call was being refused. She raised the cap and the workflow started working immediately. The lesson generalizes: when calls are being rejected, check both the credential and the account's usage limits on the provider dashboard, because a rejection does not always tell you which of the two it is.
Failure 2: Rate limit exceeded (429 Too Many Requests)
What it means: your integration is sending requests to the AI API faster than the API allows. APIs, the connection points between your tools and the AI service, cap how many requests you can make per minute or per hour.
Why it happens: you set up a workflow that processes a backlog of 200 form submissions at once, or a trigger fires more often than you expected, and the requests pile up on top of each other rather than arriving in a queue.
Fix: add a delay between processing steps. In Zapier, add a "Delay" step before the AI step. In Make, add a sleep function between modules. Most APIs throttle somewhere between 3 and 60 requests per minute depending on your plan tier, and a two to five second delay between requests usually resolves rate limit errors without noticeably slowing the workflow. If you are importing a backlog, do it in batches rather than in one push.
Failure 3: Data format mismatch
What it means: one part of your workflow sends data in a shape another part cannot read. The most common version is a date: your trigger sends "06/17/2026" and the downstream system expects "2026-06-17."
Why it happens: different software systems format dates, phone numbers, addresses and currency amounts differently, and each system assumes its own convention is obvious. An AI model that receives a badly formatted input may produce an error, or, more often, a confidently wrong output built on a misread value.
Fix: normalize the data before it reaches the AI step, using the platform's built-in formatting tools. "Format Date," "Format Number" and the general "Formatter" utilities exist for exactly this. When sending text to an AI API, state the expected format in the prompt itself: "The date will be formatted as MM/DD/YYYY. Parse it accordingly." Fixing this in a formatting step is more reliable than fixing it in a prompt, because the formatting step behaves the same way every time.
Failure 4: Empty or null input
What it means: the AI step received a blank or missing value where it expected text or data. It either produced no output at all or an error such as "invalid input."
Why it happens: a form field was optional and someone left it blank. A previous step produced no data because no records matched its filter. Or a conditional branch produced no output in a case you did not anticipate when you built it.
Fix: add a filter step before your AI step that checks the required field is non-empty, expressed as "Only continue if [field] exists and is not blank." That stops the AI step from running when the field you checked is missing. It protects only the fields you name, so if your prompt depends on several fields, check each of them rather than assuming the first one stands in for the rest.
Failure 5: AI output that downstream steps cannot parse
What it means: the AI produced a response, but the next step in your workflow, which expected a specific structure such as a JSON object or a particular keyword, could not extract what it needed from the free-form text it got.
Why it happens: AI language models produce slightly different output every time. If your prompt asks for a structured list and the model sometimes returns a paragraph instead, any downstream step that counts on the list structure will break, and it will break intermittently, which is far more confusing than breaking every time.
Fix: make the prompt more specific and ask for structured output explicitly: "Respond ONLY with a valid JSON object in this format: {summary: '...', priority: 'high/medium/low', action: '...'}. Do not include any other text." Then run the prompt 10 times and check that the format holds. Ten clean runs is evidence, not a guarantee, so where the downstream step would do something you cannot undo, keep a check on the response shape and a path for the run that comes back malformed anyway.
Failure 6: The workflow runs but produces wrong results
What it means: no error is shown anywhere. Everything reports success. But the output is not useful: summaries are too vague, categorizations are wrong, generated emails do not sound like you.
Why it happens: the prompt is not giving the AI enough context. This is not a technical failure at all. It is a prompt quality failure wearing the costume of a workflow problem, which is why people lose days to it rebuilding steps that were working correctly the whole time.
Fix: isolate the AI step. Copy the exact input it received, taken from your workflow's run log rather than from your memory of what it should have received, and paste it into a chat interface such as ChatGPT or Claude along with your prompt. Test variations there, where the feedback loop is seconds rather than a whole workflow run. Once a prompt produces good results 8 times out of 10 in the chat window, paste it back into the workflow.
| What you see | Likely failure type | First move |
|---|---|---|
| 401 Unauthorized in the run log | Authentication | Reconnect the account, then check the key and the usage limit on the provider dashboard |
| 429 Too Many Requests | Rate limit | Add a delay before the AI step and process backlogs in batches |
| Step errors on a date, number or phone field | Format mismatch | Normalize with a formatting step before the AI step |
| AI step errors with "invalid input" or returns nothing | Empty or null input | Filter for the required fields before the AI step runs |
| The step after the AI step fails intermittently | Unparseable AI output | Constrain the output format in the prompt and handle the malformed case |
| Everything succeeds but the output is poor | Prompt quality | Isolate the AI step and test the prompt in a chat interface |
The Troubleshooting Sequence
When a workflow fails, work through this sequence in order rather than jumping to the step you suspect.
- Open the run history and find the specific step that failed. Not the workflow, the step.
- Read the error code. Is it 401, meaning authentication, 429, meaning rate limit, or something else?
- If there is no error code, meaning the workflow ran but produced bad output, isolate the AI step and test the prompt directly instead of touching the workflow.
- Fix the smallest thing first. Most failures have one root cause. Do not redesign the whole workflow around a single bad run.
- Re-run the failed task manually after fixing. Most platforms have a replay or retry button for failed tasks, which also confirms the fix against the exact data that broke it.
The discipline in that list is the ordering, not the content. Every step you skip is a step you will end up doing later with less information, and the temptation is always to skip straight to number four because you have a hunch. Zainab's hunch on Tuesday morning was that her prompt was wrong, so she rewrote it twice before she opened Task History and saw that the AI step had never run at all.
Before You Call It Fixed
A workflow that passes a manual replay has proved one thing: it survives the record that broke it. That is worth having and it is not the same as being fixed. Watch the next few real runs go through the log, because the record that broke it was probably not unusual, and the failure you just repaired may be one of several instances of the same root cause sitting further back in the queue.
Then write down what happened, in one line, somewhere you will look again: what broke, what the error said, and what you changed. Integration failures repeat, particularly the credential and limit ones, which come back on the schedule that the credential or the limit dictates rather than on any schedule of yours. A note written while you still remember the details turns the next occurrence from a rediscovery into a recognition.
Anti-Patterns
- Rebuilding before reading the log. Zainab rewrote her prompt twice before discovering the AI step had never executed. The run history tells you which step failed and what it received; nothing you reason out from the outside beats that.
- Treating "no error" as "working." Failure type six shows up as a green run with a useless result. If you only monitor for errors, this is the failure you ship to customers for weeks without noticing.
- Fixing the symptom in the prompt when the cause is in the data. Asking the AI to cope with inconsistent date formats works most of the time, which is the problem. A formatting step behaves identically on every run; a prompt instruction is a request.
- Changing several things at once. Reconnect the account, add a delay, rewrite the prompt and reformat the input in one pass and you will have a working workflow and no idea why. The next failure starts from zero.
- Assuming a rejected call means a bad key. A capped spending limit, an exhausted quota and a rotated key can all present as the workflow being refused. Check the provider dashboard as well as the credential before you regenerate anything.
- Testing a prompt once and calling it stable. Models produce slightly different output each run, so one good result says the prompt can work, not that it will. Intermittent format failures are the hardest class of bug to trace.
Practice Prompts
- Interpret an error you do not recognise. "Here is the error my automation logged: [paste the full message and error code]. Tell me what category of failure this is, the most likely causes in order of probability, and what to check first. Do not suggest rebuilding the workflow."
- Isolate a bad-output problem. "This is the exact input my AI step received: [paste from the run log]. This is the prompt it ran: [paste it]. This is what it produced: [paste it]. This is what I wanted: [describe it]. Diagnose whether the problem is the prompt, the input, or my expectation, and rewrite the prompt if that is the cause."
- Constrain an output format. "Rewrite this prompt so the response is always readable by the next step in an automation: [paste your prompt]. It must return only a valid JSON object with these fields: [list them]. Include the instruction that no other text be returned, and tell me what the next step should do if the response comes back malformed anyway."
- Find the empty-input cases. "Here are the fields my AI step depends on: [list them]. Here is where each comes from: [describe]. Tell me which can legitimately arrive blank, what the AI step would do in each case, and what filter condition belongs in front of the step."
- Write your own runbook. "Turn this into a one-page troubleshooting checklist for a non-technical owner running [describe your workflow] on [name your platform]: the failure types, how each appears in the run history, and the first action for each."
Reflection
- When one of your automations fails, how do you find out, and how long does that take?
- Which of your live workflows would fail silently, producing a green run and a bad result, without anyone noticing?
- Where are your API keys and connected accounts, and do you know which of them have spending limits or usage caps set?
- Think of the last integration you troubleshot. Did you read the log first, or start from a hunch, and what did that cost?
Glossary
- Run history. The log of every execution of a workflow, showing which steps ran, what they received and what came back. Zapier calls it Task History; Make shows it in the scenario history panel.
- 401 Unauthorized. The status code meaning the service rejected the credentials presented. In practice: reconnect the account, or check the key.
- 429 Too Many Requests. The status code meaning you have exceeded the allowed request rate. In practice: slow the workflow down with a delay.
- OAuth token. The credential letting one service act on your behalf in another, such as an automation platform reaching your Gmail. It can expire through disuse.
- Rate limit. The cap on how many requests you may make to an API per minute or per hour, usually varying by plan tier.
- Delay step. A step that pauses before continuing, used to space requests out so they stay inside the rate limit.
- Formatter. A built-in utility that converts data into an expected shape, such as reformatting a date or a number, before a later step receives it.
- Structured output. A response constrained to a defined shape, such as a JSON object with named fields, so a later automated step can read it reliably.
- Replay. Re-running a failed task against the original data after a fix, confirming the repair against the record that actually broke.
Related Lessons
- Error Handling and Monitoring AI Workflows covers building the alerting that turns a silent failure into a message you receive.
- Error Recovery and Fallback Strategies goes further, into what a workflow should do instead of failing.
- Testing and Validating AI Workflows Before Launch is the preventive half of this lesson: catching these failure types before they reach a customer.
- API Basics: Connecting AI Services Directly explains keys, endpoints and limits, which are the subject of the first two failure types here.
- Common Prompting Mistakes and How to Fix Them addresses failure type six directly, where the integration is fine and the prompt is not.
Closing
Zainab lost a Tuesday to a spending cap she had set herself. The technical content of the fix was one number on a settings page; everything else she spent was the cost of not knowing where to look, and that is the gap this lesson exists to close. The habit worth building is smaller than the list of failure types. Open the log. Find the step. Read the code. Change one thing, replay, and write down what it was. That sequence turns integration failures from a category that feels like it needs a developer into a category that feels like checking why the till drawer will not open.
Key Takeaways
- Start every troubleshooting session by reading the run history, not by guessing. The error code tells you which failure type you are in, and the log shows you exactly what the failing step received.
- A 401 authentication error means a reconnect or a key refresh is needed. Check the account's usage limits on the provider dashboard too, because a capped account looks a great deal like a bad credential from the workflow's side.
- A 429 rate limit error means your workflow is sending requests too fast. Add a delay step before the AI module and process backlogs in batches rather than all at once.
- Data format mismatches are fixed with formatting steps, not prompt changes. Normalize dates, numbers and text before they reach the AI, because a formatting step behaves the same way every run and a prompt instruction does not.
- Filter for empty inputs before your AI step runs. A blank input produces an error or a useless output, and a field-exists check prevents that for every field you actually name.
- If the workflow runs but the results are wrong, it is a prompt problem, not a workflow problem. Isolate the step, take the real input from the log, and test the prompt in a chat interface where the feedback loop is seconds.
- Fix the smallest thing first, then replay and record it. Most integration failures have a single root cause you can address quickly once identified, and the note you write is what makes the second occurrence cheap.
Frequently Asked Questions
Where do I find the error in the first place?
In the platform's log of past executions. Zapier puts it in the Task History tab, Make in the scenario history panel, and a direct integration will have whatever logging your developer built, which is worth confirming exists before you need it. Open the failed run, expand the individual steps, and find the first one that is not green. Everything after that point failed because of it, not independently of it.
My workflow worked for months and then stopped. What changed?
Usually a credential or a limit rather than your workflow, because your workflow did not change. API keys get rotated, OAuth connections expire after disuse, spending caps are reached partway through a month, and plan tiers change. Find the date of the first failure in the run history, then look for something on the account side that lines up with it.
Should I fix the prompt or the workflow?
Read the run status first. If the step errored, it is a workflow or credential problem and the prompt is irrelevant to it. If every step reports success and the output is simply poor, it is a prompt problem and rebuilding the workflow will not touch it. That distinction is the single largest source of wasted troubleshooting time.
How do I stop the AI returning a different format every time?
Ask for a constrained structure explicitly, name the exact fields, and instruct the model to return nothing else. Then test repeatedly rather than once, since intermittent shape changes are what break downstream steps. Treat a run of clean tests as encouraging rather than conclusive, and keep a path for the malformed response, especially where the next step does something you cannot easily undo.
Is a rate limit error a sign I need a bigger plan?
Not usually, and not first. Rate limits are about how fast requests arrive, not how many you make in total, so the same monthly volume spread out with a short delay between calls often sits comfortably inside a limit that a burst blows straight through. Add the delay, process backlogs in batches, and see whether the problem survives that before changing plans.
How do I know a fix actually worked?
Replay the failed task against the original data, which proves the fix handles the record that broke, then watch the next few genuine runs in the log. The first confirms the repair; the second confirms it against inputs you did not choose. If the workflow matters, the more durable answer is to build notification into it, so the next failure reaches you rather than waiting to be found.
Skill.re