Prompting Inside a Workflow vs. Prompting in Chat
A prompt that worked perfectly in ChatGPT or Claude.ai will misbehave in a workflow node, and the first instinct — "the model is broken" — is almost always wrong. The model is the same. The environment is different. Three things change between chat and workflow: the implicit conversational context disappears, the role state you didn't realize you were setting evaporates, and the model loses the chat history that quietly anchored its tone. This lesson is the operator's checklist for the translation step.
Why the Same Prompt Misbehaves
You spent two hours iterating in ChatGPT. By the end you had a prompt that produced exactly the analysis you wanted on every test ticket. You copy it into the n8n "AI: Message a model" node, plug in the ticket variable, hit execute. The first run returns something half-shaped. The second run returns a sentence that begins "I think the customer is upset." The third run returns a JSON object wrapped in ```json fences. By the fifth run, you are convinced the model regressed overnight.
The model did not regress. The model was responding to a different conversation. In ChatGPT, your prompt was the last turn in a five-message conversation where you had previously said "I want you to act as a customer support analyst" and "format your response as JSON" and "be concise." Each one of those turns was setting role state, format state, and tone state in the conversation. When you copy the final prompt into a workflow, you bring the words but you leave the state behind. The model in the workflow does not have a chat history. It is reading exactly one user message with exactly one system message — and if you didn't fill in the system message, it is reading just the user message with whatever default persona the model defaults to.
The prompt is not just the text. The prompt is the text plus everything the conversation already established. In a workflow, the conversation hasn't happened. You have to write down what the conversation would have said.
This is the single most common failure mode we see in newcomer operator-builders, and we have not yet met an operator who got this right on first try. The fix is mechanical, not creative.
Gotcha One: Implicit Role State
In a ChatGPT session, you probably said "I want you to act as a senior customer support analyst" at message one or message two. Maybe you said it implicitly: "Pretend you're triaging tickets for a B2B SaaS." That single statement set role state for the entire session. Every subsequent message was interpreted through that lens. The model knew its job.
In a workflow, no such conversation happened. The default Claude or GPT model is helpful, polite, and reaches for explanations. Without an explicit role, you will get explanations where you wanted classifications, hedges where you wanted decisions, and "I'd be happy to help" preambles where you wanted machine-parseable output.
What "role state" looks like in practice
Here is a real before-and-after from a workflow port we did in February 2026. The team had a prompt that worked beautifully in Claude.ai. They pasted the user prompt into n8n with no system prompt. Output drifted within hours.
Before (the user prompt that worked in chat):
Look at this ticket and tell me whether it's a refund request, a bug report, or a question. Just one word.
What chat had that workflow didn't: Three previous messages where the user had said "I'm running customer support for a SaaS. Help me classify incoming tickets quickly. Answer in single words when I ask for one. Don't add explanations." The model knew its role, its format constraint, and its tone.
After (the workflow version that survives production):
System prompt:
You are a triage classifier for customer support tickets at a B2B SaaS. For every ticket you receive, you must respond with exactly one of:refund,bug,question, orother. Output only the single word. No punctuation, no explanation, no preamble.
User prompt:
Ticket: {{ $json.ticket_body }}
Notice four changes. The role is named explicitly ("triage classifier for customer support tickets at a B2B SaaS"). The output space is enumerated explicitly (four allowed values, including an explicit "other"). The format is banned explicitly ("no punctuation, no explanation, no preamble"). And the user prompt is reduced to the minimum dynamic input. The verbose conversational version of the prompt — "Look at this ticket and tell me..." — is replaced with a system prompt that does the role-setting and a user prompt that just supplies the data.
The role-state translation pattern
Whenever you port a chat prompt to a workflow, run this checklist:
- Open the chat conversation and read every message you sent, in order.
- For each message, ask: "Was this setting role, format, tone, or task?"
- Role and format messages go into the workflow's system prompt.
- Tone and stylistic preferences go into the system prompt as constraints (e.g., "concise, no hedging").
- The actual per-input task and data go into the user prompt.
- Anything that was implicit ("I assume you know I want JSON") must be made explicit.
Gotcha Two: Missing System Prompt
This is gotcha one's mechanical sibling. Many of the most popular workflow tools either default to an empty system prompt or hide it behind an "advanced options" panel that newcomers don't open. Zapier's wrapped "AI by Zapier" action did not even expose a system-prompt field until early 2026. n8n's "AI: Message a model" node has a System Message field that is empty by default. Make's Anthropic Claude module has a "Messages" array where the system message is its own optional element.
An empty system prompt is not neutral. Anthropic's Claude defaults to its trained persona (helpful, harmless, thorough). OpenAI's GPT defaults to its persona ("ChatGPT-style"). Google's Gemini defaults to its persona. None of these defaults are wrong, but none of them are your assistant. The result is that the same user prompt produces three different shaped responses across the three providers, and even within a single provider the model "talks like ChatGPT" when you wanted it to talk like a JSON-producing data extractor.
The minimum viable system prompt
For 90% of workflow nodes, your system prompt needs to do three things:
- Declare role. "You are a triage classifier." "You are a sales-email generator." "You are a contract-clause extractor." Specific and short.
- Declare output contract. "Respond with valid JSON matching this schema." Or "Respond with exactly one word from this list." Or "Respond with 1-3 short paragraphs separated by blank lines."
- Declare forbidden behaviors. "No preamble. No explanation. No markdown fences." This last one is the most over-looked and most necessary. Without it, chat-tuned models default to conversational helpfulness.
That is the floor. Anything more is task-specific. Anything less is a coin flip.
An empty system prompt is the model's permission slip to be ChatGPT. If you wanted ChatGPT you would still be in ChatGPT. You are in n8n because you want a classifier. Tell the model.
The "act as" debate
You will see two schools of system-prompt construction online. School one says "You are X." School two says "Act as X." Three years ago this mattered more than it does now. With Claude Sonnet 4.5 and GPT-5-class models in May 2026, either phrasing works. The thing that matters is the noun. "Act as a helpful assistant" is useless; "Act as a contract-clause extractor specialized in MSA agreements for B2B SaaS" is useful. Specificity beats grammar.
Gotcha Three: Model Behaves Differently with No Chat History
This is the subtle one. In a chat session, the model has the chat history as context — not just for content but for tone calibration. If you have been speaking informally for three messages and then ask a complex question, the model continues your informal register. If you have been speaking in technical jargon, the model matches your jargon level. In a workflow, the model gets exactly one user message and has no register to match. It defaults to the median of its training distribution, which is roughly "polite professional helper."
What this breaks
Three concrete failure modes in production:
One: tone mismatch. Your chat prompt produced terse outputs because your chat was terse. The same prompt in a workflow produces verbose outputs because the workflow gives no terse signal. We saw a sales-email generator that produced three-sentence emails in Claude.ai and five-paragraph emails in n8n with the same user prompt. The fix: a system-prompt constraint like "Each output is 80-120 words. No more, no less." With an explicit numeric anchor, the model converges.
Two: hedging language. Chat history that included previous correct answers anchors the model to confident responses. Without history, the model defaults to hedges: "I think," "It seems like," "This appears to be." These leak through into structured outputs and pollute downstream parsing or display. The fix: explicit ban in system prompt. "Make confident assertions. Do not use words like 'I think', 'seems', 'appears', 'might be'. State the classification."
Three: format drift. Chat prompts often work because the previous turn established the format. "Like before, give me a JSON object." In a workflow there is no "before." The model picks a format somewhere in its trained distribution, which on day 1 might be JSON and on day 14 might be a Markdown table. The fix: declare the format in system prompt with an exact schema, and parse defensively.
The "no chat history" smell test
When you port a prompt and the output feels off, ask: "In my original chat session, was there anything earlier in the conversation that the workflow doesn't have?" If yes, encode it explicitly in the system prompt. If you cannot articulate it, the chat session was doing more work than you realized, and you need to spend 15 minutes re-deriving the prompt with no prior context — open a brand new chat session, send your workflow's full system + user prompt as the first message, and see what comes back. That output is what your workflow gets.
The Port Checklist (Five Minutes Per Prompt)
When porting a chat prompt to a workflow node, run these in order:
- Re-derive in a clean chat session. Open a new ChatGPT/Claude.ai conversation. Send the full prompt as the first and only message. Note any quality drop. That drop is what your workflow will see — minus more, because workflows also lose ChatGPT's invisible state-tracking that may help slightly.
- Write the system prompt. Role, output contract, forbidden behaviors. Three to ten sentences.
- Reduce the user prompt to the dynamic input. Variables only. No persona, no instructions, no format reminders. The system prompt does that work.
- Add anti-drift bans. "No preamble." "No markdown fences." "No conversational wrappers." These three lines have saved more workflows than any other prompt-engineering trick we know.
- Lock the format. If the output is structured, write the JSON schema in the system prompt and include a literal example. If it's prose, give a word/sentence count.
- Run 10 historical examples. Not three. Not fresh ones. Ten real inputs that your workflow will see. Tally the failures. If any are format failures, the system prompt needs another constraint. If any are content failures, the role or task needs sharpening.
- Set temperature low. For classification and structured output: 0.0-0.3. For generation: 0.5-0.8. The chat default is often higher than you want for production.
- Pin the model version. Use
claude-sonnet-4-5(or the version-pinned form) rather than "the latest model." When the provider rolls a new version, you want to evaluate the change deliberately, not get surprised in production.
Three Port Stories
Port one: the "summarize this Slack thread" failure
A team built a Slack-to-Linear workflow. A user types /summarize in a thread, the workflow pulls the thread, summarizes, and creates a Linear issue. In Claude.ai, the prompt "Summarize this Slack thread as a bug report" produced clean three-paragraph reports. In n8n, the same prompt produced five-paragraph essays that opened with "I'd be happy to help you summarize this Slack thread!"
The fix took ten minutes. The team added a system prompt: "You are a bug-report writer. Given a Slack thread, produce exactly three short paragraphs: Reproduction Steps, Expected Behavior, Actual Behavior. Use bullet points within each paragraph. No introduction. No closing pleasantries." The user prompt became just Slack thread: {{ thread_text }}. Format consistency went from 40% to 97%.
Port two: the "extract contract clauses" disaster
A legal-ops team had a Claude.ai workflow for extracting key clauses from MSAs. The prompt was 800 words of carefully tuned instructions. They ported it to a Make scenario. Hour one: works. Hour twelve: a contract came through with non-English headers (a localized version of the same agreement). The output mixed English and French clause names and broke their downstream Airtable mapping.
The chat session had previously seen English-only contracts and the model implicitly assumed all inputs would be English. In the workflow, no such prior exposure existed. The fix was a single line in the system prompt: "Always respond in English regardless of the source contract's language. Translate clause names to standard English terminology." Hour thirteen onward: clean.
Port three: the "rewrite for the brand voice" tone mismatch
A marketing-ops team built a brand-voice rewriter. The Claude.ai prompt produced perfect on-brand copy: warm, declarative, no jargon. In the workflow it produced corporate-bland output. The chat session had three messages of back-and-forth before the prompt that established what "on-brand" meant. The workflow had none of that.
The fix was a 60-word system prompt with three brand-voice examples directly in the prompt: "On-brand voice example: 'You'll get paid in 24 hours.' Not on-brand: 'Payment will be processed within one business day.' On-brand: 'We're here when you need us.' Not on-brand: 'Our support team is available during business hours.'" Concrete examples in the system prompt did what three chat turns had done implicitly.
Few-Shot vs. Zero-Shot in Workflows
In chat, you almost always operate zero-shot — you just ask the model. The conversation history serves as a kind of in-line few-shot. In workflows, that history is gone, and the easiest replacement is explicit few-shot examples in the system prompt.
Few-shot is two to five labeled examples placed directly in the system prompt, each showing an input and the desired output. For our triage classifier:
Examples:
Input: "I want my money back, this software is terrible" →refund
Input: "Steps to reproduce the crash on Safari 17" →bug
Input: "How do I export to CSV?" →question
Input: "Hi, who handles enterprise pricing?" →other
Three to five examples is the sweet spot. Less than three gives the model too little to anchor on; more than five increases input-token cost without proportional accuracy gain. For most workflows, three examples plus an explicit "other / fallback" example covers 90% of the gain.
When to skip few-shot
You can skip few-shot when (a) the task is well-known to the model (English summarization, sentiment classification on standard categories), (b) the output schema is strict enough that the model can't drift much (single-word classification with four allowed values), and (c) you've validated 20+ historical examples and the zero-shot version produces correct outputs. Otherwise, prefer few-shot — the input-token cost (~50-200 tokens) is cheap insurance against drift.
The Temperature Question
In chat, temperature is usually invisible — most chat UIs default to 0.7 or 1.0, and you don't notice because you can re-prompt if you don't like the answer. In a workflow, the default temperature is the temperature, and every run is judged by the human reviewer the first time. For our triage classifier, temperature 0.2 produces the same classification across 20 runs of the same input. Temperature 1.0 produces three different classifications on the same input.
Rules of thumb for workflow nodes:
- Classification, extraction, structured output: 0.0 to 0.3.
- Summarization with consistent format: 0.2 to 0.5.
- Generation (emails, reports, copy): 0.5 to 0.8.
- Creative writing or ideation: 0.8 to 1.2.
If your workflow downstream parses the output, default to the low end. If a human reviews the output and the variance is desired, allow higher.
The Model Version Question
One last difference between chat and workflow. In ChatGPT or Claude.ai, when the provider rolls a new model version, you find out gracefully — UI changes, a banner appears, your next session uses the new model. In a workflow, if you used a "latest" alias, the upgrade happens silently at 3am on a Tuesday. Outputs change. You don't notice for a week. Three hundred Zendesk private notes look subtly different. Your eval set, if you have one, catches it. If you don't, your customer notices first.
The fix is two words: version pin. Use claude-sonnet-4-5-20250114 or gpt-5-2026-03-01 rather than claude-sonnet-4-5-latest. When a new version drops, you upgrade deliberately: run your eval set on the new version, compare, decide. The cost of version pinning is that you have to manually upgrade. The benefit is that you control when "different" happens.
Key Takeaways
- A chat prompt and a workflow prompt are not the same artifact. The chat prompt is the text plus the invisible state set by previous turns. The workflow prompt is only the text.
- Three gotchas: implicit role state (the persona you didn't realize you set), missing system prompt (workflow tools default to empty), and absence of chat history (which silently anchored tone and format).
- The minimum viable system prompt declares role, declares output contract (format/schema), and declares forbidden behaviors ("no preamble, no markdown fences, no hedging").
- When porting, re-derive the prompt in a fresh chat session as the first message — that is the output your workflow will see. If quality drops, you have to encode the missing context explicitly.
- Add few-shot examples (3-5) in the system prompt when the task isn't trivially well-known. The input-token cost is cheap insurance against drift.
- Set temperature low (0.0-0.3) for classification and structured output; reserve higher temperatures for generation tasks where variance is wanted.
- Pin the model version (
claude-sonnet-4-5-20250114, notlatest). Upgrade deliberately on an eval set, not silently in production. - Run 10 historical examples after every prompt change. Three is too few; demo tickets miss the real distribution.
Skill.re