←
AI Agent Builders & Citizen Developers
Capable · M16 · lesson 16 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
System Prompt vs. User Prompt in a Node-Based Tool
📖
now learning

System Prompt vs. User Prompt in a Node-Based Tool

15 min

A node-based workflow tool gives you two prompt fields, and most operators use one of them. The other field — the system prompt — is where role, contract, and tone belong; the one you've been pasting everything into is the user prompt. Getting the split right cuts token cost, stops drift, and makes Anthropic's cache_control header pay you back roughly 50%+ on every cached call. This lesson is the operator's tour of that split inside Make, n8n, and Zapier, with the cache-control pattern that turned a $4,200/month workflow into a $1,900/month workflow at one team we worked with in March 2026.

The Two Fields Most Operators Misuse

Open the n8n "AI: Message a model" node. You will see two text boxes: System Message and User Message. Open Make's Anthropic Claude module. You will see a Messages array where the first entry can be a system role and the subsequent entries are user/assistant. Open Zapier's "Send Prompt" action (the new one, post-March-2026). You will see a System Prompt field and a User Prompt field. The naming differs. The split does not.

What we see in 80% of new operator workflows: the user prompt is 600 to 1,500 words of carefully crafted instructions, role declarations, formatting rules, and at the very end a single line that interpolates the actual variable like {{ $json.ticket_body }}. The system prompt is blank, or contains "You are a helpful assistant." Both are wrong, and both are expensive.

The split is not aesthetic. The split is functional. Three things happen when you put instructions in the system field versus the user field:

  1. The model treats system content as policy and user content as data. A role declaration in the user field is much weaker than the same declaration in the system field. Prompt-injection attempts in user-supplied content are also harder to override system policy than user instructions.
  2. Token caching applies cleanly only when you mark the stable prefix. With Anthropic's cache_control, the system prompt is the natural cached prefix and the user prompt is the variable suffix. If you mix them, caching gets messy or impossible.
  3. Versioning and review get cleaner. The system prompt is the contract; the user prompt is the data envelope. Changing the contract is a deliberate event. Changing the data envelope happens every run.
The user prompt is where the variable lives. Everything else — role, format, tone, examples, bans — is the system prompt. If you cannot say "the user prompt is just data," you have a misuse.

What Belongs in the System Prompt

For a workflow node, the system prompt has five jobs. Get all five and you can stop rewriting the prompt every Tuesday.

One: Declare the role

Specificity, not grammar. "You are a triage classifier for B2B SaaS customer support tickets" beats "Act as a helpful assistant who classifies tickets" by a wide margin in practice. The model uses the role to pick a register and a vocabulary. A vague role gets a vague vocabulary. A specific role gets a specific one.

Real example from a customer-success workflow at a 60-seat SaaS we coached in April 2026:

You are an account-health risk-scorer for a B2B SaaS company that sells revenue-operations tooling to mid-market sales teams. Your inputs are the last 30 days of a customer's product usage, support tickets, and CSM call notes. Your output is a single integer risk score from 1 to 10 and a 60-word rationale.

Three things the role declaration does: it grounds the model in domain ("revenue-operations tooling," "mid-market sales teams"), it enumerates the input types ("usage, support tickets, CSM notes"), and it pre-states the output shape ("integer 1-10, 60-word rationale"). The model now has a job description.

Two: Declare the output contract

Format, schema, length, and any forbidden output values. If the downstream of your workflow parses the output, the contract is what your parser depends on. If a human reads it, the contract is what makes the output reviewable in five seconds instead of fifty.

For structured output, include the literal JSON schema and a literal example. Models follow examples more reliably than schema descriptions. The combination beats either alone:

Respond with valid JSON matching this schema:
{ "risk_score": integer 1-10, "rationale": string, "top_signal": string }
Example:
{ "risk_score": 7, "rationale": "Logins dropped 60% week-over-week and two priority tickets unresolved for 9 days. CSM call last Thursday flagged budget review.", "top_signal": "Login decline + open priority tickets" }

Three: Declare forbidden behaviors

The line every workflow operator should know by heart: "No preamble. No markdown fences. No conversational wrappers. Output the JSON object only." Without this, chat-RLHF-tuned models open with "I'd be happy to help" or wrap JSON in ```json fences. Both pollute downstream parsing. Both are fixable with one sentence in the system prompt.

The forbidden list is task-specific. Common ones we add to every classifier-style system prompt:

  • "No hedging. Do not use 'I think,' 'It seems,' 'It appears,' 'might be.' State the classification."
  • "No explanation unless the schema asks for one."
  • "No follow-up questions. Do not ask the user for clarification. If you cannot decide, use the fallback value insufficient_context."
  • "No translation. If the input is in a non-English language, classify it in the same language family rather than translating."

The last one is workflow-specific. The earlier ones generalize.

Four: Provide few-shot examples (when warranted)

If the task is non-standard, anchor with 3-5 labeled examples. Place them at the bottom of the system prompt, just above the user prompt. The model uses them for both pattern induction (input-to-output mapping) and tone induction (the register and verbosity of the rationale field, for instance).

Example fragment for our account-health scorer:

Examples:
Input: 90 days, 30% login drop, 1 open ticket aged 14 days, CSM note "considering competitor."
Output: { "risk_score": 9, "rationale": "Combination of usage decline, aged ticket, and competitive intent is a churn pattern. Engage exec sponsor this week.", "top_signal": "Competitive intent in CSM note" }

Input: 90 days, login flat, 0 open tickets, CSM note "new exec sponsor onboarded."
Output: { "risk_score": 2, "rationale": "Stable usage, no support friction, new sponsor signals continued investment. Low risk.", "top_signal": "New exec sponsor" }

Five: Declare the fallback

What does the model output when it cannot complete the task? "Insufficient context" is one good fallback. A null score is another. The point is to have a designated escape hatch so the model does not invent a plausible-but-wrong answer. We cover this in depth in Lesson 3 ("The 'Refuse if Unsure' Pattern"). The minimum here: declare an escape value in the schema, and instruct the model to use it.

What Belongs in the User Prompt

Three words: the dynamic input. That's it. The user prompt for a well-built workflow node is typically two to ten lines and contains nothing but interpolated variables and minimal labeling.

For our account-health scorer, the user prompt is:

Customer: {{ $json.customer_name }}
Time window: last 30 days
Usage data:
{{ $json.usage_summary }}
Support tickets:
{{ $json.support_tickets }}
CSM call notes:
{{ $json.csm_notes }}

Notice what's missing: there is no "please respond with JSON," no "be concise," no "remember to declare a risk score." All of that lives in the system prompt because all of that is the contract. The user prompt is the envelope that delivers this run's data.

The "label, don't instruct" pattern

The labels in the user prompt ("Customer:", "Usage data:") tell the model what each blob is. They do not instruct the model on behavior. If you find yourself writing "Customer (use this for the rationale field):" you are mixing contract into envelope. Move the instruction to the system prompt: "Use the customer name in the rationale when relevant."

Why this matters for cost

If your workflow runs 100,000 times a month and the user prompt is 1,200 words plus instructions plus the variable, you are sending 1,200 tokens of repeated instruction text every single time. At GPT-5 input rates that's roughly $30/million tokens, so 100,000 runs × 1,200 tokens × $30/million = $3,600/month in repeated instructions alone. Move those 1,200 tokens to the system prompt with cache_control turned on, and the repeated cost drops by 50% to 90% depending on cache hit rate. Same model, same outputs, half the bill or less.

The Anthropic cache_control Pattern (and Why It's Free Money)

Anthropic introduced prompt caching in mid-2024 and matured the pricing through 2025 and 2026. The mechanic is simple: you mark a portion of your prompt as cacheable, and on subsequent calls within a 5-minute (default) or 1-hour (premium) cache window, the cached tokens cost roughly 10% of the input rate for cache reads. Cache writes cost roughly 125% of the input rate. The break-even is usually after 2-3 reuses of the same prefix.

The math at the workflow scale we see most often:

  • Workflow runs 1,000 times per hour for 8 working hours.
  • System prompt is 2,000 tokens of role + contract + few-shot examples.
  • Without caching: 1,000 × 8 × 2,000 = 16M input tokens/day at $3/million (Claude Sonnet 4.5 input rate) = $48/day in system-prompt tokens alone.
  • With caching: first hour's 1,000 runs pay cache-write cost (~$3.75/million × 2,000 × 1,000 = $7.50). Remaining 7,000 runs pay cache-read cost (~$0.30/million × 2,000 × 7,000 = $4.20). Total cache cost: ~$11.70 for the day's prefix.
  • Savings: $48 - $11.70 = $36.30/day, or ~76% reduction on the prefix portion. Over a month: roughly $1,100 saved on a workflow that would have cost $1,440 without caching.

That's a real workflow we audited at a mid-market FinTech in March 2026. They were running a transaction-categorization agent. The system prompt was 2,100 tokens because they had embedded 12 few-shot examples (you can argue 12 is too many; that's a different lesson). They were spending $4,200/month on the workflow before caching. After turning on cache_control, they were spending $1,900/month. The diff bought them the rest of their AI tooling stack and then some.

How to turn it on in Make, n8n, and Zapier

n8n (community edition, v1.83+): the Anthropic Chat Model node has a "Cache Control" toggle in advanced options. Enable it. Then in the system prompt field, the entire content is automatically marked as the cached block. If you need finer control (cache only part of the system prompt), use the "HTTP Request" node and construct the payload manually with "cache_control": {"type": "ephemeral"} attached to the appropriate content block.

Make.com: the Anthropic Claude module added native cache_control support in February 2026. In the module's "Messages" configuration, the system message has a "Cache" checkbox. Check it. Make handles the header construction. Verify by looking at the response object — you will see cache_creation_input_tokens and cache_read_input_tokens in the metadata. Both should be non-zero after the second call within the cache window.

Zapier: Zapier's wrapped "AI by Zapier" action does not expose cache_control directly as of May 2026. To use caching, switch to the "Anthropic" app's direct "Send Message" action and toggle "Cache System Prompt" in advanced. The wrapper convenience costs you the caching feature; we generally recommend the direct app for any workflow exceeding 1,000 runs per month.

Manual (any HTTP-capable node): Construct the request body as:

{
  "model": "claude-sonnet-4-5-20250929",
  "max_tokens": 1024,
  "system": [
    {
      "type": "text",
      "text": "<your full system prompt>",
      "cache_control": { "type": "ephemeral" }
    }
  ],
  "messages": [ { "role": "user", "content": "<dynamic user prompt>" } ]
}

When caching does not pay off

Three cases where you should not bother:

  • Low-volume workflows (under 50 runs/day). Cache writes are slightly more expensive than non-cached writes. If you only run a few dozen times a day, the writes outpace the reads. Break-even is roughly 2-3 reads per write. At ~20 runs/day spread evenly, you cache-miss most of the time.
  • System prompt under ~1,024 tokens. Anthropic's minimum cacheable block is 1,024 tokens for Sonnet (lower for Haiku). If your system prompt is shorter, the cache_control header is ignored.
  • System prompt changes every call. If you dynamically interpolate per-run data into the system prompt (which you should not be doing), the cache key changes every call and you get zero cache hits. Keep the system prompt static across runs.

Three Real Node Walkthroughs

Walkthrough one: n8n triage classifier

A 30-seat agency, processing ~400 support emails per day. The workflow: Email arrives via Gmail trigger → n8n "AI: Message a model" node (Claude Sonnet 4.5) → routes to a Notion database based on the classification.

Before our review: User prompt was 900 words. System prompt was empty. Cost: $0.018 per email × 400/day × 30 days = $216/month. Failure rate (wrong classification): 8%.

After: System prompt is 700 words (role, output contract, 4 few-shot examples, 5 forbidden behaviors). User prompt is 4 lines (sender, subject, body, sent_at). Cache_control turned on. Cost: $0.007 per email × 400/day × 30 days = $84/month. Failure rate: 2.5%.

Net: 61% cost reduction, 69% accuracy improvement. The improvement came from the system prompt being clear, not from caching. The cost reduction came from caching.

Walkthrough two: Make lead-enrichment

A 12-seat sales-development team, processing ~2,500 new leads per week. Make scenario: HubSpot trigger → Anthropic Claude module → enrichment data written back to HubSpot.

Before: User prompt contained the full enrichment prompt plus the lead data. 1,400 tokens per call. No cache. Cost: ~$0.012 per lead × 2,500 × 4 = $120/week or $520/month.

After: System prompt holds the 1,200-token prompt (role, schema, 5 few-shot examples). User prompt holds only the lead JSON (~200 tokens). Cache_control on. Cost: ~$0.0045 per lead × 2,500 × 4 = $45/week or $195/month.

Net: 62% cost reduction. The accuracy was already fine; the win was purely on the bill.

Walkthrough three: Zapier MSA-clause extractor

An in-house legal-ops team at a 200-person company. ~30 MSAs reviewed per week via a Zapier workflow that pulls the PDF from Google Drive, extracts text, and asks Claude to extract eight key clauses.

Before: Wrapped "AI by Zapier" action. User prompt was 1,800 tokens. No cache available. Cost: ~$0.025 per MSA × 30 × 4 = $3/week. Negligible. But the team wanted multi-tenancy: each tenant gets its own clause-extractor with tenant-specific schemas and examples.

The team switched to the direct Anthropic app for Zapier, turned on cache_control, and templated the system prompt with a placeholder for tenant config (loaded from Airtable at workflow start). Per-tenant system prompts ranged from 1,500 to 2,200 tokens. The caching paid off only because 30 MSAs per tenant per week is enough to hit the cache reliably during business hours.

Across 12 tenants, cost stayed around $35/month with caching, versus a projected $90/month without. More importantly, the team had a clear architectural separation: tenant config lived in the system prompt, MSA data lived in the user prompt, and the whole thing was inspectable and roll-backable per tenant.

The System Prompt as the Versioned Artifact

Once you split correctly, a useful side effect emerges: the system prompt becomes the artifact you version, the artifact you review, and the artifact you can roll back. The user prompt template (the variable scaffolding) changes rarely. The system prompt is the policy you edit every time you want to change behavior.

This makes prompt versioning tractable. We cover it in Lesson 4, but the preview here: if your team has 14 workflows and each has a 1,500-word user prompt with embedded instructions, you have 14 places to edit when you want to add a forbidden behavior. If your team has 14 workflows and each has a clean system/user split, you edit 14 system prompts of 600-800 words each. The first scenario is unmanageable beyond a quarter. The second is auditable.

One file per system prompt

We recommend keeping system prompts as separate files in a Git repo or a prompt-management tool (PromptLayer, Langfuse Prompts, Braintrust Prompts — covered in Lesson 4). The workflow node points at the file or the prompt-management entry, not at inline text. When you need to change the prompt, you change one file and the workflow picks up the new version.

If you can't do that (because your workflow tool doesn't support external prompt references), keep a parallel Git repo with one file per prompt, and write a discipline rule: "Any prompt change happens in the repo first; the workflow gets updated from the repo." Slower, but auditable.

Five Anti-Patterns We See Weekly

  1. "Helpful assistant" system prompts. Either be specific or leave the field blank (and accept the trained-default persona). The middle ground gives you the worst of both: the model thinks it has a role, but the role is too generic to constrain behavior.
  2. Instructions in the user prompt. Move them. If you can't articulate why an instruction is system-policy and not data-envelope, it belongs in the system prompt.
  3. No fallback value declared. The model invents. Always declare an "insufficient context" or null value in the output schema and instruct the model to use it.
  4. Cache_control on a tiny prompt. Under 1,024 tokens, the cache header is ignored. Check the response metadata for cache_read_input_tokens — if it's zero, you're not caching.
  5. Dynamic content in the system prompt. Per-tenant config is fine if it's stable per tenant. Per-run data is not — it kills caching and creates a moving target for review.

The Port Routine

Whenever you build a new workflow node or audit an existing one, run this in order. It takes about 10 minutes per node.

  1. Read the current user prompt aloud. Every sentence that is not a variable interpolation is a candidate for the system prompt.
  2. Move candidates into a draft system prompt. Group them by role, contract, forbidden behaviors, examples, fallback.
  3. Reduce the user prompt to labeled variables only. Two to ten lines is typical.
  4. Validate token counts. System prompt above 1,024 tokens? Caching is worthwhile. Under? Either expand to add useful examples or skip caching.
  5. Turn on cache_control. Verify with response metadata showing non-zero cache_read_input_tokens after the second call.
  6. Run 20 historical examples. Compare against the old prompt's outputs. Same accuracy or better? Ship. Worse? Sharpen the role declaration or expand the forbidden behaviors.
  7. Lock the model version. claude-sonnet-4-5-20250929, not latest.
  8. Save the system prompt to your prompt-version store. PromptLayer, Langfuse, Braintrust, or a Git repo. Tag it with the workflow ID.

Key Takeaways

  • Workflow LLM nodes give you two prompt fields. Most operators stuff everything into the user field. The system field is where role, contract, forbidden behaviors, examples, and fallback belong. The user field is for the dynamic input variable.
  • The system prompt is policy. The user prompt is data. The split makes the model treat them differently, lets caching work, and makes versioning tractable.
  • Anthropic's cache_control reduces input-token cost on the cached prefix by roughly 90% on cache reads. Real workflows we have audited dropped from $4,200/month to $1,900/month with no other change. Break-even is roughly 2-3 reads per cache write.
  • Caching requires a minimum 1,024-token block for Sonnet, a static prefix that does not change per-run, and at least ~50 runs/day to amortize the slightly higher write cost.
  • Enable cache_control in n8n via the Anthropic Chat Model node's advanced options; in Make via the Anthropic Claude module's "Cache" checkbox; in Zapier via the direct Anthropic app (the wrapped "AI by Zapier" action does not expose caching as of May 2026).
  • The minimum viable system prompt has five components: role, output contract, forbidden behaviors, few-shot examples (when warranted), and a declared fallback value.
  • The user prompt should be 2-10 lines of labeled variables. No instructions, no persona, no formatting reminders.
  • Once you split correctly, the system prompt becomes the versioned artifact. Store it in a prompt-management tool (PromptLayer, Langfuse Prompts, Braintrust Prompts) or a Git repo with one file per prompt.
  • Five anti-patterns: "helpful assistant" placeholder, instructions in user prompt, no fallback declared, cache_control on tiny prompts, dynamic content in the system prompt.