←
AI Agent Builders & Citizen Developers
Capable · M5 · lesson 5 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Cost, Latency, and Rate Limits: The First Time You Hit Them
📖
now learning

Cost, Latency, and Rate Limits: The First Time You Hit Them

15 min

The first month of an LLM-in-the-loop workflow is a honeymoon. The second month is when the platform dashboards start mailing you. A node that ran 200 times a day in week one runs 12,000 times a day in week six because someone wired it to a new trigger. Your $40 monthly Anthropic bill turns into a $1,400 surprise. Latency that was a 2-second p50 in dev is a 9-second p95 in production. And the rate-limit error you'd never seen lands on a Tuesday morning during a marketing campaign. This lesson is how operators read the dashboard, find the offending node, and apply three standard fixes.

The Two Numbers That Matter

Forget every other metric for now. The two numbers practitioners defend in 2026 are time-to-first-token p95 and dollars per million tokens. If you cannot quote both of these for every LLM node in your stack, you do not know what you are paying for or how fast your users see the first character of an answer.

Time-to-first-token (TTFT)

This is the wall-clock time from sending the API request to receiving the first chunk of the response. Models stream, and TTFT is the latency that matters to anything resembling a user experience. A 2-second TTFT on a chat interface feels snappy. An 8-second TTFT feels broken. For a background workflow node that doesn't stream to a human, total response time matters more than TTFT — but TTFT is still the cleanest signal of API health and the first thing that drifts when a provider is under load.

Track p50 (median), p95 (the 5%-worst tail), and p99 (the 1% tail). Operators defend p95 because it's the bar that drives "this feels slow" complaints. If p95 is 4 seconds, 5% of your users wait more than 4 seconds. Most people will accept that. If p95 is 12 seconds, you have a problem.

Dollars per million tokens

The currency of LLM work in 2026 is the million-token. Anthropic Claude Sonnet 4.5: $3 input, $15 output per million. Anthropic Claude Haiku 4.5: $0.25 input, $1.25 output per million. OpenAI GPT-5: $5 input, $20 output per million (May 2026 pricing). OpenAI GPT-5 mini: $0.50 input, $2 output per million. These rates change quarterly; keep a copy of the current pricing tab open when you're modeling.

The compound metric to track is $ per workflow run. Take total monthly LLM cost, divide by total monthly runs. Compare across nodes. The one with the highest $/run is your candidate for optimization. The one with the highest variance in $/run is your candidate for input-truncation work, because variance usually means some runs are processing 10x larger inputs than others.

Reading the Usage Dashboard

Three views matter. Set up bookmarks for all three.

Anthropic Console

The Anthropic Console (console.anthropic.com, May 2026) shows a daily spend curve, a token breakdown by model, and — critically — a cache hit rate column. The cache hit rate tells you what percentage of your input tokens were served from the prompt cache versus billed at full rate. If you wrapped your system prompt in cache_control but your cache hit rate is 12%, your cache isn't warm — call volume is too sparse to stay within the 5-minute TTL window, or your system prompt is changing too often.

The Console also exposes per-API-key spend, which is the trick to attributing cost to individual workflows. Generate one API key per workflow (or per workflow family). When the bill surprises you, you can point at the offending key and walk that workflow back to its trigger.

OpenAI Usage Dashboard

OpenAI's usage page splits by model, by API key, and by date. As of March 2026, OpenAI's prompt-caching feature (released late 2024) auto-caches prefixes of 1024+ tokens that have been seen recently; cached input tokens are billed at 50% of standard input rate. You don't need to mark anything with cache_control on OpenAI — but you do need to structure your prompts so the stable prefix is at the start and the variable content is at the end. Otherwise the cache never hits.

Workflow tool dashboards

n8n's Cloud dashboard exposes per-workflow execution counts. Make's "Operations" view in scenario settings shows ops-per-day. Zapier's task usage view exposes tasks per Zap. Cross-reference these with the model dashboards: if you see 12,000 executions on a workflow but only 800 LLM API calls, you have a branching issue (or a cache). If you see 800 executions and 12,000 API calls, you have a loop or retry-storm problem.

Finding the Offending Node

Sort your nodes by descending $/month. The top one is the candidate. Ask three questions in order:

  1. Is this node necessary at this volume? Sometimes the answer is "no, the trigger is firing too often." A Zendesk webhook that fires on every ticket update (not just creation) can multiply call volume 4-8x with no business benefit. Trigger scoping is the cheapest fix.
  2. Is this node using the right model? A simple "classify into 4 buckets" task does not need Claude Sonnet 4.5. Claude Haiku 4.5 at 12x lower cost will give you 95-99% of the accuracy at a fraction of the spend.
  3. Is this node being called when nothing has changed? If your workflow runs once an hour summarizing the same Slack channel where 90% of hours have no new messages, you're paying for runs that should never have started. Add a gate.

If the answers to those three are "yes, yes, yes," now you're in the territory of the three standard fixes.

Fix One: Model Tiering (Cheaper Model for Simpler Subtasks)

Not every step needs the smartest model. A workflow that uses one LLM call to classify a ticket's intent and a second to draft a response can use different models for each. The classifier wants accuracy on 4-8 categories — Haiku 4.5 is fine. The draft generator wants longer-form coherence — Sonnet 4.5 is the right choice. Mixing the two cuts cost without hurting quality.

Real numbers from a tiering experiment

One team in February 2026 had a two-stage workflow: classify customer intent, then draft a response. Both stages were on Sonnet 4.5. Monthly cost at 8,000 runs: $384. They added a tier: Haiku for stage one (classification), Sonnet for stage two (drafting). Same monthly volume, same quality (verified on a 100-example eval set). Monthly cost: $156. Savings: 59%.

Three nuances. First, validate quality with an eval set, not vibes. A cheaper model that's 92% accurate is worse than an expensive one that's 98% accurate if your downstream depends on the classification. Run the eval before the swap. Second, watch the failure modes: smaller models often fail in specific ways — they hedge more, they ignore edge cases. The eval set should include those edge cases. Third, the model strings to use:

  • Anthropic: claude-haiku-4-5 for classification/extraction; claude-sonnet-4-5 for generation and reasoning; claude-opus-4-5 for the hardest tasks where you've validated the quality lift justifies 3-4x cost.
  • OpenAI: gpt-5-nano or gpt-5-mini for classification/extraction; gpt-5 for generation; o5 reasoning models for hardest reasoning tasks.

The tiering decision tree

  1. If the task is single-token output (classification with N choices), start with the cheapest model. Validate on 100 examples. Promote if accuracy is below your bar.
  2. If the task is structured extraction with a clear schema, start mid-tier. The cheapest models sometimes drop optional fields.
  3. If the task is open-ended generation (emails, summaries, reports), start mid-tier. Promote to the top tier if quality verification flags issues.
  4. If the task is multi-step reasoning, mathematical computation, or code generation, start top-tier. Demoting these often produces silently-wrong outputs that QA catches late.

Fix Two: Batching

A workflow that processes 500 tickets per hour can make 500 separate LLM calls (one per ticket) or batch them. Batching reduces per-call overhead, often improves cache hit rates (the system prompt is shared across the batch), and can leverage provider-specific batch discounts.

In-prompt batching

The simplest pattern: instead of calling the LLM once per ticket, accumulate 5-20 tickets and send them in one call. The user prompt becomes:

Classify the following tickets. Respond with a JSON array, one object per ticket, in the same order. Each object: {"ticket_id": "...", "intent": "...", "urgency": "..."}.

<ticket id="T-1001">...</ticket>
<ticket id="T-1002">...</ticket>
...

The system prompt is loaded once. The model's reasoning preamble is amortized. Output cost stays roughly the same (still one classification per ticket), but input cost drops because the system prompt is paid once per batch instead of once per ticket. Total per-ticket cost drops 30-60% in our measurements depending on batch size.

Caveats. Latency per ticket gets worse because you wait until 5-20 tickets accumulate. For real-time use cases this is a non-starter. For background processing it is often perfect. Batch size of 10 is a good default; above 20 you start to see model attention drop on the middle tickets in the batch, especially for smaller models.

Provider batch APIs

Anthropic Message Batches API (live since late 2024) accepts up to 100,000 requests in a single batch submission, processes them asynchronously within 24 hours, and discounts the cost by 50%. OpenAI's Batch API is similar: 50% discount, 24-hour SLA. These are perfect for non-real-time workloads — overnight CRM enrichment, weekly report summaries, end-of-day classification of all today's tickets.

The integration step in workflow tools is straightforward: instead of calling the synchronous API, submit the batch and store the batch ID, poll for completion (or set up a webhook callback), and process results when ready. In n8n this is two HTTP nodes plus a scheduled trigger. In Make and Zapier the pattern is identical.

A team that switched their nightly Salesforce summary workflow from sync calls to the Anthropic batch API in March 2026 saw their batch-processed lines drop from $89/month to $44/month. 50% off, no quality change, slightly higher latency (4 hours instead of 4 minutes) which was irrelevant for a nightly job.

Fix Three: Prompt and Result Caching

Two distinct techniques share the word "caching." Both apply.

Anthropic prompt cache (cache_control)

Anthropic introduced cache_control: {"type": "ephemeral"} as a request parameter that marks a portion of the prompt (typically the system prompt and any long context like documents or examples) as cacheable. On a cache hit, the cached portion is billed at 10% of the input rate. The cache lives 5 minutes; subsequent calls within that window hit the cache and pay the discounted rate.

To use it: structure your prompt so the stable content (system prompt, long examples, fixed context) is at the start, and add cache_control: {"type": "ephemeral"} to the last cacheable block. The first call writes the cache (full rate plus a 25% write surcharge). Subsequent calls within the window hit the cache (10% rate). Across many calls, the average input cost approaches 10% + a small write overhead.

Real numbers: for a workflow processing 30+ tickets/hour with a 2,000-token stable system prompt and a 200-token variable user content, the cache_control-enabled version cut input-token cost by 51% in our March 2026 deployment versus the same workflow with no caching. The same workflow processing 3 tickets/hour saw only 8% reduction — the cache never stayed warm because most calls landed outside the 5-minute TTL.

OpenAI prompt cache (automatic)

OpenAI's prompt cache is automatic for prefixes of 1024+ tokens that have been seen recently. There's no explicit marker. The catch is that your prompt structure has to put the stable content at the start. If your prompt opens with the variable user content and the system prompt comes second, you have no stable prefix to cache. Order matters: system prompt first, fixed context second, variable user content last.

Result caching

The other "caching" is result caching: don't call the LLM at all if you've already classified this exact input. For workflows that re-process the same content (a daily summary of an unchanged Slack channel, a re-classification of an already-classified ticket), keep a cache table keyed by a hash of the input. If the hash matches, skip the LLM call entirely.

The implementation: a Code node before the LLM that hashes the relevant input, checks a Supabase/Redis/Airtable table for a recent entry, and either returns the cached result or calls the LLM and writes the new result back. TTL depends on use case — for static content (a contract being analyzed), forever. For dynamic content (a customer's most recent ticket), 24 hours or shorter.

A team that added result caching to a "summarize all Slack channels at 9am" workflow cut LLM calls by 73% because most channels had no new activity. Same outputs, 73% fewer API calls.

Rate Limits: The Tuesday Morning Disaster

Every provider has rate limits. Anthropic in May 2026: tier-based, ranging from 50 requests/minute on the entry tier to 4,000 RPM and 1M tokens/minute on Tier 4. OpenAI: similar tier ladder. You don't think about rate limits until a marketing campaign sends 800 new signups into a workflow at 10am Tuesday, the workflow tries to summarize each new account in parallel, and the provider starts returning 429s.

The 429 response and what to do

A 429 (Too Many Requests) tells you that you've exceeded the request rate or token rate. The response includes a Retry-After header in seconds. The cheap fix: respect the retry header, back off, retry. The right fix: implement client-side rate limiting that paces your requests below the provider limit.

In n8n the pattern is the "SplitInBatches" node configured with batch size matching your provider's burst tolerance, plus a "Wait" node between batches. Make uses the Tools: Sleep module the same way. Zapier has built-in rate-limiting on premium plans via the Throttle action. All three support retry-on-429 natively.

Tier upgrades

The other lever is requesting a tier upgrade. Anthropic and OpenAI both let you request higher limits once you've hit a spend threshold. Anthropic's tier 4 (4,000 RPM, 1M TPM) requires $400 paid and 14 days of usage. OpenAI's tier 5 is similar. For workflows with predictable burst patterns (a 9am batch, a 6pm batch), the tier upgrade is often cheaper than re-architecting for low-RPM operation.

Provider failover

For genuinely high-availability workflows, a provider failover pattern matters. If Anthropic returns 529 (overloaded), fall over to OpenAI's GPT-5 or to Anthropic Haiku as a degraded backup. The system prompt needs to be portable across providers (mostly works, with minor adjustments per provider's preferences). The cost is the engineering of two prompts and the maintenance of two evals.

Real Cost Arithmetic for the Zendesk Workflow

The Zendesk private-note workflow from Lesson 1 of this chapter. Let's price it cold.

Single run, no caching

  • System prompt: 350 tokens
  • User prompt: ~1,500 tokens (truncated ticket text)
  • Total input: 1,850 tokens. At Sonnet 4.5 $3/M: $0.00555 input.
  • Output: ~125 tokens. At Sonnet 4.5 $15/M: $0.00188 output.
  • Total per run: $0.00743.
  • At 600 runs/day, 30 days: $133.74/month.

With Anthropic cache_control on the system prompt, ~80% cache hit rate

  • Cached input (system prompt, ~350 tokens, hit 80% of runs): 350 × 600 × 30 × 0.8 × $0.30/M = $1.51
  • Uncached input (system prompt for 20% miss + variable user prompt all runs): 350 × 600 × 30 × 0.2 × $3/M + 1500 × 600 × 30 × $3/M = $1.89 + $81.00 = $82.89
  • Cache write surcharge (~25% on first miss, complex math, approximate): ~$0.50/month
  • Output: ~125 × 600 × 30 × $15/M = $33.75
  • Total: ~$118.65/month. Savings: ~11%.

The cache benefit grows when the system prompt is large relative to user content. Our example has a small system prompt; for workflows with multi-thousand-token system prompts or examples, the cache savings are dramatically higher.

With model tiering: Haiku for first-pass triage, Sonnet only for complex tickets

Assume 60% of tickets are "obvious classifications" (refund / question / bug) that Haiku 4.5 handles correctly. Route those to Haiku. The remaining 40% go to Sonnet.

  • Haiku batch (60% × 600/day × 30): 10,800 runs × ($0.25/M × 1,850 + $1.25/M × 125)/run = 10,800 × ($0.00046 + $0.00016) = 10,800 × $0.00062 = $6.70
  • Sonnet batch (40% × 600/day × 30): 7,200 runs × $0.00743 = $53.50
  • Total without caching: $60.20. Down from $133.74 — savings: ~55%.

Layer caching on top of tiering: another ~10% savings. Total monthly cost: ~$54. From $134 to $54, no quality drop verified on a 100-example eval set. That is the level of cost reduction a one-afternoon refactor produces in real workflows.

The Monthly Cost Review (Half an Hour, Once a Month)

Every operator-builder we know who has cost under control runs a 30-minute review once a month. Pull the dashboards. Sort by spend. Ask the three questions. Make one change. Re-run the eval set. Ship the change.

The discipline is what matters. Cost creeps in 5% increments. A workflow that costs $40 in week one costs $200 in week ten because someone added a new trigger, the input grew, and the cache stopped warming. The monthly review catches that creep before it becomes a $1,400 surprise.

Cost arithmetic is not a one-time exercise. It is a discipline. Without the monthly review, you will rediscover the same problem after a finance-team Slack message. With it, you will be the one warning finance before they ask.

Three Anti-Patterns That Cause Cost Explosions

Anti-pattern one: the unbounded loop

A workflow that calls the LLM in a loop without a cap. We've seen one team's "agentic" workflow loop on a malformed input 47 times before a downstream timeout finally killed it. 47x cost on a single execution. The fix is a max-iteration counter checked in the workflow before each LLM call. We cover this in Level 3, but for now: if your workflow has a loop and no cap, add a cap today.

Anti-pattern two: the silent retry storm

Retries on retries. The LLM node has 3 retries. The outer workflow has 3 retries. The orchestrator has 3 retries. A single bad input triggers 27 LLM calls. Set retry policy at one layer only — usually the LLM node — and disable retries elsewhere.

Anti-pattern three: the "test in production" loop

An operator iterating on a prompt by re-running the workflow on production data 30 times an hour. Looks like 30 runs; bills like 30 runs. Use a dev environment with a fixed test fixture for iteration. Reserve production runs for "deploy and observe."

Key Takeaways

  • The two numbers operators defend in 2026: time-to-first-token p95 and dollars per million tokens. Track both per node, per month.
  • Three standard fixes when a node is too expensive or too slow: model tiering (cheaper model for simpler subtasks), batching (in-prompt batching or provider batch APIs at 50% off), and caching (Anthropic cache_control, OpenAI's automatic prefix cache, plus result-caching when inputs repeat).
  • Model tiering on the Zendesk workflow took monthly cost from $134 to $60 (~55% savings) with no quality drop, verified on a 100-example eval set.
  • Anthropic cache_control: {"type": "ephemeral"} writes the cache for 25% surcharge and reads at 10% of standard input rate; lives 5 minutes. Cache savings approach 50%+ when call volume keeps it warm.
  • OpenAI's automatic prefix cache (1024+ tokens) requires stable content first, variable content last. Re-order your prompts if you want the cache to hit.
  • Provider batch APIs (Anthropic, OpenAI) give 50% off for non-real-time workloads within a 24-hour SLA. Use them for overnight enrichment, weekly summaries, end-of-day classification.
  • Rate limits land on the day you didn't expect them. Implement client-side rate limiting with SplitInBatches/Sleep/Throttle. Request tier upgrades for predictable bursts. Add a provider failover for high-availability needs.
  • Anti-patterns to kill: unbounded loops, layered retry storms, "iterate in production." All three blow up costs faster than any model swap.
  • The 30-minute monthly cost review is the operator-builder's most under-rated habit. Pull dashboards, sort by spend, ask three questions, ship one change, re-run the eval. Cost creep dies in monthly reviews.