Monitoring & Optimization
Marcus Chen is the platform lead at Ledgerline, a hypothetical fintech that runs a five-step AI agent to process uploaded loan documents: classify the document, extract fields, validate against policy, call an external credit API, and draft a decision summary. In the demo it was fast. In production it is a black box. Everything technically works, and Marcus still cannot answer the three questions his VP keeps asking: is it fast enough, is it too expensive, and where is it breaking. This chapter is about the instrumentation and the habits that turn those questions into answers you can look up.
The Hidden Orchestration Challenges
The symptoms at Ledgerline are the ordinary ones. Support tickets say the system is "slow sometimes," which is unactionable. The monthly model bill jumped 40 percent with no obvious cause, and nobody can say which step spent the money. Last week a silent failure in the extraction step produced dozens of empty summaries before anyone noticed, because every one of those runs returned a successful response and completed on time. Tasks complete, outputs are generated, and none of that tells Marcus whether the orchestration is working well. Without observability he is flying blind while the dashboard says green.
That gap is what observability closes. Traditional monitoring answers "is it up?" Observability answers "why is it slow, where is the money going, and which step is failing?" by exposing the internal state of the system from the outputs it emits. It is a property of the system rather than a product you buy: a system is observable to the degree that you can reconstruct what happened inside it from what it emitted. For an AI orchestration that means deliberately collecting three signals at every step: metrics, which are numbers over time; logs, which are discrete events with context; and traces, which follow the causal path of one request through every step. The goal is not just an alarm when something breaks. It is understanding the system deeply enough to debug it at 2am and to optimize it on a normal Tuesday.
Key Metrics for Orchestrations
The first mistake teams make is tracking averages. An average hides the tail, and the tail is where users suffer. Marcus tracks latency as a distribution, per step and end to end, using percentiles. If the end-to-end p50 is 4 seconds but the p95 is 22 seconds, one user in twenty is having a terrible experience that the average erases entirely, and that user is the one who files the ticket saying the system is slow sometimes. The metrics worth instrumenting for any orchestration fall into four groups.
| Category | Metric | Why it matters |
|---|---|---|
| Latency | p50 / p95 / p99, per step and end to end | Percentiles expose the slow tail that averages hide; per-step tells you where time goes |
| Cost | Input + output tokens per call, cost per request, cost per successful outcome, cost by step, model selection ratio, cost trend over time | Token spend is the dominant variable cost; per-outcome catches expensive retries and the trend catches creeping bills |
| Reliability | Error rate by category, retry count, timeout rate, mean time to recovery, availability, all per step | A step with a 15% retry rate is silently tripling its own cost and latency; recovery time tells you how long users feel it |
| Throughput & quality | Requests per minute, queue depth, task success rate, human-override rate | Reveals saturation and whether "success" means the output was actually correct |
Several of these deserve unpacking, because teams routinely instrument the easy version instead of the useful one. On reliability, a bare error rate tells you something is wrong; error rate broken down by category tells you which failure is worth fixing first, since timeouts, malformed model output and downstream rejections call for entirely different remedies. Retry count matters separately because retries succeed, so they never show up as failures and yet consume the budget twice. Mean time to recovery measures how long the system stays degraded after an error, and availability records the share of time the orchestration was responsive at all. Queue depth belongs alongside them: a depth that keeps climbing is the clearest early signal of overload, well before latency degrades enough to notice.
On cost, two views change decisions. Cost per successful outcome, not cost per API call, is the honest number: if a step succeeds only after two retries on average, its real cost is triple the sticker price. Cost by step tells you which component to attack, and the model selection ratio, the share of requests routed to each model, explains changes that per-call pricing cannot. A ratio that drifts toward a more expensive model because of a fallback rule nobody reviewed is a common cause of a bill that rises while traffic is flat. Watching the cost trend rather than the current total is what would have let Marcus catch his 40 percent jump in the week it started.
On quality, task success rate has to be measured against a definition of correct, not just "returned 200 OK." Marcus's empty-summary incident returned successful HTTP responses the whole time, so every reliability metric he had was green while the product was broken. A quality metric, the percentage of summaries with all required fields populated, would have caught it in minutes instead of days. Human-override rate is the cheap companion: when reviewers start correcting outputs more often, quality has degraded even if nothing errored.
Building Your Observability Stack
The three signals do different jobs and you need all three. Metrics are cheap, aggregated and good for dashboards and alerts, but they cannot tell you what happened to one specific request. Logs are discrete events with rich context, ideal for debugging a known request but expensive to store and slow to search at volume. Traces stitch a single request's journey across every step into one timeline, which is the signal that actually localizes a problem in a multi-step orchestration. Collect metrics at the decision points of the orchestration, meaning every model invocation, every task completion and every error occurrence, so that coverage follows the structure of the workflow rather than wherever someone remembered to add a counter.
The keystone technique is the correlation ID. Marcus generates one trace ID when a document enters the system and propagates it through all five steps and into every log line and every downstream API call. Now when a user reports a bad decision at 3:14pm, he pulls the single trace ID and sees the entire causal path: which model was called, how many tokens each step used, how long each waited, and exactly where it went wrong. Without a correlation ID the same investigation means grepping five services and guessing which log lines belong together, which is why teams without one tend to conclude that the problem is unreproducible.
Structured logging is what makes those log lines usable. Emit JSON with standard fields on every event, the trace ID, step name, token counts and duration, rather than free-text prose you cannot query. A line recording that a task completed in 234ms is only valuable if you can aggregate every such line by step without writing a parser. Dashboards then turn those aggregates into visible patterns, which matters because most degradations are gradual and only look like anything when plotted against last week. Alerts sit on top, firing when a metric crosses a threshold you chose deliberately, and the standard they should be held to is catching problems before your customers do.
Practical stack guidance follows from that. Instrument with an open standard such as OpenTelemetry so you emit traces, metrics and logs in a vendor-neutral format and are not locked to one backend. Make logging asynchronous so it never blocks the request path. Sample traces intelligently: keep 100 percent of errors and slow requests, and sample the healthy majority at a low rate such as 1 to 5 percent. Set retention deliberately: high-resolution metrics for about 30 days with lower-resolution rollups kept longer, logs for around 90 days, and sampled traces retained longest because their volume is small. Instrument once at the orchestration layer so every step is covered by construction rather than bolted on later, one service at a time, with gaps exactly where the interesting failures live.
The tooling itself is a category choice more than a brand choice, and the categories map onto the three signals plus a place to look at them.
| Layer | What it does | Commonly used options |
|---|---|---|
| Metrics | Stores numeric time series for dashboards and alerts | Prometheus, InfluxDB, Datadog, New Relic |
| Logs | Ingests and indexes structured events for querying | ELK Stack, Splunk, Datadog, CloudWatch |
| Traces | Assembles per-request timelines across services | Jaeger, Zipkin, Datadog APM |
| Dashboards | Visualizes metrics so patterns become visible | Grafana, Kibana, Datadog |
Identifying and Acting on Bottlenecks
With traces in place, optimization stops being guesswork. The governing observation is lopsided: typically 80 percent of latency comes from 20 percent of your orchestration, which means the return on tuning is concentrated in a small number of places and effectively zero everywhere else. Before you look at the data, the candidate bottlenecks are worth knowing, because they recur across almost every orchestration: a single slow model component, sequential execution that should be parallel, repeated API calls that should be cached, rate limiting on a frequently used service, and insufficient resources leaving work queued behind underpowered hardware. Marcus opens a trace waterfall for a slow request and sees where the wall-clock time actually goes across the five steps. Here is the per-step breakdown with clearly hypothetical numbers for a request whose end-to-end p95 is 22 seconds.
| Step | p95 latency | Tokens | Retry rate | Notes |
|---|---|---|---|---|
| 1. Classify | 0.8 s | 600 | 1% | Small model, healthy |
| 2. Extract fields | 3.1 s | 4,200 | 2% | Large prompt, big output |
| 3. Validate policy | 0.5 s | 0 | 0% | Pure code, no model call |
| 4. Credit API | 15.6 s | 0 | 18% | External call, frequent timeouts and retries |
| 5. Draft summary | 2.0 s | 2,800 | 3% | Model call |
The waterfall makes the answer obvious: step 4, the external credit API, is 70 percent of the total latency, and its 18 percent retry rate means nearly one in five requests is paying the timeout penalty twice. The team's instinct had been to optimize the model prompts in step 2, which would have shaved maybe half a second while the real 15-second problem sat untouched. Notice that the cost bottleneck is somewhere else entirely: step 2's 4,200-token calls dominate the bill while contributing little to latency, and step 4 dominates latency while costing nothing in tokens. This is the core lesson of bottleneck analysis. Optimize the biggest measured contributor to the metric you care about, and accept that different metrics will point you at different steps.
The Optimization Toolkit
Once the waterfall has told you where to work, the available moves are a small and reusable set. Choosing among them is a matter of matching the remedy to the shape of the bottleneck rather than reaching for the technique you last enjoyed using.
- Model swaps. Replace a slow model with a faster one, accepting slightly lower accuracy where the step can tolerate it. A classification step that feeds a downstream validator usually can; a step whose output goes straight to a customer usually cannot.
- Parallelization. Run independent steps concurrently instead of in sequence. This is the highest-value structural fix when a trace shows steps waiting on each other for no reason other than the order somebody wrote them in.
- Caching. Store the results of expensive operations and repeat lookups so identical work is not paid for twice, which is particularly effective against external calls with stable answers over a short window.
- Batching. Process multiple inputs in one call where the interface supports it, trading a little per-request latency for a large reduction in overhead and cost at volume.
- Resource scaling. Add capacity when queue depth, not per-step latency, is the thing that is climbing. Scaling is the right answer to saturation and the wrong answer to a slow dependency.
- Workflow restructuring. Reorganize the orchestration to remove dependencies altogether, for example by moving a cheap validating step earlier so that expensive steps never run on inputs that were going to be rejected.
Marcus's fixes follow directly from his diagnosis. For step 4 he adds a timeout budget and a circuit breaker so a struggling external API fails fast instead of hanging, caches credit results for repeat lookups within a short window, and retries with exponential backoff instead of immediately, which stops retries from arriving while the dependency is still overloaded. That alone is projected to cut end-to-end p95 from 22 to roughly 8 seconds. Separately, because the cost dashboard shows step 2's 4,200-token calls dominating the bill, a second workstream trims the extraction prompt and moves it to a cheaper model tier, projected to cut per-request model cost by about 30 percent. Two independent optimizations, each justified by data rather than intuition, each targeting the step that actually drives the metric it improves.
Alerting on What Matters
Then you alert on what you learned. Marcus sets a small number of high-value alerts: end-to-end p95 above SLA, step-level error or retry rate above threshold, daily cost above budget, availability dropping, and the quality metric, fully populated summaries, falling below a floor. Each corresponds to an outcome someone would be upset about, which is the test an alert must pass to earn a place. He deliberately does not alert on every anomaly, because an inbox full of false alarms trains people to ignore the one that matters. Start with a handful, then refine using real incidents: every alert that fired without mattering gets tightened or deleted, and every incident nobody was warned about produces a new one.
Making Observability an Operating Rhythm
Instrumentation is worthless if no one looks at it. What turned Ledgerline around was not the tooling but the habit Marcus built on top of it. Every week the team reviews four numbers together: end-to-end p95, cost per successful outcome, the worst step by error rate, and the quality metric. Each anomaly is traced to a specific step using a correlation ID before anyone proposes a fix, so debates are settled by the waterfall rather than by seniority. Every optimization is stated as a measured before-and-after, so the team can prove the change helped and can revert it if the next week's numbers disagree.
That rhythm is what converts a dashboard into continuous improvement: observe, localize with traces, fix the biggest measured contributor, verify, repeat. It also changes what the team argues about. The question "should we optimize the prompts?" used to be answered by whoever felt most strongly; now the waterfall answers it in a minute. That is the difference between a system that merely runs and a system you understand and control.
Anti-Patterns to Avoid
Most observability failures are not missing tools. They are instrumentation that does not answer the question anyone actually has.
- Tracking averages. A p50 of 4 seconds alongside a p95 of 22 describes two completely different products, and only the percentile view shows you the second one.
- Measuring cost per call instead of per successful outcome. Retries succeed, so they never appear as failures while quietly multiplying the bill. A step that retries twice on average costs three times its sticker price.
- Treating a 200 response as success. Ledgerline's empty-summary incident returned successful responses for days. Without a quality metric, every reliability signal stays green while the product is broken.
- Logging without a correlation ID. Rich logs scattered across five services with nothing tying them together turn every investigation into grepping and guessing.
- Optimizing the fun step. The team's instinct was to tune prompts worth half a second while a 15-second external call sat untouched. Choose the biggest measured contributor, not the most interesting one.
- Alerting on everything. An inbox of false alarms trains people to ignore alerts, so the noisy configuration is worse than a small deliberate set of high-signal ones.
Practice Prompts
Work these against an orchestration you actually run.
- Define correct. Write your quality metric in one sentence, expressed as a property of the output rather than of the response code, and work out whether you could compute it today.
- Draw the waterfall. Sketch your orchestration's steps and fill in p95 latency, tokens and retry rate for each. Any cell you cannot fill is an instrumentation gap, and the gaps are the finding.
- Trace one request end to end. Reconstruct a real request's full path from your existing logs, and time yourself. The difficulty is a direct measure of whether your correlation ID works.
- Separate the two bottlenecks. Identify the step that dominates latency and the step that dominates cost, and confirm whether they are the same step. If you assumed they were, check.
- Audit your alerts. For every alert configured today, name the outcome someone would be upset about. Delete or tighten the ones with no answer.
Reflection
These questions are worth answering before your next incident rather than during it.
- If a user reported a bad output from a specific request an hour ago, how long would it take me to reconstruct exactly what happened to it?
- Which of my current metrics would have stayed green through a silent failure like the empty-summary incident?
- Am I optimizing the step that dominates the metric I care about, or the step I find most interesting to work on?
- Would I notice if a fallback rule quietly shifted my model selection ratio toward a more expensive model?
- How many of my alerts fired in the last month, and how many of those firings changed what anyone did?
Glossary
- Observability. The property of a system that lets you understand its internal state from its external outputs, as distinct from monitoring, which answers only whether the system is up.
- Correlation ID. One identifier generated when a request enters the system and propagated through every step, log line and downstream call, making a request's full path reconstructable.
- Cost per successful outcome. Spend divided by successful results rather than by calls, the only cost view that exposes the price of retries and rework.
- Model selection ratio. The share of requests routed to each model, used to explain cost changes that per-call pricing alone cannot account for.
- Mean time to recovery. How long the system stays degraded after an error before returning to normal service, which is what users actually experience during an incident.
- Trace sampling. Retaining all error and slow traces while keeping only a small percentage of healthy ones, preserving diagnostic value at a fraction of the storage cost.
Related Lessons
Monitoring sits at the end of the orchestration build and at the start of the improvement loop.
- Workflow Management & Execution covers building the orchestration whose behaviour this chapter teaches you to observe.
- Process Analysis & Opportunity Identification moves from tuning a live system to choosing which systems are worth building at all.
- Orchestration Architecture & Patterns explains the structural choices, including sequencing and parallelism, that determine where bottlenecks occur.
- Error Handling and Fallback Design for AI Workflows develops the timeout, retry and circuit-breaker patterns applied here to the failing external dependency.
- Token and Cost Economics of AI goes deeper on the token spend that dominates variable cost.
- AI Observability and Production Monitoring extends these practices to model behaviour in production beyond the orchestration layer.
Closing
Observability is the difference between orchestrations that work and orchestrations that work well. Ledgerline's system was never down, and that was exactly the problem: every signal Marcus had was designed to detect outages, and none of his failures were outages. They were a slow tail hidden inside an average, a bill rising for reasons nobody could attribute, and a broken step returning perfectly healthy responses.
Collect meaningful metrics, log structurally with a correlation ID, trace every request through every step, then build the weekly rhythm that makes someone actually look. That is what turns a black box into a system you can debug at 2am and improve on a normal Tuesday, and it is the discipline the following chapter carries forward when it asks which processes are worth automating in the first place.
Key Takeaways
- Monitoring answers whether the system is up; observability answers why it is slow, where the money goes and which step is failing.
- Track latency as percentiles per step, not averages, because the average erases the tail where your unhappiest users live.
- Measure cost per successful outcome, by step, alongside the model selection ratio and the cost trend, since retries and silent routing changes move a bill.
- Define success against a quality metric, not an HTTP status; a silent failure can return healthy responses for days.
- A correlation ID propagated through every step and log line is what makes multi-step debugging tractable.
- Roughly 80 percent of latency comes from 20 percent of the orchestration, so find the biggest measured contributor before choosing among model swaps, parallelization, caching, batching, scaling and restructuring.
- Keep alerts few and outcome-based, and hold a weekly review, because instrumentation nobody looks at changes nothing.
Frequently Asked Questions
How much overhead does observability add? Minimal when done well. Metrics are near-free. Logging is asynchronous so it never blocks execution. Trace sampling keeps overhead low while preserving visibility into the requests that matter, since you retain every error and slow request and only a small share of the healthy ones. The cost of observability is far less than the cost of optimizing blind or shipping a silent failure like the empty-summary incident.
What should I alert on? Outcomes you care about: error and retry spikes, latency past SLA, cost overages, availability drops, and quality-metric regressions. Start with a handful of high-signal alerts and refine them from real incidents. Do not alert on every wiggle in a graph, because a team that has learned to dismiss alerts will dismiss the important one too.
How long should I retain observability data? A reasonable default is high-resolution metrics for about 30 days with lower-resolution rollups kept longer, logs for around 90 days, and sampled traces retained longest since their volume is small. Tune to your trend-analysis needs and storage budget so you can see patterns without letting storage cost explode.
Should I measure cost per call or per outcome? Per successful outcome. Cost per call hides retries and rework. A step that quietly retries twice costs three times what its per-call number suggests, and only the per-outcome view reveals it.
Skill.re