Skip to content

Rishi Rai

All Writings

LLM Observability, Tracing, and Debugging

Connect model calls to retrieval, tools, policy, validation, token use, cost, privacy controls, and reproducible replay.

Technical guide

For: Engineers operating LLM applications and investigating quality, latency, cost, privacy, or reliability issues.

2026-10-02

LLM Observability for Tracing, Replay, and Debugging

Traditional service monitoring covers requests, dependencies, latency, and failures. An LLM application also needs the context shaping a probabilistic result: prompt templates, model configuration, retrieved evidence, tools, safety decisions, token usage, and validation. Without it, a slow, expensive, or wrong response is difficult to explain.

LLM observability should connect application behavior to model behavior without turning telemetry into an uncontrolled copy of sensitive conversations. The goal is an end-to-end record that lets teams locate a failure stage, compare releases, replay a case safely, and quantify operational impact. This requires deliberate trace boundaries, versioned metadata, privacy controls, and evaluations that are tied to production traces.

Model the Request as a Trace

Create one trace for the user-visible operation, not only the final model call. Include spans for authentication, routing, retrieval, reranking, prompt construction, generation, tools, validation, and persistence. Parent-child relationships locate time and errors. Trace links connect asynchronous review or queued tool results.

A typical data flow is:

  1. The gateway creates a trace identifier and records the product operation, tenant class, and request timestamp.
  2. The orchestrator starts spans for routing and policy checks, attaching stable configuration versions.
  3. The retrieval layer records query characteristics, filters, result identifiers, ranks, and timing.
  4. The prompt builder records a template version and input references, with content captured only under an approved policy.
  5. The model client records provider, model, parameters, attempts, token counts, latency, finish reason, and normalized errors.
  6. Tool and validation spans record typed inputs, outcomes, policy decisions, and bounded summaries.
  7. The application records the final disposition, such as answered, abstained, reviewed, or failed.
  8. Metrics and logs reference the same trace identifier so operators can move between aggregate symptoms and one request.

Use the tracing conventions already adopted by the wider platform where possible. LLM-specific attributes should extend existing service telemetry, not form an isolated dashboard that loses database, queue, and network context.

Capture Reproducible Metadata

Version every component that can change output: provider and exact model identifier, prompt, system policy, parameters, retrieval index, embedding and reranker, tool schemas, application release, feature flags, and validators. A friendly model alias is insufficient if a provider can redirect it to another snapshot.

Separate high-cardinality trace attributes from low-cardinality metrics. A prompt version works well as a metric dimension; a user identifier or full error message usually does not. Unbounded labels can make metrics expensive and difficult to query. Keep request-specific details in traces or structured logs with appropriate indexing and retention.

span.name = "llm.generate"
span.attributes = {
  "operation": "support_answer",
  "model.provider": provider,
  "model.id": exact_model_id,
  "prompt.version": prompt_version,
  "policy.version": policy_version,
  "temperature": temperature,
  "input.tokens": usage.input_tokens,
  "output.tokens": usage.output_tokens,
  "cache.read.tokens": usage.cache_read_tokens,
  "attempt": attempt_number,
  "finish.reason": finish_reason,
  "result.status": normalized_status
}

Do not assume provider token accounting is interchangeable. Normalize common fields but retain raw usage data for investigation. Cost should be calculated against a versioned price table because rates, caching discounts, and regional charges can change. Keep estimated and invoiced cost distinct.

Measure Tokens, Cost, Latency, and Errors

Measure queue delay, time to first token, generation time, tool wait, and complete user-visible duration. Use percentiles for interactive experiences. Track retries separately so success does not hide throttling or timeouts.

Distinguish input, output, cached input, reasoning, and provider-specific token classes when available. Relate usage to operation and result. Rising input tokens can indicate longer messages, larger retrieval context, a prompt regression, or failed truncation; trace metadata separates these causes.

Normalize errors into categories such as authentication, quota, rate limit, timeout, content policy, malformed response, context limit, tool failure, validation rejection, and internal dependency failure. Preserve the provider code in a restricted field, but do not build alerts on unstable message text. Record cancellation and client disconnects as outcomes rather than generic server errors.

Dashboards should combine traffic, success and abstention rates, errors, latency percentiles, tokens, cost, validation failures, and tool errors. Alert on user impact or budget risk, distinguishing recovered transient errors from exhausted retries.

Privacy and Security by Design

Prompts, responses, retrieved passages, and tool arguments can contain secrets, personal data, source code, or confidential records. Telemetry is often replicated into indexes, archives, and third-party services, so payload capture expands the security boundary.

Define a capture policy by environment, task, and data classification. Prefer metadata and hashes when content is not required. If payloads are needed for debugging, use field-level allowlists, redact secrets and personal identifiers before export, encrypt data in transit and at rest, enforce role-based access, and set short retention periods. Sampling should happen after mandatory security events are retained but before verbose content is exported.

Use opaque subject identifiers, never emails or names. Keep sensitive values out of labels, span names, and trace identifiers. Audit sensitive trace access and support deletion across storage, archives, and evaluation datasets. Apply threat modeling and egress controls to exporters.

Debug value does not justify collecting every payload. Capture the minimum evidence needed for a defined operational question, and make expanded capture explicit, time-bounded, and auditable.

Design Safe Replay and Debugging

Replay is most useful when a trace contains immutable references to inputs and configuration. A replay package can include sanitized input, prompt and policy versions, retrieved document identifiers, model parameters, tool schemas, and expected disposition. It should never execute production side effects by default.

Use separate replay modes. Inspection reconstructs the prompt and dependencies without a model call. Model replay invokes an original or candidate model. Comparative replay evaluates selected variants. Stub write-capable tools, payments, messages, and mutations. Route any necessary live dependency to an authorized test account.

Exact text reproduction may be impossible as models, sampling, and external data change. Compare validation status, cited sources, task completion, policy compliance, latency, and cost. Record the environment and distinguish faithful reconstruction from approximation.

Start debugging at the final disposition, find the first unexpected span, compare its inputs and versions with a known-good trace, then replay the narrowest failing stage. Changing the final prompt cannot repair defective retrieval or tool data.

Trade-Offs and Common Mistakes

Detailed traces improve diagnosis but increase storage, latency, and privacy exposure. Sampling controls volume but can miss rare failures. Tail-based sampling retains slow or failed traces but requires buffering. Redaction may remove details needed for semantic failures. Set policy by risk tier and purpose.

  • Tracing only the model call and omitting retrieval, tools, queues, and validators.
  • Recording a prompt name without an immutable template version or rendered-input references.
  • Putting full prompts, user identifiers, or raw errors into metric labels.
  • Calculating cost from a current price list when investigating historical traces.
  • Retrying errors without recording attempts, backoff, and the eventual outcome.
  • Replaying workflows against live write-capable tools.
  • Assuming the same input must produce identical text and treating variation as the only defect signal.
  • Collecting telemetry indefinitely without retention, deletion, or access-audit controls.

Testing the Observability System

Telemetry needs its own tests. In integration environments, issue known requests and assert that traces contain required spans, parent relationships, versions, usage fields, and final status. Inject timeouts, rate limits, malformed model output, failed tools, and validation rejections. Verify that error categories and retry counts are correct and that trace context survives queues and asynchronous workers.

Run fixtures with synthetic secrets and personal fields, confirming that exporters never receive prohibited values. Test roles, audit events, retention, deletion, and sampling under realistic traffic.

For release evaluation, select representative sanitized traces and replay them against proposed prompt, model, retrieval, or policy changes. Compare task-specific quality measures together with latency, errors, tokens, and estimated cost. Define acceptance thresholds before reviewing results, and retain the evaluated configuration so a regression can be investigated later.

Conclusion

LLM observability connects model experimentation to dependable operations. End-to-end traces show how routing, retrieval, prompts, models, tools, and validators produced a result. Versions make comparisons meaningful; token, cost, latency, and error signals expose consequences. Privacy-aware capture and isolated replay preserve diagnostic value without unnecessary access or repeated side effects. Debugging then becomes a controlled investigation rather than guesswork.