Skip to content

Rishi Rai

All Writings

Reducing Cost and Latency in Generative AI Systems

Optimize the complete request path through measurement, model routing, context budgets, caching, batching, streaming, and bounded fallback.

Engineering analysis

For: Engineers balancing quality, responsiveness, reliability, and operating cost in generative AI applications.

2026-10-02

Engineering Generative AI for Cost and Latency

Generative AI performance is a systems problem, not a single model benchmark. A request may pass through authentication, retrieval, prompt construction, one or more model calls, validation, and persistence before a user sees a useful result. Each stage contributes delay, cost, and failure risk. Optimizing only token price or model speed often moves the bottleneck somewhere else.

The practical goal is to deliver an acceptable answer inside an explicit service envelope. That envelope should define quality, time to first useful output, completion time, and marginal cost. Product tiers and workflows can have different envelopes. An interactive editor may value fast streaming, while an overnight analysis may favor batching and stronger reasoning.

Start With a System Model

Represent the request path as measurable stages. A typical architecture contains an API gateway, policy layer, retrieval service, prompt builder, model router, provider adapters, output validator, and cache. Keep provider-specific behavior behind adapters so routing and fallback logic do not leak into product code.

  1. Classify: identify the task, risk level, latency budget, and required capabilities.
  2. Gather: retrieve only evidence relevant to the request and authorized for the caller.
  3. Generate: select a model and parameters from policy, not scattered conditionals.
  4. Validate: check structure, citations, safety rules, and task-specific quality.
  5. Deliver: stream or return the result, then record usage and evaluation signals.

Every optimization should name the stage it changes, the metric it improves, and the quality or reliability risk it introduces.

Measure Before Tuning

Capture one trace per request with a stable correlation identifier. Record queue time, retrieval time, provider connection time, time to first token, generation duration, validation time, input and output tokens, cache outcome, model choice, retry count, and estimated cost. Report percentiles by task class and model rather than relying on averages. A healthy median can conceal a damaging p99 caused by retries or provider throttling.

Cost accounting should include embeddings, reranking, tool calls, retries, and storage, not just the final completion. Associate usage with a feature, tenant, and request class while avoiding raw sensitive prompts in telemetry. Track quality beside operational metrics. Otherwise, a shorter prompt can appear successful even when it removes evidence needed for correct answers.

  • Define a latency budget for each stage and alert on sustained budget exhaustion.
  • Measure accepted answers per unit cost, not tokens per dollar in isolation.
  • Separate time to first token from time to complete because users experience them differently.
  • Sample full traces for debugging and retain aggregate metrics for longer trend analysis.

Route Models With Explicit Policy

Model routing works when task categories have known requirements. A simple classifier can distinguish extraction, rewriting, coding, conversational lookup, and high-risk analysis. The router then chooses the least expensive model that satisfies the capability and quality threshold. Avoid routing solely by prompt length; short requests can demand difficult reasoning, and long requests can be mechanical.

function chooseRoute(request, budget) {
  const task = classifyTask(request)
  const risk = assessRisk(request)

  if (risk === "high") return routes.reviewRequired
  if (task === "extract" && budget.latencyMs < 1200) {
    return routes.fastStructured
  }
  if (task === "analysis" && budget.quality === "high") {
    return routes.reasoning
  }
  return routes.balanced
}

Version the routing policy and log the version on every trace. Shadow a proposed policy against production traffic without serving its output, then compare cost, latency, and evaluation results. Maintain per-route concurrency limits so one expensive workload cannot starve interactive requests.

Control Context as a Budget

Context windows are capacity limits, not targets. Set token allocations for system instructions, conversation history, retrieved evidence, tool results, and output. Reject or compress material when a category exceeds its allocation. Preserve recent turns and durable facts separately; blindly keeping the last N messages can discard the decision that gives later turns meaning.

Summarization can reduce cost, but it is lossy. Store the source identifier and summary version, and refresh summaries when their source changes. For large documents, retrieve small sections first and expand neighboring sections only when needed. Reserve output tokens according to the task, since an oversized maximum can increase tail latency and weaken cost controls even if it is rarely consumed.

Use Caching, Batching, and Streaming Deliberately

Caching

Cache deterministic intermediate work such as embeddings, document parsing, authorization-filtered retrieval results, and prompt prefixes supported by the provider. A completion cache is safe only when the key includes model, parameters, prompt template version, relevant data version, locale, and authorization scope. Give volatile facts short lifetimes and provide invalidation when source data changes.

Batching

Batch embeddings, evaluations, and offline generations when the provider offers lower prices or better throughput. Bound batch size by token count, not item count, and impose a maximum wait so sparse traffic does not create excessive delay. Do not batch unrelated interactive requests if one slow item delays every response or complicates tenant isolation.

Streaming

Streaming improves perceived latency but does not reduce total compute. Buffer enough output to validate encoding and basic policy before display. For structured responses, stream progress events or validated sections instead of exposing incomplete JSON. Cancellation must propagate to the provider; closing the client connection without cancelling generation still incurs cost.

Make Retrieval Earn Its Cost

Retrieval adds embedding, search, reranking, and context costs, so invoke it only for tasks that need external or current knowledge. Apply tenant and document authorization before ranking. Use metadata filters to reduce the candidate set, then rerank a small number of candidates. Include stable source identifiers in the prompt so validation can confirm that claims refer to supplied evidence.

Evaluate retrieval independently with known queries and relevant passages. Measure recall at the candidate stage and precision after reranking. If retrieval returns weak evidence, the system should ask for clarification or state that evidence is insufficient rather than filling the gap through generation.

Design Fallbacks and Quality Gates

Fallbacks should preserve semantics. Switching to a model that cannot call required tools or satisfy a schema is not graceful degradation. Define a compatibility matrix covering context size, tool support, structured output, regional availability, and safety controls. Retry only transient failures, use exponential backoff with jitter, and cap the total request deadline. Hedged requests can reduce tail latency but may nearly double cost, so reserve them for high-value operations and cancel the loser promptly.

Quality gates can combine schema validation, groundedness checks, policy rules, and lightweight task-specific assertions. A failed gate may trigger one bounded repair attempt, a compatible model fallback, or human review. Record the original and repaired outcomes for evaluation.

result = generate(route, prompt, deadline)
verdict = validate(result, evidence, contract)

if verdict.repairable and budget.allowsRepair:
    result = repair(result, verdict.errors)
    verdict = validate(result, evidence, contract)

return verdict.accepted ? result : safeFallback(verdict)

Security, Trade-offs, and Common Mistakes

Treat retrieved text, tool output, and user input as untrusted data. Keep instructions separate from evidence, enforce tool permissions outside the model, redact secrets before provider calls, and encrypt cached content. Cache keys and logs must not expose personal data. Rate limits, per-tenant budgets, and maximum token controls reduce both abuse and accidental spend.

  • Do not add retries without a shared deadline and idempotency plan.
  • Do not compress context without measuring answer quality by task class.
  • Do not share caches across authorization boundaries.
  • Do not treat streaming as proof that backend latency improved.
  • Do not optimize against a single provider snapshot without testing fallback behavior.

Test the Whole Envelope

Build a versioned evaluation set containing normal requests, long context, ambiguous queries, unsafe instructions, retrieval failures, provider timeouts, and malformed outputs. Run it for every prompt, routing, model, and retrieval change. Use deterministic assertions for schemas and permissions, rubric-based scoring for open-ended quality, and load tests for queueing and tail latency. In production, compare canary cohorts on accepted-answer rate, cost, latency, repair frequency, and human escalation.

Conclusion

Cost and latency improve sustainably when they are managed as constraints across the full request path. Instrument each stage, route by capability, budget context, cache with correct isolation, batch suitable work, stream responsibly, and require retrieval and fallbacks to pass explicit quality gates. The result is not merely a cheaper model call, but a service whose performance, quality, and risk can be explained and controlled.