Retrieval augmented generation is most reliable when it is treated as an information system, not as a prompt trick. The language model should receive a small, traceable set of passages that are likely to answer the question, and the application should preserve enough evidence to explain why those passages were selected. Reliability comes from controlling each stage, measuring it separately, and refusing to imply support that the source material does not provide.
System model and data flow
A practical RAG system has two connected paths. The indexing path converts governed source documents into searchable records. The query path turns a request into candidates, ranks them, builds context, generates an answer, and attaches citations. Keeping these paths explicit makes stale data, weak retrieval, and unsupported generation easier to diagnose.
- Ingest documents with stable identifiers, version information, ownership, access rules, and update timestamps.
- Normalize content while retaining headings, tables, code boundaries, and source offsets.
- Create chunks, embeddings, lexical fields, and metadata filters in an index.
- At query time, apply authorization and intent aware filters before retrieving candidates.
- Combine lexical and semantic results, rerank a limited pool, and remove redundant passages.
- Construct a context package with passage identifiers, then ask the model to answer only from that package.
- Validate citations and record structured traces for later evaluation.
Chunking is an information design decision
Fixed token windows are easy to implement, but they often split definitions from qualifications or combine unrelated sections. Prefer structure aware chunking for documents with meaningful headings. A chunk can include its heading path, a bounded body, and a small overlap when a sentence crosses the boundary. Code and tables usually need specialized handling because splitting them at arbitrary token counts can destroy their meaning.
Chunk size requires compromise. Small chunks can improve match precision but may omit context. Large chunks preserve context but dilute the signal and consume the model context window. Evaluate several policies against real questions. Store source offsets and a content hash so a citation can be mapped back to the exact indexed version. Do not use overlap as a substitute for coherent boundaries, because heavy overlap creates duplicate candidates and misleading confidence.
Hybrid retrieval and reranking
Semantic retrieval handles paraphrases and conceptual similarity. Lexical retrieval is strong for identifiers, error messages, product names, and exact phrases. A hybrid system uses both. Results can be merged with reciprocal rank fusion, which depends on rank rather than incomparable raw scores. Metadata filters should narrow the eligible corpus, but overly specific filters can silently eliminate the only useful passage.
Retrieve a broader candidate set than the final context requires, then rerank it. A cross encoder or carefully constrained model grader can compare the query with each passage more directly than an embedding score can. Reranking adds cost and latency, so cap candidate counts, batch requests, cache safe results, and define a timeout fallback. Diversity selection should prevent several near duplicate chunks from occupying the complete context.
Implementation guidance
Make the retrieval contract structured. Each result should carry a document identifier, chunk identifier, text, version, authorization scope, retrieval signals, and final rank. Keep prompt formatting separate from retrieval logic. This permits offline evaluation of the exact candidate list without invoking a generator, and it prevents formatting changes from obscuring retrieval regressions.
def answer(request, principal):
filters = allowed_scope(principal, request.tenant_id)
lexical = lexical_search(request.question, filters, limit=40)
semantic = vector_search(embed(request.question), filters, limit=40)
candidates = reciprocal_rank_fusion(lexical, semantic)
candidates = deduplicate(candidates)
ranked = rerank(request.question, candidates[:50], timeout_ms=800)
context = select_diverse(ranked, token_budget=5000)
if not context or ranked[0].score < MIN_RELEVANCE:
return insufficient_evidence_response()
draft = generate(
question=request.question,
passages=[passage_with_id(item) for item in context],
instruction="Use only supplied passages and cite every material claim."
)
return validate_and_attach_citations(draft, context)
The relevance threshold must be calibrated with labeled examples rather than copied from another model or index. If the reranker times out, the fallback should be observable and conservative. A system may use fused ranking directly, return fewer passages, or decline the answer depending on the risk of the use case.
Citations as verifiable data
A citation should identify the passage that supports a specific claim, not merely the document that seems related. Give the model opaque passage identifiers, require those identifiers in structured output, and reject identifiers that were not supplied. After generation, verify that cited passages still exist, belong to the authorized result set, and contain evidence relevant to the associated claim. The verifier can flag unsupported claims, but it should not silently invent replacement citations.
A fluent answer with a source list is not necessarily grounded. Grounding requires a checkable relationship between each important claim and a retrieved passage.
When evidence conflicts, preserve that conflict in the response. When evidence is absent, an explicit insufficient evidence result is more reliable than a plausible completion. Citation display should include useful source labels and version dates when available, without exposing storage paths or internal access metadata.
Evaluation and testing
Build an evaluation set from representative questions, difficult terminology, common follow ups, and known unanswerable requests. Record acceptable documents or chunks where annotation is feasible. Separate retrieval metrics from answer metrics so a generation failure is not blamed on search, and a lucky answer is not allowed to hide weak retrieval.
- Measure recall at k for whether relevant evidence entered the candidate set.
- Measure ranking quality with reciprocal rank or normalized discounted cumulative gain when graded relevance is available.
- Measure context precision, duplicate rate, citation validity, claim support, and abstention quality.
- Track latency, index freshness, empty result rate, reranker fallback rate, and token usage.
- Run deterministic tests for filters, document deletion, version replacement, citation parsing, and access boundaries.
Use human reviewers for ambiguous relevance and factual support. Model based graders can assist at scale, but they need fixed rubrics, calibration examples, and periodic agreement checks against humans. Maintain regression slices by content type, language, tenant, query complexity, and answerability. A single aggregate score can conceal a severe failure in a small but important category.
Security and privacy
Authorization belongs inside retrieval, not only in the user interface or generator prompt. Enforce tenant and document permissions before candidates enter shared caches or model context. Treat indexed documents as untrusted input because they can contain prompt injection text. The generator should never follow instructions found in retrieved content, invoke tools because a passage requests it, or reveal hidden prompts and credentials.
Minimize sensitive content in logs, encrypt indexes where appropriate, define retention rules, and propagate deletions to chunks, embeddings, caches, and evaluation snapshots. Validate file types and parser limits during ingestion. Restrict administrative reindex operations, record provenance, and protect embedding services from uncontrolled bulk extraction.
Failure modes and compromises
Common mistakes include tuning only the prompt, indexing content without versions, using vector search alone, placing every retrieved passage into context, and evaluating with a few demonstration questions. Another subtle failure is embedding access controlled content into a global index and attempting to filter it after retrieval. Even if final results are filtered, scores, caches, or traces can leak information.
More retrieval stages can improve selectivity, but they increase latency, operating cost, and the number of components that can fail. Larger contexts improve coverage while increasing distraction and generation cost. Aggressive abstention reduces unsupported answers but can frustrate users when evidence is phrased unexpectedly. These choices should be set by risk, measured on stable datasets, and monitored by slice.
Conclusion
A dependable RAG system makes evidence movement visible from source version to final citation. Structure aware chunks, hybrid retrieval, calibrated reranking, bounded context, and claim level citations provide the core controls. Separate evaluation of retrieval, grounding, and abstention reveals where quality is actually lost. With authorization, provenance, regression tests, and conservative fallbacks designed into the pipeline, RAG becomes an auditable engineering system rather than an opaque model call.