problem_kicker

Production RAG is a retrieval system with evidence, not a prompt around a vector database.

The hard part of production RAG is keeping retrieval correct, authorized, reproducible and measurable while sources, embeddings, prompts and models change. A useful architecture treats indexing, retrieval, ranking and generation as separately versioned systems.

Production RAGACL-aware retrievalHybrid retrievalRerankingMeasurable quality

DEMAND LANGUAGE / REAL-WORLD PROBLEM

Does this sound familiar?

“The RAG demo works. How do we make it reliable in production?”
“Vector DB + prompt was easy — how do we know the answers are actually getting better?”

WHAT CAUSES THIS?

Why it breaks in production

Vector-only retrieval loses exact terms and identifiers.

  • ACLs and metadata are bolted on after the retrieval model was designed.
  • Embedding or chunking migrations silently change answer behavior.
  • Prompt/model updates ship without a replayable evaluation corpus.
  • Failed indexing jobs leave partially updated knowledge states.

architecture_for PRODUCTION RAG ARCHITECTURE

engineering

We define source-of-truth semantics, authorization boundaries and evaluation queries before selecting retrieval components. Index versions, embedding versions, prompt versions and model versions are explicit so a result can be replayed and compared.

security

authority

ACL-aware retrieval should constrain candidates before protected evidence reaches generation. Service identities, index access and evaluation datasets should follow least privilege and preserve source authorization semantics.

performance

critical

Latency budgets are allocated per stage. Candidate count, fusion strategy, reranker depth, model context and cache policy are tuned against measured retrieval quality rather than optimized in isolation.

technologies

vendor

RAG · hybrid retrieval · reranking · CDC · embeddings · LLM observability

failure_kicker

anti_title

  • Rebuild the whole index for every change.
  • Mix embeddings from incompatible versions without migration state.
  • Treat a fluent answer as evidence of retrieval quality.
  • Log prompts but not retrieved evidence and version identifiers.

measure_kicker

verify_title

verify_intro

  1. Golden and adversarial query sets with expected evidence.
  2. Recall@k, MRR/nDCG and citation correctness.
  3. p50/p95/p99 latency and cost per successful answer.
  4. Replay tests across embedding, prompt and model migrations.
  5. Recovery test from interrupted or partially failed indexing.

CTO / CIO FAQ

faq_title

How do we migrate embeddings without downtime?

Use versioned indexes or versioned embedding fields, backfill incrementally, evaluate both paths and switch through an explicit compatibility gate.

Should every answer have citations?

For enterprise knowledge tasks, evidence links materially improve auditability. Citation correctness must still be evaluated; merely emitting a link is not enough.

How do we know whether RAG improved?

Use a stable evaluation set and measure retrieval, evidence correctness, answer quality, latency and cost separately.