Refonte Learning: RAG Pipelines in Production: Retrieval, Reranking, and Evaluation

RAG Pipelines in Production: Retrieval, Reranking, and Evaluation

Tue, Jul 7, 2026

RAG Pipelines in Production: Retrieval, Reranking, and Evaluation

Retrieval-augmented generation looks deceptively simple in a demo notebook: split some documents, embed them, drop them into a vector store, and ask a language model to answer questions over the top-k results. In production, the same pipeline fails in ways that feel unfair. Retrieval misses the one paragraph that mattered, the model confidently cites the wrong policy version, latency balloons when your corpus grows past a million chunks, and evaluation is a Slack thread of screenshots. This guide walks through the full stack you need to build RAG systems that hold up under real traffic, real data drift, and real accuracy requirements. You will learn how to chunk documents so they preserve meaning, how to choose and tune embedding models, how to combine dense and sparse retrieval, how rerankers rescue precision, how to assemble prompts that resist hallucination, and how to evaluate the whole thing offline before shipping.

Why RAG, and when not to use it

RAG exists because language models have three limitations: they cannot see private data, their knowledge is frozen at training time, and their context windows, while growing, still cost money per token. Retrieval-augmented generation solves all three by fetching relevant text at query time and stuffing it into the prompt. The model then reasons over that grounded context rather than relying on parametric memory. This is why RAG dominates enterprise use cases like internal knowledge bases, customer support, legal document review, and technical documentation assistants.

However, RAG is not always the right tool. If your task requires deep reasoning over an entire long document, sometimes just passing the whole document into a long-context model outperforms chunked retrieval. If your task requires stylistic mimicry, tone matching, or learning a domain-specific format, fine-tuning is often the better lever. And if your task requires structured data lookups (customer account balance, order status, inventory count), you should be calling a database or API, not embedding rows into a vector store. A useful heuristic: RAG shines when the answer is a paraphrase or synthesis of a small number of retrievable passages, and it struggles when the answer requires aggregating hundreds of documents or executing precise numerical logic.

You will also encounter hybrid architectures where RAG is one tool among several. An agent might call a SQL tool, a RAG retriever, and a calculator in sequence. If you are building those systems, spend time with the AI agents and orchestration patterns guide alongside this one, because retrieval quality directly determines whether your agent can complete multi-step tasks. And if you are still mapping the broader landscape, the Refonte Learning AI hub links every silo including agents, fine-tuning, and prompt engineering.

Finally, be honest about your accuracy floor. A legal research assistant that returns the wrong case citation 5% of the time is not shippable. A marketing brainstorm helper that misses the perfect passage 5% of the time is fine. Set your target metric before you start building, because it determines how much you invest in reranking, evaluation, and human review.

The anatomy of a production RAG pipeline

A production RAG system has more moving parts than the classic "embed and search" diagram suggests. At ingestion time, you parse documents from their source formats (PDF, HTML, Markdown, Confluence exports, Slack archives), clean and normalize the text, split it into chunks, generate embeddings, and write vectors plus metadata to a store. This ingestion pipeline runs continuously as documents change, so you need change detection, versioning, and reindexing logic.

At query time, the flow is: normalize the user query, optionally expand or rewrite it, run retrieval (often hybrid dense plus sparse), rerank the candidate set, assemble a prompt with the top results, call the language model, and post-process the response (add citations, redact sensitive data, run safety filters). Around this core sit observability, evaluation, caching, and access control layers. Each stage is a potential failure point and a potential optimization target.

The order matters. A common beginner mistake is to obsess over the language model choice while ignoring retrieval quality. If retrieval misses the relevant passage, no amount of prompt engineering or model upgrade will save you. The model can only reason over what you give it. Spend 70% of your engineering effort on retrieval and evaluation, 20% on prompt assembly, and 10% on model selection. This ratio surprises people who come from a pure prompting background, which is why the prompt engineering deep dive is best read in tandem with this page rather than before it.

A useful mental model: think of RAG as a search system with a language model at the end. Everything you know about information retrieval, indexing, and ranking applies. The language model is just a very expressive re-ranker and summarizer bolted onto the front of a search stack. Once you internalize that framing, the engineering decisions become clearer.

Document ingestion and parsing

Ingestion is where most RAG systems accumulate technical debt. The temptation is to throw everything at a PDF-to-text library and call it done. In practice, the quality of your parsed text is the ceiling on the quality of your final answers. Garbage in, confidently hallucinated garbage out.

Start by cataloging your sources. For each source type, decide whether you need structural fidelity (headings, tables, lists) or just prose. HTML from a documentation site parses cleanly with a library like BeautifulSoup, but you must strip navigation, footers, and repeated boilerplate before embedding, or every chunk will match every query. PDFs are the hardest case: text extraction quality varies wildly depending on whether the PDF was born-digital or scanned. Tools like PyMuPDF, pdfplumber, and Unstructured handle most cases, but for scanned PDFs you need OCR, and for complex layouts (multi-column academic papers, financial statements with tables) you may need layout-aware models like LayoutLM or commercial services.

Tables deserve special attention. A table rendered as flat text loses its structure, and the language model cannot recover the relationships between columns and rows. Options include: converting tables to Markdown or HTML and embedding them as-is, generating a natural language summary of each table and embedding that alongside the raw form, or extracting tables into a separate structured store and querying them via a tool call. For most knowledge base use cases, Markdown tables in the chunk work well enough. For financial or scientific documents, invest in a structured extraction path.

Preserve metadata aggressively. For every chunk, store the source document ID, title, section heading, page number, publication date, author, and any access control tags. This metadata powers filtered retrieval (only search documents this user can see, only return results from the last six months), citation rendering (show the user where the answer came from), and debugging (when a bad answer ships, you need to trace it back to the source chunk).

Finally, build for change. Documents get updated, deleted, and versioned. Your ingestion pipeline needs a stable chunk ID scheme so that reindexing a modified document replaces the old chunks rather than duplicating them. A common pattern is to hash the document ID plus chunk index, and to store a content hash so you can skip re-embedding unchanged chunks. This saves significant embedding cost at scale.

Chunking strategies that actually work

Chunking is the single most under-appreciated lever in RAG. Bad chunking guarantees bad retrieval, and no downstream tuning fixes it. The goal of chunking is to produce passages that are semantically coherent (each chunk is about one thing), appropriately sized (large enough to contain useful context, small enough that the embedding vector represents a focused topic), and independently interpretable (a chunk should make sense without its neighbors).

The naive approach is fixed-size chunking: split every N tokens, optionally with an overlap of M tokens. This works surprisingly well as a baseline, with typical values of N=512 tokens and M=50 tokens. The overlap prevents important sentences from being cut in half at chunk boundaries. Fixed-size chunking is fast, deterministic, and easy to reason about, which is why most production systems still use it.

Recursive character splitting improves on fixed-size by respecting document structure. The splitter tries to break on double newlines first, then single newlines, then sentence boundaries, then word boundaries, only falling back to character-level splits when necessary. This preserves paragraph and sentence integrity, which improves both embedding quality and readability of citations. LangChain and LlamaIndex both ship recursive splitters as defaults.

Semantic chunking uses an embedding model to find natural topic boundaries. You embed each sentence, compute similarity between adjacent sentences, and split when similarity drops below a threshold. This produces chunks that align with topic shifts rather than arbitrary token counts. It is slower and more expensive at ingestion time but can lift retrieval quality on long-form documents like reports, transcripts, and technical manuals.

Hierarchical chunking, sometimes called parent-child or small-to-big retrieval, stores chunks at multiple granularities. You embed small chunks (sentences or short paragraphs) for precise retrieval, but when you find a match, you return the larger parent chunk (section or full document excerpt) to the language model. This gets the best of both worlds: precise matching on focused text, plus enough context for the model to reason well. This pattern is worth the extra plumbing for domains where context matters, like legal, medical, and scientific documents.

Here is a quick comparison of when to use each strategy:

Strategy Best for Cost Complexity
Fixed-size Uniform corpora, quick baselines Low Low
Recursive character Mixed markdown/prose docs Low Low
Semantic Long-form articles, transcripts Medium Medium
Hierarchical (small-to-big) Legal, medical, scientific Medium High
Document-aware (headings) Technical docs with clear structure Low Medium

One rule that transfers across all strategies: never embed a chunk without prepending some context. Include the document title and the current section heading at the top of each chunk before you embed it. This dramatically improves retrieval on ambiguous queries, because the embedding now encodes not just the local text but also its position in the broader document.

Embedding models: selection and tuning

Your embedding model determines the semantic space in which retrieval happens. Two chunks with high cosine similarity in embedding space should be about similar things. If the model does not encode your domain well, retrieval fails no matter how clever your downstream stack is.

For general English text, the current strong open-source options include the BGE family (BAAI), the E5 family (Microsoft), Nomic Embed, and the GTE series. Commercial options include OpenAI's text-embedding-3 models, Cohere embed v3, and Voyage AI's models. The commercial models are typically stronger on out-of-the-box quality but lock you into an API and per-token cost. The open-source models can run on your own GPUs, which matters for cost, latency, and data residency.

Model size matters, but not as much as you might think. The jump from a 100M-parameter embedding model to a 1B-parameter model usually gives you 5-15% retrieval improvement on standard benchmarks, at 5-10x the inference cost. For most production systems, a mid-size model (300-500M parameters) hits the sweet spot. Where large embedding models pay off is on domains with heavy jargon or nuanced semantics: legal, medical, financial, scientific.

Dimensionality is a separate axis. Higher-dimensional embeddings (1024, 1536, 3072) capture more nuance but consume more storage and slow down search. Matryoshka embeddings, which are trained so that truncated prefixes remain useful, let you store full-dimension vectors and query with truncated ones for a cost-quality trade-off you tune per workload.

Multilingual and domain-specific embeddings deserve mention. If your users query in multiple languages, use a multilingual model (BGE-M3, multilingual-E5, Cohere embed multilingual) rather than translating everything to English. If you have a specialized domain with enough labeled data (query-document relevance pairs), fine-tuning your own embedding model on that data typically beats any off-the-shelf option. Fine-tuning embeddings is cheaper than fine-tuning language models and often has higher ROI. The mechanics of building labeled data and running contrastive training are covered alongside other post-training techniques in the LLM fine-tuning silo.

One evaluation habit that pays off: before committing to an embedding model, run a small labeled retrieval test on your actual data. Take 100-200 queries with known correct chunks, embed both, and measure recall@k for each candidate model. The winner on your data is often not the winner on public benchmarks.

Vector databases and index design

Once you have embeddings, you need to search them at low latency across potentially billions of vectors. This is the job of a vector database or vector index. The field has consolidated around a few strong options, each with trade-offs.

Pinecone is a fully managed vector database with strong ergonomics and predictable performance. You pay a premium for the managed experience, and you accept that your vectors live in someone else's cloud. Good default for teams that want to move fast and do not want to run infrastructure.

Weaviate is open source with a managed cloud option. Strong support for hybrid search (dense plus keyword) natively, and a rich schema model that makes metadata filtering ergonomic. Good default for teams that want hybrid search out of the box.

Qdrant is open source, written in Rust, and known for fast filtering on metadata. Excellent choice when your workload requires heavy metadata filters (access control, date ranges, source restrictions). Self-hosted or managed cloud.

Milvus is an open-source vector database designed for scale. Battle-tested at billion-vector scale, more operational complexity than Qdrant or Weaviate. Choose when you know you need scale from day one.

pgvector is a Postgres extension that adds vector search to a regular Postgres database. Best choice when your data already lives in Postgres and your scale fits in a single node (up to roughly 10M vectors comfortably, more with tuning). The operational simplicity of "just Postgres" is hard to beat.

Elasticsearch and OpenSearch now support vector fields alongside their traditional inverted indexes. Excellent choice when you already run Elastic for logs or search and want to add semantic search without a new system. Native hybrid search.

Under the hood, all of these use approximate nearest neighbor (ANN) indexes, most commonly HNSW (Hierarchical Navigable Small World). HNSW has three knobs you tune: M (the number of connections per node, higher gives better recall but more memory), ef_construction (build-time search width, higher gives better graph quality but slower ingestion), and ef_search (query-time search width, higher gives better recall but slower queries). Sensible defaults are M=16, ef_construction=200, ef_search=100, but you should tune ef_search per workload based on your latency budget and recall target.

Sharding and replication become relevant at scale. A rule of thumb: a single HNSW index fits comfortably in memory up to about 10-50 million vectors depending on dimensionality. Beyond that, shard by document ID or tenant, and use a router to fan queries across shards. Replicate for read throughput and availability. If you plan to run vector search alongside a broader ML platform, the patterns in running AI workloads on Kubernetes with MLOps pipelines translate directly to vector database deployment.

Hybrid search: combining dense and sparse retrieval

Dense retrieval (vector embeddings) is strong at semantic matching but weak at exact keyword matching. If your user asks about "error code E-4471", a dense retriever might return chunks about generic error handling rather than the specific chunk that mentions E-4471 verbatim. Sparse retrieval (BM25 or SPLADE) excels at exact matches but misses paraphrases and synonyms. Hybrid search combines both.

The simplest hybrid pattern is reciprocal rank fusion (RRF). Run both retrievers independently, get two ranked lists, and combine them by summing 1/(k + rank) across lists, where k is a small constant (typically 60). RRF is parameter-light, robust, and does not require score calibration between the two retrievers. It is the default hybrid strategy in most production systems.

A more powerful pattern uses a learned combination. You train a small model (often a gradient-boosted tree or a linear model) that takes the dense score, sparse score, and other features (freshness, source authority, click-through rate) and outputs a final ranking score. This requires labeled training data but can significantly outperform RRF when you have it.

SPLADE deserves special mention. It is a learned sparse retriever that produces sparse vectors compatible with inverted index infrastructure, but the sparsity pattern is learned by a neural model. This gives you keyword-style precision with some semantic understanding. SPLADE indexes are larger than BM25 indexes but smaller than dense indexes, and query latency is comparable to BM25. Consider SPLADE when you want stronger retrieval than BM25 but do not want the full memory footprint of dense vectors.

One practical warning: hybrid search doubles your infrastructure surface area. You now have two indexes to keep in sync, two sets of tuning parameters, and two failure modes. Only add hybrid search when your evaluation harness shows dense-only retrieval is missing important cases. Do not add it because a blog post said you should.

Query understanding and rewriting

Users write bad queries. They use pronouns without antecedents, they mix questions and commands, they use jargon inconsistently, and they ask compound questions. Query understanding sits between the raw user input and the retriever, cleaning up the input to maximize retrieval success.

Query rewriting uses a language model to transform the user's query into one or more search queries. For a follow-up question in a conversation ("what about the pricing?"), the rewriter uses conversation history to produce a self-contained query ("what is the pricing for the enterprise plan?"). For a compound question ("compare the refund policies for the US and EU"), the rewriter can split into multiple sub-queries and retrieve for each. For a vague question, the rewriter can generate several paraphrases and union the results.

HyDE (Hypothetical Document Embeddings) is a specific rewriting technique: instead of embedding the query, you ask a language model to generate a hypothetical answer to the query, then embed that hypothetical answer and use it for retrieval. This works because the hypothetical answer is stylistically closer to the target documents than the raw question is. HyDE is more expensive (one extra LLM call per query) but can meaningfully improve retrieval on question-answering workloads.

Query expansion adds related terms to the query before retrieval. Classical IR uses techniques like pseudo-relevance feedback, where the top-k results from an initial retrieval are used to identify expansion terms. In the LLM era, you can also ask a language model to list synonyms or related concepts. Expansion helps with sparse retrieval especially, where lexical mismatch is the main failure mode.

Intent classification is the last piece of query understanding. Not every query should hit your RAG pipeline. Small talk, out-of-scope questions, and structured queries (like "show me my recent orders") should route to different handlers. A lightweight classifier at the top of your pipeline improves both quality and cost by sending each query to the right backend.

Reranking: from candidates to precision

Retrieval is optimized for recall: you want the correct passage to appear somewhere in the top 20 or top 50 candidates. Reranking is optimized for precision: you want the correct passage in the top 3 or top 5, where it will actually make it into the language model's prompt. These are different problems that benefit from different models.

Cross-encoder rerankers are the workhorses of production reranking. Unlike a bi-encoder (which embeds query and document separately, enabling fast ANN search), a cross-encoder takes the query and a candidate document together as one input and outputs a relevance score. This is much slower per pair (you cannot precompute anything), but far more accurate. You use a cross-encoder only on the top-k candidates from your retriever, typically k=20 to k=100.

Strong cross-encoder options include Cohere Rerank (commercial API, best-in-class quality on English), BGE reranker (open source, self-hosted), Jina reranker, and Voyage rerankers. For most workloads, a cross-encoder rerank on the top 50 candidates, keeping the top 5 to pass to the language model, is the right recipe.

LLM-based reranking is the next step up. You ask a language model directly: given this query and these 20 candidate passages, rank them by relevance. This is expensive but can be extremely accurate, and it lets you inject task-specific instructions into the ranking ("prefer passages about pricing over passages about features"). LLM reranking is worth it for high-value queries and small candidate sets. It is not worth it for high-volume, low-margin workloads.

Metadata reranking is a cheap complement to model-based reranking. After model reranking, you can apply hard filters (correct language, correct product line, recent enough) and soft boosts (source authority, popularity, freshness). A weighted combination of model score and metadata features often outperforms pure model reranking, especially in domains where recency or authority matter (news, medical, legal).

Latency budgets constrain reranker choice. A cross-encoder rerank of 50 candidates adds roughly 100-300ms depending on the model. LLM reranking adds 500ms to several seconds. If your total latency budget for the query is 2 seconds, you have room for cross-encoder rerank but not full LLM rerank. Measure, then choose.

Prompt assembly and context management

Once you have your top passages, you assemble the prompt. This step is where many systems leak accuracy through carelessness. A well-assembled prompt makes the difference between a model that grounds every claim in the retrieved context and a model that mixes retrieval with parametric hallucination.

Start with a system message that sets the rules: "You are a customer support assistant. Answer only using the provided context. If the answer is not in the context, say so explicitly. Cite the source ID for every claim." Be specific about the behavior you want, including edge cases (what to do when the context is empty, what to do when sources conflict, how to format citations).

Structure the retrieved context clearly. Do not concatenate passages into one blob. Format each passage with a source ID, title, and any relevant metadata, so the model can cite it and so you can render clickable citations to the user. A typical format:

[Source 1] Title: Refund Policy, Version 3.2
The customer is entitled to a full refund within 30 days...

[Source 2] Title: Enterprise Terms, Section 4
Enterprise customers may negotiate custom refund windows...

Order matters. Most language models attend more strongly to the beginning and end of the prompt than the middle. Put the most relevant passage first (as ranked by your reranker) or last, and put the less certain passages in the middle. Some teams put the query at both the top (in the system message) and the bottom (right before the model responds) to keep it fresh in the model's attention.

Manage the context window carefully. If your reranked top 5 passages exceed the effective context window, you must either truncate individual passages, drop lower-ranked passages, or summarize. Truncation loses information at the ends; dropping loses whole passages; summarization loses fidelity. The right choice depends on your domain. For question-answering over policy documents, dropping lower-ranked passages is usually best because the top-ranked passage typically contains the answer. For synthesis tasks, summarization of longer passages preserves more signal.

Anti-hallucination techniques belong in the prompt. Instruct the model to quote directly from sources when possible, to distinguish between "the sources say X" and "based on the sources, X likely implies Y", and to refuse to answer when confidence is low. For deeper coverage of these patterns, work through the prompt engineering silo which covers instruction design, few-shot examples, and structured output formats in depth.

Evaluation: how you know it works

Most RAG teams do not evaluate their systems. They rely on eyeballing outputs in a Slack thread and hoping for the best. This is why most RAG systems degrade silently. Evaluation is the highest-leverage investment in your pipeline.

Start with a labeled evaluation set. Collect 100-500 real user queries (or synthetic queries that mimic real ones) and, for each, label the correct answer and the correct source chunks. This is tedious but non-negotiable. If you cannot spend the time to label 200 queries, you do not care enough about accuracy to run this in production.

Retrieval metrics measure whether the right chunks made it into the candidate set. The standard metrics are:

  • Recall@k: fraction of queries where at least one relevant chunk is in the top k
  • Mean Reciprocal Rank (MRR): average of 1/rank of first relevant chunk
  • Normalized Discounted Cumulative Gain (nDCG): weighted metric that rewards putting more relevant chunks higher

Track recall@20 for your retriever (before reranking) and MRR or nDCG at 5 for your reranker output. If recall@20 is low, no reranker will save you. If recall@20 is high but MRR@5 is low, your reranker needs work.

End-to-end metrics measure the final answer. Faithfulness (does the answer only make claims supported by the retrieved context), answer relevance (does the answer address the question), and correctness (is the answer factually right) are the standard trio. Frameworks like RAGAS, TruLens, and DeepEval automate these using LLM-as-judge scoring. LLM judges are noisy but consistent enough for tracking regressions between versions, if you fix the judge model and prompt.

Human evaluation is still the ground truth. Once per release cycle, sample 50 queries from your evaluation set (and 50 from live traffic if possible), have a domain expert rate the answers on a 1-5 scale, and compare against your automated scores. If automated and human scores diverge, your automation is misleading you.

Build evaluation into your CI. Every pull request that changes the pipeline (new chunker, new reranker, new prompt) should run the eval set and post the metrics. If retrieval recall drops by more than 2%, the PR blocks. If end-to-end faithfulness drops by more than 3%, the PR blocks. This discipline prevents the drip of small regressions that accumulate over months.

Frameworks compared: LangChain, LlamaIndex, DIY

The RAG framework landscape can feel confusing. Here is a direct comparison.

LangChain is the most popular framework and has the largest community. It provides abstractions for chains, agents, retrievers, and integrations with dozens of vector stores and LLM providers. Strengths: fast prototyping, huge integration surface, extensive documentation. Weaknesses: abstractions can leak, versioning has been unstable, production-grade features (evaluation, observability) live in a separate product (LangSmith). Best for: teams that want to move fast on prototypes and are willing to drop into lower-level code when abstractions get in the way.

LlamaIndex started as a RAG-specific framework and remains the most RAG-focused option. It has strong document loaders, hierarchical retrievers, and evaluation utilities built in. Strengths: idiomatic RAG patterns, opinionated in helpful ways, strong for advanced retrieval strategies. Weaknesses: smaller integration surface than LangChain, some overlap with LangChain that can confuse teams evaluating both. Best for: teams building RAG-heavy applications who want opinionated defaults.

Haystack by deepset is a mature framework with strong production ergonomics, particularly for pipeline definition and evaluation. It predates the current LLM boom and has heritage in classical IR. Strengths: pipeline abstraction is clean, evaluation is first class, good for teams that want a more industrial feel. Weaknesses: smaller community than LangChain, fewer bleeding-edge integrations. Best for: teams that value stability and production ergonomics over integration count.

DIY means writing your own retrieval, reranking, and prompt assembly code directly against the underlying APIs (OpenAI, your vector store, your reranker). Strengths: full control, no framework churn, easy to debug, minimal dependencies. Weaknesses: you write more code, you handle integrations yourself, you build your own evaluation harness. Best for: teams with strong engineering capacity who plan to run this system for years and want to minimize framework risk. Many production RAG systems at scale end up here after starting with a framework.

Here is a rough decision matrix:

Situation Recommendation
Prototype, less than 2 weeks to demo LangChain or LlamaIndex
Production RAG, 1-3 engineers LlamaIndex or Haystack
Production RAG, 5+ engineers, long-term ownership DIY on top of core libraries
Complex agentic workflows on top of RAG LangChain with careful version pinning
Compliance-heavy environment (finance, health) Haystack or DIY

A common trajectory: prototype with a framework, discover its abstractions do not fit your needs, gradually rewrite critical paths, end up with a hybrid where the framework provides utilities and you own the orchestration.

Cost, latency, and capacity planning

RAG systems have three cost centers: embedding, storage, and inference. Understanding each is critical to planning capacity and pricing your product.

Embedding cost scales with corpus size and update frequency. If you have 10 million chunks averaging 300 tokens each, initial embedding at $0.10 per million tokens (a typical commercial rate) costs $300. If you re-embed monthly due to content updates or model upgrades, add that recurring cost. Self-hosting embeddings on a GPU is cheaper at scale but adds operational load. The break-even is roughly 50-100 million tokens per month depending on your GPU pricing.

Storage cost scales with vector count and dimensionality. A million 1536-dimensional float32 vectors is roughly 6 GB. With HNSW graph overhead, memory usage is typically 1.5-2x the raw vector size. Managed vector databases charge per GB and per query. Self-hosted is cheaper but requires operational investment.

Inference cost dominates for most production RAG systems. Each query triggers: query embedding (cheap), retrieval (cheap), reranking (moderate), and language model generation (expensive). For a typical query with 4000 tokens of context and 500 tokens of output, using a frontier model, expect $0.02 to $0.10 per query. At 100k queries per day, that is $2000 to $10000 per day. Optimize aggressively: use smaller models when quality allows, cache aggressively, and truncate context to what actually matters.

Caching is your biggest cost lever. Cache at three levels: query embedding cache (same query text produces same embedding, cache with TTL), retrieval cache (same query returns same top-k, cache with short TTL to reflect content updates), and full response cache (identical query and context returns identical response, cache with longer TTL). A well-tuned cache stack can cut inference cost by 30-60% on workloads with repeated queries.

Latency budget planning: assume a target of 2 seconds end-to-end from user input to first token. Break that down as: query understanding 100ms, embedding 50ms, retrieval 100ms, reranking 200ms, LLM time-to-first-token 500-1500ms, network overhead 100ms. If any component exceeds its budget, either optimize it or accept a slower user experience. Streaming the LLM response to the user hides some latency because the user sees tokens as they arrive.

For teams sizing a production deployment, or negotiating vendor contracts, the AI consulting and pricing overview walks through cost modeling in more detail. For broader context on where RAG fits among other generative applications, browse the generative AI use cases catalog.

Observability, safety, and access control

A production RAG system in a business context has requirements beyond raw quality: observability so you can debug bad answers, safety so you do not leak harmful content, and access control so users see only what they are permitted to see.

Observability starts with structured logging of every query. Log the raw query, the rewritten query, the retrieved chunk IDs and scores, the reranked chunk IDs and scores, the final prompt (or a hash of it if it contains sensitive data), the model response, and any user feedback (thumbs up, thumbs down, follow-up questions). This log is the foundation for debugging, evaluation set expansion, and drift detection.

Tracing tools like LangSmith, Phoenix (Arize), and Langfuse visualize the full pipeline for each query, making it easy to spot where things went wrong. Adopt one early. Debugging a bad answer without tracing is like debugging a distributed system without logs.

Safety layers include input filtering (block prompt injection attempts, malicious queries), output filtering (block PII, block toxic content, block content that violates your terms), and output validation (verify citations point to real chunks, verify claims are grounded in retrieved context). Grounding validation deserves special attention: automated tools can score whether the response is supported by the retrieved context, flagging low-grounding responses for human review before they reach the user.

Access control is the requirement that most RAG teams underestimate. Every chunk in your vector store should carry access control metadata: which users, roles, or tenants can see it. At query time, you must filter retrieval to only include chunks the current user can see. This has to happen at the vector database level, not by post-filtering after retrieval, because post-filtering breaks recall (you might return zero results even when relevant results exist that the user is not allowed to see).

Multi-tenancy adds complexity. If you serve multiple customer organizations from one system, each tenant's data must be strictly isolated. The safest architecture is one index per tenant, but that does not scale past a few hundred tenants. For larger deployments, use metadata filtering with hard-enforced tenant IDs on every query, plus regular auditing to confirm no cross-tenant leakage. This is the kind of engineering discipline covered in the AI engineering internship and program, where trainees build production-grade systems with these concerns baked in from day one.

Advanced patterns: agentic RAG, GraphRAG, and beyond

Once your baseline RAG pipeline is solid, several advanced patterns unlock harder use cases.

Agentic RAG wraps retrieval in an agent loop. Instead of retrieving once and answering, the agent can retrieve multiple times, refining its query based on what it finds. This handles multi-hop questions where the answer requires combining information from several documents. The agent can also call other tools (calculators, APIs, code interpreters) alongside retrieval. Latency and cost go up significantly, so use agentic RAG only for queries that need it. A router at the top of your pipeline can decide whether a query needs simple RAG or agentic RAG.

GraphRAG, popularized by Microsoft Research, augments vector retrieval with a knowledge graph built from the corpus. During ingestion, an LLM extracts entities and relationships from each chunk and stores them in a graph. At query time, retrieval combines vector similarity with graph traversal, enabling questions like "what changed in the refund policy between version 2 and version 3" or "who are all the authors that have written about topic X". GraphRAG is powerful but ingestion is 5-20x more expensive than vanilla RAG, so reserve it for high-value corpora.

Contextual retrieval (a recent Anthropic pattern) preprocesses each chunk by asking an LLM to describe the chunk's context within the broader document, then prepends that description before embedding. This lifts retrieval quality by 20-40% on many workloads, at the cost of one extra LLM call per chunk during ingestion. For static corpora, the one-time cost is worth it.

Multi-vector retrieval represents each chunk with multiple embeddings, each capturing a different aspect (a summary, a hypothetical question the chunk answers, key entities). Retrieval matches against all vectors and combines the scores. ColBERT is the classic example, using one vector per token. Multi-vector retrieval improves quality at the cost of higher storage and more complex indexing.

Fine-tuned embeddings and rerankers deserve one more mention. Once your evaluation harness is solid, fine-tuning your own embedding model and reranker on your labeled data typically gives you 10-30% quality lift over off-the-shelf models. The training data comes from your evaluation set plus mined hard negatives from your production logs. This is high-ROI work that most teams skip because it feels intimidating.

For a broader view of where these patterns fit in a two-year career trajectory, the AI learning roadmap sequences them alongside foundational skills. And for the industry context of why these patterns matter, the data science trends and career strategies analysis covers the job market signal driving demand for RAG expertise.

A reference implementation walkthrough

To ground everything above, here is a step-by-step walkthrough of a production RAG pipeline for an internal company knowledge base. The corpus is roughly 200,000 documents (policies, wikis, engineering docs, meeting notes) totaling 15 million chunks after splitting.

Step 1: Ingestion. A daily job pulls changes from Confluence, Google Docs, Notion, and an internal wiki via their APIs. Each document is parsed with source-specific parsers, cleaned to remove navigation and boilerplate, and split with a recursive character splitter using a 512-token chunk size and 64-token overlap. Section headings are prepended to each chunk. Metadata captured includes source, path, author, last-modified date, and ACL tags.

Step 2: Embedding. Chunks are embedded with a self-hosted BGE-large-en-v1.5 model running on two A10 GPUs. Content hashes are checked to skip re-embedding unchanged chunks. Embeddings and metadata are written to Qdrant, sharded by document root.

Step 3: Query pipeline. User queries hit an API gateway that enforces authentication and rate limiting. The gateway calls a query understanding service that classifies intent (RAG-appropriate, structured lookup, small talk) and, for RAG queries, rewrites the query using conversation history via a small language model.

Step 4: Retrieval. The rewritten query is embedded with the same BGE model. Qdrant is queried for the top 40 candidates by cosine similarity, filtered by the user's ACL tags. In parallel, an OpenSearch BM25 query returns the top 40 candidates. The two lists are combined via reciprocal rank fusion into a top 50 set.

Step 5: Reranking. The top 50 candidates are reranked with a self-hosted BGE reranker cross-encoder. The top 8 are kept and passed forward. Latency for the rerank step is roughly 180ms on a T4 GPU.

Step 6: Prompt assembly. A system message instructs the model on behavior and citation format. Each of the top 8 passages is formatted with a source ID, title, and path. The user's original query and rewritten query are both included.

Step 7: Generation. The prompt is sent to a frontier language model with streaming enabled. The response streams back to the user with citations rendered as clickable links to the source documents.

Step 8: Logging and evaluation. Every query is logged with all intermediate signals. A nightly job samples 500 queries and scores them with an LLM judge on faithfulness, relevance, and citation accuracy. Weekly, a human reviewer samples 50 queries and rates them, comparing against the LLM judge scores. Quarterly, the evaluation set is refreshed with new queries from production logs.

Step 9: Continuous improvement. Bad answers (flagged by users or by low LLM judge scores) are triaged. Retrieval failures trigger embedding or reranker tuning. Generation failures trigger prompt updates. Ingestion failures trigger parser improvements. Every change ships through the CI evaluation harness.

This is the shape of a serious production RAG system. It is not glamorous, and most of the effort is in ingestion, evaluation, and observability rather than in the LLM itself. That is the honest truth of production ML: the model is 10% of the work.

Where to go next

You now have a full map of what a production RAG pipeline requires. The gap between reading this and shipping is the gap of practice. Build a small pipeline end-to-end on a corpus you care about, instrument it with evaluation from day one, and iterate. If you want structured practice with mentor feedback on production ML systems including RAG, retrieval, and evaluation, the AI engineering internship program at Refonte Learning builds these skills through paired projects and code review from working practitioners. Complement it with the prompt engineering deep dive for the generation side of the pipeline and the LLM fine-tuning silo for the cases where retrieval alone is not enough.

Frequently asked questions

How large does my corpus need to be before RAG is worth the complexity? There is no hard threshold, but a useful heuristic is: if your total corpus fits in the model's context window (roughly 100 pages for a 200k-token model), you probably do not need RAG. Just pass everything in. RAG becomes valuable when you have more content than fits, when content changes frequently, when you need citations to specific sources, or when you have access control requirements. Below 500 documents, prompt stuffing plus a good system message often outperforms a poorly tuned RAG pipeline.

Should I fine-tune the language model, use RAG, or both? Fine-tuning teaches the model behaviors, formats, and style. RAG gives the model access to specific facts and current information. For most enterprise use cases, RAG is the higher-ROI first step because it addresses the "how do I ground answers in my data" problem directly. Fine-tuning becomes valuable when you need consistent tone, structured outputs, or specialized reasoning patterns that prompting alone cannot elicit reliably. The two combine well: fine-tune for behavior, RAG for facts.

How do I handle documents that update frequently? Design your ingestion pipeline for incremental updates from day one. Use content hashing to skip unchanged chunks. Use stable chunk IDs based on document ID plus position so updates replace rather than duplicate. For high-churn corpora, consider a shorter cache TTL on retrieval results. If updates are near-real-time (like a news feed or ticket system), you need a streaming ingestion pipeline that writes to the vector store within minutes of a document change.

What is the biggest mistake teams make when building RAG? Under-investing in evaluation. Teams spend weeks tuning chunk sizes and testing new embedding models without a labeled evaluation set to measure whether changes actually help. The result is superstition-driven development. Build the eval set first, even if it is only 50 queries. Then every change becomes measurable and every argument becomes resolvable.

How do I know when to add a reranker? Measure your retrieval recall@20 and MRR@5. If recall@20 is high (above 90%) but MRR@5 is low (below 0.7), a reranker will help significantly. If recall@20 is low, fix retrieval first, because a reranker cannot promote a document that never made it into the candidate set. Rerankers are the second lever, not the first.

Can I use RAG with proprietary or air-gapped data? Yes, and this is a common driver for choosing self-hosted embeddings and language models. Open-source embedding models (BGE, E5, Nomic) run on a single GPU. Open-source language models (Llama, Mistral, Qwen) run on modest GPU clusters. Self-hosted vector databases (Qdrant, Weaviate, Milvus, pgvector) require no external calls. The full stack can run in an air-gapped environment. Expect quality to be somewhat lower than frontier commercial models, and expect more operational work.

How do I evaluate RAG answers that require synthesis across multiple sources? Standard question-answering benchmarks focus on single-source answers. For synthesis, you need custom evaluation: label each query with the set of required source chunks (not just one), and measure whether the answer covers all required sources without introducing unsupported claims. Faithfulness metrics from RAGAS and TruLens handle multi-source grounding reasonably well when configured for it. Human evaluation remains the ground truth for complex synthesis tasks.

Should I use LangChain or write my own? For prototypes and learning, use LangChain or LlamaIndex to move fast. For production systems you will own for years, write your own on top of the underlying libraries (a vector store SDK, an embedding client, a language model client). Framework abstractions age poorly as the field moves, and debugging framework internals costs more time than the framework saved you. Most mature production RAG systems end up as DIY code with framework-provided utilities.