EngineeringInteractive guide16 min read · 7 experiments

RAG in Production: Architecture, Chunking, Retrieval, Evaluation & Monitoring

A practical guide to building a RAG system, from preparing your documents to finding answers and checking their quality.

Author: Edward Ha
Layered paper documents flow into an archive and emerge as selected evidence.
00Overview

What is RAG?

Retrieval-augmented generation or RAG is a technique that supplies a language model with passages retrieved from a document collection. The model uses these passages alongside the question to generate an answer. Asking it to cite the passages gives readers a way to check its claims, but a citation alone does not establish that a claim is supported. 1

You have probably used it already. answers questions about the sources you upload and links each claim back to the passage it came from. does the same when it searches the web or reads a file you attach, and built its whole product around answers with numbered sources. Most customer-support bots and internal knowledge assistants work the same way.

The six stages below is all the RAG process that cover document preparation, retrieval, answer generation, and verification:

  1. 01Ingest

    Collect and parse the documents, keeping each one’s source, version, and permissions.

  2. 02Chunk

    Split them into passages small enough to search and large enough to make sense on their own.

  3. 03Embed & index

    Turn every passage into a vector and store it with its metadata.

  4. 04Retrieve

    Encode the question and find the best passages, often with keyword and vector search, then a reranker.

  5. 05Generate

    Give the model the question and passages. Ask it to cite supporting sources and say when the evidence is insufficient.

  6. 06Verify

    Check that every claim is backed by a cited passage before the answer reaches the user.

Steps 1–3 form the ingestion path and it runs whenever your sources change. Steps 4–6 form the query path and it runs on every question. The figure below shows how the two connect.

Ingestion and query processingSelect a step to explore
01 / Prepare knowledgeWhen sources change
02 / Answer a questionAt request time
Select evidence before generating.

Search eligible passages, rank candidates, and fit the useful evidence into the context budget. Exact terms may benefit from keyword or hybrid search.

Figure 1. Ingestion prepares the knowledge; retrieval selects the evidence for each answer. Evaluation informs both paths.

So what is the hard part of building a RAG system?

There are a lot of challenges when building a RAG system. Let’s look at a few problems you can run into while building one, starting with ingestion.

Some issues we’ve seen are easy to overlook: engineers forget to check whether a table was extracted correctly, formulas get lost during parsing, images containing useful information never get described, or an expired policy gets embedded and added to the index.

When the answers come back wrong, the first thing we’ve seen engineers try is switching the embedding model or adding a reranker. But neither can recover information that never made it into the index. And if an expired policy is still searchable, it can still turn up in the results.

Before changing the model, check what actually made it through ingestion.

The case below shows some example issues on the refund question from the labs: same corpus, same retriever, same question. For each one you can see what it does to the evidence, a cheap way to catch it, and the fix.

CASE FILE — THE REFUND QUESTION

Four ways the evidence breaks before retrieval

Chunking layerSame question every time: “Can I get my money back, and how long will it take?”
AS INGESTED

2of 2 boundaries land mid-sentenceat 40-word chunks, no overlap

… if they have not beenaccessed. Approved refunds are sent …
… standard purchases from our onlinestore.

A chunk that ends “if they have not been” reaches the model without the rest of its rule. Lab 05’s mock answers use complete sentences from the selected evidence and flag missing information.

AFTER THE FIX

0of 2 boundaries land mid-sentencesentence-aware, same size limit

… a receipt when contacting support.Digital downloads are refundable only …
… window require a support review.Shipping fees are not refunded …

Chunks vary in length because their boundaries follow complete sentences.

How to catch it

Print the last six words of every chunk. If chunks end on “not been” or “unless the”, the splitter ignores sentence boundaries.

The fix

Split on sentences or headings first, then pack whole sentences up to the size limit.

Cases 1–3 use excerpts or retrieval results from the demo corpus. The table example is illustrative. The displayed fixes do not change the later labs.

Besides ingestion, engineers face other challenges, for example deciding where to split a document, choosing an embedding model, keeping the index up to date, and checking whether retrieval finds enough information to answer a question. In the seven sections below, we’ll work through ingestion, chunking, embeddings, indexing, retrieval, evaluation, and the checks before shipping. Along the way, you can try the labs, change the settings, and see how those choices affect the results.

  1. 01IngestionLab 01 · Parse
  2. 02ChunkingLab 02 · Split
  3. 03EmbeddingsLab 03 · Embed
  4. 04IndexingLab 04 · Store
  5. 05RetrievalLab 05 · Try a simulated answer
  6. 06EvaluationLab 06 · Measure
  7. 07Before you shipProduction checklist
01Ingestion

What is ingestion?

Ingestion brings your documents into the search system and prepares their contents for retrieval. It helps the system find relevant information across PDFs, scanned pages, and websites while keeping enough context to interpret it correctly.

This starts with parsing: extracting content and preserving meaningful structure, such as headings, tables, and the explanations around images or formulas.

How to handling different files type

PDFs and Word files: use or to preserve reading order, headings, and tables.

Scanned pages: use an OCR tool such as to extract text from images. Check important numbers and labels against the original.

Images and formulas: Docling offers image-description and formula-extraction features. Configure these when needed, keeping captions and variable definitions with the content.

Web pages: use to extract the article and remove surrounding menus, cookie notices, and footers.

For example, “5–10 business days” in a refund table needs its payment method and column heading to make sense. The example below shows how keeping those details with each row preserves its meaning.

In the PDF
Refund arrival times
Payment methodTypical arrival
Card5–10 business days
PayPal3–5 business days
Bank transfer7–14 business days
Copied as plain text
Card 5–10 business days
No title, no column names, no source. Card what? Days for what?
Each row as its own chunk
Refund arrival timescustomer-handbook.pdf · page 4
Payment method: CardTypical arrival: 5–10 business days
Refund arrival timescustomer-handbook.pdf · page 4
Payment method: PayPalTypical arrival: 3–5 business days
Refund arrival timescustomer-handbook.pdf · page 4
Payment method: Bank transferTypical arrival: 7–14 business days
Made-up example from our pretend store. Each row keeps the table title, its source and the column names, so it still makes sense when it is the only thing retrieved.

Keeping source details

Save where each document came from, its section, version or update date, and access permissions. These details are called metadata. can carry them from a document into its smaller text pieces.

Metadata helps the system cite sources and filter by version or access level. Access permissions must also be enforced during search.

Try the parsing lab

Switch between Basic text copy and Good parser + labels below. Notice how the table stays organized, unrelated web content is removed, and source details remain attached.

LAB 01 — PARSE

What the parser gives the next step

COPIED TEXTcustomer-handbook.pdf · page 4

Example Store Customer Handbook Page 4 of 9 Refund arrival times After a refund is approved, the time it takes to arrive depends on the payment method. Payment method Typical arrival Notes Card 5–10 business days Bank holidays add delay PayPal 3–5 business days Bank transfer 7–14 business days Needs account details Internal use only · Rev. 2026-09 · 4

WHAT CHUNKING RECEIVES

27%of the words are page clutter16 of 60 words · no labels attached

LABELS (METADATA)

None. Without labels, the system cannot tell where a piece of text came from, how old it is, or who may see it.

A basic text copy keeps every word but loses the layout. The page header and footer end up in the middle of the content, and the table turns into one long line where “PayPal” sits right next to the card’s notes. Ask “how long does a PayPal refund take?” and the answer may come with “Bank holidays add delay”, and no way to tell which row that note belongs to.

Made-up pages from the same pretend store as the other labs. Crossed-out text is clutter that a good parser removes. Word counts are based on what is shown here.

02Chunking

What is chunking?

Chunking splits a document into smaller pieces that the system can search individually. This helps it find the passage needed to answer a question without sending the whole document to the model.

Each chunk needs enough context to make sense on its own. For example, “Shipping fees are not refunded unless the item arrived damaged” should stay together. Separating the exception could lead to an incomplete answer.

How to split your documents

Start with the structure preserved during ingestion: sections, paragraphs, lists, and tables. Keep headings with their content and table labels with their rows.

For general text, try LangChain’s recursive text splitter, which uses paragraph and word boundaries before splitting into smaller pieces. LlamaIndex’s SentenceSplitter is another option, with settings for chunk size and overlap.

Different documents may need different approaches:

Chunking approaches and when to consider them
ApproachHow it worksWhen to consider it
Fixed-sizeSplits at a set length, which may cut across sentences.A simple starting point for uniform text.
RecursiveTries larger text boundaries first, then smaller ones until each piece fits.Help articles, handbooks, and documentation.
SemanticSplits where the topic changes, adding processing work.Transcripts or text with few headings.
PropositionUses a model to rewrite passages as individual facts. Check that meaning is preserved.Questions targeting specific rules or facts.
Document-levelKeeps each short document as one chunk.FAQ entries and product cards.
HierarchicalSearches small chunks, then returns their larger parent section.Manuals or policies where answers need surrounding context.

How to choose size and overlap

Small chunks can lose context; large chunks can include unrelated information. Test with questions your readers would actually ask and check whether the retrieved pieces contain a complete answer.

Overlap repeats some text between neighbouring chunks to preserve context near a split. It also creates duplicate content, so increase it only where it helps. Microsoft’s chunking guide explains these trade-offs. 2

Try the chunking lab

Adjust Chunk size and Overlap below. Watch where the refund policy splits, whether conditions stay together, and how much text is duplicated.

This lab uses fixed-size chunks measured in words for readability. Production tools often measure length in tokens, the smaller text units models process.

LAB 02 — SPLIT

Find the right boundary

3chunks
97stored words
+20%duplicated
refund-policy.md81 words

Customers can request a refund within 30 days of purchase. The item must be unused and in its original packaging. Include the order number and a receipt when contacting support. Digital downloads are refundable only if they have not been accessed. Approved refunds are sent to the original payment method. Refund requests outside the 30 day window require a support review. Shipping fees are not refunded unless the item arrived damaged. These rules apply to standard purchases from our online store.

CHUNK 01words 1–40

Customers can request a refund within 30 days of purchase. The item must be unused and in its original packaging. Include the order number and a receipt when contacting support. Digital downloads are refundable only if they have not been

CHUNK 02words 33–72

are refundable only if they have not been accessed. Approved refunds are sent to the original payment method. Refund requests outside the 30 day window require a support review. Shipping fees are not refunded unless the item arrived damaged. These

CHUNK 03words 65–81

not refunded unless the item arrived damaged. These rules apply to standard purchases from our online store.

Live calculation · words are used here for readability; production splitters often count tokens. Underlined words are repeated in the next chunk.

03Embeddings

What is embedding?

Embedding is the process of turning text into numbers that represent its meaning, so the system can compare a question with your documents. An embedding model performs this conversion, producing a list of numbers called a vector.

In our search system, we embed each chunk and the user’s question. Comparing their vectors helps match “Can I get my money back?” with a passage about refunds, even though the wording differs. 3

How to choose an embedding model

Hosted models: use services such as or without running the model yourself. Your text is sent to the provider for processing.

Models you host: options such as or can run on your own infrastructure. You manage the hardware and operation.

Use the to build a shortlist, then test with your own documents and questions. Compare language support, input limits, search quality, and cost. Larger vectors also need more storage.

Use the same model and version for documents and questions, following its instructions for each input type. Switching to a different model usually means embedding your documents again.

How to prepare each chunk

Include the document title or section heading when a chunk needs more context. For example, adding “Refund policy › Customer handbook / Returns” helps identify the subject of a passage about shipping fees.

The example below compares the same passage with and without that context.

Embedded as is

Shipping fees are not refunded unless the item arrived damaged. These rules apply to standard purchases from our online store.

Which rules? Which store? This could match almost any question about rules or purchases.

Embedded with context

Refund policy › Customer handbook / ReturnsShipping fees are not refunded unless the item arrived damaged. These rules apply to standard purchases from our online store.

The title and section are added only for the embedding. The model and the reader still see the original text.

A chunk from the refund policy used in the labs below.

Remove unrelated navigation, repeated footers, and duplicate content before embedding. Keep the original passage alongside its vector so the system can show the text and cite its source.

For exact matches, such as order numbers or error codes, combine embeddings with keyword search. We cover this in the Retrieval section.

Try the embedding lab

Choose a question below and see which sources the system selects. Compare the refund question with a question about account security.

The lab uses a simplified model to illustrate the idea. Real embedding models learn their representations from data, and distances on this map are not the scores used to rank results.

LAB 03 — EMBED

A map of meaning

Refund policyPayment processingSubscription changesAccount securityDelivery guide?

“Can I get my money back, and how long will it take?”

QUESTION VECTOR
Returns
0.71
Timing
0.71
Billing
0.00
Security
0.00
Delivery
0.00
Eligibility
0.00
NEAREST SOURCES
Payment processing0.981
Refund policy0.730

Every dot is a stored chunk, coloured by source. Illustrative projection of six concept dimensions; map distances are not cosine distances. Real embedding dimensions are learned and unlabeled.

04Indexing

What is indexing?

Indexing is the process of storing and organizing your chunks so the system can find them when a question arrives. Each chunk stays connected to its embedding, original text, and metadata, such as its source, version, and access permissions. 4

For example, a refund-policy chunk should include which policy version it came from. This helps the system search the current policy and exclude an outdated one.

To store and search these embeddings, we use a vector database, or an existing database with vector-search support.

How to choose a database

Start with what your team already uses and how much infrastructure you want to manage.

Your situationConsiderWhy
You already use PostgreSQLAdds vector search alongside your existing data.
You want a dedicated vector databaseSupports vector search and metadata filtering, with hosted and self-managed options.
You prefer a managed serviceHandles the database infrastructure for you.
You want Milvus without managing the serversProvides fully managed Milvus, handling deployment and scaling.
You want to prototype locallyLets you start storing and searching embeddings locally.

Before choosing, check filtering, keyword search, updates and deletes, where the data lives, and the cost at your expected size.

How to keep the index current

Give each document a stable ID so you can find and replace its chunks when the source changes. Remove old chunks when replacing a document, and remove indexed content when its source is deleted.

Update access permissions too, and enforce them during retrieval. Saving a permission label does not prevent access by itself.

How to tune the search

An exact search compares the question’s vector with every eligible vector. An approximate search reduces that work to find results faster, but it may miss some close matches.

Start with the database defaults and compare results against exact search using your own questions. For an HNSW index, a common approximate-search method, these settings control the trade-off:

SettingWhat it changes
Search wideref_searchConsiders more candidates, which can improve results but takes longer.
Build settingsM, ef_constructionControl connections and construction effort, affecting memory, build time, and search quality.

Names vary by database. Use the comparison method your embedding model expects, such as cosine similarity. pgvector’s documentation explains these settings.

Try the indexing lab

Select a document to inspect its chunks, vectors, and version. Then toggle Current versions only to see how the archived refund policy is included or excluded.

This lab illustrates version filtering; production systems must also enforce each user’s access permissions.

LAB 04 — STORE

What belongs in the index?

SOURCES
12 chunks indexed2 archived chunks excluded
CHUNK_IDSOURCEVERSIONWORDSVECTORSTATUS
refund:0Refund policy2026-091–40current
refund:32Refund policy2026-0933–72current
refund:64Refund policy2026-0965–81current
timing:0Payment processing2026-091–40current
timing:32Payment processing2026-0933–72current
timing:64Payment processing2026-0965–82current
subscription:0Subscription changes2026-081–40current
subscription:32Subscription changes2026-0833–63current
security:0Account security2026-091–40current
security:32Account security2026-0933–59current
shipping:0Delivery guide2026-071–40current
shipping:32Delivery guide2026-0733–58current

Synthetic metadata · current-version filtering is shared with retrieval and evaluation below. In production, enforce authorization on the server before evidence reaches the model.

05Retrieval

What is retrieval?

Retrieval is the process of finding relevant chunks for a question and giving them to the model as evidence. This helps the model answer using your documents.

For example, “Can I get my money back, and how long will it take?” needs both the refund rules and the payment-processing timeline.

How to find relevant chunks

Search can match meaning, exact words, or both:

MethodWhat it doesUseful for
Vector searchCompares the question’s embedding with chunk embeddings.Different wording, such as “money back” and “refund.”
Keyword search (BM25)Ranks passages using matching words.Product names, error codes, and specific terms.
Hybrid searchCombines vector and keyword results.Questions that need both meaning and exact terms.

Hybrid search can merge results using Reciprocal Rank Fusion (RRF), which uses each passage’s position in the result lists. Microsoft’s documentation explains this approach. 5

Apply version and access filters so the search only returns eligible documents. For questions with several parts, try searching each part separately to cover the full question.

How to put the best results first

The first search finds possible matches. Reranking scores these passages against the question again to put the most relevant ones first.

Use for a hosted service, or to run the model yourself.

Searche.g. 30 candidatesFast and rough. Finds anything that might be relevant.
Reranksame 30, new orderA slower model reads the question with each candidate and scores it.
Keepe.g. top 5Only the best few go into the prompt.
Example numbers. Tune both against your evaluation set.

Fetch more candidates than you plan to use, then keep the strongest results. Reranking adds processing time, so test whether it improves answers enough to justify the delay.

How to prepare the evidence

The selected passages become the model’s context: the information supplied alongside the question.

  • Remove repetition: combine overlapping passages without losing useful details.
  • Keep source details: include each passage’s source and version so the answer can cite it.
  • Limit the context: keep enough evidence to answer the question without adding unrelated text.
  • Set clear instructions: ask the model to use the evidence and say when it is insufficient.

How to check the answer

Before generating an answer, check whether the retrieved passages cover the question. If they do not, ask for clarification or explain that the documents lack the answer.

After generation, check that each claim is supported by its cited passage. A valid link alone does not prove the claim is correct.

Check 1 · before
Is there enough evidence?

If nothing relevant was found, do not let the model guess.

Generate
Answer with citations

Instruct the model to answer from the supplied passages.

Check 2 · after
Is every claim supported?

Check each claim against its citation and confirm the source ID and version.

Both pass: send the answer.
Check 2 fails: retry with more evidence, or drop the unsupported sentence.
Check 1 fails: say the documents do not cover it, ask a question back, or hand off to a person.

Rules can detect missing citations or outdated sources. A second model can help check claim support, but its judgments need testing too. If a check fails, retrieve better evidence, revise the answer, or explain what remains unknown.

Try the retrieval lab

Select the refund question and change Evidence budget (top-k), the maximum number of chunks supplied to the model. Click Ask and compare the selected evidence and answer.

Then try the student-discount question to see what happens when the documents do not contain an answer.

The lab uses local mock data and simulated answer streaming. No AI service is called. It does not run the reranking or verification steps described above.

LAB 05 — RETRIEVE + GENERATE

Ask the knowledge base

Evidence budget (top-k)
Can I get my money back, and how long will it take?
  1. 01Question
  2. 02Encode
  3. 03Search
  4. 04Assemble
  5. 05Generate
ENCODE
Returns
Timing
Billing
Security
Delivery
Eligibility
SEARCH THE INDEX12 candidates
Refund policy—
Refund policy—
Refund policy—
Payment processing—
Payment processing—
Payment processing—
Subscription changes—
Subscription changes—
Account security—
Account security—
Delivery guide—
Delivery guide—
CONTEXT WINDOW

Selected evidence will appear here.

MOCK ANSWERwaiting

Press Ask. The demo searches the local documents and reveals a mock answer based on the selected evidence, with citations.

Exact cosine search · minimum similarity 0.15 · strongest chunk per source · similarity is not confidence. The answer and streaming are simulated locally from the selected evidence. No AI service is called.

06Evaluation

What is evaluation?

Evaluation is the process of testing whether your system finds useful evidence and produces correct answers. It helps you compare changes to chunking, search, models, or prompts before releasing them.

For example, finding the refund policy is only part of the job. The answer must also explain its rules correctly and use the right payment timeline.

What to measure

Check retrieval and answers separately so you can tell where a problem starts. 6

CheckMetricWhat it tells you
RetrievalRecallHow much of the required evidence was found.
RetrievalPrecisionHow much of the retrieved evidence was relevant.
RetrievalRank of first hitHow early a relevant result appeared.
AnswerFaithfulnessWhether the claims are supported by the supplied evidence.
AnswerRelevanceWhether the answer addresses the question.
AnswerCitation accuracyWhether each cited passage supports its claim.
AnswerCorrectnessWhether the answer agrees with a trusted reference or expert review.

Use for ready-made RAG metrics, or to run evaluation as automated tests. Some metrics use an AI model to judge answers; compare those scores with human reviews.

How to build a test set

A golden set is a collection of questions with the answers and sources you expect. Run the same set before and after a change.

Start with real questions from support tickets or search logs, with personal details removed. Include questions that need several sources, involve outdated or restricted documents, or cannot be answered from your documents.

For each question, record:

  • Expected answer: the facts a correct answer should contain.
  • Required sources: the documents needed to support it.
  • Excluded sources: outdated versions or documents the user cannot access.
One record in a golden set
question
“Can I get my money back, and how long will it take?”
type
two-part question
must find
Refund policy and Payment processing
must not use
Refund policy (archived 2024 version)
expected answer
Yes, within 30 days if the item is unused. The money usually arrives in 5–10 business days.
source of question
support ticket, personal details removed
Made-up example from our pretend store, matching the documents used in the labs.

Version the set as it grows. Keep some questions aside for final testing so you can check whether improvements work beyond the examples used for tuning.

How to use the results

  1. 01Test each change

    Change one setting at a time and compare results with the previous version.

  2. 02Review failures

    Inspect the retrieved passages and answer to understand what went wrong.

  3. 03Add real examples

    Turn reviewed production failures into new test cases.

Set pass requirements around your users’ needs. In production, also track response time, cost, and document freshness. Review results by question type because a good average can hide recurring failures.

Try the evaluation lab

Click Run evaluation to test four sample questions using your current lab settings.

Then return to the Retrieval lab, change Evidence budget (top-k) from 2 to 1, and run evaluation again. Compare Source recall, Source precision, and Cases passed.

This lab runs locally without AI calls. It checks whether the expected sources were retrieved; it does not score passage completeness or answer quality.

LAB 06 — MEASURE

Test source retrieval

40w chunks · 8w overlap · top-2
Current versions only
Source recall—Mean across 3 answerable cases
Source precision—Mean across 3 answerable cases
Cases passed— / 4All expected sources, or no results for the out-of-scope question
01
Can I get my money back, and how long will it take?
expected: refund, timing · retrieved: …
queued
02
How do I cancel my subscription?
expected: subscription · retrieved: …
queued
03
How do I protect my account login?
expected: security · retrieved: …
queued
04
Do you offer a student discount?
expected: abstain · retrieved: …
queued

For each answerable question, recall = relevant sources retrieved ÷ labeled relevant sources. Precision = relevant sources retrieved ÷ all sources retrieved (zero if nothing is retrieved). The fourth case passes only when the retriever returns no evidence. These checks measure source coverage, not passage completeness or generated-answer quality. Teaching fixture, not a held-out benchmark.

07Before you ship

Checklist before you go live

Use this checklist to review your own system before release. Each item links back to the section that explains what to check.

  1. 01IngestionCheck the parsed contentCheck reading order, table headings, OCR results, and the context around images or formulas. Remove page clutter.
  2. 02IngestionKeep source detailsKeep each chunk connected to its source, section, version or update date, and access permissions.
  3. 03ChunkingTest chunk size and overlapChoose a strategy that fits your documents. Check that retrieved chunks keep the context needed to understand rules and exceptions.
  4. 04EmbeddingsUse compatible embeddingsUse the same model and version for documents and questions, following its input instructions. Add titles or headings where context is needed.
  5. 05IndexingKeep the index currentUse stable document IDs to replace old chunks. Remove deleted content and filter out outdated versions when answering current-policy questions.
  6. 06IndexingEnforce access permissionsTest that users only retrieve documents they can access. Saving a permission label is not enough; search and cached answers must enforce it.
  7. 07RetrievalTest which evidence is selectedCheck questions with different wording and exact terms. Test whether hybrid search or reranking improves the results enough to justify the extra work.
  8. 08RetrievalCheck answers and citationsCheck claims against their cited passages. When evidence is missing, ask for clarification or explain what the documents cannot answer.
  9. 09EvaluationCompare against your test setRun the same golden set before and after each change. Review failures and meet your pass requirements before release.
  10. 10EvaluationMonitor real useTrack answer quality, response time, cost, and document freshness. Review failures by question type and add useful examples to the test set.
RELEASE RECORD

Keep a record of what you tested.

Save the source versions, parser, chunking settings, embedding model, index settings, prompt, and answer model alongside your evaluation results. This helps you repeat a test and trace what changed.