Pranav Belhekar

Visual guide

A Visual Guide to RAG Evaluation

A RAG answer can fail because retrieval missed the evidence, reranking buried it, context construction damaged it, or generation ignored it. A practical, diagram-led guide to measuring every layer and fixing the right one.

15 min read

A support assistant gives a confident answer about customer-data exports. The answer is wrong. The team rewrites the prompt, raises the temperature, lowers it again, swaps the model, and adds a stern instruction to “use only the supplied context.” Nothing fixes the failure consistently.

The relevant policy was in the index the whole time. It appeared at rank 23. The reranker only received the first 20 candidates.

This is the central problem with evaluating retrieval-augmented generation: the user sees one answer, but the answer is the output of several systems with different contracts. A single end-to-end score tells you whether the experience failed. It rarely tells you what to repair.

A useful RAG evaluation therefore works like an instrument panel. It measures whether the evidence entered the candidate set, whether the best evidence survived the context cutoff, whether the assembled context preserved the decisive details, whether the generated claims stayed inside that evidence, and whether the system knew when not to answer.

This guide builds that panel from first principles.

What’s in this guide: four contracts · evaluation records · retrieval · reranking · grounding · abstention · slices · diagnosis · a minimal harness · the operating loop · references

One answer, four contracts

RAG is often drawn as a two-box system: retriever, then generator. That is a helpful introduction and a poor debugging model. Production systems usually contain at least four contracts.

  1. Retrieval: produce a candidate set that contains the evidence needed to answer.
  2. Ranking: put the most useful evidence inside a much smaller context budget.
  3. Context construction: preserve the relevant spans, metadata, ordering, and authority signals when chunks are assembled.
  4. Generation: produce claims supported by that context, or abstain when support is missing.
ONE ANSWER, FOUR CONTRACTS User question “Can contractors download customer exports?” 01 / RETRIEVE Find candidate evidence Failure: the policy never enters the candidate set. Measure: recall@k 02 / RANK Promote useful evidence Failure: the policy is found but lands below the cutoff. Measure: nDCG / MRR 03 / CONTEXT Preserve usable evidence Failure: truncation removes the decisive condition. Measure: context coverage 04 / GENERATE Answer from evidence Failure: the model overrides the policy. Measure: faithfulness The answer is only the final symptom. If you score only the final response, every failure looks like “the model was wrong.” Score each contract and the repair becomes obvious.
One user-visible failure can originate in four different contracts. The metric under each box answers a different engineering question; no single score can replace all four.

The boxes fail differently. A retriever can miss the policy because the query says “external staff” while the document says “contractor.” A reranker can find the policy and still bury it beneath semantically similar FAQs. Context construction can cut a table away from its header or keep the permission but drop its expiry condition. The generator can receive perfect evidence and choose an answer from its parametric memory instead.

The repair follows the failed contract:

  • Low retrieval recall: improve chunking, query rewriting, metadata filters, sparse-dense fusion, or the embedding model.
  • Good recall but poor final ranking: improve the reranker, candidate count, feature set, or cutoff.
  • Good final context but incomplete evidence: fix chunk boundaries, document parsing, deduplication, and context packing.
  • Good evidence but unsupported claims: change the generation policy, prompting, claim verification, or model.

This sounds obvious written as a list. In practice, many teams record only the question and final answer. That removes the evidence needed to distinguish the four cases.

Start with an evaluation record, not a score

An evaluation set is not a spreadsheet of questions and ideal prose answers. A polished reference answer is often the least reusable part of the record: two very different responses can both be correct, and string similarity punishes the difference.

The durable unit is a small specification of what the system must know and where that knowledge lives.

AN EVALUATION CASE IS A SMALL SPECIFICATION eval_case / access-control-017 QUESTION Can contractors download customer exports? ANSWERABILITY ANSWERABLE The corpus contains an authoritative policy. EXPECTED CLAIMS • Contractors may export only with a temporary approval. • The approval expires after 24 hours. SUPPORTING EVIDENCE policy-42 §7.3 policy-42 §7.4 SLICES permission temporal policy A reference answer is optional. Reference claims and evidence are what make diagnosis possible.
The anatomy of a useful evaluation case. Expected claims support flexible answers; evidence identifiers let you test retrieval separately; answerability and slice labels turn the case into a diagnostic tool.

For each case, store:

  • Question: the language a real user would use, including ambiguity and shorthand.
  • Answerability: whether the current authorized corpus contains enough evidence to answer.
  • Expected claims: the smallest factual units a correct answer must contain.
  • Supporting evidence: stable document and span identifiers, preferably with document version.
  • Slice labels: the reason the case is difficult: paraphrase, multi-hop, temporal, permission-sensitive, unanswerable, and so on.
  • Provenance: where the case came from and who last reviewed it.

Evidence identifiers matter more than a copied reference paragraph. A copied paragraph goes stale when a policy changes. A stable source identifier and version tell you that the evaluation case itself needs review.

Start with real work. Sample questions from support logs, search misses, analyst workflows, and subject-matter experts. Add synthetic cases to widen coverage, but do not let generated questions become the entire benchmark. Synthetic data tends to inherit the vocabulary and clean structure of its source documents; real users do neither.

My starting point for a new system is 30–50 carefully reviewed cases that represent the most costly failures. That is too small to declare statistical victory and large enough to expose architecture mistakes. Grow the set from production traces after that.

Retrieval: did the evidence enter the candidate set?

The retriever’s job is not to answer the question. Its job is to make the answer possible downstream.

For a case with a known set of relevant evidence items, the basic metric is recall at a cutoff:

recall@k = relevant items found in top k / all relevant items

Suppose a question needs two policy sections and the retriever finds both in its top 50. Recall@50 is 1.0. If one section is absent, recall@50 is 0.5. For cases where any one of several passages is sufficient, record that rule explicitly and use hit rate: did at least one sufficient passage appear?

The cutoff must match the actual handoff. If the reranker receives 50 candidates, recall@10 measures an imaginary system. Measure recall at the candidate boundary that exists in production.

THE CANDIDATE SET AND THE CONTEXT SET ARE DIFFERENT TESTS Retriever output Eight candidates, ordered by retriever score rank 1 FAQ-8 distractor rank 2 POL-42 RELEVANT rank 3 DOC-11 distractor rank 4 FAQ-3 distractor rank 5 DOC-18 distractor rank 6 RUN-5 distractor rank 7 POL-43 RELEVANT rank 8 DOC-2 distractor RETRIEVER CANDIDATES · recall@8 = 2 / 2 FINAL CONTEXT · 1 of 2 relevant chunks survives Retriever question Did every required evidence item enter the pool? Reranker question Did the best evidence survive the context cutoff?
The retriever succeeds at its actual cutoff: both required policy chunks are somewhere in the eight candidates. The final context still fails because the second chunk sits below the context boundary. That is a ranking failure, not a retrieval miss.

Precision can be useful here, especially when retrieval is expensive or irrelevant documents create latency. But for the broad candidate stage, recall usually has priority. The reranker cannot promote evidence it never receives.

Three details make retrieval evaluation more honest:

Evaluate at the evidence unit the system uses. If your index stores chunks, label chunk or span identifiers rather than only document identifiers. Retrieving the right 80-page PDF is not success when the decisive paragraph never reaches the model.

Keep the authorization filter in the test. A passage retrieved without required access control is not a relevant hit. It is a security failure. Evaluation should execute the same tenant, role, and metadata filters used in production.

Record corpus version. A retrieval miss against yesterday’s index and a miss against today’s source of truth are different events. Without index and document versions, the result cannot be reproduced.

Reranking: did the right evidence reach the context?

A retriever is allowed to be generous. The final context is not. Every irrelevant chunk consumes tokens, attention, and an opportunity to distract the model.

The ranking stage therefore asks a different question: how early does useful evidence appear?

  • Precision@k measures how much of the final set is relevant.
  • Mean Reciprocal Rank (MRR) rewards placing the first relevant result early. It works well when one passage can answer the question.
  • nDCG@k handles several results with graded relevance and discounts useful evidence that appears late.
  • Context recall or coverage measures how many expected claims are supported somewhere in the final assembled context.

Do not collapse these into a universal “retrieval score.” A system can have perfect candidate recall and terrible context precision. It can also have high precision while missing one mandatory condition. For a legal or access-control answer, that missing condition can matter more than every correct sentence around it.

The cleanest experiment holds generation constant. Save the candidate set, rerank it with version A and version B, and compare the evidence metrics before running either through an LLM. This makes ranking changes cheap to test and removes generation variance from the decision.

Context construction deserves its own trace even when it is not a learned model. Store the exact text sent to the generator, in order, after deduplication and truncation. The failure may be in the glue: a markdown parser dropped a table header, two overlapping chunks repeated one policy five times, or a high-authority document lost its date.

Generation: did the answer stay inside the evidence?

Once the correct evidence reaches the prompt, the evaluation changes from search to claims.

Split the answer into atomic claims. For each claim, ask whether the supplied context supports it. Then ask whether the set of claims covers the expected answer. These are separate tests:

  • Faithfulness or supportedness: What fraction of generated claims is supported by the provided context?
  • Completeness: What fraction of expected claims appears in the answer?
  • Answer correctness: Does the response solve the user’s task, including conditions and scope?
  • Citation validity: Do the cited sources exist and correspond to retrieved evidence?
  • Citation entailment: Does each cited span actually support the claim attached to it?
A CITATION CAN EXIST AND THE ANSWER CAN STILL BE WRONG Generated answer “Yes. Contractors can export customer data whenever a manager approves it.” [Policy 42] CHECK 1 Citation validity Does Policy 42 exist, and was it actually retrieved? PASS CHECK 2 Claim support Does the cited text support “whenever” and “manager”? FAIL CHECK 3 Answer completeness Did it include the temporary approval and 24-hour limit? FAIL Validity checks the pointer. Faithfulness checks the claim. Correctness checks the task.
A citation can point to a real retrieved document and still fail the answer. The pointer is valid, but the cited text does not support the invented words “whenever” and “manager,” and the answer omits the 24-hour constraint.

Citation validity is the easiest check and the weakest guarantee. A model can cite a real policy while changing its meaning. Faithfulness evaluates the relationship between claims and evidence. Correctness evaluates the relationship between the answer and the user’s need.

LLM judges are useful here because claim support is a semantic task. They are not ground truth. Calibrate them against a small human-labeled set, keep the judging rubric and model version fixed, and sample disagreements for review. The ARES paper demonstrates this hybrid pattern: automated judges become much more useful when anchored by a few hundred human annotations and statistically corrected rather than trusted in isolation.

Also run deterministic checks wherever possible. Citation identifiers, required sections, numeric values, dates, JSON schemas, and banned unsupported sources do not need a language model to score them.

Abstention is part of correctness

A RAG system has two legitimate actions: answer from evidence or decline because sufficient evidence is unavailable. Evaluation must include both answerable and unanswerable questions or it will reward the wrong behavior.

ABSTENTION IS A CLASSIFICATION PROBLEM SYSTEM BEHAVIOR CORPUS REALITY ANSWERS ABSTAINS ANSWERABLE UNANSWERABLE CORRECT ANSWER Evidence exists and the system uses it. reward: answer quality UNNECESSARY REFUSAL The answer was available. cost: over-refusal UNSUPPORTED ANSWER The system fills the gap from prior belief. cost: fabrication CORRECT ABSTENTION The system asks, searches, or says “I don't know.” reward: calibrated trust A system that never hallucinates because it never answers is not reliable. It is unavailable.
The abstention matrix. Measuring only hallucination encourages refusal; measuring only answer rate encourages fabrication. A dependable system needs both correct answers and correct abstentions.

For unanswerable cases, test several causes:

  • The corpus genuinely lacks the information.
  • The information exists but the user is not authorized to access it.
  • Two authoritative sources conflict.
  • The question requires a newer document version than the index contains.
  • The question is ambiguous enough that a clarifying question is safer than an answer.

Track unsupported-answer rate on unanswerable cases and over-refusal rate on answerable cases. A system can drive hallucinations to zero by refusing everything. That produces a perfect safety statistic and a useless product.

The desired behavior may be more specific than “I don’t know.” In some products the correct action is to ask a clarifying question. In others it is to open a support ticket, surface the conflicting sources, or say which document is missing. Encode the acceptable action in the case.

An average score will lie to you

An aggregate is valuable for release tracking and dangerous for diagnosis. Two systems with the same 84% correctness can feel completely different: one may fail randomly across low-value questions, while the other fails every permission query from the same customer role.

THE AVERAGE IS GREEN. TWO USER JOURNEYS ARE BROKEN. ALL 240 CASES 84% answer correctness The release dashboard looks healthy. Ship? Not until you slice it. Straight lookup 96% Paraphrase 91% Multi-hop 82% Unanswerable 78% Temporal policy 61% Permissions 54% Slice by the way the system can fail, not only by department or dataset source. Averages describe a benchmark. Slices describe a product.
The overall score hides two weak slices. Temporal and permission questions are the cases where conditions, document versions, and access filters matter most; averaging them with easy lookups disguises the product risk.

Build slices around mechanisms of failure:

  • lexical lookup versus paraphrase;
  • single-hop versus multi-hop evidence;
  • stable facts versus temporal policies;
  • one source versus conflicting sources;
  • answerable versus unanswerable;
  • normal access versus permission-sensitive access;
  • short documents versus tables, scans, and long structured documents;
  • common queries versus rare, high-consequence queries.

Then set gates on important slices, not only on the global mean. A change that raises overall correctness by two points while dropping authorization-sensitive accuracy by ten should not ship.

The BEIR benchmark made a related lesson visible for information retrieval: rankings that look strong in one narrow setting do not necessarily generalize across domains and retrieval tasks. Your evaluation set needs the same heterogeneity your product faces.

How to diagnose one failed query

When a case fails, walk backward from the answer through the recorded trace.

1. Was the question answerable from the authorized corpus? If no, inspect the abstention behavior. If yes, continue.

2. Did the required evidence appear in the retriever’s candidate set? If no, the failure belongs to ingestion, parsing, indexing, query construction, filtering, or retrieval.

3. Did the evidence survive into the final context? If no, inspect reranking, deduplication, chunk expansion, ordering, and token-budget truncation.

4. Did the final context preserve every expected claim? If no, the problem is context construction even if the source document was technically present.

5. Did the generated answer use the available evidence faithfully? If no, the generator or its instructions failed.

6. Did the judge score the case correctly? Evaluation code is software. It has bugs, model drift, ambiguous rubrics, and versioning problems of its own.

This sequence prevents a common waste: changing the prompt when the evidence never reached the prompt.

A minimal evaluation harness

You do not need a large platform to start. You need stage traces, stable cases, versioned configuration, and a comparison against a baseline.

type EvalCase = {
  id: string;
  question: string;
  answerable: boolean;
  expectedClaims: string[];
  evidenceIds: string[];
  slices: string[];
};

type RagTrace = {
  query: string;
  candidates: Array<{ id: string; score: number }>;
  ranked: Array<{ id: string; score: number }>;
  context: Array<{ id: string; text: string }>;
  answer: string;
  citations: string[];
  versions: {
    corpus: string;
    retriever: string;
    reranker: string;
    prompt: string;
    generator: string;
    judge: string;
  };
};

type EvalResult = {
  candidateRecall: number;
  contextCoverage: number;
  faithfulness: number;
  completeness: number;
  abstentionCorrect: boolean;
};

Persist the trace before scoring it. A result without the retrieved items, exact final context, and component versions cannot explain a regression.

Run the same cases in three modes:

  1. Component tests for retrieval and ranking, without generation.
  2. End-to-end tests for the user-visible answer and latency.
  3. Paired comparisons between the current baseline and one proposed change.

Change one major variable at a time. If you update the chunker, embedding model, reranker, prompt, and generator together, a better final score teaches you almost nothing. Paired runs give every case the same question and corpus, making the comparison far easier to interpret.

Track cost and latency beside quality. Candidate counts, rerank depth, context length, and judge calls all have operational prices. The best configuration is the one that satisfies the quality gates within the product’s latency and cost budget, not the one that maximizes a leaderboard score without constraints.

Turn production failures into permanent tests

An evaluation suite is not a document completed before launch. It is the memory of the system’s mistakes.

THE EVALUATION LOOP IS PART OF THE SYSTEM 01 / OBSERVE Capture production traces 02 / CURATE Label cases and slices 03 / DIAGNOSE Score each contract 04 / CHANGE Fix the failing layer REGRESSION GATE Compare with the frozen baseline before shipping. Every production failure that matters should become a case that can never silently return.
The operating loop. Production traces feed reviewed cases; cases expose the failing contract; the fix is compared against a frozen baseline; every important failure becomes a permanent regression test.

The loop is straightforward:

  1. Capture traces with privacy and access controls intact.
  2. Sample failures, low-confidence answers, refusals, and user corrections.
  3. Have a domain expert label answerability, expected claims, and evidence.
  4. Add the case to the right diagnostic slices.
  5. Reproduce the failure offline and fix the responsible layer.
  6. Run the full suite against the frozen baseline before deployment.
  7. Keep the case so that failure cannot return silently.

Online feedback belongs in this loop, but a thumbs-up is not a truth label. Users reward tone, speed, and confirmation of what they already believe. Treat feedback as a signal for sampling and review, not as an automatic correctness score.

Monitor production distributions too. Offline quality can remain flat while the share of table questions, new policy versions, or unanswerable requests changes. The system did not necessarily regress; the work arriving at the system changed. Slice volume reveals that shift.

What I would build first

For a new RAG product, my first evaluation milestone is deliberately small:

  • 30–50 reviewed cases from real workflows;
  • explicit answerable and unanswerable examples;
  • stable evidence identifiers and corpus versions;
  • candidate retrieval recall at the real reranker cutoff;
  • context coverage at the real generation cutoff;
  • claim-level faithfulness and expected-claim completeness;
  • abstention and over-refusal rates;
  • exact traces for every failed case;
  • slice-level gates for the two or three highest-risk journeys;
  • one frozen baseline used for paired comparisons.

That foundation is more valuable than a dashboard with twenty opaque metrics. Add judge ensembles, synthetic generation, statistical confidence intervals, and online experiments after the traces and labels are trustworthy.

The research agrees on the central decomposition. RAGAS separates context relevance, faithfulness, and answer relevance. ARES evaluates context relevance, answer faithfulness, and answer relevance with calibrated judges. RAGChecker pushes toward finer-grained diagnosis of retriever and generator behavior. The metric names vary. The architectural lesson does not: evaluate the parts if you want to improve the system.

A short opinion

Most RAG teams do not have a model problem first. They have an observability problem.

They can show a polished answer in a demo, but they cannot reconstruct which query ran, which index version answered it, which documents were filtered out, which chunks were reranked, which text reached the model, or which claims the evidence supported. When the answer fails, the team debates prompts because prompts are the only visible component.

The highest-leverage improvement is to make the evidence path inspectable and turn its contracts into tests. Once that exists, model changes become engineering decisions rather than rituals. You can say exactly which slice improved, which layer caused it, what it cost, and what regressed.

That is also the production philosophy behind NEXUS Volume I: retrieval-augmented generation becomes reliable when every layer is measurable, replaceable, and accountable for a clear contract.

References and further reading

Get the next guide when it ships

Plus field notes from production AI, roughly twice a month.