AI Testing & Evaluations

RAG Evaluation Metrics Explained: Groundedness, Attribution and Context Relevance

A RAG system that gives a wrong answer could have retrieved the wrong documents, retrieved the right documents and ignored them, or retrieved the right documents and misread them. A single pass/fail accuracy score can’t distinguish between these — which is exactly what component-level RAG metrics are for.

01

Why RAG evaluation needs its own metrics

Retrieval-augmented generation has two distinct failure surfaces — the retrieval step (did it…

02

The four metrics worth knowing

Context relevance — of the documents or chunks retrieved, how many are actually relevant to…

03

Reading the pattern, not just the scores

The diagnostic value is in how these metrics combine.

Why RAG evaluation needs its own metrics

Retrieval-augmented generation has two distinct failure surfaces — the retrieval step (did it find the right source material?) and the generation step (did it use that material correctly?) — and a system can fail at either independently of the other. Evaluating only the final answer’s correctness conflates both failure modes into one score, which tells you something is wrong without telling you where to fix it.

Context RelevanceAre retrieved chunksactually relevant?Context SufficiencyEnough info toanswer at all?GroundednessDoes the answer traceback to the context?AttributionDoes each citation pointto the right source?
Low relevance points at retrieval; high relevance with low groundedness points at generation — the pattern across metrics, not any one score, is what tells you where to fix the system.

The four metrics worth knowing

  • Context relevance — of the documents or chunks retrieved, how many are actually relevant to the query? Low context relevance means the retrieval step itself is the problem, before generation even enters the picture.
  • Context sufficiency — even when retrieved context is relevant, does it contain enough information to actually answer the question? A system can retrieve relevant but incomplete context and still fail, through no fault of the generation step.
  • Groundedness — does the generated answer’s content actually trace back to the retrieved context, or does it include claims the context doesn’t support? Low groundedness with high context relevance points to a generation-step hallucination problem, not a retrieval problem.
  • Attribution — for systems that cite sources, does each specific claim trace to the specific source it’s attributed to? A citation that’s present but pointing to the wrong source passage is a subtler, harder-to-catch failure than a missing citation.

Reading the pattern, not just the scores

The diagnostic value is in how these metrics combine. Low context relevance with any groundedness score points at retrieval — fix the retriever, embedding model, or chunking strategy. High context relevance paired with low groundedness points at generation — the model is ignoring or misusing good source material, which is a prompting or model-selection problem, not a retrieval one. That pattern is what makes component-level evaluation worth the extra setup over a single end-to-end accuracy check.

Chunk utilization is a useful fifth signal worth tracking alongside these four: of the content actually retrieved, how much of it does the model use in generating the answer? Very low chunk utilization alongside correct answers can indicate the system is retrieving more than it needs, which doesn’t break correctness but adds unnecessary cost and latency worth tuning.

Building this into a regular evaluation cycle

These metrics are most useful run continuously against a curated test set as the system changes — a new embedding model, a different chunking strategy, an updated prompt — following the same continuous evaluation pipeline pattern used for LLM applications generally, rather than as a one-time assessment at launch that goes stale as the underlying document corpus and query patterns evolve.

Frequently asked questions

Do we need all four metrics, or can we start with just one or two?

Context relevance and groundedness are the two with the highest diagnostic value for distinguishing retrieval from generation failures — a reasonable starting pair if you need to prioritize, with context sufficiency and attribution added once the basic retrieval-versus-generation split is in place.

Can these metrics be measured automatically, or do they require human review?

Several can be measured with an LLM-as-a-judge approach (scoring groundedness and context relevance against the retrieved content) at scale, though periodic human validation of the automated scores — the same discipline that applies to LLM-as-a-judge generally — is worth building in rather than trusting the automated scores indefinitely.

How do these metrics relate to overall hallucination testing?

They’re a more granular, RAG-specific version of groundedness scoring, one of the four measurement approaches covered in our broader piece on LLM hallucination testing. These metrics add the retrieval-versus-generation diagnostic layer that general hallucination testing doesn’t always provide on its own.