In RAG, recall is the number that matters
Most RAG demos are convincing for the wrong reason. You ask a question, the answer comes back fluent and correct, and it feels solved. Then you put it in front of real documents and real questions, and it starts confidently answering from nothing, because the chunk it needed was never retrieved.
That's the failure mode that matters. Not a slightly-off phrasing, not a suboptimal prompt: the correct source simply not being in the context window.
Retrieval is the ceiling
The generation step can only work with what retrieval hands it. If the right passage isn't in the retrieved set, the model has two options: refuse, or invent. Neither is what you want. So the ceiling on answer quality is set at retrieval, not generation, which means recall is the metric to chase.
Precision (how much of what you retrieved was relevant) matters too, but it's recoverable. You can rerank a noisy set down. You cannot rerank a document that was never retrieved back into existence.
What that changes
Optimizing for recall changes concrete decisions:
- Chunk for standalone meaning. A chunk that loses its context when isolated hurts recall, because it no longer matches the queries it should.
- Retrieve more, then narrow. Pulling a wider set and reranking beats pulling a tight set and hoping.
- Measure the thing that fails. Track whether the known-correct source was retrieved at all. That's the number that predicts trust.
I wrote about the production version of this in the rag-pipeline case study. The short version: make the right document show up, and the rest of the system gets a lot easier to trust.