// the check
Before changing anything, take the question that failed and print the chunks your retriever actually returned for it. Then read them and ask one question:
That is the whole diagnostic, and it has a name. RAGAS calls it context recall; LlamaIndex's binary version is hit rate — "the fraction of queries where the correct answer is found within the top-k retrieved documents". Evaluating the retriever and the generator separately is not a folk technique; it is what every major evaluation framework is built around.
Branch one: it ranked too low
You ask about the refund window. What comes back is three paragraphs about shipping. The refund clause exists in your documents — it just ranked eleventh, and you kept five.
There is nothing in that context to extract, so the model does what models do and produces something plausible. This is FP2 in Barnett et al.'s peer-reviewed RAG failure-point taxonomy: "the answer to the question is in the document but did not rank highly enough to be returned".
The fixes all live on the retrieval side: chunking, the embedding model, hybrid search, reranking, and how many chunks you keep.
Branch two: it was right there
Now the same question, and this time the refund clause is in the retrieved context — and the answer is still wrong. Barnett calls this FP4, "Not Extracted": the answer is present in the context and the model failed to pull it out.
The best-known driver is position. The Lost in the Middle work varied where a single correct document sat in a twenty-document context and found a U-shaped curve: strong at the start, strong at the end, weakest in the middle. In the sharpest case, a model given the correct document buried mid-context did worse than being given no documents at all.
"No prompt can fix it" is half true
It is worth being precise here, because the punchy version is wrong. A prompt cannot conjure information that is not in the context. But it absolutely can change the failure mode — from inventing an answer to admitting there isn't one.
That is the documented first-party remedy on both sides of the industry. Anthropic's guidance is to explicitly give the model permission to say it doesn't know, to ground long-document work in extracted quotes, and to restrict it to the provided documents. OpenAI's cookbook uses the same instruction — answer from the context only, and say "I don't know" otherwise — and measured hallucination on deliberately unanswerable questions dropping from 47% to 25% with few-shot examples of declining.
So a retrieval bug still needs a retrieval fix. But a system that says "that isn't in my documents" is enormously more useful than one that guesses, and that part is a prompt change.
Branch three: it was never indexed
Here is the correction most versions of this advice skip. When the check comes back "the answer isn't in the chunks", that is ambiguous between two bugs with opposite fixes:
Barnett's FP1, "Missing Content" — the question simply cannot be answered from the documents you ingested — sits upstream of retrieval entirely. Follow a tidy two-stage story and you will go buy a better embedding model for a page nobody ever indexed.
Two of the seven published failure points don't fit a two-stage split at all. The other is FP3: the chunk was retrieved and still never made it into the prompt, killed by a top-k limit or context assembly. Retrieval worked, generation never saw it, and the bug is in the pipe between them.
The rule
Look at the chunks before you touch the prompt. If the answer isn't in them, the model was never going to get it right and you are debugging the wrong branch — then check whether it is in your index at all before you tune anything.
And treat this as the first thing you look at when an answer is wrong, not as a test you run once. The same paper's conclusion is that validating a RAG system is only really feasible while it is running: robustness gets discovered, not designed in.
Next: why you need an eval set at all, and how big it has to be before its numbers mean anything.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
