toolcall() ← all concepts

// concept · RAG

Your RAG is broken. Find out where.

Your system answered a question from your own documents and got it wrong. Three different bugs produce that same symptom, the fix for one does nothing for the other two, and almost everybody reaches for the prompt first. This page is the triage that runs before you pick a fix.

// the check

Before changing anything, take the question that failed and print the chunks your retriever actually returned for it. Then read them and ask one question:

# is the answer in here at all? yes → it had it and fumbled it (the model) no → then search your docs by hand: found it → ranked too low (retrieval) not there → never indexed (ingestion)

That is the whole diagnostic, and it has a name. RAGAS calls it context recall; LlamaIndex's binary version is hit rate — "the fraction of queries where the correct answer is found within the top-k retrieved documents". Evaluating the retriever and the generator separately is not a folk technique; it is what every major evaluation framework is built around.

One condition. This check needs you to already know the right answer, which is fine when you are debugging a specific failure. It is not something you can run blind on live traffic — for that, the reference-free variants (TruLens's context relevance, DeepEval's contextual relevancy) judge the chunks against the query instead.

Branch one: it ranked too low

You ask about the refund window. What comes back is three paragraphs about shipping. The refund clause exists in your documents — it just ranked eleventh, and you kept five.

query: "what's our refund window?" 1. shipping times by region 2. shipping carrier terms 3. shipping address rules — cut at top 5 — 11. refund window: 14 days

There is nothing in that context to extract, so the model does what models do and produces something plausible. This is FP2 in Barnett et al.'s peer-reviewed RAG failure-point taxonomy: "the answer to the question is in the document but did not rank highly enough to be returned".

The fixes all live on the retrieval side: chunking, the embedding model, hybrid search, reranking, and how many chunks you keep.

Branch two: it was right there

Now the same question, and this time the refund clause is in the retrieved context — and the answer is still wrong. Barnett calls this FP4, "Not Extracted": the answer is present in the context and the model failed to pull it out.

The best-known driver is position. The Lost in the Middle work varied where a single correct document sat in a twenty-document context and found a U-shaped curve: strong at the start, strong at the end, weakest in the middle. In the sharpest case, a model given the correct document buried mid-context did worse than being given no documents at all.

Date that finding. Those measurements are from 2023, on 2023 models. Position bias is still a real, documented attention effect and later work confirms it persists — but the magnitude varies by model and mitigations exist. Treat it as a reason to check position, not as a number to quote.

"No prompt can fix it" is half true

It is worth being precise here, because the punchy version is wrong. A prompt cannot conjure information that is not in the context. But it absolutely can change the failure mode — from inventing an answer to admitting there isn't one.

That is the documented first-party remedy on both sides of the industry. Anthropic's guidance is to explicitly give the model permission to say it doesn't know, to ground long-document work in extracted quotes, and to restrict it to the provided documents. OpenAI's cookbook uses the same instruction — answer from the context only, and say "I don't know" otherwise — and measured hallucination on deliberately unanswerable questions dropping from 47% to 25% with few-shot examples of declining.

# the generation-side fix Answer only from the provided context. If the context does not contain the answer, say so.

So a retrieval bug still needs a retrieval fix. But a system that says "that isn't in my documents" is enormously more useful than one that guesses, and that part is a prompt change.

Branch three: it was never indexed

Here is the correction most versions of this advice skip. When the check comes back "the answer isn't in the chunks", that is ambiguous between two bugs with opposite fixes:

not in the chunks… …but it IS in your index → rank it higher (rerank, hybrid, top-k) …and it is NOT in your index → nothing in retrieval ever helps

Barnett's FP1, "Missing Content" — the question simply cannot be answered from the documents you ingested — sits upstream of retrieval entirely. Follow a tidy two-stage story and you will go buy a better embedding model for a page nobody ever indexed.

Two of the seven published failure points don't fit a two-stage split at all. The other is FP3: the chunk was retrieved and still never made it into the prompt, killed by a top-k limit or context assembly. Retrieval worked, generation never saw it, and the bug is in the pipe between them.

The rule

Look at the chunks before you touch the prompt. If the answer isn't in them, the model was never going to get it right and you are debugging the wrong branch — then check whether it is in your index at all before you tune anything.

And treat this as the first thing you look at when an answer is wrong, not as a test you run once. The same paper's conclusion is that validating a RAG system is only really feasible while it is running: robustness gets discovered, not designed in.

Next: why you need an eval set at all, and how big it has to be before its numbers mean anything.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click