// retrieval only sees what you send it
In a conversational app the second question is usually not self-contained. "Which one is cheaper?" · "What about the annual plan?" · "Does it do that too?" The subject lives in the previous turn, not in the text you embed.
Resolve it before you retrieve
The fix is a step you insert ahead of retrieval: take the conversation so far, resolve what the question is actually about, and embed that.
This is standard, widely documented practice, and it belongs to a family of techniques — query expansion, decomposition, paraphrasing, multi-query generation, step-back prompting. The conversational case is specifically the coreference one: put the missing noun back.
The catch: a rewrite can change the question
Rewriting is not free, and more rewriting is not better. Two documented failure modes:
And the broader finding is blunter than that. Across eight conversational QA datasets, several advanced techniques failed to yield gains and could degrade performance below the no-RAG baseline — effective conversational RAG depended less on method complexity than on whether the retrieval strategy matched the data.
The rule
Resolve the question, don't redecorate it. Fill in the missing subject; do not turn a sentence into a pile of keywords. The goal is a query that means exactly what the user meant, stated in full.
Related: why your RAG retrieves garbage covers the chunking side of the same pipeline, reranking fixes ordering rather than phrasing, and working out which part is broken routes you to the right one of the three.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
