// a copy, not a view
Embedding is a one-way transform run at write time. Text goes in; a vector plus the stored chunk text comes out; the store keeps that. There is no link back to the source file, no change feed, and no invalidation.
So deleting the document does not touch the store, and neither does editing it. The only thing that removes a chunk from an index is an explicit delete against the index.
This is definitional rather than a defect, and it is true of every vector store — pgvector, Pinecone, Qdrant, Weaviate, Elasticsearch dense vectors. There is no vendor to blame and no setting that was left off.
Retrieval has no notion of "current"
Nearest-neighbour search returns the closest points. Nothing in that operation consults a timestamp, a version, or whether the source still exists.
Which means a deleted document's chunks are exactly as retrievable the day after deletion as the day before. They are the same points in the same space. The search does not know the page is gone, because nothing told it.
Editing is the worse half
Deletion is the obvious case, and most teams handle it. Editing is the one that bites the teams who did build a sync.
If the pipeline embeds the new version and writes it without deleting the old chunks first, the store now holds both. And the old ones often score better, because the wording a user quotes back to you is usually the wording they saw before the edit.
A correct sync deletes every chunk belonging to that document id, then writes the new set. Delete by document, not by chunk.
It is not hallucinating
This is the part that sends people after the wrong fix. The retrieved chunk is handed to the model as context, and a well-behaved model grounded in its context will faithfully quote it.
So the failure surfaces as the model confidently stating something the source no longer says: a price you changed, a policy you retracted, a page you took down. It is not inventing anything — it is correctly repeating what retrieval gave it.
Which is why the usual advice about hallucination does not apply here, and why "ground it in your sources" makes it worse rather than better. Grounding is the mechanism. The source is the problem.
The rule
Delete the copy, not just the file. Every write path that can remove or change a document needs a matching delete against the index, keyed by document, before anything new is written.
And be specific about which failure you are looking at, because two very different ones look similar from the outside. Changing the embedding model makes the stored numbers unreadable — same numbers, different map — and the fix is to re-embed everything. Changing the document leaves the numbers perfectly readable; they just describe text that no longer exists, and the fix is per-document delete-then-write. Related: working out which part of your RAG is broken routes on whether the right thing was retrieved at all — this is the case where retrieval worked perfectly and the text itself was wrong.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
