// a PDF is not a document
A PDF does not contain paragraphs. It contains glyphs with coordinates — instructions for putting marks on a page. Turning that back into prose is a layout problem: what is a column, what is a header, what order do humans read this in.
A naive extractor does not solve that problem. It walks the page in roughly the order the glyphs were written, which on a single-column page is fine and on a two-column page is a disaster.
Every word is real. Every sentence is fiction.
Nothing downstream can tell
This is why it survives to production. The output is still text — plausible, well-formed, UTF-8 text — so every stage after it succeeds:
There is no error to notice and no metric that dips. The first sign of trouble is a confident answer citing a sentence that is not in the document — at which point everyone goes and looks at the chunker.
Tables lose the thing that made them tables
The second failure is quieter and worse. Dense tables, multi-column spans and merged cells break naive extraction — but the numbers usually survive. What does not survive is their relationship to the row and column headers.
A lost number is obvious. A number that has quietly detached from its label is worse than lost, because it reads as a fact and retrieves as one.
The check costs a minute
There is no statistic on this page because the verdict does not need one. Take one representative document — ideally the one behind a wrong answer you have already seen — run it through your extractor, print the text, and look at it.
Nobody does this, which is exactly why it is worth doing. You will know within a minute whether you have a retrieval problem or a reading-order problem, and those have nothing in common.
If it is broken
Layout-aware parsers exist and they work differently: a vision model finds the structure of the page first — columns, headers, tables — and extraction follows that structure rather than glyph order. Docling, MinerU, Marker, LlamaParse and others are in this category, and olmOCR-Bench (roughly 1,400 documents, 7,000+ assertions) is the benchmark most of them now report against.
Related: this sits before why your RAG retrieves garbage, which is about how to split text that is already correct. And when an answer is wrong, find out where routes you between retrieval and the model — this page is the case where the document is present and indexed, and wrong anyway.
Sources: PDF parser surveys and benchmarks, 2026 · olmOCR-Bench (Ai2). Verified 2026-09-07. Deliberately omitted: a widely-quoted "40% of complex multi-column layouts" figure, which traces to a vendor blog with no methodology.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
