Skip to content

AI Industry & Models

AI Can Read a PDF Page and Still Lose the Document

Direct PDF input is now practical for many multimodal models. A test packet shows why tables, footnotes, diagrams, and cross-page references still need a parsing pipeline.

Tobias LundIndustry & Models Writer

October 5, 2026 · 7 min read

A laptop displaying a supplier PDF beside extracted table cells and a rotated scanned page.
A laptop displaying a supplier PDF beside extracted table cells and a rotated scanned page.

Current multimodal models, which accept images as well as text, have made a useful shortcut possible: upload a PDF, ask a question, and skip the extraction stack that document systems once required. For a short report with clean headings and self-contained paragraphs, that shortcut can remove an entire setup stage.

The trouble starts when “read this page” and “understand this document” stop meaning the same thing.

For this evaluation, I used a nine-page fictional supplier-review packet built around that distinction. The packet contained a table that continued onto a second page, a superscript footnote that changed the meaning of one threshold, an architecture diagram with labeled arrows, a sideways scanned appendix, and a reference from the final page back to a numbered clause near the beginning. It was designed as one ordinary workflow: deciding whether a supplier met an internal approval rule.

I ran the packet through two routes. The direct route gave the rendered PDF to a multimodal model. The structured route first used a document parser, software that extracts text while retaining elements such as headings, table cells, reading order, and page coordinates, then supplied that output alongside selected page images. Both routes received the same task: identify the approval status and show the evidence needed to reproduce it.

The direct route was faster to set up and good at describing what appeared on an individual page. The parsed route required another service, intermediate files, and more orchestration, but it preserved the links the decision depended on. That difference matters more than fluent output.

A visible table is not necessarily a usable table

The packet’s central table looked straightforward. Each supplier control occupied a row, while columns recorded the owner, evidence status, review date, and exception code. One row wrapped across several lines, and the table resumed after a page break with its column headers repeated.

Direct visual input could read individual cells when prompted with a page number and a row label. The fragile step was reconstruction. A model had to infer which wrapped text belonged to which row, recognize that the repeated header was not data, and carry the column alignment across the page boundary. A plausible answer could therefore contain the right words attached to the wrong control.

A parser represented the same table as cells with row and column positions. That made operations such as “return every control with an open exception” newly practical because the model no longer had to rebuild the grid from pixels before reasoning over it. Parsing did not guarantee correctness. Merged cells and unusual layouts still produced extraction errors, but those errors appeared in an inspectable intermediate representation rather than hiding inside the final answer.

This is the first decision point. If the workflow needs a paragraph summary of one visible table, direct input may be enough. If it filters rows, compares columns, exports values, or triggers an approval, preserve the table structure before asking the model to reason.

Footnotes fail at the link, not just the text

The packet’s threshold appeared in the main text, while a footnote narrowed its scope to suppliers handling a particular data category. The direct route could transcribe both pieces. It did not consistently treat the superscript marker as an instruction to combine them, especially when the question quoted the main sentence without naming the footnote.

That is a structural failure. The text is present, yet the answer changes because the relationship between the marker and note has been weakened. Small type can add a second problem: rendering a full page at limited resolution gives the model fewer pixels for footnotes than for body text.

The parsed route retained the footnote near its anchor and labeled both with page coordinates. Coordinates are often stored as bounding boxes, rectangular positions showing where an element appeared on the page. With that link available, the model could evaluate the threshold together with its qualification and cite both locations.

For policy, contract, research, and financial documents, footnotes should be treated as data rather than decoration. A pipeline that drops them can produce clean prose and the wrong decision.

Diagrams favor vision; rotated scans favor preprocessing

The architecture diagram reversed the result. Its meaning lived in spatial relationships: one arrow crossed a trust boundary, while another ended at a storage component. A text-only parse recovered labels and a caption but lost enough geometry to make the extracted text misleading.

Direct visual input handled the local diagram task better because the model could inspect arrow direction, proximity, and grouping. The practical setup is not to choose parsing instead of vision. It is to retain page images for elements whose meaning depends on layout, then give the model the parsed labels and the relevant crop together.

The sideways appendix exposed a different cost. A person would rotate it without thinking. The direct route sometimes recognized the orientation, but that behavior should not be the control point in a repeatable workflow. A preprocessing stage can detect rotation, correct it, and run optical character recognition, or OCR, which converts text in an image into machine-readable characters.

That adds latency before the model call and creates another place for characters to be confused. It also makes the failure observable. Teams can store the corrected image, extracted text, and confidence metadata, then route low-confidence pages for review. With direct PDF input alone, orientation handling and recognition occur inside the model request, where debugging is harder.

Cross-page references expose the real document problem

The final page of the packet referred to “the exception in Clause 2.4” without repeating its terms. Both routes had access to all nine pages, so context-window capacity was not the constraint. The issue was navigation.

A model reading page images must identify the reference, locate the clause, distinguish it from nearby numbering, and bring the relevant text back into the decision. More context does not automatically supply a document map. It can instead add competing numbers, headings, and repeated phrases.

The parsed route represented headings as a hierarchy and attached each block to a page. That allowed retrieval, the step that selects relevant material before generation, to fetch Clause 2.4 together with the final-page reference. The resulting prompt was smaller and its evidence path was explicit.

This was the supplier packet’s decisive failure mode. The approval answer depended on a table row, its footnote, and a clause cited several pages later. Direct visual input could explain each piece when asked locally, but the pipeline was more reliable at assembling the chain.

Choose according to the consequence of a wrong link

Direct input is the sensible default for occasional work on short, visually clean PDFs. It requires little engineering, preserves diagrams, and lets a person ask exploratory questions without first selecting a schema. It is also useful as a fallback when a parser mangles an unusual page.

A dedicated pipeline starts earning its cost when documents arrive repeatedly or feed another system. The cost is broader than the parser’s fee: files must be stored, page images rendered, extraction results versioned, and failures reviewed. Processing also takes longer because the model call waits for those stages.

What that extra work buys is control. A team can inspect the exact cells used in an answer, rerun extraction after changing parsers, and require citations that resolve to a page region. Those capabilities make automated comparison, audit sampling, and human approval queues practical in a way that a single opaque PDF call does not.

For the supplier workflow, I would use a hybrid pipeline. Parse every page into ordered blocks and table cells, keep the original page image, rotate scans before OCR, and send diagram crops to the multimodal model. Store page numbers and bounding boxes with every extracted element. At answer time, retrieve the relevant section rather than the nearest paragraph alone, because a section boundary often carries the footnote or definition that changes the result.

The model should return evidence before status. If it cannot point to the table row, qualifying note, and referenced clause, the workflow should withhold the approval result for human review. That rule adds friction, but it targets the exact failure the nine-page packet exposed.

Test the pipeline with relationships, not summaries

A useful acceptance test does not ask whether the model can summarize a familiar report. It asks for outputs whose correctness depends on structure: reconstruct one split table row, apply a footnote to its anchor, trace a labeled diagram edge, read a rotated page, and resolve a clause cited elsewhere.

Score each evidence link separately from the final prose. A fluent approval explanation should fail if a value came from the wrong column or if the cited clause does not support it. Run the same test after changing the model, parser, rendering resolution, or chunking rules, since any of those changes can move a boundary and break a previously correct link.

The nine-page packet is small enough to check by hand. That is the point. Before sending thousands of PDFs through either route, build one compact document that contains the structures your real decisions depend on and make every system prove it can keep those structures intact.

Questions people ask

Can a multimodal model understand a PDF without OCR?

Yes, when the PDF is rendered as page images, a multimodal model can read visible text and inspect layout without a separate OCR call. OCR remains useful for creating searchable text, correcting rotated scans before inference, attaching confidence metadata, and making extraction failures available for inspection.

When is direct PDF input good enough?

Use it for short, clean documents when the task is exploratory, the answer sits on one page, and a person will verify the result. It is particularly useful for diagrams or unusual visual layouts. It becomes risky when software will act on extracted rows, qualifications, or references without review.

Does a larger context window fix cross-page references?

No. A larger context window lets the model receive more pages, but it does not preserve heading hierarchy, footnote anchors, or table continuity. Cross-page work improves when the pipeline records document structure and retrieves the referenced section with the passage that cites it.

What should a document-parsing evaluation measure?

Measure whether the system preserves relationships, not whether its summary sounds accurate. Check table cell alignment, footnote-to-anchor links, diagram connections, scan orientation, and references between sections. Require every consequential answer to resolve to the page region and extracted element that support it.

ShareFacebook
model evaluationdeveloper toolingai at workmultimodal modelspdf parsingdocument aimodel evaluation

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next

A laptop displaying model-route logs beside a printed refund policy and a customer-support ticket.

AI Industry & Models

Model Routing Cuts AI Costs Until It Misreads One Refund

A router can send routine support work to a cheaper model. The savings disappear when a short refund request hides a policy exception, so routing and model quality need separate tests.

Tobias Lund · 7 min read

A laptop displaying a four-page PDF beside extracted text, with a chart and footnotes visible on screen.

AI Industry & Models

AI APIs Read PDFs as Text, Images or Both

OpenAI, Anthropic and Google accept PDFs, but their ingestion paths preserve different evidence. A four-page test shows when direct upload works and when preprocessing is the safer choice.

Tobias Lund · 7 min read