AI Models Read PDF Pages but Still Break the Tables
Native PDF support removes a conversion step, but it does not guarantee correct rows, footnotes, or chart labels. Use a structured parser when those relationships determine the answer.
September 5, 2026 · 7 min read

Native PDF input has made document questioning newly practical. OpenAI, Anthropic, and Google document ways for supported multimodal models to process PDFs without requiring users to convert every page manually, and some implementations combine extracted text with page images so the model receives both searchable words and visual layout.
That removes setup work. It does not turn the model into a deterministic document parser.
The distinction matters on page 7 of a supplier review, the concrete test case for this analysis. The page contains a four-column table of late orders, a footnote that changes one supplier’s second-quarter figure, and a bar chart showing total order volume. A person can connect those elements by position. A model has to reconstruct those relationships from tokens, pixels, or both.
The page 7 test
A useful evaluation starts with a controlled fixture rather than a confidential annual report whose correct answers are debatable. On the synthetic page 7, the table columns are Supplier, Q1 late orders, Q2 late orders, and Change. Northline reads 18, 9, and -9; Ibis reads 11, 8, and -3; Cedar reads 7, 10, and +3.
A superscript beside Northline’s Q2 value points to a footnote: the reported figure excludes six expedited orders. The adjacent chart uses the same supplier names but measures total order volume, not late deliveries. This is an intentionally awkward page, yet none of its parts is exotic. Procurement packs, earnings presentations, insurance schedules, and research papers routinely combine these structures.
The control question is: “Which supplier’s reported late orders fell most from Q1 to Q2, and how does the footnote change that comparison?” The expected answer has two stages. Northline has the largest reported decline at nine orders. Restoring the six excluded orders changes its Q2 figure from 9 to 15, reducing the comparable decline to three, the same decline shown for Ibis.
That answer requires more than optical character recognition, or OCR, which converts visible characters into machine-readable text. The system must keep each number attached to its row and column, follow the superscript to the correct note, apply the note to one cell, and avoid substituting values from the neighboring chart.
Run the same prompt through three input paths: the original PDF, a page image exported at a fixed resolution such as 200 dots per inch, and text extracted with a layout-preserving command such as `pdftotext -layout`. Keep the prompt and expected answer unchanged. Any difference then comes from the document representation, although vendor-side model updates can still change results between runs.
Native
PDF input sees more but hides the handoff
Native PDF input is the lowest-friction path. Vendor documentation describes systems that can use page text, rendered page imagery, or a combination, depending on the product and model. That makes questions spanning prose and figures practical without building a separate OCR pipeline first.
Page 7 exposes the tradeoff. The model may receive every visible fact and still reconstruct the table incorrectly, because seeing “18,” “9,” and “-9” does not itself encode that all three belong to Northline. PDF files often store text as positioned fragments rather than semantic rows, while a rendered page presents the same table as pixels whose relationships must be inferred from spacing and lines.
The failure can look persuasive. A response may name Northline and calculate the reported decline correctly, then overlook the footnote or apply its six orders to the Q1 value. Another answer may pull total order volume from the chart because the chart labels are visually prominent and repeat the supplier names. Fluent prose does not expose which internal association failed.
Vendor-native handling also limits observability. Unless the product returns extracted text, page references, or bounding boxes, a developer cannot easily tell whether the answer came from the PDF text layer or visual analysis. A bounding box is a set of coordinates locating an element on a page. Without that trace, debugging often means changing the input and running the question again.
Native PDF input remains the sensible first pass for summaries, document classification, and questions whose answers are stated in ordinary paragraphs. On page 7, it is a candidate generator rather than the final authority.
A page image preserves layout, not structure
Exporting page 7 as an image removes ambiguity about the PDF’s internal text order. The table, superscript, note, and chart appear exactly where a reader sees them. That can help when the source PDF contains a broken text layer, scanned pages, unusual fonts, or characters split into badly ordered fragments.
The image path introduces a different cost. Every character must be recognized visually, small footnotes receive fewer useful pixels than headings, and the model still has to infer that the second numeric column means Q2 late orders. Increasing resolution can make small type more legible, but larger images consume more input capacity and take more time to upload and process. Vendor billing methods differ, so teams should inspect the current image-token or document-pricing rules rather than assume one PDF page has a fixed cost.
Cropping is often more valuable than another prompt. Send the table and its footnote together while excluding the unrelated chart, then ask for the extracted rows before requesting analysis. This makes the model’s intermediate representation visible. If it returns Northline as `18 | 9 | -9` and links the six-order adjustment to Q2, the later comparison has a sounder base.
The crop must retain context. Cutting off the column headers, superscript, or note marker produces a cleaner image that cannot support the right answer. Page 7 therefore needs one crop covering the table and footnote, not separate crops that force the model to guess their connection.
Extracted text is cheaper to inspect and easier to break
Plain extracted text is useful because it is searchable, loggable, and easy to compare across model providers. It also avoids repeatedly sending page imagery when the task concerns prose. For large document collections, that reduction in visual processing can lower cost and latency, although the amount depends on the provider’s current token accounting.
Tables reveal the weakness immediately. A basic extractor may emit headers first, then cell values in an order determined by coordinates or internal PDF objects. Multi-column pages can interleave the footnote with the chart caption. Even `pdftotext -layout`, which uses spacing to approximate the original arrangement, represents columns with whitespace rather than explicit field names.
Page 7 can look aligned in a monospaced terminal yet lose alignment when an application collapses repeated spaces before sending the prompt. It can also preserve the superscript marker while moving the footnote several paragraphs away. The language model then receives all the necessary words but no dependable link between them.
Extracted text works when a human inspection confirms that reading order and table alignment survived. It is a poor default for automated decisions involving dense financial schedules, repeated column groups, merged cells, or notes attached to individual values.
Parse the table before asking the business question
The dependable fallback is a document parser that emits structure, such as rows in CSV, cells with coordinates, or JSON containing page numbers and header relationships. The model can then answer from explicit fields instead of reconstructing a grid during every prompt.
For page 7, the parser’s output should identify Northline as one row, label 18 as Q1 and 9 as Q2, retain the footnote marker on the Q2 cell, and store the six-order exclusion as linked note text. The application can calculate both declines in ordinary code, while the model explains the result in readable language. That division makes the arithmetic reproducible and leaves the model to handle interpretation.
Parsing adds an ingestion stage, another component to monitor, and potentially a per-page charge from a document-processing service. Complex tables still need validation. Yet the setup becomes worthwhile when the same document type arrives repeatedly, an incorrect cell would alter a payment or operational decision, or users need citations that point to a precise page region.
A practical routing rule follows from page 7. Start with native PDF input for low-stakes exploration. Use a page crop when the PDF text layer is defective or one visual region matters. Use extracted text for prose-heavy retrieval after checking reading order.
Require structured parsing when the answer depends on row alignment, merged headers, cell-level footnotes, or exact chart data.
Do not spend extra model tokens asking the same malformed input to explain itself more forcefully. Fix the representation.
Questions people ask
Can a multimodal model extract a table directly from a PDF?
Yes, and native PDF support makes one-off extraction practical without a custom pipeline. The output still needs validation when cells span rows, headers are nested, or footnotes modify individual figures, because the model can recognize every character while assigning one value to the wrong field.
Is sending a PDF page as an image more accurate?
It can be better when the PDF text layer is missing or scrambled because the image preserves visible layout. It is not inherently structured, however. Small notes can be misread, chart labels can distract the model, and higher-resolution images generally require more processing than clean extracted text.
When should I use a dedicated document parser?
Use one when table relationships affect money, compliance records, operational decisions, or repeatable calculations. A parser can return rows, coordinates, and linked notes that code can validate before a model writes the explanation, which is safer than asking the model to infer the grid on every run.
How can
I test PDF support before choosing a model?
Build a controlled page like page 7 with known rows, one cell-level footnote, and a nearby chart using similar labels. Submit the original PDF, a page image, and extracted text with the same prompt, then compare the returned structure and citations against the known answer rather than grading fluency.
One story a day
The story of the day, in your inbox
One real story about AI each morning — no hype, no alarm, just company for the road.



