Skip to content

AI Industry & Models

AI APIs Read PDFs as Text, Images or Both

OpenAI, Anthropic and Google accept PDFs, but their ingestion paths preserve different evidence. A four-page test shows when direct upload works and when preprocessing is the safer choice.

Tobias LundIndustry & Models Writer

September 27, 2026 · 7 min read

A laptop displaying a four-page PDF beside extracted text, with a chart and footnotes visible on screen.
A laptop displaying a four-page PDF beside extracted text, with a chart and footnotes visible on screen.

Uploading a PDF to a model API is now a normal request rather than a custom document-processing project. The catch sits behind the upload button: “PDF support” does not identify what the model receives, and that hidden choice determines whether columns stay ordered, chart labels remain attached to lines, scans produce words and footnotes keep their references.

I used one four-page vendor report to separate those paths. Page 1 had two text columns and a boxed sidebar. Page 2 contained a chart whose meaning depended on its legend. Page 3 was a scanned form without a usable text layer.

Page 4 paired short claims with small footnotes at the bottom. Each API received the same request: extract the claims in reading order, interpret the chart and cite the page supporting each answer.

The result was less a model ranking than a routing decision. Text extraction was cleanest and cheapest when the PDF already contained straightforward digital text. Page images preserved spatial relationships. Sending both gave the model the most evidence, but increased input use and introduced a new failure mode when the extracted text and visible page disagreed.

Three paths hide behind one file type

A born-digital PDF usually contains a text layer, the selectable characters stored inside the file, plus instructions for placing those characters on a page. A text extractor can pull out the words without rendering the page. That route is fast, searchable and compact, but reading order is inferred rather than guaranteed.

On page 1 of the test report, that distinction mattered. The two columns looked unambiguous to a person, yet an extraction-only path could place the first line of the right column after the first line of the left, then alternate between them. The words survived. The argument did not.

The boxed sidebar could also appear in the middle of the main paragraph because its coordinates did not establish a single logical sequence.

An image path renders each page into pixels and gives those images to a vision-capable model. The model can see that the report has two columns, that a legend sits beside a chart and that a footnote belongs at the bottom. It must still recognize small text, however, and higher-resolution rendering carries more image data into the request.

The combined path supplies extracted text and page images. It is the broadest option because searchable text helps with exact wording while the images preserve layout, but it can charge the context window twice for overlapping evidence and make debugging harder. If the PDF text layer says one thing while the scan visibly says another, the application needs a rule for which representation wins.

What the major APIs document

OpenAI documents PDF file inputs for models with vision support as a dual path: the API provides the model with extracted text and an image of each page. That makes direct upload practical for reports mixing prose, figures and tables, without requiring the developer to build a renderer first. It also means a PDF can consume more input than the visible word count suggests because both representations enter the model context.

Anthropic’s documented PDF support also covers textual and visual content. Claude can inspect charts and other page elements rather than treating the file as a stream of characters, which is the relevant change for page 2 of the vendor report. Limits on file size, page count and request method vary across Anthropic surfaces, so an application should validate against the current API documentation rather than assume that behavior in a chat product carries over unchanged.

Google’s Gemini API accepts PDFs for document understanding and applies native visual analysis to the pages. Files can be uploaded through the Files API when they will be referenced in later requests, avoiding repeated transfer from the client. The practical distinction is control: Gemini exposes PDF understanding as a model capability, but it does not give every application an inspectable intermediate text extraction with bounding boxes, the coordinates locating each word or block on the page.

Mistral offers a different route through its OCR API. OCR, or optical character recognition, converts visible writing into machine-readable text. The endpoint is useful as a preprocessing stage because it returns structured document content that an application can inspect, store and then send to a reasoning model. Dedicated services such as Google Document AI, Azure AI Document Intelligence and Amazon Textract go further when a workflow needs fields, tables or coordinates rather than a fluent answer.

These are not interchangeable contracts. OpenAI’s documented PDF route explicitly spends context on text and images. Anthropic and Gemini emphasize visual document understanding. An OCR or document-intelligence service makes the extraction artifact part of the application, which adds setup but lets a developer test that artifact before asking a model to reason over it.

What survived the four-page report

Page 2 exposed the strongest case for page images. A text extractor could recover the chart title, axis labels and legend terms, yet flattening those strings did not preserve which line matched which legend entry. A visual route retained that relationship, making a chart summary newly practical inside one model call. Tiny labels remained vulnerable, particularly after a low-resolution render, so the answer still needed a page citation and a check against the source.

Page 3 reversed the priorities. Because the scan had no usable text layer, an extraction-only request returned little evidence. The document needed either OCR or visual reading. Direct model vision handled the one-off question with less application code, while preprocessing produced a reusable text artifact that could be searched, corrected and compared across later runs.

The footnotes on page 4 survived text extraction as words but not always as relationships. Some extractors moved all notes to the end, detached the markers or inserted headers between the claim and its source. A page image kept the visual link available. The model could still overlook small type, which is why “read the PDF” is too loose a production prompt; asking for the claim, footnote marker, footnote text and page number made missing links visible in the output.

The combined route performed the broadest recovery across the packet, though it was not automatically the best implementation. It used more context than extracted text alone and required the model to reconcile duplicate content. For a single quarterly report, that overhead can be acceptable. Across an archive where thousands of pages are searched repeatedly, paying to reinterpret page images for every question is usually poor architecture.

Route the file before choosing the model

A useful PDF pipeline starts with a preflight. Count the pages, check whether text is selectable, detect encryption and rotation, then sample the extraction order. Those deterministic checks are cheaper to repeat than a model call and identify the important branch: clean digital text can go to a text model, while scans and layout-dependent pages need OCR, rendering or both.

For retrieval over contracts, manuals or research papers, extract and store the text once, retain page boundaries, and keep links back to the original pages. Retrieval-augmented generation, which inserts selected source passages into a model request, then sends only the relevant sections for each question. Render the cited pages on demand when a table, signature, handwritten mark or multi-column arrangement affects the answer.

For low-volume analysis, direct PDF upload is the shorter path. It removes the renderer, OCR service and storage layer, and the dual or visual route can answer questions about charts that plain extraction cannot. The tradeoff appears in cost visibility: image detail, page count and provider-specific token accounting all affect the bill, while the application may receive only the final answer rather than a separately testable extraction.

The four-page report therefore leads to a plain rule. Use direct upload when the task depends on page appearance and the document is read once. Preprocess when the same corpus will be queried repeatedly, when every extracted value needs an audit trail, or when the application must apply its own correction before model reasoning begins.

One setup step pays for itself in either design: save the original page number beside every extracted block. PDF tools sometimes renumber pages around covers or front matter, and a model citation that says “page 3” is useless if the application cannot map it back to the displayed file.

Questions people ask

Does

PDF support mean the model can read scanned documents?

Only if the API renders pages for a vision model or applies OCR. A scan may contain no selectable text, so a text-only extractor can return an empty or nearly empty result. Test one scanned page during preflight and route failures to OCR or page-image input.

Which

PDF path preserves columns and charts best?

Rendered page images preserve spatial relationships better than flattened text. A combined text-and-image path adds exact wording, which helps with citations, but it also consumes more context. For high-stakes extraction, keep the page image and an inspectable OCR result rather than relying on one generated answer.

Is uploading a

PDF more expensive than pasting its text?

It can be. APIs that send both extracted text and page images process more input than a text-only request, while image cost can vary with page count and rendering detail. Pasting clean text is usually leaner, but it discards the layout evidence needed for charts, sidebars and footnotes.

When should developers preprocess PDFs themselves?

Preprocess when documents will be queried many times, extraction must be reviewed, or downstream systems need stable fields and page coordinates. Direct upload is more practical for occasional analysis of mixed text and graphics, provided the application records citations and checks visually important answers against the source page.

ShareFacebook
developer toolingmodel evaluationai pricing and accesspdf ingestionmultimodal modelsdocument aiapi comparison

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next

A laptop displaying a supplier PDF beside extracted table cells and a rotated scanned page.

AI Industry & Models

AI Can Read a PDF Page and Still Lose the Document

Direct PDF input is now practical for many multimodal models. A test packet shows why tables, footnotes, diagrams, and cross-page references still need a parsing pipeline.

Tobias Lund · 7 min read

A laptop displaying model-route logs beside a printed refund policy and a customer-support ticket.

AI Industry & Models

Model Routing Cuts AI Costs Until It Misreads One Refund

A router can send routine support work to a cheaper model. The savings disappear when a short refund request hides a policy exception, so routing and model quality need separate tests.

Tobias Lund · 7 min read