Skip to content

Model Evaluation

Stories from people dealing with Model Evaluation — what changed, who decided it, and what the ordinary days look like now.

40 stories

A laptop showing a citation ledger with one claim marked partially supported beside its quoted source passage.

Agentic AI & Orchestration

A Second Pass Can Catch AI Citations That Do Not Fit

A research agent can check whether a cited passage supports its claim, but only after claims are split into testable units. The extra pass catches mismatches, not bad sources or missing evidence.

Mara Quintero · October 6, 2026 · 7 min

A laptop displaying a supplier PDF beside extracted table cells and a rotated scanned page.

AI Industry & Models

AI Can Read a PDF Page and Still Lose the Document

Direct PDF input is now practical for many multimodal models. A test packet shows why tables, footnotes, diagrams, and cross-page references still need a parsing pipeline.

Tobias Lund · 7 min read

A laptop displaying model-route logs beside a printed refund policy and a customer-support ticket.

AI Industry & Models

Model Routing Cuts AI Costs Until It Misreads One Refund

A router can send routine support work to a cheaper model. The savings disappear when a short refund request hides a policy exception, so routing and model quality need separate tests.

Tobias Lund · 7 min read

Account settings page with an address modal partly covered by a cookie banner in a desktop browser.

Agentic AI & Orchestration

Visual AI Agents Still Lose the Checkout Button

A moved control is the easy case. Modal windows, sticky banners, and responsive layouts show why visual browser agents need bounded tasks, state checks, and a selector-based fallback.

Mara Quintero · 8 min read

A laptop displaying a four-page PDF beside extracted text, with a chart and footnotes visible on screen.

AI Industry & Models

AI APIs Read PDFs as Text, Images or Both

OpenAI, Anthropic and Google accept PDFs, but their ingestion paths preserve different evidence. A four-page test shows when direct upload works and when preprocessing is the safer choice.

Tobias Lund · 7 min read

Laptop showing a vendor support article beside an agent tool-call log with a blocked upload request.

Agentic AI & Orchestration

Test Browser Agents Before a Support Page Hijacks Them

A vendor support page is untrusted input, even when an agent needs it to finish a task. This walkthrough tests whether page text can trigger data leaks, unsafe tool calls, or account changes.

Mara Quintero · 8 min read

A laptop showing a support-routing log beside a notebook listing cheap, strong, and escalate paths.

AI Industry & Models

AI Model Routing Cuts Costs Until the Router Guesses Wrong

Sending routine requests to cheaper models can lower inference costs. The savings survive only when the router detects mixed, unfamiliar, or high-risk requests before a bad answer triggers another call.

Tobias Lund · 8 min read

Laptop showing an AI release checklist beside a support ticket and a printed refund policy.

AI Governance & Ethics

Turn NIST’s GenAI Profile Into Evidence You Can Keep

A small company does not need a binder of AI policies. It needs retrievable records showing who approved the system, how it was tested, what counts as an incident, and what changed.

Irene Vasko · 8 min read

A PDF supplier table, footnote, and chart displayed beside extracted rows in a text editor.

AI Industry & Models

AI Models Read PDF Pages but Still Break the Tables

Native PDF support removes a conversion step, but it does not guarantee correct rows, footnotes, or chart labels. Use a structured parser when those relationships determine the answer.

Tobias Lund · 7 min read

Laptop showing side-by-side model migration logs for a support ticket with an attached screenshot.

AI Industry & Models

Model API Deprecations Need a Rehearsal Before Shutdown

Changing the model name proves that an API call still runs. A five-gate rehearsal shows whether refusals, truncation, images, token counts, and failures still fit the application.

Tobias Lund · 8 min read

A laptop showing paired production and candidate model traces for the same customer-support ticket.

AI Industry & Models

Shadow-Test a New AI Model Before It Reaches Customers

Mirror sampled production requests to a candidate model, hide its outputs, and compare each run. The method exposes task-specific regressions before a model swap reaches users.

Tobias Lund · 8 min read

A procurement PDF open beside extracted text and labeled page images on a desktop monitor.

AI Industry & Models

AI Can Read a PDF. Page Layout Still Decides Accuracy

Native PDF input removes a preprocessing step, but it does not remove layout errors. The right input format depends on whether the document contains tables, columns, scans or citation-sensitive text.

Tobias Lund · 8 min read

A laptop displaying an AI claim-evidence review ticket beside printed model and privacy documentation.

AI Governance & Ethics

Before You Publish an AI Claim, Build This Evidence File

The FTC’s familiar advertising standard already covers AI claims. Here is how to connect “unbiased,” “private,” or “more accurate” to a test, a defined scope, and recorded limits.

Irene Vasko · 8 min read

A credit decision flowchart beside a laptop showing a review queue and adverse-action notice fields.

AI Governance & Ethics

Assess the Credit Decision, Not Just the AI Model

A model card cannot explain why a customer was denied a credit limit increase. This decision-level template connects affected people, harms, controls, evidence, and appeals.

Irene Vasko · 8 min read

A recruiter’s laptop showing a ranked applicant list beside printed audit notes and a job requisition.

AI Governance & Ethics

NYC’s Hiring AI Law Requires a Narrow Bias Test

The city’s rule reaches tools that score or rank people and materially drive hiring or promotion decisions. Its required audit measures outcome disparities, not accuracy or general fairness.

Irene Vasko · 8 min read

A laptop showing two model configurations beside a support refund trace and rollback control.

AI Industry & Models

Your AI API Upgrade Needs a Rollback Plan

A replacement model can change tool calls, refusals, latency, and tone. Test it against customer-visible outcomes before a provider’s deprecation deadline forces the switch.

Tobias Lund · 7 min read

A laptop displaying a hiring model deployment log beside a printed bias audit report.

AI Governance & Ethics

A Vendor Model Update Can Outdate Your Hiring Bias Audit

New scoring inputs, thresholds or model weights can make an annual bias audit poor evidence for the tool now screening applicants. The deployment log should show whether the audited and operating systems still match.

Irene Vasko · 8 min read

Laptop showing a local test invoice beside a log of blocked browser-agent tool calls.

Agentic AI & Orchestration

Test Whether a Browser Agent Obeys Hostile Webpage Text

A local invoice page and synthetic secret can show whether webpage text redirects your browser agent. The useful comparison is between soft instructions and hard tool limits.

Mara Quintero · 8 min read

Laptop showing two model-output logs beside a refund workflow test sheet and a disabled tool-call panel.

AI Industry & Models

Your Replacement AI Model Needs a Dress Rehearsal

A vendor-designated successor can change refusals, token use, latency and tool calls. Shadow-test the workflow before moving production traffic, with rollback thresholds set in advance.

Tobias Lund · 7 min read

Laptop showing two model-response logs beside a support workflow configuration screen.

AI Industry & Models

Test the Replacement Before Your Model API Disappears

A retirement date tells you when an endpoint closes, not whether its replacement behaves the same. Shadow traffic exposes changes in refusals, tool calls, latency, and length before cutover.

Tobias Lund · 8 min read

Laptop showing a local vendor status page beside a log of blocked message and incident-update tool calls.

Agentic AI & Orchestration

Check Whether Your Browser Agent Obeys the Web Page

A controlled-page test can reveal whether an agent treats website text as evidence or as an instruction, before a connected mailbox, ticket queue, or account becomes the test environment.

Mara Quintero · 8 min read

Procurement screen beside a terminal showing a model file hash, prompt version, and deployment log fields.

AI Governance & Ethics

Open Model Weights Still Leave Auditors in the Dark

A downloadable model checkpoint reveals parameters, not where the training data came from or how a deployed system behaved. Use this checklist before accepting an AI transparency claim.

Irene Vasko · 8 min read

Other impacts

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.