Skip to content

Agentic AI & Orchestration

A Second Pass Can Catch AI Citations That Do Not Fit

A research agent can check whether a cited passage supports its claim, but only after claims are split into testable units. The extra pass catches mismatches, not bad sources or missing evidence.

Mara QuinteroAgents & Orchestration Writer

October 6, 2026 · 7 min read

A laptop showing a citation ledger with one claim marked partially supported beside its quoted source passage.
A laptop showing a citation ledger with one claim marked partially supported beside its quoted source passage.

The useful artifact from this evaluation was not the research report. It was a citation ledger: one row per factual claim, with the source URL, quoted passage, support label, and a short explanation of the verdict.

One row kept the test honest. The draft claimed that a policy covered employees and contractors, while its cited passage referred only to employees. The link worked. The page concerned the right topic.

A reader scanning the citation might accept it. Yet the passage did not establish half of the sentence.

That distinction is where ordinary citation checks tend to stop working. Link validation can confirm that a page exists, and topical relevance can show that the page discusses the same subject, but neither test establishes entailment, meaning the source passage must make the claim true without requiring an unsupported leap.

I tested a two-pass workflow built around that narrower standard. The first pass retrieved material and drafted an answer. The second received each claim beside the passage offered as evidence, then decided whether the passage supported the wording. This was a controlled evaluation with seeded mismatches rather than a benchmark, so the result shows what the setup can expose, not an accuracy rate to expect from every model or subject.

The first pass has to leave evidence behind

A research agent cannot verify a citation from a bare URL. Pages change, long documents contain conflicting sections, and a relevant title says little about the sentence being checked. The retrieval pass therefore had to save the exact text it relied on, along with enough neighboring text to preserve qualifiers.

For each claim, the ledger recorded a claim identifier, the claim text, the URL, the page title, the quoted passage, and whether retrieval succeeded. That structure required more work than asking the agent to write an answer with footnotes, but it exposed a common failure early: the model often attached one citation to a sentence containing more than one factual assertion.

The employees-and-contractors row began that way. The sentence also described the timing of a review, producing several propositions under one citation. Before verification, a claim splitter rewrote the sentence into atomic claims, each narrow enough to be supported or rejected on its own. The verifier could then accept the employee clause while rejecting the contractor clause, rather than assigning a vague score to the whole sentence.

This splitting step matters more than elaborate prompting. If the unit under review contains several facts, a single supported fragment can cause the model to approve everything around it. Shorter claims reduce that spillover, although they also increase the number of checks and can strip away context if the splitter becomes too aggressive.

The verifier did not receive the first agent’s hidden reasoning or its explanation for choosing the source. It saw the claim, the cited passage, limited surrounding text, and basic source metadata. Removing the draft agent’s rationale made it harder for the second pass to repeat a persuasive story about why the citation ought to fit.

The second pass caught relevance posing as support

The verifier used four outcomes. Supported meant the passage established the full claim. Partial meant it established only part or used narrower wording. Unsupported meant the passage did not provide the stated fact.

Unverifiable meant the agent could not inspect enough source material to decide, as can happen with inaccessible pages, truncated documents, or citations pointing to search snippets.

On the ledger’s anchor row, the result was partial support. The policy passage explicitly named employees, while the draft had expanded that group to include contractors. Nothing in the quoted text licensed the expansion. The second pass proposed narrower wording rather than searching for a justification after the fact: remove contractors, keep employees, and preserve the citation.

That behavior is materially better than a generic instruction to “double-check citations.” A broad review prompt tends to reward topical overlap. The policy page looked relevant, contained related terminology, and came from the expected source. A claim-passage comparison forced the agent to address the missing noun.

The workflow also handled numerical and causal wording more strictly. A passage describing two events does not necessarily support a claim that one caused the other, while a source saying that a value increased does not establish a draft’s precise percentage. The verifier’s instruction treated added precision, broadened scope, and stronger causal language as unsupported unless the passage supplied them.

Still, the check remained a model judgment. It was good at clear omissions in the controlled packet, especially when the draft contained a name, category, date range, or quantity absent from the citation. It was less dependable when support depended on definitions elsewhere in a document or on combining several passages. Those cases need either a larger evidence window or a human reviewer.

A correct match can still be bad research

Citation support and source quality are separate tests. A low-quality page may state a claim directly, allowing the verifier to label the citation supported even though the research should not rely on it. Conversely, a primary source may be authoritative but fail to support the particular sentence attached to it.

The ledger therefore needs another field for source type or authority. That judgment can be partly automated through domain rules and document metadata, but it should not be folded into the support label. Combining the two makes failures hard to diagnose: teams cannot tell whether the agent misread a passage or selected the wrong source in the first place.

The second pass also cannot inspect claims the drafting agent failed to register. If a paragraph contains uncited factual statements, checking only the visible footnotes creates a clean ledger with missing rows. A coverage pass must first identify externally verifiable claims in the draft and confirm that each has evidence. Opinions, transitions, and conclusions drawn transparently from prior evidence do not need citations, but the agent should not decide that merely by noticing which sentences already have links.

Retrieval errors create another blind spot. If the first pass extracts a sentence without its exception, the verifier may approve a claim against an incomplete passage. Tables, footnotes, PDF columns, and dynamically rendered pages are particularly awkward because text extraction can reorder or omit the material that changes the meaning. The fallback is unglamorous: mark the row unverifiable and retain the source for human inspection.

Using the same model for drafting and checking adds correlated error, where both passes make the same mistake for similar reasons. A separate prompt and fresh context help, but they do not create independence. Higher-stakes workflows can use a different model for verification, deterministic checks for dates and quantities, or human approval for rows marked partial or unverifiable.

The cost follows the number of claims

A two-pass agent spends more time and tokens than a research prompt that returns prose once. Retrieval must preserve passages, the splitter creates more units, and the verifier reads evidence for every factual claim. Long reports can turn a single drafting call into many checks, even when claims are batched into one request.

Batching lowers overhead but introduces a tradeoff. A verifier reading many claim-passage pairs at once may blur evidence between rows, particularly when the subject matter and wording are similar. One claim per call keeps the boundary clean and costs more. Small batches, with explicit identifiers and a machine-readable response, are a practical middle ground.

The added expense is hard to justify for disposable brainstorming notes. It makes more sense when research will reach customers, guide an operational decision, or be reused by another system that treats the output as evidence. Teams can also limit the second pass to claims involving quantities, legal or policy scope, named entities, and causal statements, though sampling leaves unchecked prose behind.

The workflow needs an enforcement action, not just a score. In this test, supported rows could remain, partial rows had to be narrowed, unsupported rows were removed unless new evidence was retrieved, and unverifiable rows went to review. Without that gate, the ledger becomes observability data that records a failure after the unsupported sentence has already shipped.

The employees-and-contractors row shows the adoption test. If the system merely reports that the link is live and the page is relevant, it has not verified the citation. If it can point to the unsupported word, rewrite the sentence to match the passage, and preserve the rejected version in a log, the second pass is doing useful work.

Questions people ask

Can a research agent verify all of its own citations?

No. It can compare a claim with retrieved evidence and reject clear mismatches, but it may miss omitted context, extraction errors, or reasoning that depends on several documents. Claims marked partial or unverifiable still need another retrieval attempt or human review.

Should the verifier use a different AI model?

A different model can reduce the chance that drafting and verification repeat the same error, though it does not guarantee independence. Fresh context, narrow claim-passage pairs, and deterministic checks for quantities or dates may matter as much as changing models.

Does a supported citation mean the source is trustworthy?

No. Support means the cited passage states enough to establish the claim. Source authority, recency, conflicts of interest, and whether a primary source is available require a separate quality check recorded independently from the support verdict.

What should happen when a citation fails the check?

The agent should narrow the claim, retrieve better evidence, remove the sentence, or send the row for review. A warning alone is insufficient if the workflow still publishes the original wording, as the unsupported claim remains attached to a plausible-looking link.

ShareFacebook
ai agentsmodel evaluationai searchai agentscitation verificationresearch agentsagent evaluation

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next

Account settings page with an address modal partly covered by a cookie banner in a desktop browser.

Agentic AI & Orchestration

Visual AI Agents Still Lose the Checkout Button

A moved control is the easy case. Modal windows, sticky banners, and responsive layouts show why visual browser agents need bounded tasks, state checks, and a selector-based fallback.

Mara Quintero · 8 min read

Laptop showing a vendor support article beside an agent tool-call log with a blocked upload request.

Agentic AI & Orchestration

Test Browser Agents Before a Support Page Hijacks Them

A vendor support page is untrusted input, even when an agent needs it to finish a task. This walkthrough tests whether page text can trigger data leaks, unsafe tool calls, or account changes.

Mara Quintero · 8 min read