Before You Publish an AI Claim, Build This Evidence File
The FTC’s familiar advertising standard already covers AI claims. Here is how to connect “unbiased,” “private,” or “more accurate” to a test, a defined scope, and recorded limits.
August 9, 2026 · 8 min read

Consider a release ticket for a candidate-screening assistant. The draft landing page says its recommendations are “unbiased,” candidate information “stays private,” and its rankings are “more accurate than professional screening tools.” Product has supplied a model card. Security has supplied an encryption diagram.
Marketing wants approval.
None of those artifacts, by itself, substantiates the three sentences. The model card may describe training data without measuring hiring outcomes; encryption protects data in a particular state but does not establish where every copy goes; an internal accuracy score means little until the team identifies the comparison tool, task, test population, and correct answer.
The practical response is a claim-evidence file attached to that release ticket. It should let a reviewer move from the exact words a customer will see to the test that supports them, including where the result stops applying. This is an operational reading of public Federal Trade Commission materials, not legal advice.
The ordinary standard already applies
The FTC’s Advertising Substantiation Policy Statement says that “advertisers and ad agencies must have a reasonable basis for advertising claims before they are disseminated.” That requirement predates current generative and predictive AI systems. The agency has not offered an exemption because a claim concerns a neural network, a proprietary model, or behavior that is difficult to measure.
A reasonable basis is not one universal study design. The expected support depends on factors including the kind of claim, the product, the consequences if the claim is false, the cost of obtaining evidence, and the evidence experts in the field would consider reasonable. A specific statement such as “testing proves” also represents that the named testing exists. Health and safety claims can demand a higher level of scientific support.
Review the advertisement’s net impression, meaning the overall message a reasonable consumer receives rather than one sentence read in isolation. An “unbiased” headline beside demographic comparison graphics may communicate measured parity across groups even if a footnote says results vary. A privacy badge can imply limits on collection or reuse that the privacy policy quietly contradicts.
FTC business guidance on AI claims applies that existing framework to performance promises. The guidance is public enforcement guidance, not a separate AI substantiation rule. FTC complaints state allegations, while final consent orders legally bind the named respondents; a proposed order has not reached that status. Other employment, civil rights, privacy, and sector-specific requirements may apply on separate tracks.
Start with the sentence customers receive
Copy each claim from the candidate-screening page into the release ticket without paraphrasing it. Add the nearby image, qualifier, button text, demonstration, and spoken sales script because they can alter the net impression. Record whether the statement is objective and testable, or vague promotional language that does not communicate a measurable fact.
Then write the narrow factual proposition that the team believes it can prove. “Unbiased recommendations” might become “the tested model produced comparable error rates across the demographic groups represented in this audit, on this job family, under this decision threshold.” That is a materially smaller claim. If marketing will not publish language consistent with the proposition, the evidence and advertisement do not match.
Each row in the file needs an owner, test artifact, model and dataset identifiers, test date or release window, applicable population, production configuration, known limitation, and expiration trigger. The trigger matters because a model update, changed prompt, new ranking threshold, or different input population can break the connection without changing the visible copy.
Make “unbiased” name an outcome
A finite audit cannot prove that a system is unbiased in every relevant sense. The team must identify the outcome it tested, such as recommendation rates or false-negative rates, and the groups over which it calculated that outcome. It must also explain how reference labels were produced, since historical recruiter decisions can reproduce the bias the test is supposed to detect.
For the screening assistant, run the production model with its real preprocessing rules and decision threshold against data that represents the intended deployment. Report results by relevant group, disclose subgroup sample sizes, and examine intersections when the available data supports that analysis. Aggregate performance can conceal a failure concentrated in a smaller population.
The evidence file should retain the evaluation code, input snapshot, model identifier, threshold settings, outputs, and reviewer notes. Document missing demographic information and groups too small for a dependable conclusion. If those gaps are substantial, replace the absolute claim with a description of the completed audit or remove it. Human review does not substantiate “unbiased” unless the team also tests the combined human-system workflow.
Trace “private” through the whole data path
“Private” is not established by showing that candidate records are encrypted in transit and at rest. The relevant test follows data from the uploaded résumé through parsing, model inference, logs, analytics, support systems, backups, external processors, deletion, and any later use for training or product improvement.
Create a data-flow map from the deployed configuration, then verify it against network traces, storage inventories, retention settings, access logs, and deletion tests. Contracts with model and cloud vendors can support claims about permitted use, but the ticket should distinguish a contractual restriction from a technical control and from observed behavior. A disabled training setting does not address telemetry that still records prompts or extracted text.
Scope the published statement to what the trace demonstrates. “We encrypt uploaded résumés and delete application content from the primary service after the configured retention period” is testable, provided backups and exceptions are described accurately. “Your data stays private” may imply no human access, no third-party disclosure, or no reuse. If the product cannot support that broader impression, the fallback is narrower copy, not a larger privacy-policy footnote.
Run the comparison the claim names
“More accurate than professional screening tools” requires a head-to-head test against an identifiable baseline performing the same task. Define accuracy first. Agreement with prior recruiter choices, correct extraction of résumé fields, and prediction of later job performance are different outcomes, with different labels and consequences.
The candidate-screening ticket should name the baseline tool and relevant configuration, hold the test inputs constant, and apply the same inclusion rules to both systems. Predefine exclusions rather than removing difficult cases after seeing results. Keep uncertainty, subgroup performance, failed inputs, and the handling of ties in the record, because a favorable average can coexist with operationally important regressions.
The production setup must match the tested setup. A benchmark using a larger model, hand-cleaned documents, or a retrieval component absent from the sold product does not support the deployed claim. Neither does a vendor benchmark covering another domain. Independent testing is not automatically required for every ordinary performance statement, but evidence gains little credibility when the test designer can quietly choose the baseline and discard failures.
Turn the file into a release gate
The reviewer should be able to answer four things from each row: what consumers are likely to understand, what evidence supports that understanding, where the evidence applies, and what would force a retest. Unresolved rows block the corresponding sentence, not necessarily the product release. Marketing can remove the claim, narrow it to the tested result, or wait for better evidence.
Preserve the version sent for approval and the public version that shipped. Screenshots, change history, test code, underlying reports, and signoffs provide the receipt later; a link to a dashboard that continuously changes does not. Route material model, prompt, policy, and data-pipeline changes back through review, while minor copy edits still need checking if they alter the net impression.
Public FTC actions show why scope matching matters. In its DoNotPay complaint, the commission alleged that claims about delivering legal services comparable to a human lawyer lacked appropriate testing. In an action involving Workado’s AI Content Detector, the FTC alleged that a broadly marketed accuracy claim rested on testing that did not support performance across the advertised uses. Those are enforcement allegations tied to particular records, but the engineering lesson is general: evidence for a narrower task does not travel automatically with broader copy.
Return to the release ticket. The model card belongs under system identity and limitations, the encryption diagram belongs under part of the privacy trace, and the accuracy report belongs under its precise test population. None should be treated as a universal claim certificate. Approval comes only after the visible sentence, deployed system, and retained evidence describe the same thing.
Questions people ask
Does the
FTC require an independent laboratory for every AI claim?
No single testing format applies to every objective claim. The required support depends on what the advertisement communicates and the consequences of error, although a claim that names independent, clinical, or expert testing must have that stated support. Higher-risk claims generally demand stronger methods and controls.
Can a company rely on its AI vendor’s benchmark?
Only when the benchmark supports the claim being made about the deployed product. Check the model version, task, population, baseline, prompts, preprocessing, exclusions, and production settings. If the vendor tested a different configuration or domain, preserve the report as background rather than treating it as substantiation.
Can a disclaimer fix an unsupported headline?
A clear, prominent qualification can narrow a claim, but it cannot reliably contradict the main message. Review the full net impression across the headline, graphics, demonstration, footnotes, and sales script. If a customer must hunt for the limitation, rewrite the headline to match the evidence.
How often should an AI claim be retested?
Set triggers rather than relying only on a calendar. Retest when the model, prompt, retrieval source, preprocessing, decision threshold, target population, baseline, or material data practice changes. The release ticket should name those triggers and disable or flag affected copy until a reviewer reconnects it to current evidence.
One story a day
The story of the day, in your inbox
One real story about AI each morning — no hype, no alarm, just company for the road.



