Skip to content

AI Governance & Ethics

Your AI Vendor’s Evaluation Set May Contain Customer Data

A support prompt retained for testing can cross from service delivery into product improvement. Approval should depend on purpose, controls, deletion coverage, and evidence.

Irene VaskoGovernance & Ethics Writer

September 5, 2026 · 8 min read

Laptop displaying evaluation-data settings beside a printed data processing agreement with marked clauses.
Laptop displaying evaluation-data settings beside a printed data processing agreement with marked clauses.

Consider one ordinary workflow: a company uses an AI service to summarize support tickets, and the vendor routes low-confidence summaries into a failed-output queue. Reviewers label the errors, engineers add selected examples to a regression set, and every later model candidate must summarize those tickets correctly before release.

That regression set, a fixed collection used to compare system versions, sounds like quality control. It may also contain names, account histories, health details, authentication fragments, or unpublished product information copied from the customer’s tickets. Once the same example tests models for other customers, supports a general release, or remains after the originating account closes, the vendor has moved beyond the narrow act of producing that customer’s summary.

The difficult word is not “evaluation.” It is “purpose.”

Service testing and model improvement are different purposes

A vendor needs some testing to deliver a reliable service. It may check whether an output is malformed, whether a safety filter fired, or whether a newly deployed model broke a customer-specific workflow. Those checks can be tightly connected to service delivery, particularly when they run automatically and discard the content after a short operational window.

The failed-output queue changes character when examples become reusable assets. A labeled ticket might help tune a prompt, select between model candidates, calibrate a safety classifier, or judge a general model release. None of those actions necessarily trains the underlying model, but they can still improve the vendor’s product.

That distinction matters because public vendor commitments often focus on training. OpenAI’s business terms, for example, state: “OpenAI will not use Customer Content to develop or improve the Services, unless you explicitly agree to this use.” Its product documentation also distinguishes business offerings, where customer content is not used for training by default, from consumer services governed by different controls. Other major vendors likewise publish separate commitments for enterprise products, consumer accounts, feedback submissions, and preview features.

A procurement review should preserve those distinctions. “We do not train on your data” does not answer whether a vendor retains prompts for evaluation, sends them to human reviewers, converts them into labeled test cases, or uses them to choose the next production model. Nor does a general statement about improving services establish whether the permission is mandatory, optional, or limited to one product surface.

The executed agreement controls the commercial relationship, while product documentation explains how the service is supposed to behave. A team needs both. This analysis is a governance framework, not legal advice, and counsel should interpret obligations that depend on jurisdiction or contract language.

Responsibility follows the decision and the control

If the vendor silently adds a ticket to a pooled evaluation set despite a contractual restriction, the vendor made the reuse decision and controls the dataset. Calling AI procurement a “shared responsibility” should not blur that fact. The vendor should be able to identify the collection rule, the permitted purpose, who accessed the example, and the mechanism that removes it.

The customer has a separate responsibility. It selected the product, configured the account, decided which records could enter the support-ticket summarizer, and accepted or rejected the relevant terms. If an administrator enabled data sharing without checking whether the setting covered API traffic, feedback buttons, or historical examples, that is a customer-side governance failure even if the vendor’s interface encouraged the mistake.

Responsibility can therefore attach to different acts. The vendor answers for collection, internal access, onward disclosure, retention, and compliance with its promises. The customer answers for data classification, product approval, account configuration, and monitoring. A subprocessor may perform labeling or storage, but that does not make the decision owner disappear; the primary vendor should disclose the subprocessor’s role and bind it to the same limits.

Silence is especially important. A visible opt-in with an account-level audit event leaves a receipt. A clause buried under “service improvement,” followed by undocumented sampling into the failed-output queue, leaves the customer unable to verify what it approved. The latter is not repaired by a broad privacy statement saying data may be used to operate the service.

Follow one ticket through the system

Before approval, ask the vendor to trace one support ticket from API request to deletion. The answer should name the records created along the way: the original prompt, generated summary, request metadata, safety scores, reviewer labels, and any copied evaluation example. Metadata, meaning information about the request rather than its main text, can still expose account identifiers, timing, feature use, and internal project names.

Next, locate the sampling rule. The failed-output queue might receive every user-reported error, outputs below a confidence threshold, or a sample chosen by an internal reviewer. Each route carries a different control. A feedback button may itself authorize sharing under separate terms, while automated sampling may run without any user action.

The vendor should also say when de-identification occurs. Removing direct identifiers after a reviewer opens the ticket is weaker than redacting them before the record leaves the production environment. Substituting names does not make a detailed incident harmless if the remaining facts can identify the person or organization.

Then inspect where the example lands. A customer-isolated regression set used only to validate that account’s configuration is easier to justify as service delivery. A pooled set available to product teams across the vendor is model improvement in practical terms, even if no gradient update ever touches model weights. Model weights are the learned numerical parameters altered during training; evaluation can guide which set of weights gets shipped without changing them itself.

This walkthrough has a cost. Isolated evaluation sets require separate storage, access policies, and release gates, while aggressive redaction can remove the context needed to reproduce a failure. The alternative is not unrestricted reuse. Vendors can use synthetic tickets, customer-authored test cases, or examples approved through a specific contribution workflow, accepting that rare production failures may take longer to diagnose.

An opt-out needs scope and a receipt

A useful control says what stops. It should cover automated sampling and manual selection, identify whether it applies across the organization or only to one workspace, and explain how long the change takes to propagate. An API account, consumer chat account, support portal, and beta feature may have separate settings even when they display the same vendor logo.

Prospective opt-out is only half the issue. If the vendor already copied tickets into the failed-output queue, changing a toggle may stop new collection while leaving old examples, reviewer labels, and derived datasets in place. Ask whether the control triggers retroactive removal, whether account closure does so, and whether previously submitted feedback follows another rule.

The receipt can be modest: an administrator event showing who changed the setting, its effective time, the products covered, and the policy version displayed. Without that event, the customer cannot prove its configuration during an audit or distinguish a vendor-side collection error from an administrator mistake.

A control that exists only through a support request deserves more scrutiny. It introduces delay, may not cover new product features, and leaves interpretation to whoever handles the ticket. For sensitive workloads, that setup may not be worth approving when a competing service offers an enforceable organization-level default.

Deletion must reach the evaluation copy

Standard deletion language often describes production content and backups without naming evaluation datasets. The procurement record should ask directly whether deletion reaches copied examples, labels, embeddings, and reviewer tools. An embedding is a numerical representation used for similarity search; deleting the original text while retaining a reversible or linkable representation may not satisfy the intended restriction.

Timing also matters. Immediate removal from an active regression set may coexist with delayed backup expiration. The vendor should distinguish those systems rather than offering one retention period that conceals several stages. If legal or security holds can override deletion, the terms should identify the category and access restrictions.

Evaluation reuse is easier to unwind than completed model training. A vendor that can identify a ticket’s dataset row should be able to exclude that row from the next evaluation run and delete its active copies. If it says removal is technically impossible, the likely problem is missing lineage, meaning the record of where data came from and where it was copied.

Return to the failed-output queue at release time. A credible record would connect the source ticket to a dataset version, document the approved purpose, show whether a reviewer saw raw content, and record deletion or exclusion. Aggregate claims about privacy do not substitute for that chain.

The approval decision can be narrow. Permit customer-isolated testing needed to operate the summarizer, prohibit pooled reuse unless an administrator opts in, and require deletion to cover active evaluation copies. If the vendor cannot separate those paths or produce a configuration receipt, keep sensitive tickets out of the service.

Questions people ask

Does evaluation count as training on customer data?

Not always. Evaluation can compare model outputs without changing model weights, but it may still support product improvement by determining which model, prompt, or safety policy gets released. A promise limited to “no training” should therefore be checked against broader language about development, improvement, quality review, and human access.

Is anonymizing an evaluation example enough?

It depends on when and how the vendor removes identifying information. Redaction before reviewer access reduces exposure, while deleting a name after copying a detailed ticket may leave the person or company recognizable. Ask what fields remain, whether records can be linked back to an account, and whether the original copy is deleted.

Should an opt-out delete examples already collected?

Do not assume it does. Many controls operate prospectively unless documentation or contract language says otherwise. The vendor should state whether the setting removes historical examples from active evaluation sets, reviewer systems, and derived copies, along with any backup delay or retention exception.

What evidence should a vendor provide after deletion?

At minimum, the customer needs confirmation that the request covered the source content and identifiable evaluation copies, plus the completion time and any stated exception. For higher-risk data, dataset lineage and an administrator audit event provide stronger evidence than a generic support message saying the account was deleted.

ShareFacebook
ai governancemodel evaluationprivacy and data rightsai governancemodel evaluationcustomer datadata retentionvendor contracts

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next

Laptop showing a declined credit application beside a policy table with the code RC-DTI-OVER-LIMIT.

AI Governance & Ethics

A Chatbot Denial Needs a Reason Code, Not More Words

A fluent explanation is useless if it cannot be traced to the rule that produced a denial. Reason codes make chatbot language reviewable before it reaches a customer.

Irene Vasko · 8 min read

A permit case file beside a laptop showing an exported AI prompt, attachment list, and redaction review log.

AI Governance & Ethics

Your Agency’s AI Prompts May Be Public Records

A permit-review prompt, its attachments, model output, and staff edits can carry different retention and disclosure duties. Agencies need a retrieval workflow before the first request arrives.

Irene Vasko · 8 min read