Skip to content

AI Governance & Ethics

Turn NIST’s GenAI Profile Into Evidence You Can Keep

A small company does not need a binder of AI policies. It needs retrievable records showing who approved the system, how it was tested, what counts as an incident, and what changed.

Irene VaskoGovernance & Ethics Writer

September 15, 2026 · 8 min read

Laptop showing an AI release checklist beside a support ticket and a printed refund policy.
Laptop showing an AI release checklist beside a support ticket and a printed refund policy.

Take one concrete system: a support assistant that reads an incoming customer ticket, retrieves passages from the company help center, drafts a response, and waits for a support representative to approve it. The assistant cannot issue a refund or send a message by itself. Its most important known failure is an unsupported refund claim, where the draft promises eligibility that the retrieved policy does not establish.

That one workflow is enough to build a right-sized evidence plan. The target is not a policy saying the company uses AI responsibly. The target is a release record that shows which version was tested, an ownership table naming the person who can stop it, and an incident rule that turns an unsupported refund promise into a reviewable event.

Start with what NIST is, and is not, asking for

NIST AI 600-1, the Generative Artificial Intelligence Profile, applies the NIST AI Risk Management Framework, or AI RMF, to risks associated with generative systems. The framework organizes work under four functions: Govern, Map, Measure, and Manage. Its profile supplies suggested actions for organizations selecting controls and evidence.

This distinction matters. The profile is voluntary guidance, not a regulation, certification scheme, or ready-made audit standard. A suggested action becomes enforceable only when another instrument gives it force, such as a contract, an internal control adopted by management, or an applicable law. A company should therefore avoid labeling every NIST suggestion a “requirement.”

NIST does, however, describe outcomes in terms that can be tested. AI RMF GOVERN 2.1 says: “Roles and responsibilities and lines of communication related to mapping, measuring, and managing AI risks are documented and are clear to individuals and teams throughout the organization.” A policy that says “the AI team owns risk” does not demonstrate that outcome.

A table naming the product owner, test owner, incident lead, privacy reviewer, and shutdown authority does.

The same translation works throughout the profile: identify the outcome, decide what record would prove it happened for this workflow, and assign someone to retain that record. Do not begin by copying every suggested action into a spreadsheet. That produces a large control catalog before the company has established which risks apply.

Write a one-page system record

Create a system record for the support assistant before drafting a broad AI policy. It should identify the business purpose, users, model provider, retrieval source, data categories, permitted actions, prohibited actions, human approval point, and fallback.

For this assistant, the permitted output is a draft support reply. The prohibited actions include sending the reply, changing an account, or approving a refund. The human approval point is the support tool’s Send button. If retrieval fails or the model cannot cite a relevant policy passage, the fallback is an empty draft with the ticket left in the normal support queue.

Record the boundaries as implemented, not merely intended. If the model’s API credentials cannot call the billing system, preserve the relevant permission configuration or deployment manifest. If the interface requires a representative to press Send, keep a screenshot and the application setting that enforces the step. A statement that people are “expected to review outputs” is weaker evidence because it does not show whether the product allows bypassing review.

The system record maps the workflow before anyone evaluates it. It also prevents scope drift. When a later release adds account lookup or automatic language translation, the owner can see that the data flow and failure modes changed rather than treating the update as a cosmetic prompt edit.

Keep six records, linked by one release ID

Give every production configuration a release ID, then attach a small evidence bundle. The identifier should cover the model name or provider identifier, system prompt, retrieval configuration, help-center corpus version, safety rules, and application code release. If a provider does not expose a fixed model version, record the alias, test date, provider, and any change notice available; the limitation belongs in the record.

| Record | What it must show for the support assistant | |---|---| | System record | Purpose, data flow, action limits, approval point, fallback, provider, and deployment owner | | Ownership table | Who approves release, runs tests, handles privacy review, investigates incidents, and can disable the feature | | Risk and test record | Applicable failure modes, test inputs, expected behavior, observed output, result, reviewer, and release ID | | Change approval | What changed, why existing evidence still applies or which tests were rerun, and who accepted the residual risk | | Monitoring record | Production signals reviewed, sampling method, user reports, access controls, retention period, and review date | | Incident record | Trigger, affected release, containment, evidence preserved, impact assessment, corrective action, and closure owner |

These records can live in an issue tracker, repository, or controlled document store. The tool matters less than linkage and retrieval. An auditor or customer assessor should be able to start with the active support-assistant release and locate its tests, approval, owner, and incident history without reconstructing the system from chat messages.

Test the failure you have named

The profile discusses risks including confabulation, where a model generates confidently stated but unsupported information, along with privacy, information security, intellectual property, human-AI configuration, and risks introduced by third-party components. The support assistant does not need an equal test program for every category. It needs documented reasoning about which risks apply and tests tied to the consequential failures.

Build a test set around the unsupported refund claim. Include tickets where the policy clearly allows a refund, clearly denies it, lacks enough information, or contains language that could lure the model away from its instructions. For each case, retain the input or a privacy-safe fixture, the retrieved passages, the final prompt, the raw output, the expected behavior, the scoring rule, and the reviewer’s decision.

A useful scoring rule is observable. The draft must not promise a refund unless the retrieved policy supports the promise; otherwise it must request missing information or route the case to a person. “Answer should be good” cannot be reproduced, and a single average quality score can hide the exact high-impact error the company meant to prevent.

Keep failed results. Deleting them turns evaluation into marketing rather than evidence. If the team changes the prompt after a failed test, save the original result, link the corrective change, and rerun the affected cases under a new release ID. Manual review costs staff time and delays release, but for a small test set it often produces more defensible evidence than buying an evaluation platform before the rubric is stable.

Make change approval narrower than change management theater

NIST’s Manage function expects organizations to address post-deployment monitoring, incident response, recovery, and change management. For a small team, that does not require a standing committee. It requires a rule identifying which changes invalidate earlier evidence.

Treat a model switch, system-prompt edit, retrieval-index change, new data source, permission expansion, or removal of the human approval step as review-triggering. Typographical interface changes need not receive the same treatment. The release owner records the change, identifies affected risks, reruns the relevant tests, and obtains approval from someone other than the person who made the change when the failure could create a customer commitment.

Return to the refund case. Updating an unrelated help-center article may need only a retrieval check. Rewriting the refund policy requires rerunning the refund test set because the source of truth changed. Allowing the assistant to send replies automatically changes the human-AI configuration and the containment strategy, so the old approval cannot cover it.

Third-party model updates are harder. A provider may change behavior behind an alias without giving the customer a fixed weight set, which means the company cannot prove that every production request used an identical model. Document that constraint, monitor provider notices, and schedule regression tests based on the workflow’s exposure. Evidence should describe the control that exists, not imply version certainty the supplier does not offer.

Define an incident before the first one

An incident criterion converts a vague concern into an operational trigger. For this workflow, define an incident as a sent customer response containing an unsupported refund promise, disclosure of another customer’s information, or instructions that bypass an account-security control. A blocked draft can remain a test or monitoring event unless repetition indicates a control failure.

The incident record should preserve the release ID, relevant prompt and retrieved material, user action, timestamps at the precision the system already records, containment decision, and corrective change. Limit access because tickets and prompts may contain personal or confidential data. Where possible, store references, hashes, redacted excerpts, and structured outcomes rather than copying every raw conversation into a governance folder.

Monitoring now has a purpose. The team can review user-reported bad drafts, samples selected under a documented method, retrieval failures, and blocked outputs tied to the named incident criteria. Logging every prompt indefinitely may increase privacy and security exposure without improving oversight. The evidence plan should state what is retained, for how long, who can access it, and what cannot be reconstructed.

A useful production entry for the anchor failure could record `outcome=blocked`, `reason=refund_claim_unsupported`, and the release ID, while the restricted test store holds the redacted material needed to investigate. That line shows the control fired. It does not prove the detector catches every unsupported claim, which is why the evaluation record remains necessary.

Questions people ask

Does following this evidence plan make a company NIST compliant?

No. The Generative AI Profile is voluntary guidance, and NIST does not turn this bundle into a certification. The records can show how a company addressed selected AI RMF outcomes, but contractual duties, internal policies, and applicable laws may require different controls or evidence.

Do we need to save every prompt and model response?

Usually not. Full prompt retention can create privacy, security, and intellectual-property risk. Keep reproducible test fixtures and enough restricted incident evidence to investigate consequential failures, then use structured production logs, sampling, redaction, access controls, and a documented retention period for the rest.

What is the minimum evidence to keep before launch?

Keep the system record, named owners, a risk-linked test record, and a signed or ticketed release approval. The support assistant should also have written incident criteria and a working shutdown path before customers use it, even if the monitoring record contains little more than the initial review schedule.

When should we rerun the tests?

Rerun affected tests when the model, prompt, retrieval corpus, data source, permissions, or human approval point changes. Also rerun them after a relevant incident or credible supplier change notice. Link the new results to a new release ID rather than overwriting the evidence for the prior configuration.

ShareFacebook
ai governanceai observabilitymodel evaluationai governancenist ai rmfgenerative aimodel evaluationaudit evidence

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next

Laptop showing a declined credit application beside a policy table with the code RC-DTI-OVER-LIMIT.

AI Governance & Ethics

A Chatbot Denial Needs a Reason Code, Not More Words

A fluent explanation is useless if it cannot be traced to the rule that produced a denial. Reason codes make chatbot language reviewable before it reaches a customer.

Irene Vasko · 8 min read

A permit case file beside a laptop showing an exported AI prompt, attachment list, and redaction review log.

AI Governance & Ethics

Your Agency’s AI Prompts May Be Public Records

A permit-review prompt, its attachments, model output, and staff edits can carry different retention and disclosure duties. Agencies need a retrieval workflow before the first request arrives.

Irene Vasko · 8 min read