The NIST AI Risk Management Framework is voluntary guidance for considering risks in the design, use and evaluation of AI systems. The practical exercise below is an AMS editorial application for a customer-service team, not a NIST certification or a substitute for reviewing the business’s obligations.
Give the test a clear reference
Choose the current, approved return policy and record its version date. Include the parts staff need: the applicable product or service, the return window, the evidence normally requested and the route for exceptions. If the policy is ambiguous, resolve that ambiguity before asking a model to interpret it for customers.
Use fictional examples rather than pasting customer messages that contain payment information, addresses or other private details. Include an ordinary eligible return, an item outside the stated window, a purchase without the usual receipt and a case the policy does not answer. These examples should test different decisions rather than repeat the same question in different words.
Check the promise, not just the tone
For each draft, mark whether the answer stays within the supplied policy, asks for only appropriate information and sends an unresolved case to a human. A response that invents a refund amount, deadline or exception fails even if its wording is friendly. Also flag a reply that presents an uncertain interpretation as a final decision.
Keep one short correction note per failure. For example: “The policy does not establish eligibility for this situation; the reply should offer staff review.” Record the expected behavior before changing the prompt so the next test has a stable comparison. Do not mark an entire workflow reliable because one revised answer looks better.
Keep authority with the person handling the case
During an initial trial, staff should review drafts before sending them. Drafting a message and authorizing a payment are separate permissions. The test does not justify giving the tool access to refund transactions or letting it decide legal rights.
Repeat the examples whenever the policy, model or connected knowledge source changes. Retain the policy version, test date and decision to continue or stop the trial. For the related data-sharing boundary, see AMS coverage of what staff may share with AI tools. A useful result is a reply the reviewer can trace to an approved rule, with a clear handoff when no rule settles the question.