Deck: An approval is a snapshot, not a lifetime license. When a prompt, connector, source, reviewer, permission, or autonomy level changes, make a small decision before the workflow touches live work again.

Suggested CTA: Download the Cortex AI Workflow Change Trigger Review Card before your next prompt, connector, data-source, reviewer, or permission change reaches live work.

Internal links: Pair this with the [AI Workflow QA & Release Rubric](../products/freebies/cortex-ai-workflow-qa-release-rubric-2026-08-19.md) for first release and the [AI Workflow Rollback Drill](../products/freebies/cortex-ai-workflow-rollback-drill-2026-09-05.md) for failure recovery.

The approval you gave last month may already be fiction

An AI workflow can be safe on Monday and unsafe on Friday without anyone changing the model.

The prompt gets edited. A connector starts returning a different field. The source document changes. A new reviewer takes over. A read-only tool gains write access. The team quietly widens the task from “draft a reply” to “send the reply.”

The workflow still has the same name. The dashboard still says approved. The original test result is still sitting in a folder.

That is how stale approval becomes operational risk.

Most teams have a launch checklist. Fewer have a change gate: a small, repeatable decision between “someone modified the workflow” and “the modified workflow touches live work.” That missing step is where AI governance quietly fails.

What the current evaluation guidance gets right

The useful lesson in current AI evaluation guidance is not a particular vendor dashboard. It is the loop.

OpenAI’s documentation describes evaluation as specifying expected behavior, testing with inputs, analyzing results, and iterating. Microsoft Foundry treats evaluation as covering both performance and safety, and supports evaluating agents against tasks, tools, user intent, and real-world traces. Anthropic’s work on long-running agents emphasizes incremental progress and clear artifacts that let the next session understand what changed.

Those ideas apply directly to business workflows. A material workflow change is not merely maintenance. It is a new evaluation event.

If the workflow’s inputs, tools, authority, or expected output changed, the old evidence proves less than the team thinks it proves.

The five-minute change gate

Before a changed AI workflow runs on live work, answer seven questions. The answers should fit on one page. The discipline matters more than the paperwork.

1. What changed?

Write the delta in plain language.

Examples: “The prompt now summarizes customer complaints into three categories.” “The CRM connector now includes account notes.” “The reviewer changed from the support lead to an untrained coordinator.” “The agent can now create a draft ticket.”

If nobody can describe the change, nobody can test it intelligently.

2. What could now be different?

Name the affected boundary, not just the edited component.

Could the change alter accuracy, privacy, permissions, tone, escalation, cost, timing, or the person accountable for the result? A small prompt edit can change which facts the system relies on. A connector update can expose a new class of sensitive record. A reviewer change can turn a real control into a rubber stamp.

3. Which test case proves the change is safe enough?

Do not rerun only the happy path. Pick the case most likely to reveal the new risk.

If the source changed, test a conflicting source. If the tool gained an action, test the stop condition. If the output became customer-facing, test an ambiguous or emotionally charged example. If the reviewer changed, test whether they can identify a wrong answer without the builder rescuing them.

The best test case is not the one that produces a beautiful demo. It is the one that can falsify the release decision.

4. What evidence must be saved?

Capture the changed version, test input, output, reviewer decision, corrections, and any exception. Keep enough context that another person can understand what was tested without reconstructing the run from memory.

This is the difference between “we checked it” and a reviewable change record.

5. Who reviews it?

Name a person with authority to reject the change. “The team” is not an owner. Neither is an automated pass score.

The reviewer should know the workflow’s boundary, the unacceptable outcomes, and the stop rule. If the change affects a regulated record, customer communication, money movement, or external action, the reviewer’s authority must match the risk.

6. What happens if it fails?

Choose the response before the test, not after a bad result creates a debate.

  • RELEASE: evidence supports returning the changed workflow to live use.
  • RETEST_REQUIRED: the change is plausible, but the evidence is incomplete or a boundary case failed.
  • FREEZE: keep the last safe version in place while the issue is repaired.
  • RETIRE: remove the workflow or changed path because the risk, value, or ownership no longer works.

The key is preserving the last safe state. A failed test should not force the team to choose between an unsafe new version and a chaotic manual scramble.

7. When does this approval expire?

Record the next trigger, not just the approval date.

Approval should reopen when the prompt, model, connector, source of truth, permission, reviewer, autonomy level, business rule, or workflow boundary changes. It should also reopen when exception patterns or correction burden show that the old evidence no longer describes live performance.

An approval with no reopening rule is a snapshot pretending to be a control system.

The operator takeaway

You do not need a committee for every typo in a prompt. You do need a proportional change gate for every change that can alter what the workflow sees, decides, sends, writes, or is allowed to do.

Start with one live workflow. List its current source, tools, permissions, reviewer, stop condition, and evidence receipt. Then define the changes that force a retest. Make the record small enough that a manager will actually use it and specific enough that a non-builder can complete it.

The goal is not to freeze useful automation. It is to prevent silent drift from being mistaken for maturity.

An AI workflow earns continued trust when it can show not only that it passed once, but also what changed, what was retested, and who decided what happened next.

Sources

  • [OpenAI: Working with evals](https://developers.openai.com/api/docs/guides/evals)
  • [Microsoft Foundry: Run evaluations for generative AI applications](https://learn.microsoft.com/en-us/azure/foundry/how-to/evaluate-generative-ai-app)
  • [Anthropic Engineering: Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents)