The most dangerous sentence in an AI workflow meeting is often: “It worked last week.”

That may be true. It may also be irrelevant.

A workflow that passed its original pilot can become unsafe or uneconomic when it starts receiving a new document type, a higher-stakes decision, a different customer context, a new source, a tighter deadline, or an exception pattern nobody tested. The system did not necessarily fail. The job changed around it.

That distinction matters because teams tend to respond to a case-mix problem with the wrong fix. They add a prompt, grant more autonomy, buy another connector, or blame the model. The better first move is to stop and ask a simpler question: Is this still the same job we approved?

If the answer is no, the workflow has not earned the new work just because the old work went well.

A passed pilot is not a permanent boundary

OpenAI’s evaluation guidance treats reliability as an iterative process: specify the task, run test inputs, analyze the results, and improve the application. That process is useful, but it only answers the question represented by the test inputs. A clean score on the old test set is not evidence for a new case class.

Anthropic’s guidance on effective agents makes a related point: use the simplest architecture that fits the task, and add complexity only when the work genuinely requires it. A new case class is therefore not an automatic reason to add an agent. It is a reason to re-check the boundary.

Microsoft’s current work-trend framing points to a world where people and AI systems share more work. That makes the boundary question more important, not less: someone must own the moment the system meets a case it was not designed or tested to handle.

The practical gap is not another launch checklist. It is a small record for the moment the live queue stops looking like the pilot queue.

What counts as a case-mix change?

Case mix is the shape of the work arriving at the workflow. It changes when any of these move:

  • Inputs: a new file type, language, data format, or missing-field pattern appears.
  • Context: the workflow now serves a different customer, department, jurisdiction, or business process.
  • Stakes: the output affects more money, a regulated decision, a customer promise, or an irreversible action.
  • Sources: a new policy, connector, data source, permission, or retrieval path enters the process.
  • Capacity: volume rises, deadlines tighten, or reviewer time falls.
  • Exceptions: a recurring failure mode appears that was absent from the original test set.
  • Authority: a draft-only workflow is being asked to send, approve, change, or write back.

None of these changes has to look dramatic. That is why they are easy to normalize. A support-triage workflow that handled internal requests may look unchanged when it starts handling customer complaints. A reporting workflow that summarized public data may look unchanged when it starts using an internal source with access restrictions. The interface stays familiar while the risk boundary moves.

The five-minute case-mix record

Before treating the new work as business as usual, record five things.

1. The original approved case mix

Write down the job the workflow was actually approved to handle: its inputs, source types, decision stakes, reviewer, excluded cases, and expected output.

If that boundary exists only in the builder’s memory, the workflow never had a reliable contract.

2. The new case class

Describe the new work in one sentence. Avoid “more complex” or “edge case.” Name the difference.

For example: “Requests from external customers now enter the same triage queue as internal requests, and the response may create a service commitment.” That sentence is more useful than “the model is seeing harder tickets.”

3. The changed stakes and likely failure

Ask what could go wrong now that was not material before. Is the likely failure a wrong classification, unsupported answer, privacy exposure, missed escalation, unauthorized action, or reviewer overload?

Also record the economic failure. A workflow can remain accurate enough while becoming unprofitable because every new case requires heavy correction or specialist review. “Accurate” is not the same as “worth keeping in production.”

4. The proof required before widening

Run at least one representative case from the new class and one boundary or unsupported case. Record the expected output, evidence checked, reviewer corrections, rework, time or cost impact, and escalation route. If the new class is consequential, increase the sample before widening access; five minutes is the trigger test, not a substitute for proportionate evaluation.

This is deliberately smaller than a full evaluation program. It is a trigger test that prevents the team from silently converting an exception into a new default.

5. The bounded decision

Choose one outcome and name the owner:

  • KEEP IN SCOPE: the case is materially the same, with the same evidence and reviewer path.
  • CONTROLLED TEST: the new class is allowed only under a named reviewer and stop condition.
  • ROUTE TO HUMAN: the workflow may collect or summarize, but a person owns the decision.
  • REDEFINE WORKFLOW: the job, evidence, permissions, or exception path changed enough to require a new contract.

If nobody has authority to choose among those outcomes, the workflow is not ready to absorb the new work.

The take

Teams do not need to retest every ordinary run. They do need a trigger for recognizing when “ordinary” has changed.

The useful control is not a promise that the workflow will never see a new case. That is impossible. It is a case-mix expansion record that makes the change visible, assigns ownership, requires proportionate proof, and preserves the decision.

The next time someone says, “It passed the pilot,” ask: Passed for which cases?

That question is often the difference between controlled expansion and accidental scope creep.

Sources

  • [OpenAI: Evals guide](https://developers.openai.com/api/docs/guides/evals)
  • [Anthropic: Building effective agents](https://www.anthropic.com/engineering/building-effective-agents)
  • [Microsoft WorkLab: The 2025 Work Trend Index](https://www.microsoft.com/en-us/worklab/work-trend-index/2025-the-year-the-frontier-firm-is-born)

Suggested CTA: Download the [Cortex AI Workflow Case-Mix Expansion Card](../products/freebies/cortex-ai-workflow-case-mix-expansion-card-2026-10-08.md) and run it the next time a live workflow receives a new class of work.

Suggested internal links: AI workflow five-case proof pack; AI workflow reviewer calibration; AI workflow maintenance schedule; AI workflow change gate.