Title: Your AI Workflow Is Not Governed Until Someone Owns the Failed Run Search intent: AI workflow exception ownership, who owns a failed AI run, AI workflow human oversight, AI workflow failure recovery

The failure mode nobody plans for

An AI workflow can have a named reviewer, a complete evidence chain, and a clean approval record, and still create invisible cleanup when the weird case shows up.

The weird case is not always dramatic. It is usually a run that completed three steps and not the fourth. A customer inquiry that produced an output the reviewer was not sure was safe to send. A report draft that pulled from a source that had changed since the workflow was approved. A transaction that partially processed before something broke.

In each case, the question is not whether the run was governed at launch. The question is whether the exception has an owner right now — someone who knows what the last safe state was, what evidence is preserved, and what the next action should be before the workflow picks up again.

Without that, the run sits. The original output quietly gets used, or quietly gets ignored, or quietly gets discovered two weeks later by someone who was not expecting to inherit it.

---

What a failed run needs in the first two minutes

The goal is not a postmortem. The goal is a bounded decision that prevents the failure from expanding.

When a run stops or produces an unsafe result, five questions need answers before anything else happens:

Who owns this run? Not "who should probably look at this" — a named person, with a backup, and an escalation deadline if neither has acted.

What is the last safe state? Which step completed safely? Which step did not? If any external action already occurred — a message sent, a record updated, a transaction initiated — that needs to be noted before anyone assumes the slate is clean.

What evidence is preserved? The input the workflow actually received, the source it used, the output it produced, and any tool or action log from the run. If that evidence disappears before the owner arrives, the failure becomes harder to recover and impossible to learn from.

What is the disposition? One of five options:

  • RESOLVE — the cause is understood, no boundary changed, and the owner can correct the run safely without reopening the broader workflow design.
  • REWORK — the workflow or input needs a change before another attempt. Someone owns the fix, and the fix needs to pass a retest before the run resumes.
  • ESCALATE — the issue touches permissions, customer commitments, financial or legal risk, or a boundary that the exception owner cannot call alone.
  • HOLD — evidence or ownership is missing. The run stays paused until the gap closes. This is not a decision to do nothing; it is a decision to do nothing unsafe until the minimum conditions exist to act.
  • RETIRE — the workflow is no longer safe, useful, or worth repairing. Document the manual replacement path and close the run.

What must be true before the workflow resumes? This is the resume gate. It should be a short, plain-language sentence: "resume only after the owner confirms the source conflict is resolved, the correction passes a normal case and the failed case, and the reviewer approves the boundary." If no one can write that sentence, the workflow is not ready to resume.

---

Why this is different from the approval record and the evidence chain

The approval record governs the workflow design. The evidence chain governs individual runs. The failed-run card governs what happens when a run stops or produces something unsafe — which is the moment both of the other layers are most likely to go silent.

Approval does not help you when the run is paused halfway. Evidence is most useful after the fact. The failed-run card is what you use in the moment: a five-question decision that preserves the last safe state, names the owner, and produces a written resume condition rather than an informal "I'll just re-run it."

"I'll just re-run it" is how a partial completion becomes a duplicate action. How a source conflict becomes a published error. How a missing field becomes a downstream cleanup bill that nobody traced back to the original run.

---

The governance test

One practical way to test whether your AI workflow is actually governed: pick the last exception or near-miss. Can you answer these five questions from the record you already have?

  • Who was the named exception owner, with a deadline?
  • What was the last safe state, in writing?
  • What evidence bundle was preserved?
  • What disposition was chosen, and who chose it?
  • What was the written resume condition?

If you cannot answer all five from a document that existed at the time of the failure — not reconstructed afterward, not assembled from memory — the workflow has an ownership gap where governance ends and hope begins.

The fix is not complicated. It is a card. One page, five sections, filled in at the moment of failure by the person responsible for the run. The card is the difference between a failure that becomes a recoverable, learnable event and a failure that becomes someone else's cleanup three weeks later.

---

Practical next step: Use the Cortex AI Workflow Failed-Run Ownership Card to define the exception owner, last safe state, disposition, and resume condition for the next AI workflow your team runs. Keep the card with the run record so the failure stays ownable even after the original operator moves on. → [Link to companion asset]

Internal links:

  • "Your AI Workflow Needs an Evidence Chain, Not Just Approval"
  • "Before You Scale the AI Workflow, Prove It Created Useful Work"
  • "Your AI Workflow Needs a Change Gate After It Goes Live"
  • "Your AI Workflow Needs a Resume State, Not Just a Handoff"