AI teams have become good at building workflows. They can connect tools, add guardrails, version prompts, run evaluations, and put an agent in front of a business process before the old project plan has finished circulating.

The operational problem is what happens next.

On June 3, OpenAI said its Agent Builder and Evals products would no longer be available on the platform after November 30, 2026, recommending the Agents SDK or Workspace Agents for different use cases. That is a normal product-lifecycle decision. It is also a useful warning for every team that treats a working AI workflow as a finished asset.

The workflow may still be “working” today. Its dependency is already on a clock.

What happened is bigger than one product sunset

OpenAI’s AgentKit announcement positioned Agent Builder, connector administration, guardrails, versioning, and evaluations as the kind of machinery teams need to build and operate agents. The later lifecycle update makes the uncomfortable part visible: even a workflow built with serious production controls may need to move when the surrounding platform changes.

That is not an argument against managed platforms. It is an argument against confusing platform features with operational ownership.

Anthropic’s guidance on effective agents makes the same point from another direction. Teams should use the simplest architecture that fits the job, because extra layers can obscure what actually happened and make systems harder to debug. A workflow that depends on several abstractions needs a clear answer for each one: what does it do, who owns it, how will we know it changed, and what test proves the workflow is still safe?

Most teams can answer the first question. Too few can answer the last three in five minutes.

The missing control is a dependency sunset test

A dependency sunset test is a small, repeatable review that runs when a model, connector, prompt runtime, evaluation layer, permission, API, or workflow platform changes.

It is not a migration project. It is the decision gate before a migration project becomes necessary.

The card needs five fields:

1. Dependency: What external component does this workflow rely on? 2. Change signal: How will the owner know it changed, was deprecated, lost a capability, or altered its limits? 3. Exposed step: Which exact workflow step is affected — retrieval, classification, tool call, approval, write-back, evaluation, or handoff? 4. Revalidation test: What last-known-good case must run again, with what evidence, before the workflow resumes? 5. Release decision: Does the owner keep, patch, narrow, pause, or retire the workflow?

The absence of a named owner or a last-known-good case should be a pause trigger. If nobody knows what to rerun, nobody has proved that the workflow is safe to keep running.

Why versioning is not enough

Versioning tells you what changed inside your repository or builder. It does not automatically tell you what changed in the world around it.

A connector can change its permissions. A source system can alter a field. A model can become more cautious or more eager. An evaluation dataset can stop representing the real work. A vendor can retire the builder that created the workflow. A policy can change the set of actions the agent is allowed to take.

The workflow’s visible version may be identical while its effective behavior is different.

This is why “we have logs” is not the same as “we have revalidation.” Logs explain what happened after a run. A revalidation test helps decide whether the next run should be allowed to happen.

The practical test for operators

Pick one workflow that matters but is still bounded — for example, customer-request triage, internal knowledge retrieval, or a draft-first reporting process.

Then ask:

  • Can we list every external dependency in one place?
  • Can the owner name the first signal that a dependency changed?
  • Can a non-builder identify the step most likely to fail?
  • Do we have one normal case, one boundary case, and one prohibited case saved as test inputs?
  • Is the release decision recorded by someone with authority to pause the workflow?

If the answer is “no,” the workflow is not ready for more autonomy. It is ready for a dependency card.

The point is not to predict every vendor change. That is impossible. The point is to make change survivable: visible enough to catch, bounded enough to test, and owned enough to stop. A dependency register that nobody uses is paperwork; a sunset test tied to a release decision is control.

The take

The next phase of AI operations will not be won by teams that add the most features to a workflow. It will be won by teams that can absorb change without losing control of the work.

Before adding another connector, agent, or orchestration layer, give every live workflow a dependency sunset test. Name the dependency. Map the exposed step. Save the last-known-good case. Assign the decision owner. Require evidence before resume.

That is the difference between a workflow that merely runs and one the business can continue to trust.

Sources

  • [OpenAI: Introducing AgentKit](https://openai.com/index/introducing-agentkit/)
  • [Anthropic: Building effective agents](https://www.anthropic.com/engineering/building-effective-agents)

Suggested CTA: Download the Cortex Dependency Sunset & Revalidation Card, then run it on one live workflow before expanding its permissions or volume.

Suggested internal links: AI workflow change gates; AI workflow maintenance schedules; AI workflow rollback drills; AI workflow evidence chains.