The first credible AI project is not a strategy deck. It is a bounded workflow with a baseline, a reviewer, and a decision rule.

Most companies do not have an AI adoption problem. They have a proof problem.

They have attended the workshop, bought the tools, circulated the policy, and collected a small museum of prompts and automation templates. Then someone asks the only question that matters: which workflow earned the right to continue?

The answer is usually vague. “The team is experimenting.” “People are saving time.” “The pilot is going well.” Those are activity reports, not operating evidence.

The first AI project should not be another strategy deck or broad readiness assessment. It should be one bounded workflow that produces a reviewable decision in five working days.

AI readiness is too broad to be useful

The AI shelf is crowded with readiness checklists, prompt packs, workflow bundles, and automation demos. They are not useless. They are simply too easy to buy and too hard to use as proof.

The serious conversation around workplace AI has moved toward repeatable workflows, defined inputs, checkpoints, human review, and the tradeoff between quality, speed, and cost. Lean teams are not looking for another abstract explanation of what AI could do. They need a practical way to decide what it should be allowed to do next.

That creates a gap between “we started using AI” and “this workflow earned another dollar of trust.”

Most pilots die in that gap. Nobody selected a narrow enough use case. Nobody named the reviewer. Nobody recorded the baseline. Nobody tested the weird case. When the first bad output arrived, the team had no decision rule beyond optimism or panic.

The five-day workflow-proof sprint

A proof sprint is deliberately small. It does not make a company AI-ready in a week. It answers a narrower question:

Can one real workflow be used safely, reviewed consistently, and released narrowly enough to create useful work?

Day 1: Select one workflow

Choose a recurring workflow with a clear owner and visible output.

Good candidates include drafting a meeting recap from approved notes, preparing a first-pass internal research memo, creating an initial SOP, or drafting customer follow-up from a defined source pack.

Avoid “improve operations with AI.” That is a department aspiration, not a testable workflow.

Write down the workflow name, who performs it, what the output must contain, and what acceptable work looked like before AI. If the team cannot describe the baseline, it cannot honestly claim improvement.

Day 2: Bound the workflow

Before anyone writes a clever prompt, define the input and judgment contract.

What sources may the system see? What tool or model will it use? What may it produce? What must a human decide? What is it never allowed to decide? What happens when a source is missing, contradictory, or outside the approved scope?

This is where an AI experiment becomes an operating process. The boundaries matter more than the wording of the prompt.

Name the reviewer now, not after the first questionable result. Name the stop and escalation rule too.

Day 3: Test normal cases and one weird case

Run at least three ordinary examples and one case designed to expose the workflow’s limits.

The normal cases show whether the process works when the inputs behave. The weird case shows whether the process knows what to do when they do not.

For a meeting-recap workflow, the awkward case might include incomplete notes, conflicting action items, or a decision that was discussed but never approved. The correct result is not a confident paragraph that smooths over the conflict. It is a visible flag for human review.

A workflow that performs well only on the happy path has not proved itself. It has performed a demonstration.

Day 4: Review the review

Record what the human reviewer caught: unsupported claims, missing context, formatting problems, wrong decisions, unnecessary edits, or cases that should have stopped the run.

This correction log is more useful than a satisfaction survey. It tells the team whether the review burden is light and repeatable or whether a supposedly efficient workflow has simply moved the work downstream.

Price the hidden labour. Review, cleanup, correction passes, reruns, manager coaching, and support questions all count. If the first draft is faster but the approved output takes longer, gross time saved is a misleading metric.

The real question is whether the workflow creates net useful work after review and cleanup.

Day 5: Make a release decision

End with a decision card, not a celebratory adjective.

There are three honest outcomes:

  • Go narrow: release the workflow to a named role, with the current controls and review process.
  • Revise: change the input contract, test set, ownership, or review rule and run the proof again.
  • Stop: the quality risk, cleanup burden, or support cost is not worth the claimed benefit.

“Pilot successful” is not a decision. It is what people write when they have not decided how much trust to grant.

A worked example: internal research memos

Imagine a three-person operations team that spends several hours each week turning approved source material into internal research memos.

On Day 1, the owner defines the workflow: produce a one-page memo from a named folder of source documents. The baseline is the existing average completion time and the quality standard: every material claim must be traceable to an approved source.

On Day 2, the team specifies that the system may summarize and structure the supplied documents but may not invent facts, browse for unapproved evidence, or make the recommendation itself. A manager is the reviewer.

On Day 3, the team runs three ordinary memos and one deliberately awkward source pack with a missing date and two conflicting figures. The workflow produces useful drafts on the ordinary cases and correctly flags the conflict on the weird case.

On Day 4, the reviewer finds that the drafts are faster to shape but still require a source-traceability check and one recurring correction to the executive-summary format. Those are manageable controls. If the reviewer had to rewrite every memo, the answer would be different.

On Day 5, the team chooses “go narrow”: one role, one source folder, one reviewer, and a two-week measurement period. It has not proved that AI can transform research. It has proved something more valuable: this specific workflow can be trusted within a defined lane, subject to evidence.

What to measure

Do not measure only minutes saved. Track eligible cases, AI-assisted cases, gross time saved, review and cleanup time, useful outputs kept, quality failures, support burden, and repeating correction patterns.

Then ask where the reclaimed capacity goes. Time saved but left to evaporate into more low-value busywork is weak business value. A workflow earns wider trust when it creates net useful work, not when it generates an impressive demo.

The better first question

Stop asking whether your company is “AI-ready.” It is too broad to guide a decision and too flattering to expose a weak process.

Ask whether one real workflow is specific enough to test, bounded enough to control, and useful enough to release.

In five days, a lean team can have a named owner, a reviewer, a baseline, a weird-case result, a correction log, and a go/revise/stop decision. That is not the whole AI strategy. It is the first piece of evidence a credible strategy should contain.

The next workflow should have to earn its turn.

---

Editorial source notes: Lean-Team Workflow Proof Gap Market Readout (2026-08-08); Cortex AI Workflow Control Starter Pack Improvement (2026-08-10); Cortex AI Useful Work Per Dollar Proof Sheet (2026-07-30).