The awkward part of an AI pilot comes at the end. The tool did useful things, made mistakes, and still needed a person to check the output. Do not decide based on the best demo or the worst miss. Treat the pilot as evidence.
The rule
Keep when required safeguards pass and the workflow creates useful net value. Fix when the problem is specific and correctable. Stop when the workflow misses a must-have safeguard, quality floor, or value test.
If the workflow itself is not defined, start with the Automation Fit Test and the AI readiness checklist.
The Pilot Evidence Sheet
Grade the workflow from input to reviewed result, not how impressive the model feels. Record:
- Test volume: number of representative cases.
- Baseline time: average handling time before the pilot.
- Pilot time: average time including review, corrections, and exceptions.
- Quality: usable as-is, minor edit, major edit, or unusable.
- Exceptions: cases that left the normal path.
- Safeguard misses: cases where a required rule did not hold.
- Real cost: software, usage, maintenance, integration, and human effort.
1. Start with the safeguard gate
Some requirements are not tradeoffs. If every outgoing reply must be reviewed before sending, a path that skips review is a failed safeguard even if most drafts are excellent. Do not hide that inside an average score.
For sensitive information, use the AI data-handling guide before continuing.
2. Measure quality in four buckets
- Usable as-is: no meaningful correction.
- Minor edit: quick cleanup; core result is correct.
- Major edit: substantial correction or rewriting.
- Unusable: starting over would be faster or safer.
Quality math
Usable-output rate = (usable as-is + minor edit) ÷ total cases
Major-correction rate = (major edit + unusable) ÷ total cases
There is no universal passing percentage. Set the minimum for the workflow before you look at the final numbers.
3. Measure net time, not generation speed
Include loading information, reviewing output, correcting mistakes, and handling exceptions.
Net time saved
Baseline handling time − full pilot handling time = net minutes saved per task
If the old process took eight minutes and AI generated a draft in seconds, that does not mean you saved eight minutes. If review and correction take six, the useful saving is about two. Feed the measured number into the AI automation ROI worksheet.
4. Count exceptions separately
Averages hide operational pain. Track missing inputs, ambiguous requests, failed integrations, specialist handoffs, and cases where the system should refuse to continue. A clear handoff is often better than a plausible guess.
Worked example: inquiry reply drafts
This is an illustrative example, not a real company or first-hand test.
A home-service business tests 40 inquiry replies. A person reviews every AI draft before sending.
- Baseline: 8 minutes per manual reply.
- Pilot time: 3.3 minutes including review and correction.
- Quality: 20 usable as-is, 12 minor edits, 6 major edits, 2 unusable.
- Usable-output rate: 32 ÷ 40 = 80%.
- Net time saved: 8 − 3.3 = 4.7 minutes per inquiry.
- Safeguard issue: one draft included a detail that was not present in the source; the reviewer corrected it.
The workflow saves time, but the pilot does not justify removing review. At 120 inquiries per month, the measured saving is about 564 minutes, or 9.4 hours, before software and maintenance costs.
A sensible decision is: keep the drafting assistant, keep human review, tighten the source rule, and retest the weak cases.
Keep, Fix, or Stop
Keep
- Must-have safeguards passed.
- The quality floor was met on representative work.
- Net time or another chosen outcome improved.
- Exceptions are manageable.
- Realistic value justifies ongoing cost.
Fix
- The use case still appears valuable.
- The failure is concentrated in a known part of the workflow.
- A specific change can be tested: better source data, tighter instructions, clearer handoff, stronger review, or narrower scope.
Stop
- A must-have safeguard cannot be made reliable enough.
- Quality remains below the minimum.
- Review and correction erase the value.
- The workflow works only on cherry-picked easy cases.
Stopping is not a failed pilot. One purpose of a pilot is to discover cheaply that a workflow should not become permanent.
Copy this decision memo
Workflow: What job did we test?
Baseline: What did it require before?
Sample: Were the cases representative?
Quality: What were the usable and major-correction rates?
Time: What was full handling time?
Exceptions: What failed or needed a handoff?
Safeguards: Did every must-have rule hold?
Cost: What is the realistic ongoing cost?
Decision: Keep, fix, or stop—and why?
Limits
This framework works best for bounded operational workflows where inputs, outputs, review time, and misses can be observed. It is weaker when outcomes take months to appear, samples are tiny, or consequences are difficult to reverse. Higher-stakes work may need specialist review and stricter safeguards before any pilot expands.
If the pilot ran only on clean examples or with the workflow builder standing nearby, retest under realistic conditions before expanding scope.
What to do next
If the pilot clears the keep threshold, calculate the business case with real pilot numbers using the ROI worksheet. If the product is the problem, return to the AI tool buying framework and compare alternatives against the same job and test set.
The standard
A pilot earns the right to continue when the evidence shows a better workflow—not when the AI produces an impressive moment.