The awkward part of an AI pilot comes at the end. The tool did useful things, made mistakes, and still needed a person to check the output. Do not decide based on the best demo or the worst miss. Treat the pilot as evidence.

The rule

Keep when required safeguards pass and the workflow creates useful net value. Fix when the problem is specific and correctable. Stop when the workflow misses a must-have safeguard, quality floor, or value test.

If the workflow itself is not defined, start with the Automation Fit Test and the AI readiness checklist.

The Pilot Evidence Sheet

Grade the workflow from input to reviewed result, not how impressive the model feels. Record:

  • Test volume: number of representative cases.
  • Baseline time: average handling time before the pilot.
  • Pilot time: average time including review, corrections, and exceptions.
  • Quality: usable as-is, minor edit, major edit, or unusable.
  • Exceptions: cases that left the normal path.
  • Safeguard misses: cases where a required rule did not hold.
  • Real cost: software, usage, maintenance, integration, and human effort.

1. Start with the safeguard gate

Some requirements are not tradeoffs. If every outgoing reply must be reviewed before sending, a path that skips review is a failed safeguard even if most drafts are excellent. Do not hide that inside an average score.

For sensitive information, use the AI data-handling guide before continuing.

2. Measure quality in four buckets

Quality math

Usable-output rate = (usable as-is + minor edit) ÷ total cases
Major-correction rate = (major edit + unusable) ÷ total cases

There is no universal passing percentage. Set the minimum for the workflow before you look at the final numbers.

3. Measure net time, not generation speed

Include loading information, reviewing output, correcting mistakes, and handling exceptions.

Net time saved

Baseline handling time − full pilot handling time = net minutes saved per task

If the old process took eight minutes and AI generated a draft in seconds, that does not mean you saved eight minutes. If review and correction take six, the useful saving is about two. Feed the measured number into the AI automation ROI worksheet.

4. Count exceptions separately

Averages hide operational pain. Track missing inputs, ambiguous requests, failed integrations, specialist handoffs, and cases where the system should refuse to continue. A clear handoff is often better than a plausible guess.

Worked example: inquiry reply drafts

This is an illustrative example, not a real company or first-hand test.

A home-service business tests 40 inquiry replies. A person reviews every AI draft before sending.

The workflow saves time, but the pilot does not justify removing review. At 120 inquiries per month, the measured saving is about 564 minutes, or 9.4 hours, before software and maintenance costs.

A sensible decision is: keep the drafting assistant, keep human review, tighten the source rule, and retest the weak cases.

Keep, Fix, or Stop

Keep

  • Must-have safeguards passed.
  • The quality floor was met on representative work.
  • Net time or another chosen outcome improved.
  • Exceptions are manageable.
  • Realistic value justifies ongoing cost.

Fix

  • The use case still appears valuable.
  • The failure is concentrated in a known part of the workflow.
  • A specific change can be tested: better source data, tighter instructions, clearer handoff, stronger review, or narrower scope.

Stop

  • A must-have safeguard cannot be made reliable enough.
  • Quality remains below the minimum.
  • Review and correction erase the value.
  • The workflow works only on cherry-picked easy cases.

Stopping is not a failed pilot. One purpose of a pilot is to discover cheaply that a workflow should not become permanent.

Copy this decision memo

Workflow: What job did we test?
Baseline: What did it require before?
Sample: Were the cases representative?
Quality: What were the usable and major-correction rates?
Time: What was full handling time?
Exceptions: What failed or needed a handoff?
Safeguards: Did every must-have rule hold?
Cost: What is the realistic ongoing cost?
Decision: Keep, fix, or stop—and why?

Limits

This framework works best for bounded operational workflows where inputs, outputs, review time, and misses can be observed. It is weaker when outcomes take months to appear, samples are tiny, or consequences are difficult to reverse. Higher-stakes work may need specialist review and stricter safeguards before any pilot expands.

If the pilot ran only on clean examples or with the workflow builder standing nearby, retest under realistic conditions before expanding scope.

What to do next

If the pilot clears the keep threshold, calculate the business case with real pilot numbers using the ROI worksheet. If the product is the problem, return to the AI tool buying framework and compare alternatives against the same job and test set.

The standard

A pilot earns the right to continue when the evidence shows a better workflow—not when the AI produces an impressive moment.