Skip to content
A broad vermilion form separates charcoal and pale gray printed masses across an opaque canvas.

Adoption and productivity measurementArticle

When should a pilot be stopped, extended or scaled?

Decide whether to stop an AI pilot, fund a bounded extension or approve expansion with clear ownership, operating support and a review date.

Jump to a section

Stop an AI pilot when the intended use no longer has a credible path to worthwhile, acceptable operation. Extend it when a specific uncertainty could change the decision and a bounded test can resolve it. Scale when the evidence supports a defined expansion and the people who will run it have accepted the work, cost and controls.

Those decisions concern the next commitment. A promising result does not oblige a company to fund a permanent service, and an inconclusive result does not justify another month by itself.

This article is for the business sponsor reviewing an enterprise pilot. If the comparison, quality assessment or missing results are still unclear, start with how to run a credible AI pilot experiment. Here, the question is what the company should authorize after reading that evidence.

Name the commitment you are being asked to approve

“Scale” can mean adding colleagues to the same assisted workflow, opening a new department, connecting more data, or allowing software to act without someone approving each output. These changes create different obligations. Write down the proposed change before judging whether the pilot supports it.

Suppose Jamie Lee manages the order desk at a distributor. The pilot drafts explanations for delivery-date changes using order records and warehouse updates. An employee checks every draft before sending it to a customer. The team now wants access for the evening shift.

That request leaves human approval in place. It still raises a new question: can evening staff check the drafts and resolve conflicting warehouse updates with the support available then? Successful daytime use does not answer that automatically. Nor does it establish that messages can be sent without approval.

Ask the sponsor to specify:

  • Who and what changes. Name the additional people, work, data and permitted actions.
  • What continues. Preserve the checks and support that made the tested workflow acceptable.
  • What remains unproven. Identify the conditions the pilot did not cover.
  • What the decision commits. State the funding, staff time and operating responsibility being requested.

Keep the original acceptance criteria beside the findings. If priorities have changed, explain the change and reassess the evidence against it. Quietly lowering the bar after seeing the result makes the decision harder to defend.

Stop when the next investment has no credible purpose

A pilot can teach the company something useful and still deserve to close. The question is whether another investment has a plausible route to an outcome worth maintaining.

Michael Sharp's NIST discussion of industrial AI assessment asks whether added complexity brings enough benefit, and whether a simpler solution can meet the need. That is a useful challenge for Jamie: perhaps correcting the warehouse update process would resolve the customer confusion more directly than generating better explanations of inconsistent records.

We recommend stopping the proposed use when one of these conditions holds:

  • The business need has disappeared or a simpler approach meets it adequately.
  • The benefit is too small once review, correction and ongoing support are included.
  • A material failure has no credible remedy within the intended scope and resources.
  • Nobody with the necessary authority and capacity can take responsibility for operation.

A fixable gap may justify a revised test. Name the fix and its cost before making that exception. Past spending alone does not tell you whether the next stage is worthwhile.

Close the pilot deliberately. Tell participants when access ends and how to finish open work. Have the responsible owners handle retained records, access removal and supplier commitments under the company's existing policies. Preserve the findings, including what would need to change before the idea is reconsidered. A stop decision should leave colleagues with a working route for tomorrow's orders.

Extend only to answer a question that matters

An extension should buy information that can change the decision. “The team needs more time” is incomplete until someone can say what that time will reveal.

For Jamie, suppose the pilot missed the peak dispatch period, when warehouse updates arrive late and the order desk is busiest. A further test could be useful if that workload is the main unresolved concern. Repeating quiet daytime work would not answer it.

Write the extension around four commitments:

  1. The question. Can the evening shift check delivery explanations during peak dispatch without building an unacceptable queue?
  2. The test. Include that workload and the staff who would actually operate it. Record review effort, unresolved records and the customer-service outcome.
  3. The limit. Set a workload boundary, allocated support, spending cap and review date. Keep human approval and an available manual route.
  4. The consequence. State what evidence would support expansion, require a narrower use, or end the test.

The analyst should judge whether the proposed work can supply enough evidence, using the pilot's comparison and analysis plan. A calendar deadline cannot make a sparse sample informative. Equally, more observations of the wrong conditions do not resolve the missing question.

Sometimes the gap is an unfinished implementation task rather than uncertainty. If nobody has arranged evening support, waiting another month will not test support capacity. Assign and fund the work, or keep the deployment within the boundary the company can support.

Scale a service that someone is ready to run

Two order-desk colleagues sit side by side reviewing a delivery record, with the sheet and laptop facing them.
The receiving team needs to understand the real records, exceptions and support work before accepting the handover.

The GSA guide to moving from pilot to production identifies ownership, an implementation plan and future retirement evaluation as production considerations. It also stresses looking beyond the limited problem a pilot addressed to the full pipeline.

For Jamie's evening shift, that means an operating agreement with concrete answers:

ResponsibilityWhat the receiving team needs to accept
Customer outcome
Who owns correct, useful delivery explanations and reviews whether service improves?
Exceptions
Who resolves a conflict between an order record and the latest warehouse update?
Capacity and cost
Who funds checking, support and usage at the proposed workload?
Service problems
Who can restrict use, restore the manual route and contact the supplier?
Continued evidence
Who reviews quality and benefit, and what prompts a fresh decision?

Do not assume a pilot team's close attention will continue indefinitely. In a Reddit discussion about production handoffs, a user described the tension between a project team eager to hand over and operations staff concerned about support and capacity. They suggested keeping the builders involved for a short period after launch. This is an anonymous account, but its practical question is useful: who helps the receiving team when the first unfamiliar case arrives?

Name that support period and its exit conditions. For example, the evening supervisor should know whom to call when the warehouse record is stale and what to do if no specialist is available. If every difficult case still depends on one pilot engineer, the handover is unfinished.

Expansion can be bounded while improvements continue. In their March 2026 account of GOV.UK Chat testing, Sam Dub and Sharon McDonald described an app-first expansion plan alongside remaining work on answer speed and questions the service could not answer. That dated account illustrates a limited next step with acknowledged gaps; it does not establish that another organization's gaps are acceptable.

For your decision, distinguish an improvement you can safely schedule from a condition that must hold before anyone new uses the service. Record who accepts the remaining uncertainty and the evidence behind that judgment.

Leave the meeting with a decision people can execute

Keep the record short enough that the operating team will use it. It should contain:

  • The decision and its exact scope, including what remains excluded.
  • The findings that justify it and the important limitations.
  • The accountable owner, resources and work still required.
  • The next review date, evidence to collect and reasons to reconsider sooner.

For the distributor, an executable decision might be to extend the assisted drafting pilot through the next peak dispatch cycle, with evening staff, a named escalation contact and a fixed review date. Access stays within that test until the review. A later expansion would separately approve the staffing and support needed to keep it running.

Use the wider AI value framework to connect the decision to business outcomes, and net time measurement to keep checking and rework visible. The useful result of the meeting is a commitment whose boundaries everyone can explain.

If the funding conversation keeps returning to a vague claim that the pilot is not ready, identify the unresolved acceptance condition. Establish whether another test could answer it or whether a team must make a commitment.

Questions about stopping, extending and scaling AI pilots

How long should an AI pilot extension last?

An AI pilot extension should last long enough to observe the conditions that could change the decision, within an approved resource limit. A delivery-support pilot may need a peak dispatch cycle; another workflow may need enough completed cases for a useful comparison. There is no universal number of days that establishes readiness. Define the question, ask the analyst what evidence is feasible, and agree a review date and consequence using the extension commitments above.

Can a successful AI pilot still be stopped?

Yes. A pilot may meet its technical or task-level target while the proposed service costs too much to operate, solves a lower-priority problem, or lacks an accountable operating team. Explain which result passed and which business condition prevents continuation. Preserve the useful findings and give participants a clear route for unfinished work. The stop criteria help separate that decision from a repair worth testing.

Can we scale with one team while another keeps testing?

Yes, when the approved team's evidence and operating arrangements support its defined use. Keep the other team's access and responsibilities within its own test boundary. Different shifts, records or review capacity can change the result, so do not present one team's success as approval for all. Name the owner and controls for each boundary, then use the operating handover questions before expanding it.

Updated

aiready

A home for your company’s AI community.

Share what works and help each other put AI into practice.

  • Real use cases

  • Practical guides

  • Company policies

  • Shared experience

Explore aiready
Explore the blog