A credible AI pilot starts with a comparison you can defend. Define the work, the change being tested and the outcome that would justify expansion before giving people access. Where feasible, assign access or rollout timing randomly. Track unsuccessful attempts and missing results alongside completed work, and keep the conclusion within the conditions you tested.
For an adoption lead, the aim is a decision: expand a particular workflow, change it, stop it or gather better evidence. A popular tool and a successful demonstration can both be useful without proving that the intervention improved ordinary work. The broader guide to measuring AI value connects that evidence to operational and financial outcomes.
Write the decision before opening the pilot
Suppose Claire Bennett leads software support and wants to test AI-drafted replies to customers whose CSV files will not import. Agents receive an error message, a sanitized sample file and approved help documentation. They must send instructions that identify the problem and give valid next steps without inventing product features.
“Does AI help support?” is too broad. Claire's question is whether offering an approved drafting workflow reduces the team's effort to resolve these cases while meeting the existing quality standard.
Write a short protocol that Claire, an analyst and the reviewers can use:
- Eligible work. Define the import problems included, required inputs and exceptions that stay outside the pilot. Apply those rules before seeing a draft.
- Intervention. Specify the tool, version where available, access, training, source material and required review. The tested change is this whole package.
- Comparison. Describe how the same work is handled today, including existing tools and support. Do not make normal service worse to create a contrast.
- Primary outcome. Choose the measure that answers the decision, such as total human effort per accepted resolution. Define acceptance and the follow-up window for reopened cases.
- Guardrails. Name unacceptable outcomes and who can pause the pilot. Here, invented troubleshooting instructions would require investigation even if replies became faster.
- Decision rule. Agree what improvement would be worth the recurring cost and what uncertainty would prevent expansion. Have the analyst check whether the design can distinguish that improvement.
If the AI group also gets better documentation and extra coaching, the pilot evaluates that combination. It cannot isolate the model's contribution without a design that separates those changes. Use the workflow redesign guide to make the changed process explicit.
Choose a comparison that fits how colleagues work
Collect a baseline before rollout: the mix of import problems, review effort, resolution outcomes and relevant differences between teams. A baseline describes the starting point. A concurrent comparison helps address what would have happened during the pilot without the intervention.
HM Treasury's AI impact-evaluation guidance describes randomizing access, the timing of access or encouragement to use a tool. It also discusses nonrandom comparison methods when random assignment is infeasible. Its government setting differs from a company pilot, but the central question transfers: what exactly is the intervention being compared with?
| Design | What it can help answer | Main condition to check |
|---|---|---|
Randomly assign access | What changes when eligible people or teams are offered the workflow? | Assignment is respected and both groups have comparable outcome records |
Randomly phase in access | What changes for earlier groups while others continue normal work? | Timing is genuinely randomized and the analysis accounts for time |
Use a nonrandom comparison group | How do outcomes change relative to a relevant unaffected group? | An analyst can justify the comparison and its assumptions |
Observe one group before and after | What changed after introduction? | Other changes remain possible explanations; the result alone cannot isolate an AI effect |
Choose the assignment unit before choosing the spreadsheet layout. If agents routinely share draft replies and troubleshooting techniques, individual assignment may not keep the groups separate. Assigning whole teams might fit better, provided there are enough teams for a useful comparison. The OHID guide to randomized trials explains this distinction between individual and grouped assignment.
For Claire, that means mapping who shares work with whom. An AI-assisted reply copied into the comparison team's shared answer library changes what that team receives. Record such sharing rather than pretending the comparison remained untouched. Hundreds of tickets from two teams do not become hundreds of independently assigned teams.
Have an analyst choose the sample size and analysis for the assignment unit. Do not split a tiny team in half and assume randomization guarantees a precise answer. Nor should agents repeat the same import problem twice and treat the second attempt as fresh work: they may remember its solution.
Keep people and tasks from disappearing from the result
A comparison can become misleading after assignment. People may decline to use the tool, choose only easy cases, switch methods or stop recording difficult work.
METR encountered a version of this problem in its February 2026 developer-study update. Some developers were reluctant to join a study that might require working without AI, and some tasks assigned to the AI-disallowed condition were not completed. The researchers judged their newer productivity estimates unreliable. Their LinkedIn discussion of the update described the design difficulties, not a settled new speed estimate.
The lesson for Claire is to keep a record of eligible cases and assigned groups, not just AI-assisted successes. For a randomized offer-of-access question, keep people in their assigned groups for the primary analysis, including those who rarely use the tool. Show actual uptake separately. Comparing enthusiastic users with nonusers answers a different question and can reintroduce selection bias.
This follows the intention-to-treat principle described in the MHRA's statistical guidance. That guidance concerns clinical investigations; the applicable methodological point is to preserve the randomized comparison. It also warns that missing outcomes can bias results. An intention-to-treat label does not recover data you never collected.
Ask the analyst to predefine how missing records will be handled. Report how many are missing in each group and why, and check whether plausible alternative outcomes would change the decision. Do not silently treat an unrecorded resolution as either a success or a failure.
Keep these events visible in the pilot log:
- An eligible case was excluded, with the reason and when that decision was made.
- An assigned agent did not use AI, or a comparison agent used it.
- An AI attempt was abandoned and the case finished manually.
- A colleague supplied extra checking, correction or advice.
- A case had no recorded outcome by the agreed follow-up date.
A Reddit user describing a support pilot reported a lead quietly checking supposedly resolved tickets and senior agents handling difficult cases in Slack. The account is unverified, but it raises a concrete question worth asking: whose work keeps the pilot looking successful without appearing in its records?
Review the outputs before interpreting the average

For the import-support pilot, score each reply against three requirements:
- It identifies a cause supported by the error message and file sample.
- Its troubleshooting steps are available in the product.
- It tells the customer what to do if those steps fail.
Fluent wording alone is not acceptance. Use the same criteria and sampling approach in both groups; separate this evaluation from the review needed to protect customers during normal service.
Where practical, remove tool and group labels before evaluation. Reviewers may still infer the method from the writing, so do not claim perfect blinding. Resolve scoring disagreements using examples before interpreting a small difference between groups.
Pair the quality results with total effort through an accepted outcome. Report attempted and accepted counts with the effort measure. A lower average among completed cases can conceal more abandoned cases or a harder queue left for other people.
Ask for the estimated difference and its uncertainty, not only a percentage or a statement that a result is statistically significant. The analyst should account for the assignment design and repeated work from the same people. A wide interval may leave both a worthwhile benefit and no useful improvement compatible with the evidence.
Inspect any planned breakdowns that matter to the decision, such as common import errors versus complicated formatting problems. Treat patterns discovered afterward as leads for a further test. Avoid searching many small subgroups until one looks successful.
Make a bounded decision and record what changed
Do not let the final meeting turn into a vote on whether people like AI. Bring the original decision rule, the comparison, quality outcomes, missing records and operating cost together.
| Finding | Appropriate next step |
|---|---|
Useful improvement, acceptable quality and a credible comparison | Expand within the tested scope, with ongoing monitoring |
A fixable workflow problem obscures the result | Change that part and evaluate the revised workflow |
Unacceptable harm or effort with no compensating benefit | Stop the tested use and retain what was learned |
Too little information or a compromised comparison | Report uncertainty and repair the evaluation before claiming an effect |
Document changes during the pilot, including tool updates, new documentation, staffing and ticket-routing rules. If the intervention changes substantially, separate the periods instead of blending them into one unexplained average. A result for this version, support queue and review arrangement is not automatically a result for every department.
The next action is small and specific: write the comparison and decision rule for one workflow, then ask the analyst and the people doing the work what could make that comparison misleading. Resolve those objections before the pilot starts.
For the next funding and operating commitment, use the decision guide to stopping, extending or scaling an AI pilot. It covers bounded extensions and the responsibilities a receiving team needs to accept.
Questions about credible AI pilots
Can we run a useful pilot without a control group?
Yes. A pilot without a control group can reveal usability problems, failure modes, review effort and whether the workflow is feasible. A before-and-after comparison alone usually cannot establish how much change AI caused. If that causal question matters to the investment, consider a concurrent comparison or randomized rollout with an analyst. The comparison options above explain their different limits.
How long should an AI pilot run?
Run the pilot long enough to observe the work and follow-up outcomes required by its decision, using a sample plan agreed in advance. Duration depends on task frequency, variation, learning time and the improvement you need to detect. A fixed number of weeks is not evidence of adequacy. For the import-support example, the plan must allow time to identify reopened cases as well as initial replies.
Should people who barely use AI be excluded?
Not from the primary comparison of a randomized offer of access. Keep participants in their assigned groups and report uptake separately, so low use remains part of the result of offering that workflow. Missing outcomes require an explicit analysis plan; they should not be silently dropped. Use interviews to understand low use and the case log described above to distinguish it from missing data.
What can a small pilot actually prove?
A small pilot may identify specific failures and show whether a workflow can operate under the tested conditions. It may be too imprecise to establish a modest productivity improvement, and successful cases do not establish safety for rare failures or other settings. Ask an analyst what effect the available sample could reasonably detect. Use an uncertain result to plan the next test rather than presenting it as proof of no benefit.



