Keep work moving by deciding in advance which tasks can wait, which have an approved fallback and which must stop when AI becomes unavailable or unreliable. Retain access to the source records, give someone authority to switch modes, and test whether the alternative can handle the work that matters. A fallback is usable only when people can complete and check the task with it.
For a service owner or operational manager, the aim is to preserve an acceptable level of service. That may mean processing fewer orders or delaying an internal report. It does not mean maintaining every AI-assisted activity at its normal speed. This is part of owning your tools and workflows, alongside access, support and maintenance.
Recognize failure even when the service responds
An outage is the visible case. In its June 2025 incident report, OpenAI attributed elevated errors to a routine operating-system update that disrupted network connectivity on servers. Recovery happened in stages across services. A supplier's overall recovery announcement therefore needs a local check of the workflow you actually use.
The less obvious case is an answer that arrives but no longer meets your requirements. Anthropic's September 2025 postmortem described three infrastructure bugs that intermittently degraded responses. Its evaluations had not captured the problems users reported. That incident shows why availability and output quality need separate checks; it does not explain every disappointing response from an AI tool.
In a September 2025 Reddit discussion, a user describing a company with a few hundred employees said their team had spent much of July correcting unfamiliar errors before evaluating alternatives and changing its setup. The account is unverified, and its July timeline precedes the bugs in Anthropic's report. Other participants described better performance. The useful lesson is to preserve examples of what changed in your own work rather than treating community agreement as the test.
Give employees a clear way to report the affected task, time, tool and example of an unacceptable result. Keep confidential material in the approved support channel. An ongoing support owner can investigate whether the cause lies in the model, source information, permissions or another part of the workflow. You can hold questionable work while that investigation continues.
Choose the minimum service you will maintain
Start with the consequence of delay or error. The same department may need several responses to the same outage.
Suppose a warehouse team uses AI to summarize delivery exceptions and draft customer updates. The order-management system remains the authoritative record. The dispatch manager, Jamie Lee, needs to decide what the team can still promise when the assistant stops working.
| Work during the interruption | Chosen response | What must remain available |
|---|---|---|
Internal weekly summary with no immediate deadline | Wait, with an owner and a revised completion time | The underlying records and a visible queue |
Customer update about a confirmed delivery change | Use the approved manual template | Current order details, a trained coordinator and the usual review |
Message that would promise an unconfirmed replacement delivery | Hold until the responsible person confirms it | An escalation route and a record of what is still unknown |
These are decisions for this example, not universal classifications. A normally routine message can become urgent if a customer is waiting to receive a shipment. Review the fallback when the circumstances change.
The UK Government's AI Playbook includes fallback processes for maintaining critical functions when an AI change must be reverted or a system terminated. That is government guidance rather than a general private-sector requirement. It is a useful prompt to name the service you intend to preserve, rather than merely naming a spare tool.
Record the decision in the existing operating instructions. Include the person authorized to switch, the affected work, the approved alternative, the point at which capacity or uncertainty requires escalation, and who sends the next update. Keep these instructions somewhere accessible without the failed assistant.
Rehearse the fallback against real constraints
Run a controlled exercise on approved sample work before relying on the fallback. Simulate the loss of the AI step without disrupting live customer commitments. Ask the people who would cover the interruption to perform the task, including finding records, checking the output and handing it over.
For Jamie's warehouse team, a rehearsal might reveal that the manual template is available but a second coordinator lacks access to the current order record. That is a continuity gap even if the AI supplier has excellent uptime. Fix the access or assign an authorized alternative, then repeat the affected part of the exercise.

Measure the work the fallback adds. Include preparation, review and recording, not only the time spent drafting.
The calculation assumes comparable cases and work that can be divided between the coordinators. Difficult cases, interruptions and incoming demand can reduce the capacity further. Use the exercise to expose these assumptions, then agree a manageable service level with the people responsible for it.
Keep a short record of the exercise:
- Task and interruption. What AI step was unavailable, and what work still had to finish?
- Evidence. Which source records, instructions and permissions were needed?
- Observed result. What finished correctly, how long did it take, and what remained queued?
- Gap and owner. Who will repair each problem, and when will the affected step be tested again?
Repeat when staffing, source access or the workflow changes materially. A document showing how last year's team worked is not evidence that today's backup person can do it.
Check an alternative model before depending on it
A second AI service can be part of the fallback if it is approved for the information and task. Check its output, permissions, capacity and dependencies before the incident. A model that answers quickly but invents a delivery promise is not a successful replacement.
Consultant CTO Lee Crossley described investigating overlapping provider incidents in a September 2026 LinkedIn article. He could not establish the cascading-failure explanation he initially suspected. That restraint matters: several services failing near one another does not prove a shared cause, and several supplier names do not prove independent recovery paths.
Ask the technical owner which components the primary and fallback routes share, including identity, source retrieval, gateways and infrastructure. Where independence is unknown, record it as unknown and retain a route that does not depend on either model. For an ordinary internal report, waiting may be more practical than building a second system.
Changes also need checks when nothing is down. A model replacement, revised prompt or changed source connection can alter an approved task. In a 2024 study of older GPT-3.5 models on toxicity classification, researchers found that model and prompt changes needed attention together, with different behavior across groups of examples. It is a narrow study, not a current ranking of tools. It supports testing your task rather than inferring suitability from an overall model score.
Retain approved examples that cover ordinary work and consequential exceptions. For the warehouse, include a confirmed delay, conflicting dates and a missing replacement date. Define the checks before running them: source dates must be preserved, conflicts flagged and unconfirmed promises withheld. Record the model or service version where available, instructions, source version and results. Repeat uncertain cases because generated answers can vary.
An attractive average must not hide a failure that would make the task unacceptable. The responsible owner decides whether to release, restrict or keep the change in review. Match the depth of testing to the consequences; a few examples cannot establish safety for a high-impact system.
Restore work deliberately
Before returning to the normal route, confirm that the specific workflow works again and reconcile what happened during the interruption. A supplier's green status indicator is useful evidence, but it does not account for your queued cases or partially completed work.
- Check recovery on representative work. Include the behavior that failed and confirm source access, output quality and the intended review step.
- Reconcile the queue. Identify completed, pending and uncertain cases. Check whether an action already succeeded before retrying it, so a customer does not receive duplicate updates.
- Authorize the return. The named owner records which tasks can resume and any remaining restrictions.
- Tell the team what changed. Explain the active route, outstanding backlog and next update. Preserve the incident evidence and assign follow-up repairs.
If the cause remains uncertain, say so in the internal incident record. Do not turn a temporary improvement into a confirmed diagnosis. Continue the restricted mode when the checks do not support normal operation.
For your next continuity review, choose one important AI-assisted task and rehearse its alternative with the person who would actually use it. A missing permission, an unworkable queue or an unclear restart decision is a useful finding while you still have time to fix it.
Questions about AI service continuity
Should we switch to another AI tool during an outage?
Switch only to a tool already approved for the task and information, with a fallback route that has been checked. Verify the alternative's output and dependencies; opening another model does not establish that it is suitable or independent. If no approved alternative exists, use the agreed manual route, wait or hold the task according to its consequences. Start with the minimum-service decision.
How many examples should we use to test a model change?
There is no universal test count for an AI workflow. Choose examples that cover its important inputs, known failures and consequences, and define acceptable results before testing. Repeat uncertain cases and examine important groups separately rather than relying only on an overall score. Higher-impact uses need more extensive evaluation and specialist oversight. The warehouse checks illustrate a starting set of conditions, not proof that three cases are sufficient.
When can we stop using the fallback?
Return when the named owner has evidence that the affected workflow meets its requirements again and the interruption's backlog is understood. Test the failed behavior, confirm source access and review pending or partially completed actions before restarting them. Communicate the approved route and any restrictions to the team. The recovery sequence keeps supplier recovery and your own restart decision connected.



