Fallback chains: designing AI that fails gracefully.
Models time out, providers go down and outputs fail validation. A production system needs an answer for each of those moments before they happen.
Every external dependency fails eventually, and model APIs fail more often than most. Requests time out under load, providers have outages, rate limits arrive at the worst moment, and sometimes the call succeeds but the output is unusable: malformed JSON, a refusal, an answer that ignores the documents it was given.
A demo can ignore all of that. A production system needs a planned response to each failure, decided in advance and tested. We call that plan a fallback chain.
Failure is a when, not an if
Start by listing the ways a model call can fail, because they need different responses:
| Failure | Typical cause | Response |
|---|---|---|
| Timeout | Provider load, long outputs | Retry once with a tight deadline, then move down the chain |
| Rate limit or outage | Quota, provider incident | Skip the provider while it’s failing and use the next rung |
| Invalid output | Broken JSON, missing fields, a refusal | Repair it if the fix is trivial; otherwise retry or move down |
| Ungrounded answer | Weak retrieval, the model ignoring sources | Abstain or hand off; retrying usually repeats the failure |
| Everything down | Rare, but it happens | A deterministic answer, a cached result, or a person |
The chain
A fallback chain is an ordered list of ways to produce an acceptable answer. A typical one looks like this:
- the primary model, with a timeout sized to your latency budget;
- one retry for transient errors, with backoff and jitter;
- a secondary model from a different provider, with prompts tested for it specifically;
- a deterministic answer: a cached response to a known question, a rules-based reply or a template;
- a person, with the context they need to pick up the request.
Each rung is tested on its own against the eval set. A fallback model that hasn’t been evaluated isn’t a fallback; it’s a second, unknown failure mode.
Validate before you trust
A response that arrives isn’t necessarily one you can use. Every rung’s output goes through the same validation before it’s accepted: schema checks for structured output, groundedness checks for answers that should cite sources, and policy checks for anything a customer will see. A response that fails validation counts as a failure of that rung, and the chain moves on.
interface Rung<T> {
name: string;
timeoutMs: number;
call(signal: AbortSignal): Promise<T>;
}
type Attempt = { rung: string; outcome: "ok" | "invalid" | "error"; ms: number; error?: unknown };
export class AllRungsFailed extends Error {}
/**
* Try each rung in order. A rung succeeds only if it answers in time and the
* answer passes validation. Every attempt is logged, so fallback rates show up
* on a dashboard instead of in a post-mortem.
*/
export async function withFallbacks<T>(
rungs: Rung<T>[],
isValid: (value: T) => boolean,
log: (attempt: Attempt) => void,
): Promise<{ value: T; rung: string }> {
for (const rung of rungs) {
const started = performance.now();
try {
const value = await rung.call(AbortSignal.timeout(rung.timeoutMs));
const ms = performance.now() - started;
if (isValid(value)) {
log({ rung: rung.name, outcome: "ok", ms });
return { value, rung: rung.name };
}
log({ rung: rung.name, outcome: "invalid", ms });
} catch (error) {
log({ rung: rung.name, outcome: "error", ms: performance.now() - started, error });
}
}
throw new AllRungsFailed(`No rung produced a valid answer (${rungs.map((r) => r.name).join(" → ")})`);
}
The last rung, a person, sits outside this function. When AllRungsFailed is thrown, the caller routes the request to the handoff queue instead of showing an error.
Circuit breakers
When a provider is down, sending every request to it first makes things worse. Each user waits for a timeout before the chain moves on, and the retries add load to a service that’s already struggling.
A circuit breaker tracks recent failures for each provider. After enough failures in a short window it opens, and requests skip that rung entirely for a cooldown period. Then it lets a few requests through to test whether the provider has recovered, and closes again if they succeed. With a breaker in place, an outage costs a handful of slow responses rather than thousands.
Tell the user the truth
Degraded doesn’t have to mean broken, but it should be honest. If the system has fallen back to a simpler mode, it says so briefly: “I can answer general questions right now, but I can’t look up your order. A colleague will follow up within the hour.”
Users forgive a limitation they’re told about. They don’t forgive a confident answer that turns out to be wrong.
Measure every rung
The attempt log feeds three numbers we watch on every system: the share of requests answered by each rung, the validation failure rate for each rung, and end-to-end latency including fallbacks. A rising share for the secondary model is often the first sign of a provider problem, or of a prompt that has started failing validation, and it shows up well before users start to complain.