AI Agent Fallback Chains: Always Serve a Response
The circuit breaker opened. After four retries with exponential backoff, your agent decided the LLM API isn't responding. The retry logic is working exactly as designed.
Now what?
If your agent throws an exception and the UI shows a generic error, you solved the technical problem but destroyed the experience. The circuit breaker solves when to stop trying. The fallback chain solves what you serve afterward.
These are two distinct patterns. This article covers the second one.
Why Fallback Is Not "Show an Error"
There's a difference between a system that fails and a system that degrades.
A system that fails returns an error when it can't do what you asked. A system that degrades returns the best possible response given the current context — even if that response is less precise, slower, or generated by a different mechanism.
In production, an agent that degrades gracefully has a radically better SLA than one without fallback. Not because it fails less, but because when it does fail, it still serves something useful.
Most teams building AI agents in production have implemented retry and circuit breaker (or something close to it), but rarely have the fallback hierarchy defined. That's what we build here.
The Structure of a Fallback Chain
A fallback chain is a sequence of mechanisms ordered by descending quality and ascending availability. When the first fails, you try the second. When the second fails, the third. The last one always responds.
type FallbackTier = {
name: string;
execute: (input: AgentInput) => Promise<AgentOutput>;
isAvailable: () => Promise<boolean>;
};
async function withFallbackChain(
input: AgentInput,
tiers: FallbackTier[]
): Promise<AgentOutput & { tier: string }> {
for (const tier of tiers) {
const available = await tier.isAvailable().catch(() => false);
if (!available) continue;
try {
const result = await tier.execute(input);
return { ...result, tier: tier.name };
} catch {
continue;
}
}
// Last tier never fails: minimal response always available
return buildGracefulError(input);
}
The typical chain for a support or extraction agent has four levels:
Tier 1 — Primary model: The best model you have. Maximum quality, higher cost, higher latency. claude-opus-5-5 or equivalent.
Tier 2 — Fast model: A smaller, cheaper model from the same provider or an alternative. claude-haiku-4-5-20251001 or gpt-4o-mini. Responds in ~400ms and is usually available even when the primary model is overloaded.
Tier 3 — Semantic cache: Previous high-quality responses searched by vector similarity. If someone asks something very similar to what the agent already answered well, serve that response without calling any LLM.
Tier 4 — Deterministic logic: Rules, regex, templates, database queries. For cases with a clear, stable answer, the LLM is unnecessary overhead. A billing agent can extract an order number with regex; a FAQ agent can do keyword search.
Tier 5 — Contextual error: Never a generic 500. A message that explains what happened, what the user can do, and when to expect resolution.
Real Implementation with All Four Tiers
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
// Tier 1: Primary model
async function callPrimaryModel(input: AgentInput): Promise<AgentOutput> {
const response = await client.messages.create({
model: "claude-opus-5-5",
max_tokens: 1024,
messages: [{ role: "user", content: input.query }],
system: input.systemPrompt,
});
return parseResponse(response);
}
// Tier 2: Fast model from same provider
async function callFastModel(input: AgentInput): Promise<AgentOutput> {
const response = await client.messages.create({
model: "claude-haiku-4-5-20251001",
max_tokens: 512,
messages: [{ role: "user", content: input.query }],
system: input.systemPrompt,
});
return parseResponse(response);
}
// Tier 3: Semantic cache search
async function searchSemanticCache(
input: AgentInput
): Promise<AgentOutput | null> {
const embedding = await getEmbedding(input.query);
const similar = await vectorDB.search({
vector: embedding,
threshold: 0.92, // Only if very similar
limit: 1,
});
if (!similar.length) return null;
return { ...similar[0].cachedResponse, cached: true };
}
// Tier 4: Deterministic extraction
function deterministicFallback(input: AgentInput): AgentOutput | null {
// Example: support agent with indexed FAQs
const keyword = extractMainKeyword(input.query);
const faqMatch = faqIndex.findByKeyword(keyword);
if (faqMatch) return { answer: faqMatch.answer, source: "faq" };
return null;
}
// Full chain
async function agentWithFallback(input: AgentInput) {
// Tier 1
if (circuitBreaker.isHealthy("claude-opus")) {
try {
return await withTimeout(callPrimaryModel(input), 8000);
} catch {
circuitBreaker.recordFailure("claude-opus");
}
}
// Tier 2
if (circuitBreaker.isHealthy("claude-haiku")) {
try {
return await withTimeout(callFastModel(input), 3000);
} catch {
circuitBreaker.recordFailure("claude-haiku");
}
}
// Tier 3
const cached = await searchSemanticCache(input).catch(() => null);
if (cached) return cached;
// Tier 4
const deterministic = deterministicFallback(input);
if (deterministic) return deterministic;
// Tier 5: always responds
return {
answer: degradedMessage(input.context),
degraded: true,
retryAfter: 60,
};
}
The degradedMessage Function Matters More Than You Think
The error message isn't a UX detail. It's your system's last communication with the user during a failure. What you write there directly affects how users perceive product reliability.
function degradedMessage(context: AgentContext): string {
const messages: Record<AgentContext, string> = {
"invoice-extraction":
"Automatic processing is temporarily unavailable. " +
"You can upload the file manually in the invoices panel. " +
"The system will recover within the next few minutes.",
"customer-support":
"Our assistant is temporarily unavailable. " +
"Your query has been logged and you'll receive a response within 2 hours.",
"data-analysis":
"Automatic analysis is paused for maintenance. " +
"Your data is saved and will be processed automatically once the system recovers.",
};
return messages[context] ?? "Service temporarily unavailable. Please try again in a few minutes.";
}
When Each Tier Saves Your SLA
Circuit breakers and retries solve transient failures (a momentary 429, a network timeout). The fallback chain solves sustained failures.
In practice, these are the scenarios where each tier intervenes:
Tier 2 kicks in when the primary model is overloaded (529 overloaded) or when latency exceeds your defined timeout. This happens several times a month in any production system. Users notice nothing except perhaps a slightly shorter response.
Tier 3 kicks in during provider incidents (30-60 minute outages) and in cases of high concurrent demand. Semantically cached responses with high similarity are indistinguishable for the user from real-time generated ones.
Tier 4 kicks in during long incidents or when both LLM tiers are down. It covers the percentage of queries with a known structured answer — in many domains, between 20% and 40% of traffic.
Tier 5 kicks in only when everything else has failed. In well-designed systems, it covers less than 1% of traffic.
The Mistake That Breaks the Chain
The most common failure in fallback chain implementations isn't technical. It's a design issue: tiers aren't correctly isolated.
If Tier 3 (semantic cache) calls the same embeddings API that Tier 1 uses internally, and that API is down, your "fallback" depends on the system you just declared unavailable. Same problem if your deterministic logic (Tier 4) queries a database with the same credentials as the primary agent.
Each tier must have its own dependencies. The semantic cache can use an in-memory index for the most frequent hits. Deterministic rules should run with data in memory or in a separate database.
If the AI integration systems you build share infrastructure between tiers, the chain will collapse exactly when you need it most: when your primary infrastructure is under pressure.
Observability: Which Tier Responded and Why
A fallback chain without metrics is an opaque system. You need to know, for each response:
- Which tier responded
- Why previous tiers failed
- How long each tier has been in "degraded mode"
// Tag every response with its tier
const result = await agentWithFallback(input);
await metrics.record({
tier: result.tier,
latency: Date.now() - startTime,
cached: result.cached ?? false,
degraded: result.degraded ?? false,
context: input.context,
});
if (result.degraded) {
await alerts.notify({
level: "warn",
message: `Agent serving degraded responses for context: ${input.context}`,
duration: degradedSince,
});
}
An alert that fires when Tier 2 has been responding for more than 5 minutes — instead of Tier 1 — gives you time to act before users notice.
Real Results
In AI agent integration projects where we've implemented four-tier fallback chains, provider incidents that previously caused 100% error rates during their duration now result in 0-2% of visibly degraded responses for users.
The change doesn't come from providers failing less. It comes from the system knowing what to do when they do.
If you're building a production AI agent and haven't defined your fallback chain yet, you'll find out about the next provider incident through user error reports.