Provider fallback keeps an LLM application alive after a failure. Model routing chooses the right healthy model before that failure happens.
On September 11, 2026, OpenAI's status history recorded elevated API errors for GPT-5.6 Sol. A provider incident is not unusual: any LLM API can slow down, enforce a rate limit, or become temporarily unavailable.
The important question is what your application does next.
The smallest possible answer often looks like this:
try:
return await openai_call(request)
except TransientProviderError:
return await anthropic_call(request)
That is provider fallback. It is useful, but it is not model routing.
Fallback asks: the provider failed; where should this request go now?
Routing asks: which model should receive this request in the first place?
The distinction matters as soon as reliability, latency, quality, and cost affect a production system.
Not every failure needs another provider.
A timeout, network error, 429, or 503 may be temporary. Google's Gemini API documentation recommends bounded exponential backoff for retryable errors such as 429 RESOURCE_EXHAUSTED and 503 UNAVAILABLE. Its guidance is equally important on what not to retry: a bad request, invalid credential, unsupported parameter, or broken tool schema needs a code or configuration fix, not another attempt.
A basic policy looks like this:
Retries are not free. If a provider is degraded for several minutes, three long retries can turn a fast error into a bad user experience. Set a timeout and a small retry budget for the task, then fail over only when the error is genuinely transient.
Fallback should protect against infrastructure failures. It should not conceal application bugs.
Once an application can call more than one provider, a better question appears: why wait for the preferred provider to fail?
Different workloads need different things:
They do not automatically need the same model.
Hardcoding the choice pushes that decision into every caller:
await generate(
model="gpt-6-sol",
messages=messages,
)
Instead, make the application describe the job:
await llm.generate(
task="rag_answer",
messages=messages,
requirements={
"structured_output": True,
"max_latency_ms": 5000,
},
)
The application says what it needs. The routing layer decides which available model can meet those requirements right now.
This does not need to be an AI-powered router. A good first version is deliberately boring:
One component owns the policy. Without that boundary, provider-specific conditions slowly spread across API handlers, agents, background workers, and scheduled jobs.
Published token prices show why a single default model is rarely ideal. The table below is a pricing snapshot from September 29, 2026 for standard API usage and short-context requests. It excludes cached-input, batch, long-context, tool, and regional-processing charges.
| Model | Input / 1M tokens | Output / 1M tokens | | ---------------- | ----------------- | ------------------ | | Gemini 3.8 Flash | $0.75 | $3.75 | | GPT-6 Sol | $2.00 | $10.00 | | Claude Sonnet 4 | $3.00 | $15.00 |
For one million input tokens and 250,000 output tokens, that works out to approximately:
| Model | Calculated token cost | | ---------------- | --------------------- | | Gemini 3.8 Flash | $1.69 | | GPT-6 Sol | $4.50 | | Claude Sonnet 4 | $6.75 |
This is a pricing comparison, not a quality ranking. The models are not capability-equivalent, and the cheapest route is not necessarily the correct one.
The useful question is not which model is cheapest? It is: which eligible model is the least expensive one that still meets this task's quality and latency requirements?
That order matters. First remove models that cannot produce the required result. Then remove unhealthy targets. Only then rank the remaining candidates by latency and cost.
Static configuration is not enough. A model can be the normal first choice at 9:00 a.m. and a poor choice at 9:05 because latency has jumped or errors are increasing.
Track at least:
429) rateA circuit breaker prevents the router from repeatedly sending normal traffic to a target that is already failing:
In the closed state, requests use the target normally. In the open state, the router skips it. After the cooldown, the half-open state sends a small number of probe requests. Successful probes close the circuit; failures open it again.
Libraries such as LiteLLM provide building blocks for retries, fallbacks, cooldowns, and load balancing. They are useful infrastructure, but the routing policy still belongs to the application: it is the application that knows whether a task can tolerate a cheaper model, a different tool-calling format, or a slower response.
An HTTP 200 from a fallback model is not enough.
Changing models can affect:
Every model eligible for a route should therefore appear in that route's evaluation set.
For a RAG route, measure correctness, faithfulness, citation or structured-output validity, latency, and cost. For an agent route, measure task completion, tool selection, invalid tool calls, number of steps, total tokens, and cost.
The selection pipeline should be simple and explicit:
Quality comes before cost optimization. Otherwise routing quietly converts a reliable system into one that is merely cheap.
A normal LLM log might say this:
request completed in 2.8s
With routing, that is not enough to explain an outcome. Record the decision as well:
{
"task": "rag_answer",
"provider": "anthropic",
"model": "claude-sonnet-4",
"routing_reason": "primary_target_cooldown",
"attempt": 2,
"previous_error": "rate_limit",
"latency_ms": 2814
}
If answer quality changes, this tells you whether the likely cause is retrieval, the prompt, the model, or the routing policy. Without it, dynamic routing creates another black box.
You do not need a complex optimizer to get value from routing. Start with explicit task routes, two providers, bounded retries, a circuit breaker, and enough telemetry to explain every decision.
Then collect production evidence before making the policy smarter.
Fallback is reliability: it keeps the application running when the selected provider has a problem.
Routing is architecture: it decides which provider and model should handle the request before that problem happens.
The question is not which LLM is best. It is which eligible model is good enough for this specific task, right now, without paying or waiting more than necessary.