A business workflow tied to one LLM provider can lose a critical dependency when that service goes down. Customer support, code generation, and internal knowledge bases can all depend on a small group of providers, sometimes without a tested fallback.

Other parts of production infrastructure have established approaches to redundancy: multi-region database replication, cross-cloud object storage, and BGP failover for network connectivity. LLM-dependent workflows need similar planning, but switching models introduces a different kind of compatibility problem.

A second endpoint doesn't guarantee a working service

Cloudflare and AWS outages have familiar failure patterns and documented mitigations. When Claude or ChatGPT becomes unavailable, recovery isn't necessarily as simple as pointing traffic at a secondary service.

Prompts can behave differently across models. Output formats can drift, and function-calling syntax can vary. An application might reach the backup provider successfully and still receive a response it can't use.

An OpenRouter-style abstraction can provide a common interface, but it doesn't remove every dependency on the model underneath. Carefully tuned prompts and evaluations, the tests used to judge model responses, may still be coupled to one provider's behavior. The fallback is a different product, even when its interface looks familiar.

Designing a fallback that can carry traffic

Disaster recovery as a service, or DRaaS, could become a useful category for LLM applications. Its value would depend on whether it can keep workflows functioning across providers, rather than simply offering access to another model.

Self-hosting Llama 3.3 on a couple of H100s is one possible option, but the hardware decision doesn't settle the architectural questions:

  • How portable are the prompts and tool integrations across providers?
  • Can a local model stay ready to handle reduced-function service without consuming full serving capacity around the clock?
  • At what service-level agreement, or SLA, does the cost of running dedicated inference become worthwhile?

Those decisions involve both capacity and acceptable behavior. A degraded mode needs a clear definition of which tasks can continue and which outputs remain usable. Otherwise, a backup may be reachable without being useful.

Accepting the outage risk is also a possible decision. It should be an explicit one. Treating a second provider as tested failover leaves the compatibility work unfinished until the primary service is already unavailable.