SRE Lab · failover

Plan B

Pick a made-up incident and ask an AI model for the first three triage steps. Then break the primary model on purpose and watch a backup model, from a different provider, answer instead. Each attempt shows how it ended and how long it took.

1 · Incident (fictional)
2 · Primary model

Up to 10 runs per 10 minutes per visitor · shared daily cap

Result

Results appear here: each attempt, which model answered, and the answer.

How it works

Request path The page posts to a Pages Function, which calls AI Gateway with two steps. Step 0 goes to Workers AI. If it errors or times out, the gateway falls back to step 1, Anthropic. Everything except Anthropic runs on Cloudflare. runs on Cloudflare This pageincident + mode, no keys Pages Function/api/planb · limits · caps AI Gatewaysteps [0, 1] · no cache step 0 Workers AIllama-3.2-3b-instruct error / timeout step 1 Anthropicclaude-haiku-5-5

The page sends only an incident number and a mode. A Pages Function checks the per-visitor limit and the daily caps, then sends one request to Cloudflare AI Gateway with two steps. The gateway tries step 0 (Llama 3.2 3B on Workers AI). If that returns an error or runs past its timeout, the gateway retries the same question on step 1 (Claude Haiku 5.5 on Anthropic). There are no retries on step 0, so the fallback is visible.

The two-step request uses AI Gateway's Universal Endpoint format, which Cloudflare now marks as deprecated but still supported. Its replacement, dynamic routing, configures the same kind of fallback in the gateway itself.

What this does and doesn't show

It shows recovery from a failed model request: a bad model id, a provider error or a slow response. The function and the gateway both run on Cloudflare, so if Cloudflare itself had a wide outage, this page, the function and the fallback would all be down together. Plan B is a backup model, not a backup platform.

Surviving a real platform outage takes more:

Guardrails and cost