Plan B
Pick a made-up incident and ask an AI model for the first three triage steps. Then break the primary model on purpose and watch a backup model, from a different provider, answer instead. Each attempt shows how it ended and how long it took.
Result
Results appear here: each attempt, which model answered, and the answer.
How it works
The page sends only an incident number and a mode. A Pages Function checks the per-visitor limit and the daily caps, then sends one request to Cloudflare AI Gateway with two steps. The gateway tries step 0 (Llama 3.2 3B on Workers AI). If that returns an error or runs past its timeout, the gateway retries the same question on step 1 (Claude Haiku 5.5 on Anthropic). There are no retries on step 0, so the fallback is visible.
- Break it: down asks the gateway for a model that doesn't exist, so step 0 fails with a provider error.
- Break it: slow gives step 0 a 1 ms timeout, so it times out before the model can answer.
- Which model answered comes from the gateway's
cf-aig-stepresponse header when it's there, and otherwise from the shape of the answer (Anthropic and Workers AI reply in different formats). The result says which one it used. - The response cache is skipped, so every run is a real request. The API key for the backup stays on the server.
The two-step request uses AI Gateway's Universal Endpoint format, which Cloudflare now marks as deprecated but still supported. Its replacement, dynamic routing, configures the same kind of fallback in the gateway itself.
What this does and doesn't show
It shows recovery from a failed model request: a bad model id, a provider error or a slow response. The function and the gateway both run on Cloudflare, so if Cloudflare itself had a wide outage, this page, the function and the fallback would all be down together. Plan B is a backup model, not a backup platform.
Surviving a real platform outage takes more:
- Failover outside the failing platform: a client or SDK that can call a second provider directly, or two independent edges (different CDN or cloud) behind DNS with health checks.
- Health-checked routing: synthetic probes that take a provider or region out of rotation before users notice, and put it back slowly.
- Degraded modes: a cached or rule-based answer, a queue for later, or a clear "AI help is unavailable" state, so the product still works without the model.
- Budgets on the backup: the fallback path gets real traffic only during incidents, so give it its own rate limits and spend caps, and test it regularly the way this page does.
Guardrails and cost
- Fixed, fictional incidents and a fixed prompt: no free text from visitors, so nothing to inject.
- About 10 runs per 10 minutes per visitor (a salted, daily-rotating hash of the IP; raw IPs are never stored).
- Daily caps: about 300 runs and 100 backup calls. When a cap is hit, Plan B rests until midnight UTC.
- Answers are capped at 200 tokens. Workers AI runs inside the free daily allowance; a backup call to Claude Haiku 5.5 costs a small fraction of a cent.