Sally Roth

SRE Lab

Small, working tools and explainers for reliability engineering: setting reliability targets, knowing when something is really broken, and learning from the failures that keep coming back. More on the way: incident management people can follow, metric cardinality, Kubernetes drawn out, and the same patterns in Python, Go and Ruby.

SLO Burn SimulatorSee how fast an error budget drains, which burn-rate alerts fire and when, and get matching Prometheus rules. Down or MeUptime checks with an AI classifier that decides whether a failure is real, and only alerts when it's sure. The Incident Files310 real postmortems, searchable by failure mode, plus the nine failure patterns that keep coming back and how companies warded them off.

Coming next