Sally Roth

SRE Lab

Small, working tools and step-through explainers for reliability engineering. The tools cover reliability targets, telling real outages from blips, and the failures that keep coming back. The explainers cover how the systems actually work, from Kubernetes and Linux to infrastructure as code, delivery, observability, security, incident response and agentic engineering, and what breaks once you run them at scale.

SLO Burn SimulatorSee how fast an error budget drains, which burn-rate alerts fire and when, and get matching Prometheus rules.updated Down or MeUptime checks with an AI classifier that decides whether a failure is real, and alerts only when both of its confidence thresholds are met.updated Plan BBreak an AI model on purpose and watch a backup from a different provider take over through Cloudflare AI Gateway, with which model answered and how long each attempt took.updated The Incident Files310 real postmortems, searchable by failure mode, plus the nine failure patterns that keep coming back and how companies warded them off.updated Kubernetes, DrawnHow Kubernetes works in 18 step-through diagrams, from kubectl apply to DNS, storage and upgrades, each with a 30-second answer, gotchas, a quiz and what breaks at scale.updated Linux, DrawnWhat a Linux node really does with memory and packets: page cache, the OOM killer, cgroups, a packet's path, conntrack, TCP queues and the first 60 seconds of performance analysis.updated Sane Incident ManagementIncident response people can follow at 3am: severity, roles, the lifecycle, communication, sustainable on-call, blameless reviews, action items that land, and readiness.updated Taming CardinalityHow one label multiplies series and cost, where Prometheus pays for them, how to find the culprit, and how to drop, aggregate and govern before the bill arrives.updated Terraform, Five Years OnWhat changed since 1.0: moved, import and removed blocks, checks and tests, the license change and OpenTofu, ephemeral values, and splitting state at scale.updated Delivery, DrawnFrom commit to production: CI, build once and promote, GitOps with Argo CD, canaries, feature flags, safe schema migrations, supply-chain signing, and merge queues at scale.updated Observability, DrawnHow telemetry flows: OpenTelemetry and the Collector, Prometheus scraping and PromQL, tracing and sampling, Loki logs, and alerting that wakes the right person.updated Security for SREs, DrawnHow AWS decides allow or deny, federation instead of keys, TLS 1.3 and mTLS, secrets that rotate, network segmentation, least privilege and audit trails.updated One Pattern, Three LanguagesEveryday SRE tasks (logs, APIs, AWS cleanup, Kubernetes, Prometheus, fleet checks, exporters) in Python, Go and Ruby, to compare how each language structures the same job, and what breaks at scale.updated Infrastructure as Code, DrawnRunning IaC well on a team and across an org: remote state and locking, modules, environments, plan and apply in CI with policy checks, drift, safety rails, testing, and scale.updated Agentic Engineering, DrawnHow Stripe, Spotify, Shopify and others run hundreds of coding agents and ship safely: harnesses the model can't skip, verifiers it can't edit, no route to prod, risk-tiered review, and Lauren Tan's pstack.updated Ownership at ScaleResearch on how very large companies decide who owns code and services: OWNERS files and CODEOWNERS, service registries synced from HR, routing pages, incidents, vulnerabilities and costs to owners, keeping it accurate through reorgs, and fleet-wide changes that route around owners.updated