Sally Roth

SRE Lab

Small, working tools and step-through explainers for reliability engineering. The tools cover reliability targets, telling real outages from blips, and the failures that keep coming back. The explainers cover how the systems actually work, from Kubernetes and Linux to infrastructure as code, delivery, observability, security, incident response and agentic engineering, and what breaks once you run them at scale.

SLO Burn SimulatorSee how fast an error budget drains, which burn-rate alerts fire and when, and get matching Prometheus rules. Down or MeUptime checks with an AI classifier that decides whether a failure is real, and only alerts when it's sure. Plan BBreak an AI model on purpose and watch a backup from a different provider take over through Cloudflare AI Gateway, with which model answered and how long each attempt took. The Incident Files310 real postmortems, searchable by failure mode, plus the nine failure patterns that keep coming back and how companies warded them off. Kubernetes, DrawnHow Kubernetes works in 18 step-through diagrams, from kubectl apply to DNS, storage and upgrades, each with a 30-second answer, gotchas, a quiz and what breaks at scale. Linux, DrawnWhat a Linux node really does with memory and packets: page cache, the OOM killer, cgroups, a packet's path, conntrack, TCP queues and the first 60 seconds of performance analysis. Sane Incident ManagementIncident response people can follow at 3am: severity, roles, the lifecycle, communication, sustainable on-call, blameless reviews, action items that land, and readiness. Taming CardinalityHow one label multiplies series and cost, where Prometheus pays for them, how to find the culprit, and how to drop, aggregate and govern before the bill arrives. Terraform, Five Years OnWhat changed since 1.0: moved, import and removed blocks, checks and tests, the license change and OpenTofu, ephemeral values, and splitting state at scale. Delivery, DrawnFrom commit to production: CI, build once and promote, GitOps with Argo CD, canaries, feature flags, safe schema migrations, supply-chain signing, and merge queues at scale. Observability, DrawnHow telemetry flows: OpenTelemetry and the Collector, Prometheus scraping and PromQL, tracing and sampling, Loki logs, and alerting that wakes the right person. Security for SREs, DrawnHow AWS decides allow or deny, federation instead of keys, TLS 1.3 and mTLS, secrets that rotate, network segmentation, least privilege and audit trails. One Pattern, Three LanguagesEveryday SRE tasks (logs, APIs, AWS cleanup, Kubernetes, Prometheus, fleet checks, exporters) in Python, Go and Ruby, to compare how each language structures the same job, and what breaks at scale. Infrastructure as Code, DrawnRunning IaC well on a team and across an org: remote state and locking, modules, environments, plan and apply in CI with policy checks, drift, safety rails, testing, and scale. Agentic Engineering, DrawnSoftware development at scale in the agentic age: the agent loop, context, verification-first workflows (Lauren Tan's pstack), orchestrating parallel agents, review, safety, agents in ops, and org-wide practice.