Site Reliability Engineering
Listed inSLOs and Error BudgetsObservability & Operationson
The full Google SRE book, free online. The chapters on SLOs, error budgets, and alerting philosophy are the ones to read first.
Running what you built: knowing it is healthy, finding out why it is not, and responding without heroics.
13 articles
Listed inSLOs and Error BudgetsObservability & Operationson
The full Google SRE book, free online. The chapters on SLOs, error budgets, and alerting philosophy are the ones to read first.
Listed inBlameless PostmortemsObservability & Operationson
How blameless review works in practice, what a good postmortem document contains, and how to keep action items from rotting.
Listed inDistributed TracingObservability & Operationson
Spans, context propagation, sampling, and reconstructing one request across a dozen services.
Listed inIncident ResponseObservability & Operationson
Roles, communication, mitigate-before-diagnose, and the timeline you write while it is fresh.
Listed inCost AwarenessObservability & Operationson
Where cloud spend actually goes, attributing it to a team, and the query that costs more than the feature.
Listed inMetricsObservability & Operationson
Counters, gauges, and histograms, cardinality as the thing that breaks your bill, and RED and USE.
Listed inDebugging in ProductionObservability & Operationson
Forming a hypothesis, narrowing with the tools you have, and resisting the fix you cannot explain.
Listed inOn-CallObservability & Operationson
Runbooks, escalation, handover, and a rotation that does not burn the team down.
Listed inLoggingObservability & Operationson
Structured events over prose, levels that mean something, and what must never appear in a log line.
Listed inSLOs and Error BudgetsObservability & Operationson
Choosing indicators users feel, setting a target below 100%, and spending the budget deliberately.
Listed inCapacity PlanningObservability & Operationson
Load testing, headroom, autoscaling limits, and the dependency that does not scale with you.
Listed inAlertingObservability & Operationson
Alerting on symptoms rather than causes, and the pager that only fires when a human must act.
Listed inBlameless PostmortemsObservability & Operationson
Contributing factors instead of a root cause, and action items that someone actually owns.