What Is Site Reliability Engineering (SRE)?
Site reliability engineering applies software engineering to operations — error budgets, SLOs, and automating incident response at scale.
Site reliability engineering (SRE) is a discipline that applies software engineering practices to operations — treating reliability as a problem you solve with code, automation, and measurable targets, rather than with manual runbooks and after-the-fact firefighting. The term and the practice originated at Google, where the founding idea was simple: an operations team staffed by software engineers will, by disposition, automate their way out of repetitive toil instead of accumulating it.
Error budgets: the central idea
The concept that distinguishes SRE from generic “ops” work is the error budget. A team sets a service level objective (SLO) — say, 99.9% of requests succeed over a rolling 30 days — which implies an error budget of the remaining 0.1%: the amount of unreliability the business has explicitly agreed is acceptable.
That budget becomes a shared resource product and engineering both draw from. Plenty of budget left means the team can ship features aggressively, take risks with a rollout, or defer some hardening work. A budget that’s nearly exhausted means the team stops shipping new features and shifts entirely to reliability work — an automatic, pre-agreed circuit breaker that removes the usual tension between “ship fast” and “keep it up,” because both sides agreed on the tradeoff before an incident forced the conversation.
What SRE teams actually do
In practice, SRE work spans a few recurring categories:
- Defining and tracking SLOs. Deciding what “reliable” means for a given service in measurable terms, and building the observability — metrics, logs, and traces — needed to know whether it’s being met.
- Automating toil away. Toil is manual, repetitive operational work that scales linearly with traffic or team size and doesn’t require human judgment — restarting a stuck process, manually rotating a credential, hand-editing a config for every new customer. SRE teams treat toil as a bug: something to automate out of existence, not accept as the cost of running a service.
- Incident response and review. Being on call for production incidents, and — just as important — running a structured, blameless review after each one to find the systemic cause rather than a person to blame.
- Capacity planning and load testing. Making sure a service has the headroom to handle predictable growth and unpredictable spikes before either one becomes an outage.
- Building the platform underneath other teams. Load balancers, deployment pipelines, chaos engineering practices — the infrastructure that makes reliability a property of the platform rather than something every team has to re-solve.
SRE vs traditional operations vs DevOps
Traditional operations teams are often measured by uptime alone and staffed to react to incidents as they happen — a valuable function, but one that tends toward manual intervention as the default fix. SRE reframes the same responsibility around a target (the SLO) and a budget for missing it, and staffs the team with engineers who write code to close the gap between current reliability and the target.
SRE is frequently described as “a specific, prescriptive implementation of DevOps” — DevOps is the broader cultural idea that development and operations shouldn’t be separate silos with a wall between them; SRE is one concrete way to run that combined team, with error budgets, blameless postmortems, and a defined cap on how much time goes to operational toil versus engineering work. A team practicing GitOps or running on Kubernetes with well-defined liveness and readiness probes is applying the same underlying instinct: make reliability an automated, observable property of the system instead of a person’s vigilance.
Runbooks, on-call, and reducing toil
SRE teams still write runbooks — a runbook and an automated fix aren’t in competition, they’re stages of the same process. The typical lifecycle is: an incident happens, it’s diagnosed manually, the fix is documented as a runbook so the next responder doesn’t start from zero, and if the same runbook gets used often enough, it’s a strong signal that step should be automated instead of repeated by a human. A growing shelf of runbooks that never shrinks is itself a symptom worth investigating — it usually means the underlying automation work keeps losing to feature work.
On-call rotations in SRE are also explicitly bounded: a common practice caps the fraction of an engineer’s time spent on operational load (on-call plus toil) at around half, with the rest protected for the engineering work that actually reduces future incidents. Uncapped on-call load is one of the more reliable predictors of burnout and attrition on infrastructure teams.
The takeaway
Site reliability engineering turns “keep it running” into an engineering discipline with a measurable target — the SLO — and a budget for missing it, staffed by engineers who automate away the repetitive work rather than absorbing it manually forever. The core shift from traditional ops isn’t tooling, it’s treating reliability the same way you’d treat any other engineering problem: define what “good” means, measure it continuously, and write code to close the gap instead of doing it by hand every time.
Tagged
Keep reading
Chisato · · 4 min read Incident Severity Levels Explained (SEV1-SEV4)
Incident severity levels rank outages by impact so teams respond proportionally. What SEV1 through SEV4 typically mean and how to set the scale.
Chisato · · 4 min read What Is OpenTelemetry? Traces, Metrics, and Logs
OpenTelemetry is a vendor-neutral standard for instrumenting apps with traces, metrics, and logs — one API, any observability backend.
Chisato · · 5 min read On-Call Rotations and Incident Response Explained
On-call rotations spread responsibility for production incidents across a team on a schedule, paired with a defined incident response process for when alerts fire.