FIELD NOTE / 2026.09.213 MIN READ / 5 SOURCES

The Minds Behind Site Reliability Engineering – 7 People Redefining Software

Seven practitioners helped turn large-scale operations into an engineering discipline centered on reliability targets, automation, service-level objectives, and learning from failure.

TL;DR

Seven practitioners helped turn large-scale operations into an engineering discipline centered on reliability targets, automation, service-level objectives, and learning from failure. [1][2]

Why you should read it anyway

SRE changed operations by giving reliability a budget. Instead of demanding impossible perfection, teams define service objectives, measure user-visible reliability, automate repetitive work, and decide how much change velocity a service can safely absorb.

Imagine where Site Reliability Engineering would be without them

Without Google’s SRE lineage, automation and production engineering would still improve, but the widely shared concepts of error budgets, SLOs, toil reduction, blameless learning, and software-engineering approaches to operations would have spread later.

Time Estimate of how many years we would be hindered without them for human progress

Editorial counterfactual estimate: 4–8 years. This is not a measured historical fact. It is an editorial estimate of how much slower the field might have matured without this cluster of people, institutions, practices, and tools.

The 7 people behind Site Reliability Engineering

1. Ben Treynor Sloss

Why they matter: originated the term Site Reliability Engineering at Google and built the organization around the principle that operations problems should be approached with software engineering. SRE introduced explicit reliability goals, automation, and engineering work as alternatives to endlessly scaling manual operations.[1]

2. Niall Richard Murphy

Why they matter: co-edited Google’s Site Reliability Engineering book and helped articulate SRE practices for a broad technical audience. His work helped translate internal operating experience into reusable concepts around reliability, capacity, monitoring, incident response, and organizational design.[2]

3. Betsy Beyer

Why they matter: edited Google’s SRE books and helped turn a large body of distributed operational experience into a coherent discipline that could be learned outside Google. Her editorial and program work helped institutionalize the vocabulary of SRE.[3]

4. Jennifer Petoff

Why they matter: co-edited Site Reliability Engineering and helped document the practices, processes, and management systems surrounding large production services. This was essential because SRE is not only a set of technical mechanisms; it also depends on incentives and team structure.[4]

5. Chris Jones

Why they matter: co-edited the first SRE book and contributed to the documentation of Google’s approach to production engineering. The book connected reliability principles with concrete operating practices across the software lifecycle.[5]

6. Todd Underwood

Why they matter: became a prominent Google SRE leader and educator, especially around reliability culture, human factors, incident response, and learning from failure. His work helped broaden SRE beyond tools into organizational behavior and decision making.[1]

7. Liz Fong-Jones

Why they matter: worked as an SRE at Google and later became a major public educator on SRE, observability, service-level objectives, and production engineering. She helped translate reliability ideas across company boundaries and into the modern cloud-native community.[2]

How they each differ from one another

Treynor Sloss created the organizational model; Murphy, Beyer, Petoff, and Jones codified it in the canonical book; Underwood emphasized reliability culture and human systems; Fong-Jones carried SRE and observability ideas into the broader industry.

Final Take

SRE made reliability negotiable in the best sense: measurable, explicit, and connected to product decisions. That turned operations from an emergency function into a continuous engineering process.

RESEARCH / PROVENANCE

Works Cited

5 SOURCES
  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

CodeHistory is a living archive. Citations document the evidence used for this edition; later evidence may refine the account.

Contribute / Corrections

Improve the record.

Use this moderated submission form to suggest a correction, provide a source, challenge a priority claim or identify a missing contributor. Submissions are treated as research leads, not automatically published comments.

Submit a research lead

Please do not submit confidential material or claims you cannot support.