The Site Reliability Workbook
by Betsy Beyer, Niall Richard Murphy, David K. Rensin, Kent Kawahara, Stephen Thorne · Betsy Beyer, Niall Richard Murphy, David K. Rensin, Kent Kawahara, Stephen Thorne
O'Reilly follow-up to Google's SRE book, focused on implementation: setting SLOs and error budgets, alerting on symptoms, running blameless postmortems, and managing toil and on-call, with case studies from companies adopting SRE outside Google.
This link may earn us a small commission at no extra cost to you. Affiliate disclosure
More resources on Site Reliability Engineering
sre.google
Google's official Site Reliability Engineering resource hub, offering core SRE principles, best practices, and real-world case studies on reliability at scale, including monitoring, incident response, capacity planning, and error budgets.
r/sre Reddit Community
Reddit community where site reliability engineers discuss on-call practice, SLOs, incident response, tooling choices, and career questions. Useful for gauging how teams actually run reliability work and for crowd-sourced opinions on books, certifications, and job moves.
SRE Workbook
Google's companion volume to the SRE book, free to read online, showing how SLOs, error budgets, alerting, and incident response are implemented in practice through worked case studies from Google and other companies.
SRE Book
The full text of Google's Site Reliability Engineering, readable online at no cost. Chapters from Google engineers cover service level objectives, error budgets, monitoring, release engineering, postmortems, and how reliability responsibilities are split with product teams.
Building Secure and Reliable Systems
Google engineers describe how security and reliability requirements interact across a system's life cycle, covering threat modeling, design for least privilege and recovery, secure coding practices, testing, and incident response. Free PDF also available from Google.