sre.google
Unknown
Google's official Site Reliability Engineering resource hub, offering core SRE principles, best practices, and real-world case studies on reliability at scale, including monitoring, incident response, capacity planning, and error budgets.
More resources on Site Reliability Engineering
r/sre Reddit Community
Reddit community where site reliability engineers discuss on-call practice, SLOs, incident response, tooling choices, and career questions. Useful for gauging how teams actually run reliability work and for crowd-sourced opinions on books, certifications, and job moves.
SRE Workbook
Google's companion volume to the SRE book, free to read online, showing how SLOs, error budgets, alerting, and incident response are implemented in practice through worked case studies from Google and other companies.
SRE Book
The full text of Google's Site Reliability Engineering, readable online at no cost. Chapters from Google engineers cover service level objectives, error budgets, monitoring, release engineering, postmortems, and how reliability responsibilities are split with product teams.
The Site Reliability Workbook
O'Reilly follow-up to Google's SRE book, focused on implementation: setting SLOs and error budgets, alerting on symptoms, running blameless postmortems, and managing toil and on-call, with case studies from companies adopting SRE outside Google.
Building Secure and Reliable Systems
Google engineers describe how security and reliability requirements interact across a system's life cycle, covering threat modeling, design for least privilege and recovery, secure coding practices, testing, and incident response. Free PDF also available from Google.