Site Reliability Engineering (SRE) Foundation introduces the principles and practices organizations use to scale critical services reliably and economically. Introducing a site-reliability dimension requires organizational re-alignment, a new focus on engineering and automation, and the adoption of new working paradigms.
Day one contrasts SRE with DevOps and works through core principles and practices before turning to Service Level Objectives, error budgets, and error budget policies. Attendees then tackle toil directly: what it is, why it erodes team capacity, and how to reduce it. The day closes with monitoring, Service Level Indicators, and observability as the instruments that make reliability measurable.
Day two covers SRE tools and automation, from defining automation and its focus areas through a hierarchy of automation types and secure automation practices. A module on anti-fragility examines why teams should learn from failure and how to shift organizational balance toward it, including on-call practices and blameless post-mortems. The course closes by comparing SRE with other operational frameworks and looking at where the discipline is headed.
The course traces the evolution of SRE and its future direction throughout. It equips participants with the practices, methods, and tools to engage people across the organization around reliability and stability. Real-life scenarios and case stories bring these ideas to life. Upon completion, participants take away practical skills, including how to understand, set, and track Service Level Objectives (SLO’s).
RX-M built the course by drawing on key SRE sources and engaging SRE thought leaders. The team also worked with organizations that have embraced SRE to extract real-life best practices. It teaches the key principles and practices needed to start SRE adoption.
This course also prepares learners to complete the SRE Foundation certification exam.
Who Should Attend
Business Managers, Business Stakeholders, Change Agents, Consultants, DevOps Practitioners, IT Directors, IT Managers, IT Team Leaders, Product Owners, Scrum Masters, Software Engineers, Site Reliability Engineers (SREs), System Integrators, Tool Providers
What Attendees Will Learn
Upon completing Site Reliability Engineering (SRE) Foundation, attendees will be able to:
- Explain SRE principles and practices and how they differ from DevOps
- Define and track Service Level Objectives (SLOs) and error budgets
- Identify and reduce operational toil
- Apply monitoring and Service Level Indicators (SLIs) to service reliability
- Apply SRE tools and automation, including secure automation practices
- Apply anti-fragility principles and learn from failure through blameless post-mortems
- Assess the organizational impact of SRE adoption
- Compare SRE with other operational frameworks
- Prepare for the SRE Foundation certification exam
Prerequisites
Attendees should have an understanding of common DevOps terminology and concepts, along with related work experience.