Site Reliability Engineering (SRE) Practitioner equips participants with the practices, methods, and tools to engage people across the organization around reliability. Real-life scenarios and case stories bring these ideas to life.
Day one opens by naming common SRE anti-patterns, from rebranding Ops as SRE to mob-style incident response and point fixing, so teams can recognize and avoid them. Attendees then define SLOs and SLIs that measure reliability from a user’s perspective, set system boundaries in distributed ecosystems, and use error budgets to make data-driven decisions, including how to handle error thresholds when third-party services are involved. The day closes with building secure and reliable systems: designing for changing architecture, fault tolerance, security, resiliency, scalability, and data privacy.
Day two dives into full-stack observability, covering the pillars of observability, synthetic and end-user monitoring, distributed tracing, and instrumenting applications with libraries and agents. A platform engineering and AIOps module shows how a platform-centric view addresses fragmentation and inconsistency at scale and how AIOps and DataOps improve resiliency. The day finishes with SRE incident response management, applying the OODA loop, closed-loop remediation, swarming, and AI/ML-assisted incident management.
Day three turns to chaos engineering: its origins with Chaos Monkey, designing chaos experiments and GameDay exercises, and applying security chaos engineering. A closing module frames SRE as the purest form of DevOps, covering key principles, metrics for success, and an execution model for scaling reliability across the product spectrum, reinforced with an SRE case study.
Upon completion, participants leave with tangible takeaways, including how to implement SRE models that fit their organizational context, build advanced observability in distributed systems, design for resiliency, and run effective incident responses using SRE practices.
RX-M built the course by drawing on key SRE sources and engaging SRE thought leaders. The team also worked with organizations that have embraced SRE to extract real-life best practices. It teaches the key principles and practices needed to start SRE adoption.
This course also prepares learners to complete the SRE Practitioner certification exam.
Who Should Attend
Business Managers, Business Stakeholders, Change Agents, Consultants, DevOps Practitioners, IT Directors, IT Managers, IT Team Leaders, Product Owners, Scrum Masters, Software Engineers, Site Reliability Engineers (SREs), System Integrators, Tool Providers
What Attendees Will Learn
Upon completing Site Reliability Engineering (SRE) Practitioner, attendees will be able to:
- Identify and avoid common SRE anti-patterns
- Define SLOs and SLIs that meaningfully measure service reliability
- Apply error budgets to make data-driven reliability decisions
- Design secure, fault-tolerant, and resilient systems
- Apply full-stack observability practices, including distributed tracing
- Apply platform engineering and AIOps practices to improve resiliency
- Apply SRE incident response management practices
- Apply chaos engineering practices to test system resilience
- Prepare for the SRE Practitioner certification exam
Prerequisites
RX-M highly recommends that students attend the SRE Foundation course or have equivalent knowledge. Students should also understand common SRE terminology and concepts, along with related work experience.