Site Reliability Engineering (SRE) Practitioner

}
3 Days

Available On-Site

Available Virtually

Open Enrollments Available
f

Customizable

Site Reliability Engineering (SRE) Practitioner equips participants with the practices, methods, and tools to engage people across the organization around reliability. Real-life scenarios and case stories bring these ideas to life.

Day one opens by naming common SRE anti-patterns, from rebranding Ops as SRE to mob-style incident response and point fixing, so teams can recognize and avoid them. Attendees then define SLOs and SLIs that measure reliability from a user’s perspective, set system boundaries in distributed ecosystems, and use error budgets to make data-driven decisions, including how to handle error thresholds when third-party services are involved. The day closes with building secure and reliable systems: designing for changing architecture, fault tolerance, security, resiliency, scalability, and data privacy.

Day two dives into full-stack observability, covering the pillars of observability, synthetic and end-user monitoring, distributed tracing, and instrumenting applications with libraries and agents. A platform engineering and AIOps module shows how a platform-centric view addresses fragmentation and inconsistency at scale and how AIOps and DataOps improve resiliency. The day finishes with SRE incident response management, applying the OODA loop, closed-loop remediation, swarming, and AI/ML-assisted incident management.

Day three turns to chaos engineering: its origins with Chaos Monkey, designing chaos experiments and GameDay exercises, and applying security chaos engineering. A closing module frames SRE as the purest form of DevOps, covering key principles, metrics for success, and an execution model for scaling reliability across the product spectrum, reinforced with an SRE case study.

Upon completion, participants leave with tangible takeaways, including how to implement SRE models that fit their organizational context, build advanced observability in distributed systems, design for resiliency, and run effective incident responses using SRE practices.

RX-M built the course by drawing on key SRE sources and engaging SRE thought leaders. The team also worked with organizations that have embraced SRE to extract real-life best practices. It teaches the key principles and practices needed to start SRE adoption.

This course also prepares learners to complete the SRE Practitioner certification exam.

Who Should Attend

Business Managers, Business Stakeholders, Change Agents, Consultants, DevOps Practitioners, IT Directors, IT Managers, IT Team Leaders, Product Owners, Scrum Masters, Software Engineers, Site Reliability Engineers (SREs), System Integrators, Tool Providers

What Attendees Will Learn

Upon completing Site Reliability Engineering (SRE) Practitioner, attendees will be able to:

  • Identify and avoid common SRE anti-patterns
  • Define SLOs and SLIs that meaningfully measure service reliability
  • Apply error budgets to make data-driven reliability decisions
  • Design secure, fault-tolerant, and resilient systems
  • Apply full-stack observability practices, including distributed tracing
  • Apply platform engineering and AIOps practices to improve resiliency
  • Apply SRE incident response management practices
  • Apply chaos engineering practices to test system resilience
  • Prepare for the SRE Practitioner certification exam

Prerequisites

RX-M highly recommends that students attend the SRE Foundation course or have equivalent knowledge. Students should also understand common SRE terminology and concepts, along with related work experience.

Delivery

Available for Instructor-Led (ILT) in-person/onsite training or Virtual Instructor-Led training (VILT) delivery.

Each attendee will require the ability to ssh into a cloud hosted virtual machine (provided with the course). In environments where SSH is not possible, local lab VMs or browser accessible lab systems can be provided. For web-based delivery, participants require an Internet-connected computer capable of teleconferencing.

Related Instructor-Led (ILT & VILT) Training Courses

DevOps Leader (2 Days)

JIRA (1 Day)

Systems Thinking (2 Days)

If you are interested in other Cloud Native, AI, programming, or other courses, check out the full course list or search our entire catalog:

Secret Link