Kubernetes Day 2 Operations gives Kubernetes operators a deeper look at the platform and the tasks involved in maintaining and running a production cluster. The course runs across three hands-on days. It picks up where Kubernetes Foundation leaves off, building attendees’ day-two operations knowledge through lecture, demonstrations, and extensive lab work. Day one opens with an architecture review from a day-two perspective, and attendees provision a cluster of their own. They then practice expanding and shrinking both the Kubernetes cluster and its underlying etcd cluster. This involves manually adding and removing members while the cluster stays online.
Day two turns to observability, starting with tracing activity using OpenTracing and Jaeger to follow a request as it moves across services. The day then moves into monitoring with OpenMetrics and Prometheus to track cluster and workload health over time. It closes with log management, where attendees assemble a logging pipeline from Elasticsearch, FluentBit, and Kibana. This lets them search and correlate events across every node in a cluster.
Day three closes out the course with cluster upgrades, walking through upgrading worker and master nodes with minimal downtime. Attendees then practice backing up and restoring clusters, replicating a production environment to test changes safely before they reach live traffic. The final module puts everything together in troubleshooting scenarios that mirror real production incidents. These give attendees practice diagnosing problems under realistic time pressure. By the end of the course, attendees can keep production Kubernetes clusters healthy, observable, and recoverable. They can do this no matter what a given incident throws at them.
Who Should Attend
K8s operators, IT and SRE Staff
What Attendees Will Learn
Upon completing Kubernetes Day 2 Operations, attendees will be able to:
- Expand and shrink Kubernetes and etcd clusters safely
- Implement observability and distributed tracing for production clusters
- Monitor Kubernetes with OpenMetrics and Prometheus
- Manage Kubernetes logs with Elasticsearch, FluentBit, and Kibana
- Upgrade worker and master nodes with minimal downtime
- Back up and restore Kubernetes clusters for disaster recovery
- Troubleshoot common production Kubernetes issues
Prerequisites
The RX-M “Docker Foundation” and “Kubernetes Foundation” courses or equivalent knowledge.