SRE-201
Intermediate
⏱ 6 weeks
SRE Fundamentals
Run reliable services: SLIs, SLOs and error budgets, monitoring with Prometheus and Grafana, and structured incident response.
Prerequisites: LNX-201 Linux System Administration
Syllabus
Module 1 — Reliability engineering
- SRE vs DevOps
- SLIs, SLOs, SLAs
- Error budgets and toil
Module 2 — Monitoring
- Prometheus architecture
- PromQL
- Grafana dashboards
- Alertmanager
Module 3 — Incident management
- On-call fundamentals
- Incident response process
- Blameless postmortems
Module 4 — Practical reliability
- Health checks and runbooks
- Capacity basics
- Chaos engineering intro