Site Reliability Engineering Foundations
0 of 18 lessons complete (0%)
Exit Course
Stage One — What the Discipline Actually Is
Reliability as an Engineering Problem
Preview
Toil and Why It Is the Central Metric
How the Function Fits the Organisation
3 lessons
Stage Two — Service Level Objectives and Error Budgets
Indicators Objectives and Agreements
Choosing Indicators That Reflect User Experience
Error Budgets and Using Them
3 lessons
Stage Three — Observability
Metrics Logs and Traces
Instrumentation and Cardinality
Dashboards and Alerts People Trust
3 lessons
Stage Four — Incident Response
Roles and Running an Incident
On-Call That Is Sustainable
Blameless Review and Follow-Through
3 lessons
Stage Five — Eliminating Toil
Finding and Prioritising Automation
Runbooks and Self-Healing
Change Management and Safe Deployment
3 lessons
Stage Six — Capacity Resilience and Making It Stick
Capacity Planning and Load Testing
Resilience Patterns and Failure Testing
Embedding the Practice and Career Path
3 lessons
Stage Four — Incident Response
Blameless Review and Follow-Through
You don’t have access to this lesson
Please register or sign in to access the course content.
Take course
Sign in
Previous
Next