Cloud-Ops-Workshop
Hands-on session on observability and incident response for growing platform teams.
What you will learn
- How to design useful SLOs without drowning in metrics
- Building a practical runbook for the first 30 minutes of an incident
- Pairing alerting with ownership so noise stays low
Agenda
- 10:00 — Welcome & context
- 10:20 — Observability patterns that scale
- 11:10 — Break
- 11:20 — Incident response lab (live scenario)
- 12:30 — Retro & next steps
Prerequisites
Basic familiarity with cloud hosting (any major provider) and at least one monitoring tool. No Kubernetes deep-dive required.
Who this is for
SREs, platform engineers, and engineering managers responsible for reliability.