Michael Mattis

Cloud-Ops-Workshop

2026-10-09T10:00:00+00:00 · Online

Hands-on session on observability and incident response for growing platform teams.

What you will learn

  • How to design useful SLOs without drowning in metrics
  • Building a practical runbook for the first 30 minutes of an incident
  • Pairing alerting with ownership so noise stays low

Agenda

  • 10:00 — Welcome & context
  • 10:20 — Observability patterns that scale
  • 11:10 — Break
  • 11:20 — Incident response lab (live scenario)
  • 12:30 — Retro & next steps

Prerequisites

Basic familiarity with cloud hosting (any major provider) and at least one monitoring tool. No Kubernetes deep-dive required.

Who this is for

SREs, platform engineers, and engineering managers responsible for reliability.