Complete each sentence with ONE phrase from the box.
Ask the teacher: Write down anything you want to remember from this section. Your notes are saved automatically.
Rewrite each sentence in the passive so the SYSTEM is the subject.
Ask the teacher: Write down anything you want to remember from this section. Your notes are saved automatically.
Rewrite each sentence using 'have something done' to signal ownership without claiming the keyboard work.
Ask the teacher: Write down anything you want to remember from this section. Your notes are saved automatically.
Decide whether each incident line is BLAMELESS (describes the system) or BLAME-LOADED (names a villain).
'A configuration change was pushed at 14:02 that removed the health-check.'
'Ravi broke the auth service again.'
'The failover was never tested in production — we caught that in the post-mortem.'
'This wouldn't have happened if ops actually cared.'
'The second-reviewer requirement had been waived for launch week.'
'Whoever wrote this alert should be fired.'
Ask the teacher: Write down anything you want to remember from this section. Your notes are saved automatically.
Rewrite each blame-loaded line as a blameless incident sentence using the passive and, where useful, a role instead of a name.
Ask the teacher: Write down anything you want to remember from this section. Your notes are saved automatically.
You'll hear three engineering leads walk through the same outage. Decide who should own the incident write-up going forward.
🎙️ Three post-mortems — same 47-minute outage — multi-voice
Bianca (Diego):This is really on Ravi. He pushed the config at 14:02 without waiting for a second reviewer. Everyone was under launch pressure, sure, but you don't push to prod on launch day without a review. Action items: written warning for Ravi, more training for the whole team.
Will (Ines):Timeline. At 14:02 a configuration change was pushed that removed the /healthz endpoint. At 14:06 the on-call engineer was paged. At 14:14 the change was rolled back. Service was fully restored at 14:49. Why. WHY 1: the change had been approved with only one reviewer. WHY 2: the second-reviewer requirement had been waived for launch week. WHY 3: the runbook did not classify auth-adjacent changes as always requiring two reviewers. Three action items. One, we'll have the runbook updated — SRE lead, by Wednesday. Two, we'll have the /healthz alerting reconfigured to page within 60 seconds — on-call lead, by Friday. Three, we'll have a monthly failover drill scheduled — Head of Platform, first drill by end of month.
Vera (Kwame):Timeline is what Ines said. We should be more careful with deploys, honestly. We should probably test the failover more often. And, um, we should look at the runbook at some point too. Yeah.
Ask the teacher: Write down anything you want to remember from this section. Your notes are saved automatically.
Pick a real incident from the last 6 months. Write the 4-sentence summary (max 80 words). Sentence 1: timeline in the passive. Sentence 2: root cause, at least three whys deep. Sentence 3: ONE 'have something done' action item with owner and date. Sentence 4: what changes in the RUNBOOK, not what a person will 'try to do better'.
Ask the teacher: Write down anything you want to remember from this section. Your notes are saved automatically.