1. Half the internet goes down. The status page stays green.
February 28, 2017. Amazon's cloud storage service — the one quietly holding files for a huge chunk of the internet — breaks for four hours. Slack uploads die. Trello dies. Smart lightbulbs stop listening to their owners.
Best detail: Amazon's own status dashboard showed green the entire time. The dashboard's icons were stored on the thing that was down.
An engineer was following an approved, step-by-step playbook to take a few servers offline for debugging. One value in the command was mistyped, and the tooling cheerfully removed a far bigger set of servers than intended — including ones running two core systems that hadn't been fully restarted in years. They needed a long, careful cold start, and the internet got a four-hour nap.
An authorized human. An approved playbook. One wrong keystroke.
If one typo can take out a thousand times more than you intended, the typo isn't the problem — the missing guardrail is. Cap how much any single command can remove. And never host your status page on the thing it's supposed to report on.









.webp)

.webp)


.webp)

.webp)

.webp)


















