1. Half the internet goes down. The status page stays green.
February 28, 2017. Amazon's cloud storage service — the one quietly holding files for a huge chunk of the internet — breaks for four hours. Slack uploads die. Trello dies. Smart lightbulbs stop listening to their owners.
Best detail: Amazon's own status dashboard showed green the entire time. The dashboard's icons were stored on the thing that was down.
An engineer was following an approved, step-by-step playbook to take a few servers offline for debugging. One value in the command was mistyped, and the tooling cheerfully removed a far bigger set of servers than intended — including ones running two core systems that hadn't been fully restarted in years. They needed a long, careful cold start, and the internet got a four-hour nap.
An authorized human. An approved playbook. One wrong keystroke.
If one typo can take out a thousand times more than you intended, the typo isn't the problem — the missing guardrail is. Cap how much any single command can remove. And never host your status page on the thing it's supposed to report on.









.webp)


.webp)



.webp)

.webp)
















