Alarm design = deciding which signals are allowed to interrupt a person, and in what order.
Most product teams think they do not have an alarm system. They have a dashboard, a Slack channel and a pager rota. That is an alarm system. Nobody designed it, and it ranks nothing.
Three Mile Island is not the origin story of good alarm design. It is proof that a solved problem can sit unused for thirty five years.
1. The indicator that told a true lie
What happened: At about 4am on 28 March 1979, a relief valve on the Unit 2 reactor stuck open. The light in the control room said it was shut.
Why it failed: The light was wired to the signal sent to the valve, not to the valve itself. It was honest about the command and silent about the world.
Ship instead: Go through your status indicators and ask what each one measures. "Deploy succeeded" often means a request returned 200, not that the thing is running.
2. A hundred alarms, none of them ranked
What happened: The President's Commission found that more than 100 alarms went off in the first few minutes, with no system for suppressing the unimportant ones so operators could concentrate on the significant ones.
Why it failed: Every tile looked like every other tile. An operator on shift said he would have liked to throw the alarm panel away, because it was not giving them useful information.
Ship instead: If two alerts cannot be told apart in one second, you do not have two alerts. You have one, firing twice.
3. The fix was already thirty years old
What happened: In the 1940s Alphonse Chapanis worked out why B-17 pilots kept retracting the landing gear after landing. The flap and gear levers sat close together and felt the same. He put a wheel on one and a triangle on the other. The mistakes stopped.
Why it failed: Shape coding was not invented after Three Mile Island. It was decades old and already regulated in civil aircraft. It had simply never crossed into nuclear control rooms.
Ship instead: Before inventing a way to rank your alerts, read what a neighbouring industry already settled. Aviation, medicine and process control have all published theirs.
4. Your dashboard is failing the same test
NeuBird surveyed 1,039 SRE, DevOps and IT operations people in February 2026. 77% get at least ten alerts a day. 57% say fewer than 30% of them are worth acting on. 83% ignore or dismiss alerts at least sometimes. And 44% had an outage in the past year linked to an alert somebody had suppressed or ignored.
That is the 1979 control room, rebuilt in software, by people who all know the story.
5. "Too many" has a published number
EEMUA 191 and ISA-18.2 put a figure on it. In normal running an operator should get no more than one alarm every ten minutes, roughly 150 a day. In a serious upset, no more than ten in the first ten minutes. Your pager probably breaks both. The useful part is that someone wrote a number down, so the argument stops being about taste.
| The 1979 failure | Your version of it | The fix |
|---|---|---|
| Valve light showed the command | "Deploy succeeded" means a 200 | Show the state, not the request |
| A hundred identical tiles | Forty identical Slack lines | Rank it before it fires |
| No way to suppress noise | Everything pages | Write down what waits until Monday |
| No agreed limit | No agreed limit | One alert per ten minutes |
Run this audit
Take last month's alerts and count how many led to somebody changing something. If it is under a third, you are standing in that control room. Then take your five loudest alerts and write, beside each one, what a person is supposed to do when it fires. The ones where you cannot finish the sentence are the ones to delete.
Get the PDF