Most founders rehearse for the single disaster. The database goes down. The payment processor hiccups. A deploy goes sideways on a Friday afternoon. You imagine one thing breaking, you imagine yourself fixing it, and you move on feeling prepared.
What you don't rehearse is the week when three things break at once and none of them are related.
That week happened to us. And the interesting part wasn't the failures. It was the silence.
On Monday morning, a third-party dependency we rely on for a data pipeline started returning errors intermittently. Not a full outage — just enough to make results unreliable. The kind of failure that doesn't set off alarms but slowly poisons downstream work.
On Tuesday, a configuration change deployed the previous week surfaced a subtle issue in how we handled an edge case during tenant onboarding. Two customers noticed before we did. That stings every time.
On Wednesday, an infrastructure provider had a regional disruption. Partial, not total. The worst kind, because partial failures are harder to reason about than clean outages.
Three unrelated problems. Three different systems. One week.
No all-hands war room. No Slack channels full of capital letters. No engineer sleeping under a desk. No founder sending a breathless update to investors at midnight.
Instead, the response looked boring. Each issue was picked up by the person closest to it. They followed the runbook — not because someone told them to, but because the runbook existed and they'd practiced using it. Status updates went to the right channels at the right cadence. Affected customers got specific, honest communication within the hour. Not a template apology — an actual description of what was happening and when we expected resolution.
By Thursday, all three issues were resolved or mitigated. The week continued.
There's a version of this story that sounds heroic. A leader rallies the troops, engineers pull all-nighters, someone has a breakthrough at 2 AM, and the team emerges exhausted but bonded. Founders love telling that story. Investors love hearing it.
But that story is a failure mode.
If your response to simultaneous failures requires heroism, your systems are fragile. Heroes burn out. Heroes make mistakes at 3 AM that they wouldn't make at 10 AM. Heroes create bottlenecks because only they understand the context. A culture that depends on heroes is one bad vacation away from a real outage.
What we saw instead was the result of choices made months earlier, when nothing was on fire. Write the runbook now, before you need it. Define who owns what before there's an incident. Make sure alerts route to the right person, not to everyone. Practice communicating with customers about problems when the stakes are low, so you're not drafting your first incident message during an actual incident.
None of those choices felt important when we made them. That's the point.
You don't know your ops culture is real until it absorbs a shock without drama.
You can write values on a wall. You can run retros and postmortems. You can build dashboards. But the only honest test of operational maturity is an unplanned one — and ideally, several unplanned things at once.
A single failure tests your system. Simultaneous failures test your culture. Do people know their roles without being told? Do they communicate outward without being asked? Do they trust each other enough to stay in their lane instead of piling into one problem?
Our week of failures told us something we couldn't have learned any other way: the process we'd been building quietly, without fanfare, actually held weight.
We still did postmortems. We still found things to improve. The Monday pipeline issue revealed a monitoring gap. The Tuesday onboarding bug pointed to a testing scenario we'd missed. The Wednesday infrastructure event led us to tighten our partial-failure detection.
But the changes were incremental. We were tuning, not rebuilding. That's the difference between a system that failed and a system that bent.
If you're a founder reading this, the takeaway isn't a checklist. It's a question: if three unrelated things broke in your company this week, overlapping — would the response be quiet and competent? Or would it require someone to be a hero?
If the answer is heroes, you have work to do. And the work isn't technical. It's cultural. Writing things down before they're urgent. Practicing response when nothing is wrong. Giving people ownership and trusting them to use it.
The week everything broke was the week we learned the boring work had compounded into something real. Not because we planned the test. Because we couldn't have.
Be the first to comment.
0 comments
Loading comments...