Most bad weeks have a single root cause. You trace the thread, find the knot, untie it, move on. Our worst week was not that kind of week. Three unrelated problems arrived inside five days, each compounding the last, none sharing a root cause.
We are not publishing a postmortem here. We already did that privately with affected customers, and that is where the technical detail belongs. This post is about the human side: what it felt like, where our instincts were wrong, and what we built afterward to make sure the next bad week is just bad — not compounding.
Monday morning, alerts fired for elevated error rates on a subset of tenant workloads. Not dramatic. Not pager-worthy at first glance. The kind of thing that looks like a flicker until it doesn't resolve.
By noon we knew it was real. A configuration change from the previous Friday had introduced a subtle regression. The damage was narrow — a handful of accounts — but it had been accumulating over the weekend.
Here is the part that stung: we had a review step for that class of change. It was documented. It existed. Nobody skipped it. The check simply didn't catch this failure mode, because nobody had imagined it.
Wednesday brought a completely separate issue — an infrastructure-level disruption outside our direct control that degraded availability for a wider set of customers. We had redundancy in place, but the failover behavior didn't perform the way our runbooks assumed it would.
Two incidents open at the same time. Different owners. Different channels. Customers seeing our status page update twice in one week and drawing the reasonable conclusion that something systemic was wrong.
That conclusion was incorrect, but it didn't matter. Perception compounds faster than root causes.
Friday, running on short sleep and long Slack threads, a team member pushed a remediation for the Monday issue that briefly broke something else. Caught in minutes. Rolled back in minutes. But the damage to internal morale was immediate. The person who shipped it went quiet.
That silence was the most dangerous thing that happened all week. Not the outages.
There is a specific fatigue that hits when problems overlap. Each incident on its own is manageable. You have a playbook, you have people, you work the problem. When they stack, the emotional math changes. Confidence drops. Decision-making slows. People second-guess things they normally trust.
We noticed three failure patterns in ourselves:
Over-communication disguised as action. The volume of internal updates went up while the rate of decisions went down. We were narrating the problem instead of solving it.
Hero impulse. A couple of people tried to work both incidents simultaneously. They made neither better.
Blame avoidance as a reflex. Nobody was pointing fingers, but everyone was pre-defending their decisions. That is a sign the team feels unsafe, even when leadership hasn't done anything threatening.
We made four changes. None were technical fixes for the specific failures — those got addressed in the first week. These were structural decisions about how we operate under stress.
A circuit breaker on concurrent incidents. If two unrelated incidents are open, a designated coordinator steps in whose only job is resource allocation. They don't work either problem. They make sure the right people are on the right problem and nobody is splitting attention.
A pre-written "bad week" communication template. When multiple issues hit the same window, customers deserve a framing message: these are separate, here is what we know, here is what we don't. We now have that draft ready before we need it.
Explicit stand-down rules. After 14 hours of incident work, you are off the rotation. Not optional. Not heroic. Sleep is an operational input.
A blameless debrief with a feeling check. We already did blameless postmortems on the technical side. We added a section at the start where each person says one sentence about how the week affected them. That sounds soft. It is the hardest part and the most useful. The Friday silence taught us that.
Founders love the mythology of grit. Push through. Stay up. Ship the fix. That story works for a single hard night. It does not work for a compounding week.
Resilience is not about toughness. It is about having practices already in place so that when the second problem hits, you don't collapse the response to the first.
You cannot install these practices during the crisis. That is the whole point. You install them now, when things are calm, and you feel a little silly writing a playbook for a scenario that seems unlikely.
The week will come. The only question is whether you've built the structure or you're building it live.
Be the first to comment.
0 comments
Loading comments...