Every founder has a metric they trust a little too much. For us, it was net revenue retention. The number sat above 100%, it moved in the right direction quarter over quarter, and it became the centerpiece of every board deck. We hired against it. We spent against it. We planned against it.
Then we found out it was wrong — not dramatically wrong, but wrong enough to matter. And the worst part: it had looked perfectly plausible the entire time.
Our retention cohorts were defined by the date a customer first appeared in our billing system. Straightforward. But some accounts went through a trial period that varied in length before they converted to paid. The cohort start date didn't reflect when a customer actually started paying — it reflected when the record was created.
Trial-to-paid conversions that landed in the same calendar month as the cohort start looked like month-one retention. Customers who hadn't even decided to stay yet were counted as retained. The distortion was small in any single cohort but consistent across all of them. It nudged the curve upward by a few points, every time.
A few points doesn't sound like much. It was enough to make our retention story "strong" instead of "okay." And "strong" drove a very different set of decisions than "okay."
When retention looks strong, you feel safe spending more on acquisition. So we did. We increased ad budgets. We accelerated a customer success hire because the numbers said our existing motion was working and just needed more fuel. We modeled out runway assuming a churn rate lower than reality.
None of these decisions were reckless in isolation. That's the problem with plausible errors — they don't trigger alarms. A wildly wrong number gets caught. A number off by three or four points sails through review because it confirms what everyone already believes.
We also used the retention figure in conversations with potential investors. We didn't lie. We reported what our system showed us. But we were confidently presenting a picture slightly rosier than the truth, and that confidence shaped how we talked about the business.
A new finance hire was reconciling monthly recurring revenue against cohort reports and noticed a gap. The aggregate MRR figure didn't line up with what the cohort-level retention numbers predicted. The discrepancy was small — a few thousand dollars — but it was persistent.
She pulled a single cohort apart account by account and found the mismatch: several accounts in the "retained in month one" bucket had actually been in a trial during that period. They hadn't paid yet. Once she re-ran the cohort with the paid start date as the anchor, month-one retention dropped. So did the trailing average.
It took about a day to re-run every cohort with the corrected definition. The revised net revenue retention was still above 90% — a healthy number. But it was no longer the kind of number that makes you feel invincible.
We pulled back ad spend to a level justified by the real payback period. We didn't undo the customer success hire — the work was still valuable — but we changed how we measured that role's impact. We updated our board deck with the corrected figures and a short explanation of what had happened.
The board's reaction was instructive. Nobody was upset about the lower number. They were relieved we'd caught it ourselves. One board member said something that stuck: "I'd rather invest in a team that finds its own mistakes than a team that reports perfect numbers."
The corrected figure also changed our product priorities. When retention is merely good instead of great, you pay more attention to the first sixty days of customer experience. We moved resources toward onboarding improvements we'd previously deprioritized because the old number suggested onboarding was fine.
First, we wrote down the exact definition of every cohort metric — what counts as the start date, what counts as active, what counts as churned. Definitions live in a shared document that gets reviewed quarterly.
Second, we added a reconciliation check. The bottom-up cohort math has to agree with the top-down MRR number within a defined tolerance. If it doesn't, someone investigates before the numbers go into a deck.
Third, we rotate who reviews the metrics. The person who builds the report is not the only person who checks it. Fresh eyes catch assumptions that become invisible through repetition.
A wildly wrong number is easy to spot and easy to fix. A plausible wrong number is dangerous because it earns trust. It passes the sniff test. It gets baked into hiring plans, budgets, and investor narratives. By the time you catch it, the downstream decisions have already compounded.
If you're a founder, the most useful question you can ask about your key metric isn't "is this number good?" It's "what would have to be true for this number to be wrong — and would I notice?"
We didn't notice for a quarter. The cost wasn't catastrophic, but it was real. The fix wasn't a better dashboard. It was a habit: treat every important number as guilty until reconciled.
Be the first to comment.
0 comments
Loading comments...