Fewzon logoFewzon
10 min readAnalyticsProduct

Why Most Metrics Are Lying to You

Dashboards feel objective, but they are aggressively editorial. A tour of the specific ways metrics mislead the teams that trust them most.

Metrics have a reputation for objectivity that they do not deserve. They look like numbers. They compare cleanly. They fit in a dashboard. All of these properties invite a specific mistake: treating the metric as a description of reality rather than a compressed opinion about it.

Every metric is an opinion. The choice of what to count is an opinion. The choice of what to ignore is a bigger opinion. The aggregation function is an opinion. The time window is an opinion. By the time a number reaches the dashboard, it has passed through a dozen editorial decisions, and it presents itself as if none of them happened.

This is not an argument against measurement. It is an argument for reading metrics the way you would read an essay — with attention to the author's angle, the author's omissions, and the author's assumptions.

The composition problem

The most common way metrics lie is by composition. An average that looks stable can hide two subpopulations moving in opposite directions. A weekly active user count that grows can hide a churning base offset by aggressive acquisition. A conversion rate that improves can hide a shrinking top of funnel.

None of these are edge cases. All of them are the default behavior of aggregate metrics, and all of them are invisible unless you deliberately decompose. The dashboard will not decompose for you. It will show you the calm surface and let you infer the depth.

The habit that prevents this is asking, of every headline number, "what is the distribution behind this?" If the answer is a shrug, you do not know what the metric means. You know what it looks like.

The proxy problem

Most metrics are proxies for something you actually care about. You care about customer satisfaction; you measure NPS. You care about product quality; you measure crash rate. You care about team health; you measure retention.

The proxy is never the thing. Sometimes the proxy tracks the thing closely, and sometimes it drifts, and the drift is almost always invisible until it matters. NPS can climb while satisfaction falls, because the survey happens to be catching a different slice of users. Crash rate can drop while quality falls, because the code that used to crash now silently corrupts. Retention can hold while team health collapses, because the people staying are the ones with the fewest options.

The proxy is a lens. Lenses distort. Teams that trust their proxies without periodically checking them against the underlying reality get, eventually, a very sharp view of something that isn't quite the thing.

The Goodhart problem

The moment a metric becomes a target, it stops being a measurement. This is Goodhart's Law, and it is more universally true than most people appreciate.

A team measured on ticket close rate will close tickets. Whether they resolve the underlying issues is a separate question, and one the metric will not answer. A team measured on velocity points will hit their velocity points. Whether the points correspond to shipped value is a separate question. A team measured on uptime will report uptime. Whether the uptime is meaningful — whether the system is doing what users needed it to do — is a separate question.

None of this means metrics are useless. It means that any metric you use to evaluate people will become a game they are playing, and you should design the game with that in mind. The best metrics for evaluation are the ones that are hardest to game without also producing the underlying value you cared about in the first place.

The context collapse problem

Dashboards are context collapse machines. They present numbers stripped of the situations that produced them. A 20% drop in a metric can mean disaster or it can mean a competitor launched a promotion or it can mean a holiday. The number, alone, tells you nothing about which.

Good metric practice restores context. Annotations on time series. Notes on anomalies. A written record of what the team believes is happening and why. This is more work than the dashboard invites you to do. It is also the difference between numbers that inform decisions and numbers that decorate them.

What we do at Fewzon

We are, at this stage, small enough that we could pretend we don't need metrics. We haven't. We measure a short list of things that we believe map closely to reality, and we have written down, explicitly, what we believe each metric would look like if it started to drift from the thing it's proxying.

This is more valuable than the metrics themselves. It is a hedge against our own future confidence. When the number moves, we can check whether the underlying reality moved with it, or whether the proxy has quietly broken and the number is just a shadow of itself.