Thinkr

Craft

Vanity Metrics in Product Specs

A vanity metric is not a fake number. It is a real one that can only go up — which is why it passes review and settles nothing afterwards.

Galang Aulia · 5 min read
Craft

A vanity metric is not a fake number. That is why it survives review.

It is a real, correctly measured number that happens to be incapable of telling you whether your work succeeded — usually because nothing that could realistically happen would make it fall. It goes up when the feature works, up when it does not, and up when nobody notices it shipped.

The test

One question, applied to any metric in a spec:

What would this number do if the feature failed completely?

If the honest answer is "go up slightly slower," you have a vanity metric.

Cumulative signups rise because time passes. Total page views rise because marketing is running. Time-in-product rises when people get value and also when they get lost. None of these has a failure case you could point at in six weeks, which means none of them can settle the question they were put in the document to settle.

How they get in

Not stupidity. Four mechanisms, and knowing which one you are in changes the fix.

It is the number you already have. Instrumenting the real metric is work; the vanity one is already on a dashboard. Under deadline, available beats correct.

It explains easily upward. "Signups are up 20%" lands in a leadership review. "Median time-to-first-critique fell from 7.4 to 3.1 minutes for workspaces created after the change" needs a sentence of setup, and the person presenting knows it.

It was in the template. The field said Success metric and somebody filled it with the most obvious number the product produces.

It is genuinely correlated with success. The hardest case. Engagement really does rise when products get better. It also rises for six other reasons, which is why the correlation cannot be converted into evidence about your change.

The usual suspects

VanityWhat it hidesReplace with
Total / cumulative signupsTime passingActivated users in a launch cohort
Page viewsTraffic source changesTask completion rate
Time in productValue and confusion, identicallyTime to the outcome, measured downward
Feature clicksOne-time curiosityRepeat use within 30 days
Users who "tried" itNothing about whether it helpedUsers who completed the job it exists for
NPS on 40 responsesSampling noiseA specific behaviour you predicted would change

The right column is not more rigorous by being more complicated. It is more rigorous by having a direction it can move that would mean this did not work.

The seductive one

Engagement time deserves its own paragraph because it is the metric most often defended.

More time in your product can mean people are getting value. It can equally mean they cannot find what they need. The number is identical in both cases, and which story gets told depends on who is presenting it.

For most tools — and certainly for anything people use to get a job done — reduced time to outcome is the win. A spec that targets "increased session length" for a document-review tool has written down a goal that is satisfied by making the tool slower to use.

If engagement genuinely is your success condition, say why in the spec. The cases where it is legitimate exist, and they are rarer than the metric's popularity suggests.

Why they pass review

Two reasons, both social rather than analytical.

They are unfalsifiable, so there is nothing to argue with. A reviewer who cannot construct a scenario in which the metric fails also cannot construct an objection. The section passes untouched, and passing untouched reads as agreement.

They flatter. Challenging a number that is going up requires saying, out loud, that the good news is not evidence. That is an uncomfortable thing to be the person who says, especially about somebody else's launch, and especially in a room that has already decided it went well.

This is why the question to ask in review has to be mechanical rather than judgmental. "What would this do if it failed?" is a question about the metric. "Is this a vanity metric?" is a question about the person who wrote it, and gets answered defensively.

Activity versus outcome

The shortest way to separate them:

Activity metrics describe what the product did. Views, clicks, sessions, time, volume.

Outcome metrics describe what changed for someone. They complete something, faster, more often, with fewer attempts, or they stop needing help with it.

Almost every vanity metric is an activity metric that got promoted. The replacement is usually the same event, re-anchored to the user's goal rather than the product's surface — not "views of the critique screen" but "critiques completed within a day of starting a spec."

When it is fine

Tracking a vanity metric is not a sin. Making it the success criterion is.

Total signups is a perfectly good health signal — you want to know if it falls off a cliff. Engagement is a reasonable guardrail for a feature you expect not to reduce it. And a leading indicator is legitimate provided it is labelled as one: "clicks on the new entry point, as an early read while we wait for 30-day completion data."

What fails review is the same number in the Success Metric field with a target attached, where its only real property is that it was going to rise anyway.

FAQ

What is a vanity metric? A real number with no failure case — it rises whether or not your work succeeded.

How do I spot one? Ask what it would do if the feature failed completely. "Rise more slowly" is the tell.

Is tracking one ever fine? Yes, as a health signal or a labelled leading indicator. Not as the success criterion.

What replaces engagement time? Usually time-to-outcome, measured downward. More time is equally consistent with value and confusion.

New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing