Thinkr

Trends

The benchmark utility gap has three named failure modes

A paper on AI evaluation names three ways a rising score hides a flat outcome. All three have exact equivalents in the metrics section of a PRD.

Galang Aulia · 4 min read
Trends

A paper published in May makes an argument about AI evaluation that turns out to be an argument about product metrics generally, and it supplies something the product conversation has been missing: names for the three specific ways a rising number can accompany a flat outcome.

What the paper says

Ishani Mondal and Shweta Bhardwaj identify what they call a benchmark utility gap — generative AI systems achieving strong benchmark performance while failing to deliver real-world usefulness — across 28 deployment cases spanning education, healthcare, software engineering and law.

Their diagnosis is that this is not primarily a model problem. It is an evaluation-practice problem, and it reduces to three recurring failures:

Proxy displacement. The measurement stands in for something it cannot fully capture, and then the proxy becomes the target. Optimisation continues, the proxy keeps improving, and the link to the thing it was standing in for quietly breaks.

Temporal collapse. The outcome is a trajectory — capability developing through sustained use — and the evaluation is a snapshot. A single measurement at one moment cannot see whether anything compounded.

Distributional concealment. Aggregate performance conceals the distribution underneath it. The average is fine; some group the system consistently fails is invisible inside it.

Their proposed alternative centres on what they call the missing construct — utility, defined as "the change in a stakeholder's capability induced through sustained interaction with an AI system within a deployment context" — operationalised through a four-stage framework, SCU-GenEval: stakeholder-goal mapping, construct-indicator specification, mechanism modeling, and longitudinal utility measurement.

What it is and is not

This is a position paper, and worth reading as one. It analyses 28 deployment cases and proposes a framework; it does not run that framework and measure whether it produces better decisions. There is no controlled comparison and no effect size, because that is not what the paper set out to do.

It is also an arXiv preprint (v2, revised 11 May 2026), not peer-reviewed work.

The three failure modes are the durable contribution. The framework is a proposal, and the honest way to read it is as an argument for a change in practice rather than evidence that the change works.

The part that transfers

Here is why this matters outside AI evaluation. Take the three names and apply them to the metrics section of an ordinary product spec:

Proxy displacement is the vanity metric. You cannot measure "users understood the feature," so you measure clicks on it. Then clicks become the target, the team optimises the entry point, clicks rise, and nobody learns anything about understanding. The proxy was reasonable when it was chosen. It stopped being reasonable the moment it became the goal.

Temporal collapse is the single-point measurement. "Adoption at 30 days" is a snapshot of something that is actually a curve. A feature that 40% of users try once and abandon and a feature that 40% adopt permanently produce the same number on the same day.

Distributional concealment is the denominator problem. Median time-to-outcome improved 20% — for whom? Enterprise workspaces with 200 members may have got dramatically faster while small teams got slower, and the aggregate reports success. This is the one that most often survives a launch review intact, because the number that was promised is the number that was delivered.

Three failure modes, all of which appear in specs that nobody would describe as badly written.

What to do with it

The vocabulary is the useful part, and vocabulary is not nothing — it is much easier to object to "this might be temporal collapse" than to "I don't know, it feels like this metric won't tell us much." The first is a specific claim about a mechanism. The second sounds like reluctance.

Three questions that fall out of the three names, usable on any metrics section:

  • Is this the outcome, or a stand-in for it? If a stand-in, what would tell you the link had broken?
  • Is this a moment or a trajectory? If the value compounds with use, a 30-day snapshot cannot see it.
  • Whose result does this average contain? Which cohort could be getting worse inside a number that is getting better?

None of these requires the paper's framework, which is a substantial thing to adopt. They require noticing that "the number went up" and "the thing we wanted happened" are different claims, and that most specs only make the first one checkable.

The broader point

The AI field is currently discovering, with unusual rigour and at considerable expense, something product teams have known informally for a decade: a measurement that is easy to compute will beat a measurement that is correct, unless somebody actively defends the second one.

What is genuinely new here is the precision. "Vanity metric" has been a useful accusation and a vague one. Proxy displacement, temporal collapse and distributional concealment are three separable mechanisms with different remedies — and a metrics section can fail any one of them independently.

That is a better tool than the word it replaces.

FAQ

What is the benchmark utility gap? Strong benchmark scores alongside weak real-world usefulness, identified across 28 deployment cases.

What are the three failure modes? Proxy displacement, temporal collapse, and distributional concealment.

Is it empirical? No — a position paper proposing a framework, on arXiv, not peer-reviewed.

Why should a PM care? All three describe ordinary metrics sections, not just AI benchmarks.

Sources

  1. Benchmarked Yet Not Measured — Generative AI Should be Evaluated Against Real-World UtilityarXiv (Mondal & Bhardwaj) ·
New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing