Thinkr

AI

The best model found 47% of what the experts found

A benchmark of ten models against expert-derived ground truth. Necessity and correctness — the two criteria needing judgment — were almost never detected.

Galang Aulia · 5 min read
AI

We published a post two days ago arguing that automated review catches structural defects and misses anything requiring knowledge of the world. A benchmark published this month measures roughly that, and the number is not flattering to anyone in this category, including us.

The study

Researchers at Virginia Tech and the University of Arizona built expert-derived ground truth for requirement quality using INCOSE criteria, then tested ten models across two families (OpenAI and Anthropic), five generations each, over 100 independent runs, two requirement sets and five sampling temperatures.

The headline result:

The best-performing Anthropic model detects a median of only 47% of expert-identified issues while false-flagging 11%.

The error profile is asymmetric, and the direction matters. These models are comparatively good at not crying wolf and comparatively bad at finding real problems. For a review tool that is the worse of the two failure modes, because a quiet result reads as a clean one.

The part that lines up exactly

Three findings, and the second is the one worth stopping on.

Performance degrades where judgment is required. Specifically, necessity and correctness issues are almost always missed — the authors describe near-zero detection and name both as priority targets for future work.

Those are not arbitrary categories. Necessary asks whether the requirement should exist at all. Correct asks whether it accurately reflects the real need. Both require knowing something the document cannot contain. Every criterion that can be settled by inspecting the text is detected far better than these two.

That is the same line we drew on Sunday from the other direction: an automated reviewer checks the document against itself, and the questions that need the world stay human. It is more useful to have it measured than asserted, and more uncomfortable.

Generational progress is non-monotonic. Newer models were not reliably better at this task. The usual assumption — wait a generation and the gap closes — does not hold in this data.

The variation is characteristic, not stochastic. Error behaviour shifted only modestly and non-monotonically across temperatures, which the authors read as evidence of consistent model deficiencies rather than randomness you could average away.

One consequence they state plainly: single-execution evaluations are statistically unreliable. Run the same review once and you have sampled from a distribution you cannot see.

The warning to anyone building agents

Worth quoting because it cuts against the dominant architectural fashion:

Orchestrating these LLM modules in specialized architectures risks compounding these deficiencies rather than correcting them.

If each module misses the same judgment-dependent issues, chaining eleven of them does not triangulate toward the truth — it produces eleven confident silences about the same blind spot. Specialisation helps where failures are independent. It does not help where they are shared.

What the authors say the study is not

They are notably careful, and the caveats bite.

The results are explicitly a lower bound. The study used off-the-shelf models with no prompt engineering, no few-shot examples grounded in the quality criteria, and no retrieval of context — all of which they name as interventions worth trying. Optimised setups may do considerably better.

Quality assessment was operationalised as binary classification, which does not capture severity, rationale quality, or partial correctness. A model that identifies an issue but explains it poorly scores the same as one that nails it, and vice versa.

The model set is not exhaustive, the temperature range is limited, and locally hosted or organisation-specific deployments are excluded. The authors say the specific values "should not be interpreted as universally representative."

The work was funded by the National Nuclear Security Administration through the Systems Engineering Research Center — which explains the INCOSE framing and is worth knowing. The underlying data is available on request rather than openly.

The caveat we should be loudest about

These are systems-engineering requirements, not product specs.

An INCOSE-style requirement is a single atomic statement judged against formal criteria. A PRD is a document with narrative, context, flows and trade-offs. The tasks are related but not the same, and nobody should transplant 47% onto product spec review as though it were the same measurement.

What transfers is not the number. It is the shape: the criteria requiring judgment about the world were the ones missed, the misses outnumbered the false alarms, and more recent models did not fix it. That shape is structural, and it is the part we would expect to hold.

What we are changing

Reporting this honestly means saying what it implies for our own work rather than only for the field.

Two things are actionable. The variability finding is the sharper one: if single-pass review is statistically unreliable, then the reliability of any review tool is a property of how many looks it takes and how it aggregates them, not of the underlying model. That is a design constraint we should be explicit about rather than quietly benefit from.

The second is about framing. If necessity and correctness are near-zero across the board, then any tool in this space — ours included — should say what it does not check, prominently, rather than letting a clean result imply approval. We wrote a post about that failure two days ago. The paper is a reminder that writing the post is the easy part.

The authors' conclusion is the right one and it is not a hedge:

Their defensible near-term role is human-in-the-loop decision support.

FAQ

The headline number? A median 47% of expert-identified issues detected by the best model, with 11% false-flagged.

What got missed? Necessity and correctness — near-zero detection, and the two criteria that need judgment rather than inspection.

Do newer models fix it? Not reliably; progress across five generations was non-monotonic.

Biggest caveat? These are systems-engineering requirements, not PRDs, and the results are an explicit lower bound on unoptimised off-the-shelf models.

Sources

  1. Two Truths and A Lie? Benchmarking Off-the-Shelf LLMs for Requirements Quality Assessment: Performance, False Alarms, and MissesarXiv
New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing