Thinkr

AI

Vaguer requirements tended to produce worse AI code, but not conclusively

A controlled study injected defects into requirements before code generation. Pass rates often fell, yet no effect survived correction, and the authors say so.

Galang Aulia · 6 min read
AI

The intuitive claim is easy to state: give an AI model vague requirements and you get worse code. A new controlled experiment set out to measure it, and the most useful thing about the result is how carefully its authors refuse to overstate it.

The study

The paper comes from a group of requirements-engineering researchers across fortiss, TU Munich, Chalmers, Trinity College Dublin, the University of Duisburg-Essen and Blekinge Institute of Technology. It builds on an earlier study by some of the same authors that looked at how "requirement smells" affect automated traceability between requirements and code.

A requirement smell is a defect in how a requirement is written. The paper groups them into three kinds:

  • Semantic — the meaning becomes ambiguous or context-dependent.
  • Syntactic — the structure obscures things, such as passive voice hiding who acts.
  • Lexical — a word choice loosens precision.

The examples are recognisable to anyone who has reviewed a spec. In one, a precise threshold — points going above 11 — becomes points being "sufficient to win." Still grammatical, still plausible, now open to interpretation.

The setup:

  • Four small game applications — Dice, Arkanoid, Snake and Scopa — with between 14 and 25 requirements each, a fixed Java skeleton, and a unit test per requirement.
  • Smells injected at five increasing densities, plus separate sets for each smell category.
  • Two models, GPT-4o and DeepSeek-V3, at temperature 0, five runs per condition.

What happened

Clean requirements worked well, but not perfectly. With no injected smells, both models passed most requirement-level tests. They still failed some consistently. GPT-4o's average on Scopa was 72.5%, with a 15-point standard deviation across runs; DeepSeek-V3 reached 86.3%.

More smells often meant lower pass rates — often, not always. DeepSeek-V3 on Dice fell from 88.0% to 77.6% at the highest density. On Snake, GPT-4o slipped from 71.4% to 67.1% and DeepSeek-V3 from 64.3% to 52.9%. GPT-4o on Scopa fell from 72.5% to 57.5%, though not in a straight line.

But GPT-4o on Dice fluctuated between 74.4% and 80.0% with no clear pattern. Both models stayed high on Arkanoid at every density. DeepSeek-V3 on Scopa moved between 81.3% and 90.0% without a consistent decline.

No smell category was reliably worse than the others.

And the sentence that governs everything above it, from the authors' own answer to their second research question:

None of the correlations remained statistically significant after correction for multiple comparisons.

What the authors say it means

They call the findings exploratory evidence that requirement smells may influence code generation quality. Not a demonstrated effect.

They are equally careful in the other direction. The study had limited statistical power — only five density levels, and small numbers of requirements per category — so, in their words, a non-significant result should not be read as evidence that there's no effect. The honest summary is suggestive, unconfirmed.

They also note that the smells had more visible descriptive effects on code generation than the earlier study found for traceability, suggesting quality effects may depend on the task. That comparison is flagged as exploratory too.

One thing worth flagging for anyone who reads only the summary: the abstract doesn't mention the significance result. It says higher smell density was "generally associated" with lower correctness. The qualification lives in the results, discussion and threats-to-validity sections. A write-up based on the abstract alone will overstate the finding — which is how most research reaches product teams.

How this squares with Orchid

Readers of this briefing will remember a benchmark we covered in September, Orchid, which found that every model tested degraded on ambiguous requirements, with GPT-4 dropping by more than 30%. This study found declines it couldn't confirm statistically. That looks like a contradiction. It probably isn't.

The designs differ in the way that matters most for significance: scale. Orchid built 5,216 ambiguous variants of 1,304 function-level tasks. This study used four applications with 14–25 requirements each, and its authors say outright that its statistical power was limited. A small experiment that sees the same direction without reaching significance is weak confirmation, not a refutation.

They also measure different things. Orchid tests isolated Python functions; this study tests whole small applications built on a fixed skeleton, where one requirement's ambiguity can be absorbed — or amplified — by the structure around it. That difference is itself a reason the effect might be noisier here.

The fair reading across both: ambiguity probably costs correctness, the size of the cost depends heavily on the task, and the evidence is stronger for small isolated units than for whole systems.

The caveats, from the paper itself

The threats-to-validity section is unusually thorough, and several items bite:

  • Two models only. GPT-4o and DeepSeek-V3 are both older than today's frontier models. The authors note that newer models may resolve ambiguity better, and call that an open question.
  • Small, well-scoped requirements. Games with 14–25 short requirements are not an industrial spec.
  • Injected smells, not natural ones. Designed to resemble real defects, but controlled manipulations nonetheless.
  • Possible training-data exposure. These games are common programming exercises with public implementations, which may have inflated baseline performance.
  • One prompt template. Iterative prompting, reasoning prompts or agent workflows may behave differently.
  • Tests may miss valid alternatives. A correct implementation that takes a different route can fail a unit test written for the expected structure.

The finding worth keeping

The headline question — do vague requirements break AI code? — ended inconclusive. The detail underneath it didn't.

Clean requirements did not guarantee correct code, and the failures weren't random. The authors note they recurred around requirements involving shared functions, ordered dependencies and dynamic state — and that the descriptive decline was most visible in the two games with the most interconnected behaviour. They are clear that task complexity wasn't manipulated, so that link is a hypothesis for now.

For anyone writing specs that a model will build from, that points somewhere specific. The defects the study injected are the ones a line-by-line review catches: an unclear word, a missing actor, a vague threshold. The failures that persisted even without them live between requirements — this one must happen before that one, these two share a rule, this state changes what that action does.

That matches the pattern in a lot of recent work. Clarity of individual sentences is the part of a spec that tools and reviewers check well. Relationships between requirements — sequence, shared state, dependency — are where unnamed states hide and where dependencies go unwritten. A spec can pass a sentence-level review and still leave the model guessing at how its parts connect.

What to take from it

Not "vague requirements don't matter." The study couldn't show that, and says it couldn't.

Two narrower points hold up:

  1. Write precise thresholds anyway. "Above 11" and "sufficient to win" produce different code often enough that the cost of precision is trivially worth it — even if this experiment can't put a significant number on the difference.
  2. Spend review time on how requirements relate, not just on how each one reads. Order, shared rules and state transitions were where both models kept failing, on clean input.

And one point about the paper itself: a study that finds a plausible trend, can't confirm it, and says so plainly in its results, discussion and limitations is rarer than it should be. The field's confidence about AI and requirements quality is running well ahead of its evidence. This paper is a useful reminder of how far.

FAQ

What did it test? Whether injected requirement defects reduce the correctness of LLM-generated code, across four small games and two models.

Did vaguer requirements produce worse code? Often, descriptively. Not significantly after correction.

So quality doesn't matter? The study was underpowered. It couldn't confirm an effect, which isn't the same as ruling one out.

The useful finding? Clean requirements still failed where requirements depend on each other.

Sources

  1. On the Impact of Requirement Smells in LLM-Based Code GenerationarXiv ·
New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing