Thinkr

AI

For requirements work, the task decided more than the model did

A two-part study tested LLMs on five requirements tasks. No model or prompt won across all of them, and generating requirements was where fabrication showed up.

Galang Aulia · 5 min read
AI

Every team adopting AI for spec work ends up asking some version of the same question: is a general-purpose assistant enough, or does this need something more specialised? A paper submitted to arXiv on 21 August 2026 tested language models across five requirements tasks and declined to name a winner — which turns out to be the useful answer.

The study

The authors — Jacek Dąbrowski, Manjeshwar Aniruddh Mallya, Alessio Ferrari, Mohammad Amin Zadenoori and Yijun Yu, from Lero at the University of Limerick, Trinity College Dublin, CNR, the University of Padova and the Open University — combine two studies.

Study I was reported earlier by part of the same group; this paper adds Study II to it. Five small open-weight models (Llama 2, Llama 3, Mistral, Gemma and Phi-3 Mini), run on a workstation with a 6 GB GPU, worked on app-store reviews: classifying each review by request type (feature request, bug report, other) and by non-functional requirement type, then turning 90 sampled reviews into a structured requirements specification.

Study II is new. GPT-5.5 and DeepSeek-R1 worked on the Rust language project's public goals, tracking issues and developer chat — nine cases and 80,642 Zulip messages — finding which discussions relate to which goal or issue, and explaining why. Most of that study was carried out during a research visit to Huawei's Ireland research centre. (The paper's result tables label GPT-5.5 as "ChatGPT".)

What they found

Sorting feedback was moderate to good. Classifying reviews by request type reached an average F1 of 0.59–0.68 depending on the prompt, with the best single result 0.74 (Llama 3 with examples in the prompt). Classifying by quality attribute — usability, reliability, performance — was harder: averages of 0.47–0.51 and a best of 0.55. The authors attribute the gap to implicit qualities being harder to read than explicit requests.

Generating requirements was the weak spot, in a specific way. The specifications averaged 3.1 out of 5. They scored well on clarity and completeness (averages of 3.8 and 4.0) and poorly on fidelity to the source feedback (3.0) and conciseness (2.3). The authors describe the text as fluent, but note the models occasionally fabricated requirements that were not in the original feedback — plausible content that wasn't grounded. Prompts that imposed explicit constraints helped little.

Finding links across documents worked well with the larger models. The best result was a precision of 0.91 in the top five candidates, from GPT-5.5 with step-by-step prompting on issue-to-discussion links. Step-by-step prompting raised average top-five precision for the harder goal-to-discussion links from 0.59 to 0.73. The explanations of those links scored between 3.94 and 4.35 out of 5.

No model or prompting strategy won everywhere. The paper's summary table names four different best models across the five tasks, and a different best prompt depending on how much reasoning the task needed.

The part that bears on "general or specialised"

Two findings speak directly to the question.

First, the practical one. In a pilot, the small local models that did reasonably on short review texts were impractical for the long-context traceability work — limited context windows and slow execution. The authors reason that organisations are unlikely to maintain a specialised model for one task, so general-purpose frontier models that cover many tasks "may currently offer the more practical solution" — while flagging vendor lock-in, privacy and running costs as the trade-off.

Second, the one that applies whichever you pick. Across both studies, the tasks that asked a model to generate new requirements text, or to synthesise across many documents, were the ones most prone to unsupported or incomplete output. The authors' recommendation is to treat every output — classifications, links, specifications, explanations — as decision support to be validated, regardless of model family.

The caveats, from the paper

The authors are explicit that this is exploratory work.

  • One evaluator for the specifications. Study I's specification scores came from a single evaluator, and the reliability of those ratings was not validated.
  • No inferential tests in Study I, and a small sample for the generation task, so differences between models there are descriptive.
  • Small models only in Study I. The authors note larger models may behave differently.
  • One ecosystem in Study II. Every case comes from the Rust project, each configuration was run once (72 runs in total), and with no complete ground truth, recall could not be measured. A second author checked a random 10% of judgements, with agreement of 0.83 on links and 0.71 on explanation ratings. The authors' own line: "We do not claim direct generalisability."

One limitation we'd add, which the paper's framing doesn't foreground: no model ran all five tasks. Small models did the classification and generation; frontier models did the traceability. So "performance is task-dependent" is partly confounded with "different classes of model were used for different tasks". The cross-task comparison is suggestive, not a controlled test.

What it means for picking a tool

The question in our post on general assistants versus dedicated PRD tools is usually framed as one decision. This study suggests it is several, split by job.

Finding and sorting — which feedback is a bug, which discussion relates to which goal — produced generally reliable output that the authors say can reduce manual effort, though it still needs checking.

Writing requirements is where fluency and accuracy came apart. A specification that reads cleanly and covers the ground can still contain a requirement nobody asked for. That is the job that needs a reviewer whatever produced the draft, and it is exactly what a seven-check test of any generator is designed to catch: count what was invented, not how polished it looks.

It also matches the other side of the problem. When models are asked to review requirements rather than write them, the judgement-heavy issues are the ones they miss. Generation and review fail in the same place: wherever the right answer depends on facts the text doesn't contain.

So "ChatGPT or a dedicated tool?" is a fair question, but a narrower one than it sounds. The more useful question for each step is: if this output is wrong in a way that reads well, who notices?

FAQ

What did it compare? Language models on five requirements tasks — two kinds of feedback classification, specification generation, traceability links and link explanations.

Which model was best? None across all tasks; the authors recommend choosing models and prompts per task.

Where did models struggle most? Generating requirements — clear-looking output with lower fidelity and occasional fabricated requirements.

How strong is the evidence? Exploratory: a single evaluator for specifications, one ecosystem for traceability, single runs, and different models on different tasks.

Sources

  1. Large Language Models for Requirements Engineering: A Cross-Task Empirical EvaluationarXiv ·
New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing