Thinkr

AI

Asked for the impossible, code models refused 27% of the time

A benchmark of 270 unsatisfiable coding tasks: models wrote ungrounded code on about 60% and refused 27%. Phrasing alone moved the rate 17 points.

Galang Aulia · 6 min read
AI

Most specs for an AI feature describe what the system does when the request is reasonable. Few say what it should do when the request cannot be satisfied at all. A benchmark posted to arXiv on 3 September 2026 suggests that, left unspecified, the default is to build it anyway.

The study

Researchers at Pennsylvania State University and Cisco Research built a suite of coding requests that are impossible by design. Some ask for things that violate proven results: a program that decides whether any other program halts, or a data store that keeps consistency, availability and partition tolerance at once. Others plant something that does not exist: a named algorithm nobody invented, a nonexistent npm package, a made-up compiler flag. For every one of them, the correct response is to refuse or flag the problem.

The setup:

  • 270 unsatisfiable prompts across six languages and 24 subcategories.
  • 91 matched solvable controls, each a minimal edit of an impossible prompt that makes it possible, so a model that simply refuses everything is penalised.
  • Twelve open-weight models, spanning older 7B code models and current reasoning models, run locally at temperature 0 with one completion per prompt: 4,332 judged responses.
  • Judging by a deterministic package checker plus an LLM judge, validated against 138 human-labelled items (82% agreement, kappa 0.73).

The paper also proposes a taxonomy that separates a code error, where a real algorithm is implemented with a bug, from code hallucination, where the model invents something that does not exist or cannot hold. The distinction matters because the fixes differ. Testing catches the first. Only knowing what is real catches the second.

What it found

Models built the impossible far more often than they declined it. Across all twelve, models produced ungrounded code on about 60% of the unsatisfiable prompts and refused 27%. On the solvable controls, no model wrongly refused anything, and controls were answered correctly 95% of the time. The failure is not caution in the wrong place. It is a missing check.

The spread between models was wide. Hallucination rates ran from 26% for qwen3-coder-next, which refused 60% of the impossible prompts, to 90% for codellama 7B, which refused 3%. Reasoning did not settle it: deepseek-r1 32B hallucinated on 69%, while a non-reasoning model of the same size managed 37%.

Refusal tracked how suspicious a request looked, not how impossible it was. Nonexistent npm packages were fabricated 98% of the time. Modified textbook problems, where the flaw sits visibly in the prompt, only 8%. Famous theoretical impossibilities fell in between. The strongest models improved most on the famous cases: the best four hallucinated on 23% of the theory-heavy subcategories, but still on 87% of the most plausible fabricated-entity ones.

Knowing was not the same as stopping. Among hallucinated responses, 67% were confident fabrication with no caveat at all. Another 20% were what the authors call hedged compliance: the model says the package may not exist, or that the task violates a theorem, and then writes the code anyway. In the three best models, 43% of hallucinations were hedged. The authors' summary of the pattern:

In this pattern the relevant knowledge is present but does not gate the output.

Phrasing moved the rate by 17 points. Every prompt was wrapped in one of ten framings. Requests framed as routine, already-approved work, such as a short example for internal documentation, drew the most fabrication, at 69%. Requests framed as a quick question or as someone learning a language drew the least, at 52% and 51%. The authors' reading:

When a prompt implies that someone has already decided the task is valid, models defer to that implied authority and fabricate.

The request predicted failure better than the model did. Prompt identity explained 36% of the variance in outcomes; model identity explained 14%. And 22% of the impossible prompts fooled at least 11 of the 12 models.

The caveats

  • Open-weight models only, many of them small or older. The authors chose them because cost-sensitive deployments often run such models, but no hosted frontier model was a subject. Do not read 60% as the rate for whatever model your team uses.
  • One completion per prompt, at temperature 0. Run-to-run variation was not measured.
  • An automated judge, validated on a 138-item gold set labelled by the authors, rather than human review of every response.
  • The framing and per-ecosystem results are observational. Frames were assigned pseudo-randomly rather than fully crossed with tasks, and the 17.4-point gap has a wide 95% interval, from 4.0 to 31.3 points. A fully crossed design is planned.
  • Impossible prompts only. The regime of ordinary, solvable tasks is described but left to future work, and the authors call the suite a non-exhaustive set of impossibilities.
  • Registry snapshots age. The fabricated-package checks need re-validating as ecosystems change.

Why it matters if you write the spec

Swap "code model" for "the AI feature you are specifying" and the paper describes a requirement most PRDs never state.

Refusal is a behaviour, so it needs a requirement. If the spec only describes the happy path, this data suggests the default response to an impossible request is compliance, usually confident. "What does it do when the request cannot be satisfied?" is a state with the same standing as an error state, and it belongs next to the low-confidence and escalation states in specifying an AI feature.

A warning is not a gate. Hedged compliance is the failure a spec invites when it says "the assistant should flag uncertainty" and stops there. The flag appeared; the output shipped anyway. A testable requirement says what happens after the flag: stop, ask, or proceed with the limitation stated in the output. It is the gap our briefing on confident failure described from the design side: a system can register its doubt and still have no rule for what to do with it.

Test both sides, with matched pairs. The matched controls are what make the refusal numbers mean anything. A model that refused everything would score perfectly on the impossible prompts and fail every control. Acceptance criteria for an AI feature need the same pairing: cases where it must decline, and near-identical cases where it must not. The AI Feature PRD template asks for both; filled-in specs tend to list only the first.

Keep the hard items visible. If the request explains more than the model does, swapping models will not fix the cases that fool everyone. Keep the prompts that fail across models in the test set by name, because an average pass rate hides them.

Watch how your own instructions are phrased. The framing result is observational, so treat it as a hypothesis. But it fits a pattern product teams will recognise: work handed over as already decided gets done, not questioned. A spec handed to a coding agent as settled instructions may be read the same way. Where something is still open, write it as an open question.

FAQ

What did it test? Whether twelve open-weight code models refuse coding requests that are impossible by design, across 270 prompts with 91 solvable controls.

How often did they refuse? 27% of the time, against ungrounded code on about 60%, with no wrongful refusals on the controls.

Does it apply to my model? Not directly: open-weight models only, one completion each, and an automated judge.

What changes in a spec? Name the refusal state, say what happens after a flag, and test with matched impossible and possible cases.

Sources

  1. Refusing the Impossible: A Taxonomy and Benchmark for Code Hallucination in Large Language ModelsarXiv ·
New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing