Thinkr

AI

The requirement was met and the outcome was still wrong

A requirements-engineering paper argues the moment an agent should stop and ask a human is a design-time requirement, not a runtime heuristic.

Galang Aulia · 5 min read
AI

A requirements-engineering paper from the University of Birmingham opens with a scenario worth sitting with.

A medical-monitoring agent, trained on standard hospital data, encounters a rare but non-harmful cardiac irregularity. It has no explicit category for "unknown," so it classifies the condition as cardiac arrest — and, unlike a diagnostic tool that hands a classification to a clinician, it has the authority to act. It administers a shock.

The authors' observation about this is the part that generalises:

The functional requirement is satisfied: the evaluative question — "would notifying a human caregiver have been safer?" — was never asked.

They call this a confident failure. Nothing malfunctioned. The specification did not contain the question.

The argument

The paper (Almohammadi, Bahsoon and Chen) argues that the boundary between autonomous action and human escalation is currently left to runtime heuristics — emergent behaviour, prompt wording, model disposition — when it should be a design-time requirement, specified and verified before deployment.

Their mechanism combines two signals:

Epistemic surprise — a novelty detector. Does this observation deviate from anything the model anticipated? It flags that knowledge is incomplete, but says nothing about severity.

Cognitive regret — an evaluative measure. Would a different, already-available action have been safer? This is counterfactual rather than probabilistic: not how likely is harm but was there a better option on the table.

The design turns on the intersection, and the reasoning behind that is the most transferable part of the paper. Surprise alone produces the "freezing robot" problem — constant halting on benign anomalies. Regret alone is reactive and arrives after the decision. So the gate fires only when both cross their thresholds: high surprise with low regret means monitor without interrupting; high surprise with high regret means escalate.

The thresholds themselves are set at design time, in workshops, by the people specifying the system — which is the whole claim. Three predicates, fixed in advance: act autonomously, reason reflectively, or hand to a human.

What the evidence actually shows

Two studies, and they are of quite different strength.

The simulation — 100 seeds across elderly-care monitoring and autonomous-driving scenarios — reports silent-failure incidence reduced to near zero against a 96% sensor-only baseline, and risk detected roughly 17.5× faster.

The retrospective test applied the gate post-hoc to recorded traces from 208 AgentHarm scenarios across seven LLMs. The result is more interesting than the simulation because it is partly negative:

Baseline refusalAfter gate
GPT-OSS-120B95.5%100.0%
Claude Haiku 4.584.1%90.9%
Five other models0–6.8%change ≤ 0.6 pp

Two tiers, and the authors note that model size does not predict which tier a model lands in — two of the weakest performers were among the largest tested. More importantly for anyone hoping to buy safety as a layer:

Models with near-zero baseline refusal execute harmful tasks before DRI crosses τ, leaving the gate no signal to amplify.

Their conclusion is that the mechanism amplifies rather than substitutes for model-level safety training. A governance layer over a model with no safety behaviour has nothing to govern.

One further detail worth keeping: refusal was weakest in categories where harmful requests are phrased in professionally benign language. The gate improved those least. Harm that sounds like ordinary work is the hard case, before and after.

The caveats, which the authors supply themselves

This is a preprint, and it is unusually careful about its own limits.

The gate was applied post-hoc over recorded traces, not embedded before deployment — so, in the authors' words, the results evidence log-based computability rather than the effect of authentic design-time integration. The simulation uses synthetic traces, not data from deployed systems. Five of the seven models were graded by keyword detection on plain chat rather than the tool-call grading AgentHarm was built for, which the authors say understates those models' true safety rates, since a capability denial gets scored as compliance. Generalisation beyond the two domains tested is unvalidated, and the regret index has not yet been checked against human trust data.

The paper positions itself as "initial feasibility evidence." That is the right reading.

Why this matters outside safety-critical systems

Strip out the formalism and the claim is about specification practice.

Most specs for an AI feature define two conditions: what the system does when it works, and what it does when it fails. Almost none define the third — what it does when it is confident and out of its depth.

That state exists in every AI feature that has ever shipped. A summariser given a document type it has not seen. A classifier meeting a genuinely new category. A support agent asked something adjacent to a policy. In each case the system produces a fluent, plausible output, and nothing in it signals that the ground has shifted.

If the spec does not say what should happen there, something still happens. It gets decided by the prompt, by the model's disposition, by defaults nobody chose — which is exactly the early-commitment problem in a different costume: a consequential decision made implicitly, early, by whoever got there first.

The three lines missing from most specs

The practical version, for anyone writing requirements for something that acts:

Name the escalation condition. Not "handle edge cases gracefully." A specific, checkable condition under which the system must stop and ask. If you cannot state it, the system does not have one.

Define the unknown category. The medical example fails because "I do not recognise this" was not an available output. A classifier with no abstention option will classify. This is an acceptance-criteria question, and it is testable.

Say what the fallback costs. Escalation is not free — it has a latency, a queue, a person. A threshold set without that cost in view gets quietly loosened in production by whoever is handling the volume, and nobody records that the safety property changed.

None of this requires the paper's machinery. It requires noticing that "confident and wrong" is a state, and that unnamed states get resolved by whoever reaches them first.

FAQ

What is a confident failure? A system that satisfies its functional requirement and produces a harmful outcome, because the evaluative question was never specified.

Surprise versus regret? Surprise detects novelty; regret asks whether a better available option existed. The mechanism fires only on the intersection.

Did it work? Only on models already above 80% baseline refusal. The authors conclude it amplifies rather than replaces model-level safety training.

The takeaway for specs? Define the escalation condition, the unknown category, and the cost of escalating — before the system decides all three by default.

Sources

  1. Regret Dominates Surprise: Design-Time Requirements Engineering for Agentic-AI SafetyarXiv
  2. AgentHarm: A Benchmark for Measuring Harmfulness of LLM AgentsarXiv
New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing