Thinkr

Craft

Specifying an AI Feature: The States Nobody Names

Most AI feature specs define success and error. The states in between (low confidence, abstention, refusal, escalation) decide how the feature fails.

Galang Aulia · 6 min read
Craft

A spec for a deterministic feature can get away with two states for most requirements: it worked, or it failed with an error. Both are visible. The error throws, the screen shows something, somebody gets paged.

A model breaks that. Its most common failure is not an error. It is a fluent, well-formatted answer that happens to be wrong, delivered in exactly the same voice as a right one. And between "right" and "wrong" sit several states that most AI feature specs never name: the model is unsure, the model does not know, the model must not answer, a person needs to take over.

Unnamed states still happen. They just get decided by the prompt, the model's defaults, or whoever writes the handler. This post names them and says what a spec should contain for each.

Why the usual state list does not cover it

The error-state discipline still applies to the deterministic parts of an AI feature: validation, permissions, timeouts. What it misses is that a model can succeed technically and fail substantively. The API returned 200. The output parsed. The answer was wrong.

A requirements paper we covered calls the extreme version a confident failure: the system satisfies its functional requirement and produces a harmful outcome anyway, because the specification never asked what it should do when it was confident and out of its depth. The paper's example is a medical monitor with no "unknown" category. It classified a rare, harmless irregularity as cardiac arrest, and acted on it.

Most product features are nowhere near that stakes level. The mechanism is the same.

The states

1 · Low confidence. The model has an answer and weak grounds for it. The spec has to decide whether to show the answer marked as uncertain, show alternatives, or suppress it. Most importantly it has to name the signal: a classifier score, a retrieval match, a second-pass check. "When the model is unsure" is not a trigger unless something computes it.

2 · Abstention. The model declines to guess. This only exists if the spec makes "I don't know" a valid output. A classifier with no unknown category will always pick a category, so the absence of this state is itself a design decision. Say what the user sees, and what they do next, because an abstention with no next step is a dead end.

3 · Refusal. The model could answer and must not: the request is out of scope, against policy, or asks for a commitment only a person can make. Abstention is "cannot", refusal is "will not", and they need different messages. Also specify the opposite failure, over-refusal, where the model declines ordinary requests because they resemble forbidden ones.

4 · Escalation. The output goes to a person instead of, or before, the user. The spec needs the trigger, who receives it, how fast they must respond, and what the user sees while waiting. Escalation is not free. It has a queue, a latency and a salary behind it, and a threshold set without that cost in view tends to get loosened quietly once the volume arrives.

5 · Unavailable. The model API is down, slow or rate-limited. This one is an ordinary dependency failure and belongs with your other error states. The AI-specific question is the fallback: does the feature degrade to a non-AI path, or disappear?

6 · Confidently wrong. The state nothing at runtime can detect, by definition. You cannot specify a trigger for it. You specify recovery instead: how a user corrects the output, whether the correction is captured, which outputs get sampled for human review, and who reads that review.

What to write for each

The same five parts as any error state, with one change: the trigger must name a signal the system can actually compute.

StateTriggerUser seesNext actionWho finds out
Low confidenceNamed signal below a stated valueAnswer, visibly markedAccept, edit or discardLogged with the score
AbstentionNamed signal, or input outside supported types"Couldn't determine" plus reasonManual pathCounted weekly
RefusalRequest matches the refusal listWhat it won't do, and who canRoute to a personLogged by category
EscalationStated conditionPending state with expected timeWait, or continue manuallyNamed queue and owner
UnavailableTimeout or error from the model APIFeature degrades to manualContinue without AIAlert to on-call
Confidently wrongNone at runtimeThe wrong answerCorrect itSampled review, correction log

One feature, specified

An illustration: a feature that suggests a category and priority for each incoming support ticket. The thresholds below are placeholders to show the shape, not recommendations.

Low confidence. If the classifier score is under 0.7 (uncalibrated: validate against labelled tickets before launch), show the suggestion greyed out with "low confidence". The agent must confirm before it applies.

Abstention. Tickets in languages the model was not evaluated on, or with no body text, get no suggestion and a "categorise manually" label. Abstention rate is reported weekly.

Refusal. The model never assigns the "legal" or "security incident" categories. Tickets that look like either are routed to the named queue with no suggested priority.

Escalation. Any ticket suggested as "urgent" from an account on an enterprise contract goes to the duty lead for confirmation within 15 minutes, before the customer-facing SLA starts.

Unavailable. If the model does not respond within 3 seconds, the ticket enters the normal queue with no suggestion. No retries block ticket creation.

Confidently wrong. Every agent change to a suggested category is logged with the original. A weekly sample of 50 accepted suggestions is reviewed by the support lead.

Notice what the spec had to decide that no model could decide for it: which categories are too consequential to automate, which customers get a human check, how long a person has to respond, and who reads the correction log. Those are product decisions. Left out, they become prompt-engineering decisions, made later, by someone without the context.

Three decisions the spec cannot skip

The signal behind every threshold. "Confidence" is not a property every model exposes. Name what you are reading, and mark the value as uncalibrated until it has been tested. A threshold with no stated source is a guess that looks like a decision.

The cost of the fallback. Every abstention and escalation lands on a person. Estimate the volume, name the queue, and say who owns it. Otherwise the first busy week produces a quiet change to the threshold, and nobody records that the behaviour changed.

How often each state is acceptable. A feature that abstains on half its inputs may not be worth shipping. A feature that never abstains is probably guessing. State the range you expect, so launch has something to compare against.

Where it goes in the document

The AI Feature PRD template has two sections built for this: a behaviour definition, which is where refusal and escalation rules belong, and a guardrails table with one row per failure mode. The states above are the rows most often missing from that table.

The same idea applies when the thing reading your spec is a coding agent rather than a user: the most useful line is the one that says when to stop and hand back instead of guessing. Our critique includes an AI-Specific Readiness pass for documents that describe model-driven behaviour, and the generator asks about the brief before drafting it.

FAQ

Which states does an AI feature need? Low confidence, abstention, refusal, escalation, unavailable and confidently wrong, each with a trigger and a next step.

Abstention or refusal? Abstention is "cannot", refusal is "will not". Different messages, different next steps.

How do I specify a threshold? Name the signal, the value and both sides of it, and mark it uncalibrated until tested.

What about confidently wrong? Specify recovery: correction, capture, sampled review and an owner.

New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing