Thinkr

Craft

The Six Ways Specs Actually Fail

A synthesis of thirty posts in thirty days. Every spec failure we covered was one of six patterns — and two are about order, not content.

Galang Aulia · 8 min read
Craft

Thirty posts in thirty days on specification practice. This is the pass that reads them from the end.

One thing worth saying before anything else: this is a synthesis of writing, not a study. There is no corpus of reviewed documents behind it and no measurement. What there is, is thirty attempts to describe a specific failure — and the discovery, reading them together, that they were describing six things rather than thirty.

1. The unnamed state

The most common failure, and the most mechanical.

Anything a spec does not name gets decided anyway — during implementation, by whoever reaches it first, under time pressure, with no review and no record. The empty list, the expired link, the state that occurs once and never again, the failure the user sees, the row limit nobody checked.

The pattern showed up well outside product work. A safety paper we covered this month describes a monitoring agent that harms a patient because "I do not recognise this" was not an available output — the same failure, with the stakes raised. A classifier with no abstention option will classify.

It also showed up in documentation itself: a superseded spec that nobody marked as superseded is an unnamed state in the filing system.

The fix is enumeration, and it is boring. List the states. Every requirement gets a failure behaviour. Every limit gets a number.

2. The well-formed wrong thing

The expensive one, because nothing catches it.

A spec can be internally flawless — consistent, complete, every term defined, every criterion testable — and describe something nobody needs. Coherence and correctness are unrelated properties.

This is the pattern behind vanity metrics, which are well-formed by construction; that is what distinguishes them from errors. It is behind the difference between a good and a bad spec being something other than polish. And it is the hard limit on automated review: an AI reviewer checks the document against itself, so a confident coherent falsehood looks exactly like a good spec from inside the document.

Ambiguity announces itself. This does not. The only detector is somebody who knows the domain asking whether the premise is true — which is why the score is a starting point rather than a verdict.

The fix is a question, not a check: what would have to be true for this to be a mistake?

3. The decision that left no trace

Something gets decided. Nobody writes it down. Six months later the reasoning is gone and the decision looks arbitrary — so somebody relitigates it, or worse, reverses it without knowing what it was protecting.

The out-of-scope list is the clearest case. It gets agreed at the exact moment everyone is paying attention and the trade is explicit, and then lost in the transition from pitch to spec. Rebuilding that agreement three weeks later, from memory, mid-sprint, is not the same act.

Three of the month's posts were really about this one problem from different angles: decision memory as a capability, the log as the artifact, and the decision record as the format. The recurring finding is that specs do not serve this purpose — they get superseded and archived, and the reasoning goes with them.

The fix is a separate, append-only place, and the discipline that amendments land in the document the same day rather than in the thread where they were decided.

4. The measurement taken where it flatters

Measure at the step where the gain appears, and the gain is real. The cost lands somewhere else, for someone else, with no instrument pointed at it.

This was the month's most consistent finding, and the one with the most external support. Most teams measure AI by time saved — a genuine improvement at the generation step, taken before the verification tax reallocates that time to checking. A spec drafted in ten minutes that produces five clarifying questions and one wrong implementation has not saved forty minutes; it has moved them into three calendars and recorded a win.

The same shape governs success metrics that lack a baseline, and the benchmark that stops tracking the thing it stood for.

The fix is end-to-end rather than step-level. A metric worth having is one that can return bad news.

5. The constraint that arrived too late

Two of the six are about order rather than content, and this is the one that surprised us most.

A limit stated in paragraph nine arrives after an approach has been chosen. A permission rule appended at the end arrives after the data model is settled. Neither is late in the document's own terms. Both are late in the reader's.

The research support here was the strongest of the month. A planning paper describes agents making early commitments that are systematically amplified and difficult to recover from — which is also a fair description of how most specs get written: each section reasonable on its own, in order, with no pass that looks backwards from the end.

Hence scope appearing third rather than eighth, and dependencies named before the work rather than discovered during it. Accessibility is the sharpest instance: EN 301 549 and WCAG 2.2 obligations written in after the flows are designed are obligations that will be retrofitted expensively.

The fix is front-loading the constraints that close options, and marking which are genuinely fixed rather than merely habitual.

6. The document mistaken for the agreement

Writing something down is not the same as anyone agreeing to it. A spec can faithfully record a consensus that never existed and read perfectly while doing it.

This is why the handoff produces so much — it is often the first moment anyone reads the document adversarially, and the questions that surface were always there. It is why who writes the spec matters less than who owns the decisions in it, and why engineering readiness is a property of shared understanding rather than of document length.

The question count at handoff is the cheapest available instrument for this, and almost nobody tracks it.

The fix is reading the document as the person who has to live with it, which is a different act from proofreading it.

Starting from the symptom

Patterns are only useful backwards. In practice you notice a symptom, so:

What you observeProbably patternWhere to look
Same five questions every handoff1 or 6handoff, edge cases
Built as specified, wrong thing shipped2good vs bad
"Why does it work this way?" — nobody knows3decision log
The metric improves, nothing feels better4vanity metrics
Late discovery that forces rework5dependencies
Agreement in the room, disagreement in the build6who owns it
Estimates that collapse under questioning6engineering readiness
Scope creep nobody authorised3 or 5scope

Two symptoms map to more than one pattern, and that is not imprecision — a recurring question at handoff is either something the spec never named, or something it named that nobody agreed to. Those have different fixes, and telling them apart takes one question: was it in the document?

What changed over the month

Two positions shifted, and it is worth recording which.

Tooling is not the bottleneck, but it is not neutral. We had argued the choice of tool barely matters. Working through it properly, that is too strong: each tool makes a specific failure easy, and knowing which one you are exposed to is more useful than knowing which tool other teams picked.

The "spec as source of truth for AI" story is less settled than we presented it. A study of real agent sessions found agents spent most of their documentation attention on their own instruction files rather than human-authored documentation, with no verification sequence following consultation. That does not refute the spec-driven premise, but it does mean the mechanism is assumed rather than observed — and we had written about it as though observed. The practical consequence is that the context you give a model matters more than the prose quality of the document it is drawn from.

The weakest bets

In the interest of not only reporting what worked: the definitional posts underperformed. Content built to answer "what is X" attracts readers who are not going to do anything, and the format posts on fidelity and on length are useful but were never going to move much.

The posts that earned their place were the ones naming a specific failure the reader had already experienced and could not articulate. That is a more reliable test than keyword volume, and it is the one we would use again.

If you only do one thing

Across all six patterns, one act does the most work: write down what you are deliberately not doing, before you write the requirements.

It closes unnamed states by forcing the boundary to be explicit. It records the trade at the moment it was agreed. It makes the constraint arrive early enough to shape the approach. And it surfaces disagreement while disagreement is still cheap — because the exclusions are where people discover they were not, in fact, agreeing.

Everything else in thirty posts is elaboration on that.

FAQ

The six? Unnamed states, well-formed wrong things, untraced decisions, flattering measurement, late constraints, and documents mistaken for agreements.

The worst? The well-formed wrong thing — nothing detects it, because there is no defect to detect.

The cheapest fix? Write the exclusions first.

Is this data? No. It is a synthesis of a thirty-post series, not a study of a measured corpus.

New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing