Thinkr

Craft

Writing Requirements an Agent Can Build From

A coding agent reads early, commits early, does not ask and does not check its work against your spec. Six changes to how you write the requirements.

Galang Aulia · 6 min read
Craft

A spec written for engineers relies on something nobody writes down: the reader will ask. An engineer who hits "filtered by the threshold" sends a message and finds out which way. That question is a working error-detection system, and most specs lean on it more than their authors realise.

A coding agent removes it. Five studies we have covered in our briefings, read together, describe this reader fairly precisely. It reads early, commits early, does not ask, and does not come back to check its work against your document.

None of that makes the spec less important. It makes the spec the only place your intent survives. Here is what the research says about the reader, and six changes to how you write for it.

What the research says about the reader

It mostly reads files written for it. A study of 557 real agent sessions found agent instruction files and working notes made up 60.5% of documentation interactions, against 10.6% for classical technical documentation. The authors note their document classifier was not human-validated, so treat the split as indicative.

It reads at the start, not when stuck. In the same data, consultation was self-initiated in 70.2% of cases and failure-driven in 7.5%. No documentation-driven verification was observed: agents did not read the spec and then check their output against it.

It commits early. A planning paper describes step-by-step reasoning as producing early myopic commitments that are amplified over time and hard to recover from. The results come from deterministic, fully structured environments, but the shape is familiar: the damaging choice is made first, when the least is known.

It does not ask. A benchmark of ambiguous coding requirements describes models as forced into determinism: they collapse an unclear requirement into one implementation, and two runs of the same requirement can produce functionally opposite code.

Clean requirements are not enough. A controlled study found that even with no injected defects, models failed some tests, and the failures clustered around shared functions, ordered dependencies and changing state. Its main effect did not reach statistical significance, and the authors say so.

Six changes to how you write

1. Put constraints first, and where they will be read

If the agent reads at the start and commits early, a limit in paragraph nine arrives after the approach is chosen. Lead with the things that close options: limits, permissions, data boundaries, what must not change.

Then repeat the non-negotiables in the agent's instruction file, whatever your tool calls it. That is where the agent's attention demonstrably goes. The spec stays canonical; the instruction file carries the handful of rules that must not be missed.

2. Mark what is fixed and what is open

An agent, like a new hire, cannot tell a regulatory obligation from a habit. Both appear as the status quo. One word per constraint fixes it: fixed or open. Without it, the agent treats your preference as a law and your law as a preference, and you find out which in review. This is the same distinction that makes context worth writing down at all.

3. Resolve ambiguity in the spec, because nothing downstream will

"Filtered by the threshold." "Recent items." "Handle appropriately." Each used to produce a question. Now each produces working code that implements one reading. Replace every adjective and every relational word with the thing it stands for: above or below, the last 30 days, a named behaviour. The same discipline as testable acceptance criteria, applied to every line rather than only the criteria.

4. Write the relationships, not just the lines

The failures that persisted on clean input sat between requirements: this must happen before that, these two share a rule, this state changes what that action does. A line-by-line review does not catch those, and neither does a line-by-line spec. Say the order. Name the shared rule once and reference it. List the states that change behaviour. The dependencies you forgot to write down matter more when the reader cannot ask about them.

5. Turn every must-hold rule into a check

The agent was not observed checking its work against documentation. So a requirement that has to hold needs a test, not a sentence. Write acceptance criteria as conditions that can fail, and ask for the test to be written alongside the code. If a rule matters and cannot be tested, say so explicitly and flag it for human review. That is the line a PRD test plan depends on.

6. Say when to stop

The most important line is usually missing: the condition under which the agent should stop and hand back to a person instead of choosing. A requirements paper on agent safety argues the escalation boundary is a design-time requirement, not something to leave to the prompt. For a coding agent it is usually mundane. If the change needs a schema migration, stop. If two requirements conflict, stop and quote both. If the permission rule does not cover a case, stop rather than guess.

One requirement, rewritten

An illustration. The version written for a human reader:

Add CSV export for the members list. It should handle large workspaces and respect permissions.

Every phrase in it needs a question: how large, which permissions, what goes in the file. A person asks. An agent decides. The version written for an agent:

Constraints (fixed): only workspace owners and admins can export. The file must not include members removed from the workspace. No schema changes. Constraints (open): file naming, column order. Behaviour: exports up to 50,000 rows synchronously. Above that, queue the export and email a link to the requester. Order: check permission before counting rows, so non-admins never learn the member count. Checks: a test per fixed constraint, including a non-admin receiving a permission error and a removed member absent from the file. Stop and ask if: the members table lacks a field needed for the export, or the email service is not available in this environment.

It is longer by a few lines. Most of the additions are answers to questions an engineer would have asked in the first hour. The agent will not ask them, so the spec has to.

What this does not fix

Two honest limits.

The evidence is early. All five studies are arXiv papers. The trace study is observational, and in its own section on implications its data do not support, it reports that actionability and verifiability, two properties widely recommended for agent-friendly documentation, lacked consistent behavioural support. The practices here are consistent with the evidence. They are not proven by it.

A precise spec for the wrong feature is still wrong. Everything above makes an agent more likely to build what you wrote. None of it checks whether what you wrote is worth building. That judgement stays with the people reading the spec before the agent does.

Where a generator helps and where it does not

Most of the six changes are about the brief, not the prose. A generator that drafts immediately from a thin brief fills every gap an agent would otherwise fill, just earlier. Ours scores the brief on six dimensions and asks about what is missing before it drafts, and reads your workspace's past decisions and standards rather than starting blank. That is the design intent. The stop conditions and the fixed-or-open labels still have to come from you, because they are decisions, not text.

FAQ

How do you write a spec for a coding agent? Constraints first, fixed versus open, no ambiguous words, explicit relationships, a test per must-hold rule, and stop conditions.

Do agents read the spec? Less than assumed. One study found agent-facing files were 60.5% of documentation interactions.

Will it ask when something is unclear? Do not count on it. It picks a reading and builds it.

Is the evidence solid? Early: arXiv papers, mostly observational or controlled. Treat these as practices consistent with it.

New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing