Thinkr

Craft

What AI Still Can't Catch in a PRD

An AI reviewer checks the document against itself. Checking it against the world stays human — and that gap is structural, not a capability we are waiting on.

Galang Aulia · 7 min read
Craft

We build an AI PRD reviewer, so this post is written against our own interest. It is also the most useful thing we know about the tool.

There is a clean line between what an automated review can and cannot do, and it is not the line most people assume. It is not about model quality, and most of it will not be fixed by a better model.

An AI reviewer checks the document against itself. A human reviewer checks the document against the world.

Everything follows from that.

What it genuinely does well

Worth stating plainly, because the limits only matter if the strengths are real.

Omissions against a known shape. Specs have a recognisable structure. A missing success metric, an absent out-of-scope section, a user flow with no failure path — these are detectable because the absence is visible against the pattern.

Internal contradictions. Section 2 says admins only, section 7 describes an editor doing it. Humans miss this constantly, because we read for intent and quietly repair inconsistencies as we go. Machines do not perform that repair.

Unnamed states. The empty state, the loading state, the state that only occurs once — enumerable, and therefore checkable.

Undefined terms and unquantified limits. "Recent," "fast," "a lot of," "the user." Each of these gets resolved by whoever implements it, and the resolution is invisible until it ships wrong.

Missing acceptance criteria. Whether a requirement can be tested at all is close to a mechanical property.

And one thing no human review can match: consistency. The same standard on the twentieth spec of the quarter as the first, at 6pm on a Friday, for a project the reviewer has no stake in. Human review quality degrades with fatigue, familiarity and politics. Automated review does not.

That is a real category and it is most of what goes wrong in practice.

Category one: out of reach in principle

These need facts that are not in the document. No model can retrieve them, because they are not written down anywhere.

Whether the problem is real. A spec can be internally perfect and describe something nobody wants. Coherence and demand are unrelated properties. The most dangerous specs in any company are the well-written ones for features that should not exist — they pass every check precisely because the writer was competent.

Whether the evidence says what you claim. A reviewer can see that you cited research. It cannot see that the research was six interviews with your friendliest customers, or that the quote was about a different problem. The citation is in the document; its quality is not.

Whether the metric is the one that matters. It can verify that a metric has a baseline and a target. It cannot know you picked that metric because it was the one already trending up. Vanity metrics are well-formed by construction — that is what makes them vanity metrics rather than errors.

The unwritten constraints. Billing is mid-migration. That integration is contractually frozen until March. The VP who owns this surface has strong views nobody has documented. These decide whether the plan is possible and appear in no spec.

Whether the right people agreed. Consensus is a social fact. A document can record an agreement that never happened, and read perfectly while doing it.

Whether now is the right time. Sequencing is a judgement about everything else in flight. A reviewer sees one document.

Category two: out of reach in practice

Different failure. Here the information is present but the judgement is not reliably automatable.

Whether the scope is right. A reviewer can check that scope exists and is specific. Whether cutting that particular thing guts the feature is a product call that depends on who the user is and what they are actually trying to do.

Whether the estimate is plausible. This requires knowing the codebase, the accumulated debt, who is on leave and how the last three projects went.

Whether the risk is worth it. Risk appetite is set by the company's position, not by the document. The same risk is correct at a startup and reckless at a bank.

Second-order effects. What this does to support volume, to the other team's roadmap, to the mental model users already have. Occasionally inferable, never dependably.

The failure worth naming

The most expensive spec is not the incomplete one. It is the confidently wrong one — internally consistent, fully specified, and built on a premise that is false.

Ambiguity gets caught, because ambiguity looks like ambiguity. A confident, coherent, wrong spec looks exactly like a good one from inside the document. There is no structural signal to detect, because there is no structural defect. Only someone who knows the domain can see it.

This is why "the AI found nothing" carries almost no information about whether the spec is good. It means the spec is well-formed. Well-formed is the cheapest property a spec has.

What this looks like in one spec

Concretely, because the abstract version is easy to nod along with.

A spec for bulk CSV export. It names the entry point, the permission rule, the empty state, the row limit, the failure behaviour when the export times out, and the acceptance criteria for each. Every term is defined. Nothing contradicts anything. A mechanical review returns clean, correctly.

Five things are wrong with it, and none is visible from inside the document:

  • The three customers who asked for this wanted a scheduled export, not a manual one. They said so in the calls; the spec records the feature, not the ask.
  • The 50,000-row limit was chosen because it was in the last spec. Nobody checked the distribution. Sixty percent of the accounts that need this exceed it.
  • The success metric is exports generated. It will go up whether or not anyone can use the file.
  • The data team is deprecating the table this reads from next quarter. It is on their roadmap and not on yours.
  • Support was not consulted and will absorb every question about why the file opens wrong in Excel.

Each of these is a fact about the world. The document is a faithful, well-formed description of a plan that will waste a quarter — and it is precisely because it is well-formed that it will sail through review and into a sprint.

The risk nobody mentions

A clean review feels like approval.

That is the real hazard, and it is a human one rather than a technical one. You run the review, the issues come back, you fix them, you run it again, it comes back clean. The feeling produced by that sequence is done. And the feeling is doing work that the check did not earn — because none of it touched whether the thing should be built.

It is the same shape as the verification tax: the saving is real and visible at the step where it happens, and the cost lands somewhere nobody has an instrument pointed at. Here the cost is a skipped question.

The defence is a habit rather than a tool. After the mechanical review comes back clean, one human asks three things:

  1. Is the problem real, and how do we know?
  2. What would have to be true for this to be a mistake?
  3. Who has not seen this who will have to live with it?

None of those is in the document. That is the point of them.

How to actually split the work

The division that works follows the line at the top.

Give the machine the exhaustive pass. Every state, every term, every requirement checked for testability, every section checked for presence. This is the work humans do badly, resent, and skip when busy. It is also the work where consistency beats insight.

Keep the world-facing questions with people. Is this worth doing, is the evidence sound, does this survive contact with the other team, is the timing right. These need someone who was in the room.

The sequencing matters more than it sounds. Run the mechanical pass first, so that human review time is not spent on undefined terms and missing states. Reviewers who spend their attention on formatting defects never reach the questions only they can answer — and they leave the review believing they reviewed it.

That is the honest case for automated review, and it is narrower than the marketing in this category usually suggests: not that it catches what humans miss, but that it clears the floor so humans reach the part that was always theirs.

FAQ

What can't it catch? Anything requiring facts outside the document — whether the problem is real, whether the evidence holds, whether the right people agreed, whether now is the time.

What is it good at? Mechanical completeness, internal contradictions, unnamed states, undefined terms, untestable requirements — consistently, every time.

Will better models close the gap? Not the first category. A reviewer that cannot observe your organisation cannot know what nobody wrote down.

The main risk? That clean feels like approved. Well-formed is the cheapest property a spec has.

New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing