Thinkr

Craft

How to Evaluate an AI PRD Generator

Judge a PRD generator by what it does when it lacks information: ask, mark the gap, or invent. A seven-check test you can run on any tool, ours included.

Galang Aulia · 7 min read
Craft

We build a PRD generator, so read this with that in mind. It is also the test we would want someone to run on ours before trusting it.

Most comparisons of AI PRD generators rank them on what a demo shows: how fast, how long, how polished, how many templates. Those are the wrong axes, because they measure the draft the way its quickest reader sees it — you, skimming — instead of the way its slowest reader will: an engineer building from it three weeks later.

There is one question underneath every useful check.

What does the tool do when it doesn't know something?

A generator always has less information than a good spec needs. You gave it two paragraphs; the spec needs a baseline, a limit, a permission rule, an empty state and a decision about what's out of scope. The gap is unavoidable. What the tool does with it is the entire evaluation, and there are only three options: ask, mark it, or invent.

Set up the test properly

Do not evaluate on a made-up feature. You will not be able to tell a good guess from a wrong one.

Use one real brief from something you have already shipped — ideally one that went wrong in an interesting way. You know the real answers, the constraints nobody wrote down, and the edge case that caused the incident. That knowledge is your answer key.

Then give every tool the same thin input: the two or three paragraphs you actually had at the start, not the finished spec. A generator given a finished spec is only reformatting it.

The seven checks

1. Does it ask before it writes? Two sentences in, 2,000 words out, is a warning rather than a feature. The best tools stop and ask about the things that change the spec — who the user is, what the constraint is, what success means. Count the questions. Zero is a result.

2. Count the invented numbers. Go through the output and list every number: targets, baselines, limits, percentages, timeouts. For each one, ask whether it came from your input. An invented baseline is worse than a blank one, because it looks like a decision. Somebody will build to it.

3. Are assumptions marked as assumptions? A good draft separates you told me this from I assumed this. If the tool writes its guesses in the same voice as your facts, you have to re-derive which is which — and you won't, for all of them.

4. Did it name the states? Before running the test, write down the states you know this feature has: empty, loading, error, no permission, the second visit, the one that caused the incident. Count how many the tool named without prompting. This is the most common gap in specs written by people, so a generator that closes it is doing real work.

5. Is the out-of-scope list specific? Most generators produce one. The test is whether it belongs to this feature or to every feature. "Mobile is out of scope" is boilerplate. "Bulk edit is out of scope; single-row edit ships first" is a decision someone can hold you to.

6. Could QA write a test from each acceptance criterion? "The export should be fast" fails. "Exports of up to 50,000 rows complete within 30 seconds" passes — provided 50,000 and 30 came from you, which is check 2 again. Testable criteria are nearly mechanical to judge, so judge every one.

7. Run it twice. Same brief, same tool, two runs. If the drafts disagree on scope, on the metric, or on a limit, the tool is making consequential choices at random — and you would have read either version as considered. A recent benchmark on requirements review found that single-run evaluations are statistically unreliable; the same applies to single-run generation.

A scoring sheet

CheckPassesFails
Asks firstAsks about user, constraint, success before draftingDrafts immediately from anything
Invented numbersEvery number traces to your input, or is markedPlausible targets and baselines from nowhere
AssumptionsLabelled, separate from your factsSame voice as your facts
StatesNames most of your known states unpromptedHappy path only
ExclusionsSpecific to this featureGeneric boilerplate, or absent
CriteriaEach one testable as writtenAdjectives: fast, easy, intuitive
Two runsSame decisions both timesDifferent scope, metric or limits

Score each check pass, partial or fail. Don't weight them. A tool that fails check 2 badly should lose regardless of the rest, and you will notice that without a formula.

What a failure looks like

An illustration. The brief, as thin as real ones are:

Let admins export the member list to CSV. Customers keep asking; support sees a few tickets a week.

A failing draft comes back long and confident. Somewhere in it:

  • Success metric: 40% of admins use export within 30 days. Where did 40% come from? There's no baseline, and nobody chose it. Check 2.
  • Exports are limited to 10,000 rows. A reasonable-sounding number. If your largest accounts have 60,000 members, this limit quietly excludes the customers who asked. Check 2 again.
  • Out of scope: mobile export. True of nearly every feature, so it tells nobody anything. Check 5.
  • Nothing about who counts as an admin in an account with several workspaces, what happens when an export fails halfway, or whether email addresses belong in the file — which is a privacy decision, not a formatting one. Check 4.

A passing draft is shorter and less impressive at first glance. Before drafting, it asks three things: who can export, whether there's an existing row limit, and whether the file should include personal data. Where you don't answer, it writes Assumption: workspace admins only — confirm rather than deciding for you. Its metric section says Baseline needed: current export-related tickets per week.

The second draft is less finished. It is also the one you can build from, because every open question is visible instead of answered by someone who wasn't there.

What not to judge it on

Length. From the same thin brief, a longer draft almost always means more invented content, not more thinking. The best first draft is often the one with the most honest blanks.

Polish. Formatting, headings and tone are the cheapest things a model does. They are also what makes a draft feel finished, which is the problem in the next section.

Speed. Every tool in this category is fast enough. Seconds saved on drafting are not where the time goes.

Template count. Forty templates that invent numbers are forty ways to ship a confident guess.

The trap is fluency

A generator's best feature is also its main risk. Fluent, well-structured prose reads as considered, and a generated spec has never been wrestled with by anyone. Nobody argued about the limit. Nobody noticed the permission rule was missing. It arrives looking like the output of a process that didn't happen.

That matters more than it sounds, because a clean-looking document feels like an approved one. A human first draft usually looks rough, and the roughness tells readers it isn't finished. A generated first draft looks finished on arrival.

So the question isn't whether to use one. It's whether the tool helps you see where the draft is thin — or hides it.

Where ours sits, stated plainly

Thinkr's generator is built around the first three checks. It scores the brief before drafting, asks about what's missing, and reads your workspace's past decisions and standards rather than starting from a blank prompt — context is the thing a model can't retrieve for you.

That's the design intent, not a result. Run the seven checks on it with the same brief you use for everything else. Like every tool here, it can't know facts you didn't give it, and it won't catch a well-written spec for the wrong feature. If it fails a check, that's worth knowing before you rely on it, and worth telling us.

The comparison page lists the other tools in the category, with dated prices.

After you've picked one

The generator is the start. What makes the spec safe to build from is still the review — a pass that reads the draft against a fixed checklist, and a human who asks whether the premise is true. A good generator shortens the first part. Nothing shortens the second.

FAQ

The most important test? What it does when it doesn't know something: ask, mark, or invent.

How to compare fairly? One real, already-shipped brief, the same thin input for every tool, the same seven checks.

Judge on length or polish? No. Both reward invention.

Does Thinkr generate too? Yes. Run the same test on ours.

New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing