Thinkr

AI

Specs written for agents state the task and rarely the safeguards

A study of 1,248 GitHub Agentic Workflow files: in a labelled sample, task, output and constraints appeared in over 93%, prompt-injection defences in 9.4%.

Galang Aulia · 6 min read
AI

A spec written for an engineer can leave things out, because the engineer will ask. A spec written for an agent that runs on a schedule, with nobody watching each run, cannot. A study posted to arXiv on 23 September 2026 looks at more than a thousand such specs written by practitioners, and counts what they put in and what they left out.

What was studied

GitHub Agentic Workflows, or gh-aw, lets developers describe a recurring repository job in a Markdown file. The top of the file is configuration: what triggers a run, and which permissions and tools the agent gets. The body is prose instructions for the agent. A compiler turns the file into a GitHub Actions workflow, and an agent then carries out the job, such as triaging new issues, investigating failing CI or updating documentation, without a person starting each run.

The authors, Jasem Khelifi, Issam Oukhay, Ali Ouni and Mohammed Sayagh of École de technologie supérieure in Montreal and Mohamed Aymen Saied of Université Laval, treat those bodies as what they are: operational specifications. Their data:

  • 1,248 gh-aw Markdown files from 276 repositories, all from projects with at least 10 GitHub stars, with 20,841 file-level changes in their commit histories.
  • A sample of 294 files, hand-coded by two authors into a taxonomy of ten categories and 42 subcategories of instruction. Six could not be labelled, leaving 288. Before resolving disagreements, the coders agreed 83.61% of the time, a Cohen's kappa of 0.61.
  • 348 files with at least 120 days of activity, from 43 projects, used to track how the specs were maintained over their first four months.

They also tested whether LLMs could apply the taxonomy automatically (best F1 score: 0.818), which matters more to researchers than to spec writers.

What developers put in

These are not short prompts. The median file body was 556.5 words over 104 lines, with a median of nine headings. Code blocks appeared in 62.1% of files. The most frequent heading words were step, issue and phase.

The core of a spec is nearly always there. In the labelled sample, task instructions appeared in 97.6% of files, output instructions in 95.8%, constraints in 93.8% and process instructions in 93.4%. The most common specific patterns were a required sequence of steps (88.2%), decision branches covering conditions, fallbacks, early stops and error handling (87.2%), scope boundaries on which files, actions or changes are allowed (76.7%), and quality standards (74.7%). The paper's opening example, an auto-triage workflow, has sections headed When to Run, Your Task and Do Nothing If.

The thin categories are the protective ones. Safety instructions of any kind appeared in 25.0% of files. Explicit defence against prompt injection, meaning instructions that treat issue text, comments or linked content as data rather than as commands, or limit which sources can direct the agent, appeared in 9.4%. Resource budgets on time, retries, tokens or cost appeared in 38.2%, checks on the credibility of evidence in 11.1%, and worked examples of acceptable and unacceptable cases in 19.4%.

The specs keep changing. Among the 348 long-lived files, 78.2% were still being updated in their fourth month. Median project-level churn fell sharply after the first month, from 31.2 changed lines per 100 to 5.6, but the share of repository commits spent maintaining the workflows did not change significantly from month to month. A third of content-changing edits left the line count unchanged, so file size alone cannot show whether the instructions moved. And 108 of 365 projects (29.6%) generated at least one workflow from a Markdown file hosted in another repository, so changes its owner made could reach them the next time they recompiled.

The caveats

The authors are explicit about what the study does not show.

  • Description, not effect. The taxonomy records what developers wrote. It does not test whether any type of instruction improves task success, cost or safety. The paper cites earlier work in which repository context files did not generally improve task success and added inference cost, as a warning against reading instructions as evidence that they work.
  • Missing from the text is not missing from the system. Coding covered only the Markdown body, not the configuration. The low safety and budget counts, in the authors' words, identify "areas for closer review without establishing that corresponding runtime protections are absent."
  • Coding is interpretation. Agreement before resolution was substantial, not perfect, and one subcategory, use of the local workspace, had a kappa of only 0.238.
  • A particular slice. The corpus is public repositories with at least 10 stars that GitHub's code search could reach. The maintenance analysis deliberately kept files active for 120 days or more, so it describes specs that were maintained, not how many are written once and left alone.
  • Not product specs. These are recurring repository chores, not feature requirements. The habits transfer; the numbers describe this setting.

Why it matters if you write specs for agents

Our companion post on writing requirements an agent can build from argues for constraints first, explicit relationships, a check for every must-hold rule and a stated point where the agent stops. The practitioners in this corpus, writing for agents that run unattended, mostly arrived at the same shape: ordered steps, conditional branches with early stops, and scope limits are the norm, not the exception. Three findings add to that.

Write down what the agent must not obey. An agent working from a spec also reads things that are not the spec: tickets, comments, customer text, linked pages. Fewer than one file in ten addressed which of those sources may direct the agent. An illustration: if a requirement reads "summarise new customer feedback into the issue", the spec needs a line saying the feedback is material to summarise, not instructions to follow. In this corpus, that kind of line is one of the least common.

Put firm limits where the system enforces them. The paper's sharpest recommendation for developers:

Developers should use supported configuration controls for firm limits, while using body instructions to explain task-specific conditions and expectations.

Their example is a triage workflow whose instructions tell the agent to post no more than one comment, while its configuration caps the comment handler at one. The prose explains; the configuration holds. It is the same conclusion as the earlier finding that agents do not check their work against documentation, so a rule that has to hold needs a test, not a sentence. If a limit matters, do not leave it only in text the agent weighs against other text.

Treat the spec as something that changes, and review the changes. Most of the long-lived files were still being edited in their fourth month. When an agent runs from the spec, an edit changes behaviour on every later run, and no person starts or watches those runs. The authors recommend pinning and reviewing shared workflow files like any other dependency, and checking whether a new version changes permissions, tools, cost or how much authority the agent has. For a product spec feeding an agent, the equivalent is a visible record of what changed, reviewed before the agent picks it up.

The paper's motivating example shows why. In the .NET runtime repository, after a rejected attempt to change a platform-specific cryptography test, a feedback workflow proposed a new rule: send such cases to the domain owners. It went in through a reviewed pull request. That is a spec learning when the agent should act, hold back or ask a person for help.

What to take from it

Not that safeguards are missing everywhere: the study cannot see protections set in configuration. The narrower point stands. When people write specs for agents that will run without them, they reliably write the task, the steps and the scope. The lines that tell the agent what to distrust, what it may spend and what it must keep private are the ones a reviewer has to ask for.

FAQ

What was studied? 1,248 GitHub Agentic Workflow files from 276 repositories, with 288 hand-labelled for instruction types.

What do these specs usually contain? Task, output, constraint and process instructions, each in over 93% of labelled files.

What do they usually lack? Explicit prompt-injection defence (9.4%), resource budgets (38.2%) and safety instructions in general (25.0%).

Does it prove what works? No. It describes what was written, not its effect.

Sources

  1. Specifying and Maintaining Agentic Workflows: An Empirical Study of GitHub Agentic WorkflowsarXiv ·
New posts and release notes. No spam, unsubscribe anytime.

Stop shipping foggy PRDs.
Start the critique loop.

Three minutes to sign up. No credit card. Cancel by closing the tab.

Start freeSee pricing