← All posts A vague instruction expands into many branches of generated code, which narrow through a verification gate into one production-ready release

Coding agents can turn a loose goal into plausible code within minutes. That is not the same as delivering software faster. Reviewers still have to reconstruct the intended behaviour, find the missing cases and decide whether the result is safe to release.

A longer specification does not automatically solve this. It can contain contradictions, stale decisions and details already expressed more accurately in the code.

The question is:

What does the agent need to finish the work, and what evidence will let us verify it?

A useful specification prevents misunderstandings that would be expensive to discover in code. It also tells the agent when to stop and ask.

What does the research establish?

No controlled study tells us how many requirements, examples or acceptance criteria a coding agent needs. The “right amount” is still engineering judgement, not a measured formula.

The evidence supports four narrower conclusions.

1. AI accelerates code production more reliably than delivery

Reinvently’s systematic review of 116 empirical studies found that controlled studies of conventional coding assistance generally report roughly 20–30% gains at the coding stage. Larger gains appear in bounded spec-driven and agent-native cases. Those estimates carry less evidence weight and bundle the model with changes to requirements, repositories, testing, review and team practice.

The strongest direct study of the journey from code to delivery followed more than 100,000 developers. Commit activity rose much faster than release output as tools progressed from autocomplete to interactive and autonomous agents. The largest observed increase—180% more commits—became 30% more releases and no measured increase in application usage (NBER working paper). This was an observational matched event study, not a randomised trial. It measures a more consequential outcome than code volume.

Specification matters in this picture because it can reduce avoidable rework and give verification a stable target. It should not be credited with the full productivity result. Most high-performing cases change several parts of the system at once.

2. Clear tests can improve human judgement of generated code

The most direct evidence is about executable clarification rather than long prose. In a 15-programmer study, TiCoder generated tests to clarify intent. Participants became better at judging whether generated code was correct. They also reported lower task-induced mental effort (IEEE Transactions on Software Engineering). The study did not estimate end-to-end delivery throughput.

“Handle failed payments gracefully” sounds like a requirement, but it leaves the important decisions to the reviewer. Which failures? What should the customer see? Should the payment be retried? Tests, examples and named failure states turn those decisions into something that can be checked repeatedly.

Not every requirement belongs in a test. Use executable checks for behaviour that is stable and important. Leave questions that genuinely require judgement with a named person.

3. Bounded delegation can outperform conversational assistance

In a 24-developer brownfield-onboarding experiment, Copilot Agent reduced mean completion time by 61.7% relative to Copilot Ask. It also reduced reported workload, without a statistically significant correctness improvement (PACIS 2026 paper). The workflow shifted from active collaboration towards supervision.

Industrial cases point in the same direction but require more caution. An expert-led workflow on the 1.52-million-line PicoScenes system reported 68.3% less implementation time, lower mean cyclomatic complexity and fewer defects (ICSE 2026). It is a comparison on one system using a bundled method, not a transferable estimate of what specification alone will produce.

These studies support giving agents bounded work and room to execute it. They do not show that writing a longer specification caused the improvement: the agent mode, tools and working practices changed as well.

What they provide is evidence for bounded delegation—not evidence for a particular specification template.

4. Verification must scale with generation

More autonomous systems produce larger changes and more review work. A study of 567 Claude Code pull requests found that 83.8% were merged. Yet 45.1% of merged changes still required human revision (On the Use of Agentic Coding). Another study examined 12,433 agent-authored pull requests. Specification mismatch and logic defects were the leading visible functional reasons for rejection (Coding Agents in the Wild). These repository studies are observational. They show why merge rate is not the same as zero-cost acceptance.

Automated review helps, but it is not a substitute for an oracle. A year-long Atlassian evaluation covered more than 1,900 repositories. It associated an integrated review agent with a 30.8% reduction in pull-request cycle time. The tool’s developer conducted this observational study (ICSE 2026 paper). Other review systems produce many comments that developers reject or ignore.

Use machines for checks they can repeat. Bring people in as the consequences or need for judgement increase:

Tests and static analysisReject defects that can be checked the same way every time.
Automated reviewCatch routine issues before they consume a person’s attention.
Accountable human reviewOwn decisions where the consequences are serious or the answer depends on interpretation.
The useful optimum minimises total delivery cost A conceptual U-shaped total-cost curve falls as specification reduces ambiguity and rises as specification overhead grows. A variable green band marks the minimum sufficient contract. The useful optimum minimises total delivery cost Ambiguity + rework Specification overhead Total delivery cost The optimum shifts with risk and uncertainty Minimum sufficient contract Specification effort and formality → Relative total cost → Conceptual model—not an experimentally estimated curve.
Write enough specification to prevent expensive misunderstandings, but stop when more detail adds more work than it removes. The chart illustrates that trade-off; research has not measured a universal sweet spot.

The right amount of spec is a risk allocation decision

An agent does not need every fact about the product. It needs the facts that would otherwise cause an expensive or dangerous misunderstanding.

Four variables determine what the contract needs to contain.

VariableLess specification can work when…More specification is justified when…
UncertaintyThe task is exploratory and the desired outcome is still being discovered.The expected behaviour is stable and disagreement would create rework.
Consequence of errorThe output is disposable, isolated and easy to reverse.The change touches money, personal data, security, safety or regulated decisions.
Delegation distanceOne engineer works with one agent in a tight feedback loop.Work crosses sessions, people, agents, repositories or service boundaries.
ObservabilityA human can inspect the result cheaply and failures are obvious.Correct-looking output can conceal logic, integration, performance or security defects.
Turn project conditions into the right form of specification Costly errors, more handoffs and hard-to-detect failures point to a stronger contract. Uncertain outcomes point to a boundary-led contract. Together they define the minimum sufficient contract. Match the contract to the project WHEN THEN Production risk is higher • mistakes are costly • work crosses more handoffs • failures are hard to detect MORE RIGOUR Write a stronger contract • tighter constraints • executable acceptance checks • named owners and approval gates The outcome is uncertain Do not write a falsely fixed answer. Define the search space instead. CLEAR BOUNDARIES Write a boundary-led contract • non-goals and constraints • evidence and stop-and-ask points Both belong in the minimum sufficient contract.
Costly, distributed or hard-to-check work needs stronger controls. Exploratory work needs clear boundaries if your agent is to know when it has finished the loop.

In short, the level of specification should rise with the level of project risk. It is perfectly acceptable for an exploratory spike to be a one-shot exercise. Work in a highly regulated domain such as financial services needs greater rigour, with specifications and approval records forming part of the required audit trail. That becomes more important as the work moves closer to production.

More complex production setups also call for more detailed specifications. A solo developer building small web apps in their bedroom with a single agent sits at the opposite end of the spectrum from a medical-device company developing software with an agentic swarm.

The practical target is a minimum sufficient contract. It gives the agent enough information to finish the job and gives you enough evidence to decide whether it succeeded.

“Minimum” does not mean short. A reversible prototype may need only a goal, boundaries and a few checks. Production work involving money, personal data or several agents may need schemas, failure behaviour, rollback and named approval. Every extra requirement should earn its place.

What belongs in a minimum sufficient contract?

For most bounded feature work, a useful agent-facing specification has seven parts.

  1. Outcome. What observable change should exist when the work is done?
  2. Context. Which current behaviour, domain terms and repository conventions matter?
  3. Constraints. What must remain true, including security, compatibility and performance boundaries?
  4. Non-goals. What plausible adjacent work is deliberately outside scope?
  5. Acceptance checks. Which outcomes can be tested or inspected repeatably?
  6. Failure behaviour. What should happen on invalid input, partial failure, timeout, retry or rollback?
  7. Open decisions. Where must the agent stop and ask rather than infer?

The contract fixes the outcome and its boundaries, not the implementation.

Put each requirement where the agent can use it and the team can maintain it. Stable behaviour belongs in tests, types and schemas. Rationale and non-goals usually belong in prose. Do not copy an interface into three documents when the code already expresses it clearly; the copies will drift.

How should specification change by task type?

Exploratory work: specify boundaries and evidence

Do not ask an exploratory agent to deliver an answer you have not yet discovered. Tell it what question to investigate, which directions are out of bounds, what evidence to collect and when to stop.

The output is learning. Any code it produces is a probe until it passes a separate production review.

Bounded feature work: specify outcomes and checks

For a small feature or familiar integration, the product requirements document (PRD) just needs to define the visible outcome, important constraints, non-functional requirements and acceptance criteria. Include examples wherever two engineers could reasonably interpret the requirement differently.

Keep the change small enough for a person to understand and review.

Deterministic or consequential work: make the contract executable

Financial transactions, migrations, security controls and data transformations justify more precision in the PRD. Specify invariants, failure handling and rollback. Alongside the PRD, start with a set of fixtures, integration tests and unit tests.

Multi-agent work: specify the handoffs

Once several agents are working in parallel, vague boundaries become overlapping edits, incompatible assumptions and merge conflicts.

Define each handoff using SIPOC: supplier, input, process, output and customer. State who provides the input, what the agent receives and does, what it returns and who consumes the result. Then define how the output will be checked and what “good” looks like, so each handoff has a strong quality gate.

Review the spec before paying to implement it

A polished specification can still be wrong. Before an agent writes code, run a short adversarial pass:

An agent can look for gaps, but a person must settle any ambiguity that changes the outcome or risk. Do this before implementation, when changing the specification is cheaper than changing the code.

Retire prose when stronger artefacts replace it

A specification should not remain the main source of truth once a requirement exists as a test, type or schema. Keep the decision log, but remove implementation instructions that the code has made obsolete to keep your agent context clean.

Code is truth, but it does not explain why a constraint exists, and it can implement the wrong behaviour perfectly. Keep each fact in the place where it is easiest to maintain and hardest to misunderstand.

Measure the whole loop, not prompt-to-code time

If a team wants to discover its own right amount of specification, it should measure comparable work at different levels of structure. Useful measures include:

Do not optimise for the first diff. A workflow that produces code in ten minutes and consumes two hours of review may be worse than one that spends 30 minutes clarifying intent and passes review once.

Do not ask how much code the agent produced. Ask whether a verified change reached production with less effort and without causing more failures.

From principle to operating model

Once you know how rigorous the contract must be, decide how much of the workflow you want the framework to control.

That places the four frameworks in different parts of the landscape:

The question is not which framework is most comprehensive. It is which missing control you need it to provide.

Choose a framework by specification rigour and scope of control A two-axis matrix positions OpenSpec towards focused change control, GSD towards rigorous execution control, GitHub Spec Kit towards a broader standardised workflow, and BMAD towards comprehensive lifecycle control. Match the framework to the control you need Required specification rigour Targeted Comprehensive Desired scope of control One part of the workflow Requirements to release OpenSpec Governed change record GSD Planning and execution GitHub Spec Kit Portable spec workflow BMAD Lifecycle and roles Editorial orientation based on framework scope—not an empirical ranking.
Move up the chart as project risk demands more rigorous specifications; move right as the desired control expands from one part of the coding workflow towards the full delivery lifecycle.

For the practical choice, see the full comparison of GSD, BMAD, OpenSpec and GitHub Spec Kit. It covers where these starting points overlap, how much ceremony they introduce, and which project settings suit each one.

The practical rule

Coding agents have made implementation cheaper. Deciding what to build—and proving that the result is safe to release—has not become cheaper by the same amount.

The agent needs enough direction to finish the loop. You need enough evidence to trust the result. That’s the sweet spot.


Evidence note. This article draws on Reinvently’s version 1.28.2 systematic review of AI-assisted software development. Numerical claims retain the review’s distinctions among controlled effects, observational associations and lower-weight organisational cases. The conceptual cost curve and decision model apply that evidence. They are not validated measurement instruments. Markus Eisele’s O’Reilly Radar essay on specification in agentic development prompted the question.

Frequently asked questions

How much specification does a coding agent need?

Enough to tell the agent what result to produce, what must not change, how the result will be checked and when it should ask for help. Small reversible tasks may need only a goal and acceptance checks; consequential or multi-agent work needs stronger contracts, failure handling and approval gates.

Does more specification always improve AI-generated code?

It is a classic trade-off. Too little specification can create correction loops, but excessive or stale prose can add cost and conflicting instructions. Research has not established a universal optimum, but it does give us some useful guidelines.

What should an agent-facing specification contain?

For bounded feature work, define the outcome, relevant context, constraints, non-goals, acceptance checks, failure behaviour and decisions that require human input.

Which spec-driven development framework should a team use?

Choose according to the missing workflow control. GSD focuses on execution and context, BMAD on lifecycle and roles, OpenSpec on governed change, and GitHub Spec Kit on a portable workflow across coding tools. The full framework comparison covers their practical trade-offs.


New research and engineering guides land here first. Sign up for email updates.

← All posts