← All posts Engineering

Spec-Driven Development for AI Coding Agent Teams

<article> <p>For an engineering manager or senior engineer trying to make agent-written code reviewable, <strong>spec-driven development for AI coding agents</strong> gives you the control plane: put intent, constraints, tasks, and verification evidence into durable repo-visible artifacts before the agent writes code. The answer is not a longer prompt. It is a workflow where the team can review what the agent is about to do, what it must not touch, how it will prove the change, and how the spec stays current after merge.</p>

<p>The practical shift is simple. Agents make implementation cheaper, so ambiguity becomes more expensive. If a team hands an agent a vague request, the agent can still produce a coherent diff. That diff may compile, pass adjusted tests, and still be wrong for the product, the architecture, or the repository's rules.</p>

<p>Spec-driven development changes the handoff. Instead of asking reviewers to reconstruct the prompt from a pull request, the team writes down the agreement first. The agent implements against that agreement. Reviewers compare code to the spec, not to a chat transcript that may be incomplete, stale, or inaccessible.</p>

<p>That review target connects directly to an <a href="/guides/ai-coding-agent-pr-reviewability-checklist">AI coding agent PR reviewability checklist</a>: reviewers need scope, evidence, and risk notes before they judge the diff.</p>

<h2>Spec-Driven Development for AI Coding Agents Starts With Intent</h2>

<p>Picture a common team workflow. A product engineer asks an agent to add a new account setting. The agent scans the repo, edits a form, updates a service, changes a permission check, rewrites a test fixture, and opens a pull request. The code looks plausible. The reviewer now has to answer basic questions before reviewing the code itself:</p>

<ul> <li>What user behavior was supposed to change?</li> <li>Which files were intentionally in scope?</li> <li>Which existing patterns should the agent have followed?</li> <li>Which security and compliance rules applied?</li> <li>Which tests prove the requested behavior rather than the agent's interpretation?</li> </ul>

<p>If those answers live only in the agent conversation, review slows down. If they live in a small spec committed with the change, review has a stable target.</p>

<p>This is why the strongest pattern in the brief is not one tool. GitHub Spec Kit, OpenSpec, Superpowers-style procedures, AGENTS.md files, CLAUDE.md files, skills, checklists, and PR gates all point at the same operating model: standardize the upstream engineering artifacts that agents consume.</p>

<h2>The Team Story: From Prompt Handoff to Reviewable Work</h2>

<p>A useful spec-driven workflow feels less like ceremony and more like a disciplined ticket turning into a reviewable pull request.</p>

<p>First, the engineer writes a short spec. It states the user problem, expected behavior, non-goals, constraints, affected capability, and definition of done. For small work, this may be a compact plan in the issue or branch. For production features, it can be a full Spec Kit style path with constitution, specify, plan, tasks, implement, and converge phases.</p>

<p>Second, the agent grounds the plan in the actual repository. It reads the relevant code, tests, project rules, and architecture notes. The Spec Kit Agents paper in the brief reported 128 runs across 32 features and 5 repositories, with context-grounding hooks improving judged quality from 3.51 to 3.66 in the full workflow. Treat that as a directional signal, not a universal guarantee: grounding matters because agents can otherwise make clean plans for the wrong codebase.</p>

<p>Third, the team reviews the plan before broad implementation. This is the cheap correction point. A senior engineer can reject a risky abstraction, narrow the task, require a migration plan, or ask for a different test strategy before the agent generates a large diff.</p>

<p>Fourth, the agent implements task by task. Each task should be small enough to inspect. The agent should leave evidence: commands run, tests added, manual checks, known gaps, and any files it intentionally avoided.</p>

<p>Finally, the pull request links back to the spec and includes verification evidence. Reviewers check whether the implementation matches the agreed behavior and constraints. After merge, the team updates durable specs so the next agent session inherits the truth that was just created.</p>

<h2>Use Three Lanes, Not One Big Process</h2>

<p>The failure mode of spec-driven development is Markdown bureaucracy. Martin Fowler's SDD tools analysis, as summarized in the brief, warns that SDD tools can overproduce artifacts for small tasks. Addy Osmani makes the same practical point from another angle: throwing a massive spec at an agent does not work. Agents need focused, structured, testable, bounded context.</p>

<p>In practice, use lanes by change size and risk.</p>

<table> <thead> <tr> <th>Lane</th> <th>Use it for</th> <th>Minimum artifact</th> <th>Review gate</th> </tr> </thead> <tbody> <tr> <td>Tiny fix</td> <td>Small defects, copy changes, narrow refactors, local test fixes.</td> <td>Issue comment or plan with scope, non-goals, and test command.</td> <td>Human pre-review plus passing targeted checks.</td> </tr> <tr> <td>Feature spec</td> <td>User-visible behavior, new flows, API changes, shared components.</td> <td>Spec, plan, task list, acceptance criteria, and verification checklist.</td> <td>Plan review before implementation and PR review against the spec.</td> </tr> <tr> <td>Governed change</td> <td>Security, compliance, cross-repo work, migrations, platform contracts.</td> <td>Constitution rules, design note, task traceability, risk notes, rollout and rollback plan.</td> <td>Owner approval, explicit verification, and post-merge spec update.</td> </tr> </tbody> </table>

<p>This keeps the process proportional. A one-line bug should not need a constitutional convention. A permission model rewrite should not be driven by a chat prompt and a hopeful test run.</p>

<h2>What Belongs in the Spec</h2>

<p>A reviewable agent task needs five concrete pieces of context. Andrew Crookston's field report in the brief describes the same pattern: each task needs what to build, what not to touch, how to prove it works, constraints, and definition of done.</p>

<ul> <li><strong>Intent:</strong> the user problem, expected behavior, and acceptance criteria.</li> <li><strong>Boundaries:</strong> files, modules, APIs, tests, and systems the agent should not touch.</li> <li><strong>Constraints:</strong> security, compliance, design system, performance, migration, compatibility, and integration rules.</li> <li><strong>Plan:</strong> ordered tasks that can be reviewed before implementation.</li> <li><strong>Verification:</strong> tests, commands, manual checks, and evidence required before the work is done.</li> </ul>

<p>Keep the spec boring and specific. "Improve settings" is not a spec. "Add an email notification toggle to account settings, persist it through the existing preferences API, do not modify billing settings, and prove it with the existing form test plus one API persistence test" is closer to something a reviewer can use.</p>

<p>The point is not to predict every line of code. The point is to remove the avoidable choices an agent should not be making alone.</p>

<p>This is especially important for <a href="/guides/prevent-ai-coding-agents-rewriting-tests">preventing AI coding agents from rewriting tests</a> to match their own assumptions.</p>

<h2>Repo-Visible Context Beats Disposable Chat</h2>

<p>The brief's strongest governance lesson is that context should survive the chat session. OpenSpec emphasizes durable repo storage, where specs live by capability and archived change deltas update the source-of-truth spec library. Its website puts the point plainly: context should not disappear when a chat session ends.</p>

<p>That matters for review. A pull request can link to a repo spec. A teammate can diff it. A later agent can read it. A platform team can standardize templates around it. A chat transcript is useful evidence, but it is a weak source of truth for team delivery.</p>

<p>Use durable layers:</p>

<ul> <li><strong>AGENTS.md or CLAUDE.md:</strong> always-loaded project rules, repository map, commands, ownership expectations, and "do not touch" areas.</li> <li><strong>Feature specs:</strong> requirements and acceptance criteria for meaningful changes.</li> <li><strong>Task lists:</strong> implementation steps the agent can execute and humans can audit.</li> <li><strong>Skills:</strong> reusable packages of instructions, resources, and scripts for repeatable workflows.</li> <li><strong>Checklists:</strong> review, verification, security, release, and post-merge update gates.</li> </ul>

<p>Vercel's narrow Next.js eval in the brief is a useful warning about context loading. An 8KB AGENTS.md docs index reached a 100% pass rate, compared with 79% for explicitly instructed skills and 53% baseline. Do not overgeneralize that number across every stack. The practical lesson is narrower and stronger: concise always-loaded repo context can matter more than optional context the model may or may not invoke reliably.</p>

<h2>Plan Review Is Where Senior Engineers Save Time</h2>

<p>The best time to catch an agent going wrong is before it writes the broad diff. A plan review can be short, but it should be real.</p>

<p>Ask these questions before implementation starts:</p>

<ul> <li>Does the plan solve the stated behavior, or did it quietly widen the problem?</li> <li>Does it follow existing repository patterns?</li> <li>Does it touch owned or sensitive files?</li> <li>Does it separate mechanical changes from behavior changes?</li> <li>Does the verification plan prove the actual requirement?</li> <li>Does the task list create reviewable chunks?</li> </ul>

<p>This is where the human role shifts. Staff engineers and platform teams spend less time typing boilerplate and more time designing artifacts, gates, permissions, and escalation paths. Crookston's field report captures the change well: code becomes the byproduct.</p>

<h2>Context Drift Is the Main Failure Mode</h2>

<p>Context drift happens when the agent's working model of the task separates from the real repository, the reviewed plan, or the team's constraints. It is subtle because the output can still look coherent.</p>

<p>Common drift patterns include:</p>

<ul> <li>The agent follows a stale architecture note instead of current code.</li> <li>The implementation satisfies the prompt but violates a repo rule.</li> <li>The agent edits tests to match its assumption rather than the requirement.</li> <li>The plan says one module, but the diff spreads across unrelated ownership areas.</li> <li>The agent creates a new abstraction where the codebase already has one.</li> <li>The PR summary describes the intended change but omits the risky changed file.</li> </ul>

<p>Use explicit grounding to reduce this. Require the agent to cite local files it inspected in the plan. Require it to name existing patterns it is following. Require changed-file summaries in the PR. Require a gap list when verification is incomplete. These are not decorations. They are review handles.</p>

<h2>Compare the Workflow Families Without Picking a Religion</h2>

<p>The brief identifies three useful families. Treat them as tools for different jobs.</p>

<p><strong>GitHub Spec Kit</strong> is the full workflow. It frames spec-driven development as defining what to build before building it, with phases for constitution, specify, plan, tasks, implement, and converge. Its quick start also distinguishes shorter workflows for smaller features from fuller workflows with clarify, checklist, and analyze gates for production work.</p>

<p><strong>OpenSpec</strong> is a lighter brownfield path. Its workflow centers on explore, propose, apply, and archive. The artifact set is practical for existing systems: proposal, design, tasks, and spec deltas. Its durable capability-based spec library is especially useful when teams need the source of truth to evolve after each merged change.</p>

<p><strong>Superpowers-style methodology</strong> focuses less on a requirements format and more on agent procedure: brainstorming, worktrees, plans, subagents, test-driven development, review, and branch finishing. That is valuable when the team's problem is not missing requirements, but inconsistent agent operating behavior.</p>

<p>A mature team may combine them. Use AGENTS.md as the base layer, a Spec Kit style path for larger features, OpenSpec style deltas for brownfield capability updates, and skills for repeatable team workflows.</p>

<p>For repo-rule placement, compare <a href="/guides/agents-md-vs-claude-md-vs-cursor-rules">AGENTS.md vs CLAUDE.md vs Cursor rules</a> and choose the always-loaded layer your tools actually read.</p>

<h2>Make Verification Part of the Contract</h2>

<p>An agent task is not done when code exists. It is done when the agreed verification evidence exists.</p>

<p>For each spec, define the evidence before implementation starts:</p>

<ul> <li>Required unit, integration, end-to-end, or manual checks.</li> <li>Commands the agent must run and report.</li> <li>Known checks the agent cannot run locally.</li> <li>Expected before-and-after behavior.</li> <li>Review focus areas for humans.</li> <li>Post-merge spec updates.</li> </ul>

<p>Be especially strict when the agent edits tests. Reviewers should know whether a test was added to capture an existing requirement, changed because behavior intentionally moved, or modified because the agent needed the suite to pass. Those are different situations.</p>

<p>Verification should also include negative space. If the agent did not test a migration, did not run a browser flow, or could not verify a third-party integration, the PR should say that plainly. Hidden gaps create false confidence.</p>

<h2>Governance Without Waterfall</h2>

<p>Spec-driven work can sound like a return to waterfall. It should not be. The point is not to freeze design before learning anything. The point is to make changes to intent visible and reviewable.</p>

<p>Use governance where it pays for itself:</p>

<ul> <li><strong>Constitutions:</strong> non-negotiable engineering rules, such as security, compliance, design system, ownership, and deployment constraints.</li> <li><strong>Clarify gates:</strong> stop the agent when requirements are ambiguous instead of letting it guess.</li> <li><strong>Analyze gates:</strong> compare plan, tasks, implementation, and spec before merge.</li> <li><strong>Owner gates:</strong> require review from teams responsible for sensitive paths.</li> <li><strong>Archive gates:</strong> update the durable spec library after the change lands.</li> </ul>

<p>The discipline is in deciding which gate applies. A tiny fix may need only scope and test evidence. A cross-repo authentication change needs written constraints, owner review, rollback notes, and updated source-of-truth docs.</p>

<h2>A Practical Adoption Playbook</h2>

<p>Start with one team and one class of work. Do not begin by asking every engineer to write long specs for every prompt. Begin where review pain is already visible: agent PRs that are too broad, unclear, hard to test, or hard to route.</p>

<ol> <li>Create a short AGENTS.md or equivalent repo rule file with commands, architecture map, ownership notes, and hard constraints.</li> <li>Add a small spec template with intent, boundaries, constraints, task plan, and verification.</li> <li>Define three lanes: tiny fix, feature spec, governed change.</li> <li>Require plan review for feature and governed lanes before broad implementation.</li> <li>Require PRs to link the spec and summarize changed files, evidence, risks, and gaps.</li> <li>Track why agent PRs bounce: unclear scope, missing tests, drift, owner concerns, or spec-code mismatch.</li> <li>After merge, update the durable spec or archive the change delta so future agents inherit the decision.</li> </ol>

<p>This is enough to change behavior. The team no longer debates whether an agent's code "looks fine" in isolation. It asks whether the implementation satisfies a reviewed contract.</p>

<h2>Spec-Driven Development for AI Coding Agents Checklist</h2>

<ul> <li>Put intent, non-goals, constraints, tasks, and verification in repo-visible artifacts.</li> <li>Use lightweight and heavyweight lanes based on change size and risk.</li> <li>Keep always-loaded project context concise and current.</li> <li>Ground plans in repository evidence before implementation.</li> <li>Review plans before agents generate large diffs.</li> <li>Require pull requests to link the spec and show verification evidence.</li> <li>Watch for context drift between spec, plan, code, tests, and PR summary.</li> <li>Update durable specs after merge so the next agent starts from current truth.</li> </ul>

<p>The practical goal is not prettier documentation. It is reviewable agent work. When specs, repo rules, skills, checklists, and gates are treated as execution infrastructure, the agent has less room to guess and the reviewer has a clear standard to enforce.</p> </article>

Get started

Deploy your fleet.

Put a fleet of sandboxed agents to work on your own infrastructure, provisioned in seconds and watched live from one console.

Get started →

Admin-provisioned · Self-host in one command · Your data never leaves your VM