← All posts

Repo Token Budget for AI Coding Agents: Team Guide

A repo token budget for AI coding agents is not a badge that proves your whole monorepo fits in a model window. It is a practical readiness check: can the agent fit the task's likely working set, instructions, tool schemas, logs, diffs, conversation history, and output reserve without losing the signal it needs?

For staff and platform engineers, that distinction matters. Whole-repo token counts are useful as a bloat signal. They are a poor operating metric by themselves. Real agent work is usually bounded by a package, service, feature area, or recent pull request shape.

Related reading: AI coding agent repo readiness.

What Goes Into a Repo Token Budget for AI Coding Agents

A useful budget includes more than source files. Count the files the agent may read or edit, generated maps or summaries, AGENTS.md, CLAUDE.md, Copilot instructions, tool and MCP schemas, test output, diffs, PR metadata, conversation history, and room for the final patch and explanation.

This is why "full repo fits" is usually the wrong bar for monorepos. The better question is whether the agent can identify and load a bounded working set: the owner service, public interfaces, relevant tests, build config, local instructions, and recent related changes.

Measure Three Budgets

1. Whole-Repo Budget

Measure the repository after excluding obvious noise: generated files, build output, coverage folders, vendored dependencies, snapshots, and large fixtures. Tools such as Repo Tokens can count codebase tokens with tiktoken and publish badges against a configured context window. Configure that window explicitly because defaults can change across versions.

Use this number as a trend. If it jumps, look for committed artifacts, verbose docs, migrations, duplicated fixtures, or generated code drifting into normal traversal paths.

2. Package or Service Budget

Measure each package, service, app, or major module independently. This is much closer to how agents work. A clear service boundary with a small interface is easier for an agent than a flat directory where every change requires broad search.

Package-level counts also tell you where to add local instructions, interface summaries, or architecture indexes.

3. Historical Working-Set Budget

Measure files edited in real work. Count files changed in the last 30, 60, and 90 days. Count touched-file bundles from the last N merged PRs. This gives you a practical estimate of what agents will need for ordinary tasks.

If historical working sets are large, the fix may be architectural. Weak module boundaries cost humans review time and agents context.

Use Simple Scoring

Pick the target context size for the models your team uses, then score typical working sets:

  • Green: under 20 to 30% of target context
  • Yellow: 30 to 60% of target context
  • Red: over 60% of target context

The remaining room matters. Agents need space for instructions, tools, retrieval misses, failed test output, diffs, and the response. If code alone consumes the window, the agent has little room left to work.

Track Overhead Separately

Do not hide every cost inside one number. Measure instruction files, tool schemas, repo maps, logs, diffs, and PR metadata separately. VS Code and GitHub Copilot engineering identify repeated prompt components such as system instructions, tool definitions, repository context, and conversation history. Those fixed costs add up before the agent reads much code.

VS Code reported tool-search median token reductions of 9.81% for GPT-5.4, 8.61% for GPT-5.5, and about 18% for Anthropic models. The lesson is practical: expose capability when needed, not by default. An all-tools MCP config can spend context before the task starts.

Tools That Help

Repo Tokens is useful for a visible whole-repo signal. Repomix can report file and repository counts, respect ignore and include patterns, filter secrets, and compress code using Tree-sitter signatures plus comment and empty-line removal. Aider's repository map takes a different approach by sending a compact symbol map, with a default --map-tokens budget of 1k.

Compression is not automatically a win. Removing comments and examples may save tokens while deleting the only explanation of a business invariant. Use compression where it preserves meaning, and test it on real agent tasks.

Related reading: context engineering for coding agents.

How to Reduce the Budget

Start with low-risk exclusions: generated files, coverage, dist, build directories, vendored code, snapshots, and large fixtures. Then improve retrieval quality: package-level instructions, architecture indexes, interface summaries, repo maps, AST summaries, and quieter tests.

Instruction files deserve special care. Evidence on AGENTS.md is mixed. One study found more than 20% cost increase and lower success. Another found 28.64% lower median runtime and 16.58% lower output tokens. The practical rule is simple: scoped instructions help, bloated root instructions tax every task.

Anti-Patterns to Fix First

  • A giant root AGENTS.md that tries to teach the whole company
  • Generated files committed where agents naturally search
  • Weak service boundaries
  • Repetitive docs copied across packages
  • Huge test logs and stack traces
  • Snapshots and fixtures included in default context
  • Every MCP server and tool exposed to every task

These are not only token problems. They make the right path harder to find.

A Starter Report

Build a weekly report with whole-repo count, package counts, recent touched-file budgets, last N PR working sets, instruction file sizes, tool-schema overhead, largest files, fastest-growing directories, and green/yellow/red scoring.

The standard is not "the whole repo fits." The standard is that an agent can fit the right files, the right instructions, the right tools, and enough feedback to finish safely. Count the whole repo to catch bloat. Budget working sets to improve real agent work.

Get started

Deploy your fleet.

Put a fleet of sandboxed agents to work on your own infrastructure, provisioned in seconds and watched live from one console.

Get started →

Admin-provisioned · Self-host in one command · Your data never leaves your VM