← All posts

Context Engineering for Coding Agents in Large Repos

The Short Answer

Context engineering for coding agents is the practice of deciding what the agent sees, when it sees it, how that context is refreshed, and where its limits are. In a large repo, the useful goal is not to squeeze the whole codebase into a prompt. The useful goal is to make the agent's work reviewable: which instructions applied, which files were retrieved, which commands were run, which evidence supported the change, and which controls remained outside the model.

For platform teams, the operating model is simple. Treat context as infrastructure. Version it, scope it, lint it, evaluate it, and audit it. Keep broad guidance short. Use path-scoped rules for local conventions. Prefer exact search for known symbols and errors. Add codebase indexes and prior experience only when selection quality is measured. Put hard constraints in permissions, hooks, CI, code owners, sandboxing, and review gates, not in prompt text.

The trade-off is measurable. Repository instruction files can help when they route the agent to the right files and commands. They can also raise cost, increase latency, and mislead the agent when they are bloated, stale, or vague. Public findings are mixed: one study reports no general success improvement and over 20% higher inference cost from repository context files, while another reports 28.64% lower median runtime and 16.58% lower output token use with AGENTS.md. The difference is not magic syntax. It is context quality.

Related reading: coding agent context management.

The Failure Mode: The Agent Loses the Thread

The familiar failure starts with a reasonable ticket. A developer asks an agent to change behavior in a service inside a monorepo. The repo has multiple packages, generated files, old tests, service-specific commands, and conventions that live in code review history rather than in one document. The agent reads a root overview, opens a plausible file, edits the wrong layer, and then runs a generic test command that does not exercise the changed package.

The patch may even look competent. That is the danger. The reviewer now has to reconstruct what the agent missed: the local fixture, the ownership boundary, the service-specific test, the generated file rule, the security exception, and the reason a nearby implementation was not the right one. Review becomes archaeological work.

Context engineering changes the shape of that work. Instead of asking the model to remember everything, the platform gives it a staged path: global rules, repo rules, local package rules, exact search, codebase index, call sites, tests, logs, and task evidence. The agent still has to reason. The difference is that its reasoning can be inspected against concrete inputs.

What Context Engineering for Coding Agents Includes

A practical context stack has several layers, each with a different owner and failure mode. Mixing them together in one always-loaded file creates cost and drift.

LayerWhat belongs thereMain risk
Global or org contextSecurity policy, approved tools, workflow defaults, credential constraintsPolicy text mistaken for enforcement
Repo-root contextArchitecture map, build and test commands, service boundaries, ownership mapBroad summaries that go stale
Path-scoped contextPackage commands, local conventions, fixtures, domain gotchasConflicting instructions across nested files
Task contextAcceptance criteria, relevant logs, known files, user constraintsOne-off details leaking into durable memory
Retrieved evidenceSymbols, imports, errors, call sites, tests, recent commits, docsNoisy retrieval that hides the useful file
Durable memoryRecurring repo lessons with provenance and review rulesVague memories that never expire
Skills and playbooksLong procedures that load only when relevantSkill leakage into unrelated tasks

This separation matters because coding agents do not need every layer on every task. Anthropic distinguishes always-loaded memory from skills that load only when relevant. Its Claude Code docs also recommend concise, specific, structured instructions, with a target under 200 lines per CLAUDE.md. OpenAI's AGENTS.md docs state that Codex reads AGENTS.md files before work and uses layered guidance, with a default combined project-doc maximum of 32 KiB. Those limits are not administrative trivia. They are a signal that always-loaded context is expensive.

Instructions Are Context, Not Controls

AGENTS.md, CLAUDE.md, and similar files shape model behavior. They do not enforce safety or correctness. That distinction should be explicit in any platform design.

Use instruction files for facts the agent needs to do ordinary work: exact commands, package boundaries, generated-code rules, ownership hints, sensitive paths, test expectations, and review evidence. Use enforcement layers for actions the agent must not be able to bypass: permissions, CI, hooks, sandboxing, code owners, branch protection, secret scanning, and human review.

The security case is concrete. Backslash Security describes AGENTS.md as an indirect prompt-injection surface, especially in non-interactive execution modes with ambient developer credentials. A repository instruction file is part of the software supply chain for an agent. Treat it like code: review changes, limit who can edit it, and assume malicious or stale text can influence behavior.

Related reading: coding agent security controls.

What AGENTS.md Should Do in Large Codebases

AGENTS.md is useful when it gives the agent a compact operating brief. The open AGENTS.md format recommends a root file plus nested files for monorepos, with the nearest file taking precedence. Suggested content includes overview, build and test commands, style, testing, and security. Codex concatenates global, project, and nested guidance, with closer files appearing later and able to override earlier guidance.

In practice, the root file should answer five questions. What is this repo? Which commands prove a change? Where are the service boundaries? Which files are generated or sensitive? What evidence should appear in the final response or PR? The nested file should answer the local version of the same questions for a package, service, or domain.

A large monorepo does not benefit from a root file that tries to summarize every subsystem. That pattern burns tokens on information unrelated to the ticket. It also creates a maintenance problem: once the overview is wrong, the agent follows a confident map to the wrong place.

Static Overview Versus Live Exploration

The strongest pattern in the brief is the tension between static guidance and live retrieval. Broad repo overviews are easy to write, but precise task-triggered retrieval often matters more.

Cursor's search docs still emphasize exact search for function names, variables, errors, and regexes. Cursor Agent uses grep automatically for specific symbols and Instant Grep for large codebases. That should not be surprising to staff engineers. When the ticket names an error, symbol, migration, or API, lexical search is often the highest-signal first move.

Semantic and shared indexes still matter. Cursor reports that naive indexing of large repos can take hours, while teammate index reuse can cut large-repo time-to-first-query from hours to seconds. Its 2026 indexing post also reports average 92% similarity between clones inside organizations. The platform implication is practical: index reuse can reduce startup friction, but privacy, access control, deletion, hashing, and retention rules need to be part of the design.

Related reading: large codebase setup for coding agents.

The Evidence Is Mixed, So Measure the Workflow

The research brief does not support a blanket rule that repository instruction files always help. It supports a narrower, more useful conclusion: context files help when they improve selection, routing, and verification enough to offset their cost.

FindingReported resultPlatform reading
Gloaguen et al. on repository context filesNo general task-success improvement, inference cost up by over 20%Do not treat root overviews as free performance gains
Lulla et al. on AGENTS.md efficiency28.64% lower median runtime and 16.58% lower output token use, with comparable completion behaviorGood instructions may reduce wasted work
Shepard and Albrecht on probe-and-refine guidance33.0% mean resolve rate, compared with 28.3% static knowledge base and 25.5% unguided baselineGenerated and refined guidance can help when it improves coverage and reaches correct files
SWE Context BenchCorrectly selected summarized prior experience improves accuracy and reduces runtime and token cost, while unfiltered or wrong prior context can hurtSelection quality is the control point
Configuration-smells studyLint Leakage in 62%, Context Bloat in 42%, Skill Leakage in 35% across 100 popular reposInstruction files need maintenance, not only adoption

These results point in the same operational direction. Measure context as a system, not as a document. Track whether the agent selected the right files, whether it ran the right tests, how many tokens it spent, how long it took, how much review correction it needed, and whether the patch stayed maintainable.

Exploration Quality Is the Hidden Metric

Large-codebase work is often lost before the first edit. SWE-Explore isolates repository exploration across 848 issues, 10 languages, and 203 repos. The average codebase in that benchmark has 759 non-test files and 179.6K non-test lines of code, while the ground truth averages 4.3 files, 4.7 regions, and 1,578 lines. The signal is sparse.

That distribution explains why "more context" can hurt. The correct answer usually lives in a small slice of the repo. The agent needs a way to find that slice, cite it, and keep unrelated architecture text out of the working set.

Other benchmarks extend the concern beyond patch success. SWE-Bench Pro includes 1,865 problems from 41 active repos, with tasks that can require hours to days for a professional engineer and often span multiple files. SWE-Cycle reports solve rates dropping sharply when agents move from isolated tasks to full issue lifecycle in bare repos. SWE-CI evaluates long-term maintainability with tasks averaging 233 days and 71 consecutive commits. A context strategy that helps a benchmark patch but increases review burden or architectural drift is not a platform win.

A Platform Pattern That Works

For a platform team, the practical design is a ranked context plan. The agent should start from task context, load only the relevant instruction layers, search for exact symbols and errors, expand into semantic or codebase index results when needed, and cite the evidence it used. Long procedures should live in skills or playbooks that load on demand. Durable memory should include provenance, review status, and expiry rules.

The context plan should also have budgets. A useful budget names maximum instruction size, maximum retrieved file count before summarization, maximum log length, maximum prior-memory items, and required evidence for a final patch. Without budgets, context grows until the reviewer becomes the filter.

Recommended metrics should cover five surfaces: retrieval quality, patch quality, cost and latency, review burden, and safety. Retrieval quality asks whether the agent reached the correct files and regions. Patch quality asks whether tests and behavior line up. Cost and latency track runtime, output tokens, and index wait time. Review burden tracks review comments, rework, and diff size. Safety tracks denied actions, prompt-injection exposure, and changes to sensitive files.

Related reading: private repo evals for coding agents.

Maintenance Rules for Context Files

A context file should have an owner, a review path, and a deletion path. Add a rule when the agent repeatedly makes the same mistake or when a convention is expensive to rediscover. Remove a rule when CI, a linter, a type checker, or a hook already enforces it. Rewrite a rule when it names old commands, obsolete owners, or broad claims without a file path.

Scoping is the cheapest improvement. Put global norms in global instructions, repo facts in the root file, service facts in nested files, and long procedures in skills. A path-scoped rule that tells the agent the correct test command for one package is more useful than a root paragraph listing every test framework in the company.

Conflict handling should be documented. If closer files override root files, say so. If a tool has different load order semantics, note the difference in the tool-specific bridge file. GitHub announced Copilot coding agent support for AGENTS.md custom instructions on 2025-08-28, while Claude Code uses CLAUDE.md and can import AGENTS.md. Portability is improving, but tool-specific semantics still matter.

The Decision Frame

There are two common options. The first is a document-centered approach: create one root instruction file, keep adding lessons, and hope larger context windows absorb the growth. The second is an infrastructure approach: split context by scope, retrieve evidence by task, measure selection quality, and enforce hard rules outside the prompt.

Comparatively, the document-centered approach is faster to start and easier to explain. Its cost appears later: stale summaries, context bloat, contradictory instructions, and review work that shifts from code quality to agent forensics. The infrastructure approach takes more platform design, but it gives teams a way to ask better questions: which context changed the outcome, which context was noise, and which constraint should never have been left to the model.

The practical conclusion is conditional. If your agents mostly handle small, local changes, a concise root AGENTS.md plus exact commands may be enough. If they work across large codebases, monorepos, generated code, private APIs, or unattended workflows, context engineering becomes part of the delivery platform. The work is not to make the agent omniscient. The work is to make its limited view intentional, current, and reviewable.

Get started

Deploy your fleet.

Put a fleet of sandboxed agents to work on your own infrastructure, provisioned in seconds and watched live from one console.

Get started →

Admin-provisioned · Self-host in one command · Your data never leaves your VM