← All posts

Private Repo Evals for Coding Agents That Matter

Private repo evals answer the question public leaderboards cannot: will this coding agent make safe, useful, reviewable changes in our repository, under our constraints, without leaking source code or learning the answer key?

The practical answer is to build SWE-bench style evals from completed internal work, then harden them for private code. Start from real issues and pull requests. Freeze clean base commits. Strip access to Git history and merged diffs. Add held-out tests. Run everything in containers. Keep benchmark exports private. Score both test results and maintainer judgment.

That last part matters. A coding agent that passes a hidden test can still produce a patch your maintainers would reject. For platform teams, the adoption question is not "which model has the highest public resolve rate?" It is "which agent workflow earns trust inside our repos at an acceptable cost and risk level?"

Why Public Benchmarks Are Not Enough

SWE-bench changed coding-agent evaluation by moving from toy functions to real GitHub issues, repository context, patch generation, and test-based grading. The original benchmark includes 2,294 software engineering problems from 12 Python repositories. SWE-bench Verified narrowed that to a human-validated subset of 500 tasks.

Those benchmarks are useful market signals. They are not rollout proof for your engineering organization.

Your repositories have different languages, build systems, service boundaries, CI costs, flaky tests, legacy modules, ownership rules, and review standards. Public benchmark tasks also become familiar over time. Vendors, prompts, harnesses, and agent workflows can optimize against them, intentionally or not.

The gap shows up in research. SWE-Bench Pro, which moves closer to enterprise-like long-horizon work, reports 1,865 problems from 41 actively maintained repositories, with frontier-agent Pass@1 still below 25% in the paper's unified scaffold. SWE-Bench+ found benchmark quality problems such as solution leakage, weak tests, and possible training-data leakage. METR found that roughly half of test-passing SWE-bench Verified pull requests written by mid-2024 to mid or late-2025 agents would not be merged by maintainers.

The lesson is direct: public scores can help you shortlist systems, but private coding agent benchmark results should decide production access.

What Private Repo Evals Should Measure

A good private eval measures the agent, the harness, and the operating model together. The model matters, but so do repository search tools, context limits, retry loops, approval settings, sandbox policy, and the way tasks are written.

At minimum, track these signals:

Signal What it tells you
Resolve rate Whether the agent can pass the required validation for representative tasks.
Maintainer acceptance Whether experienced reviewers would merge the patch.
CI pass and regression rate Whether the change works beyond the narrow hidden test.
Diff size Whether the agent solves narrowly or creates review debt.
Runtime and cost Whether the workflow is economically viable after model, runner, and CI cost.
Retry count Whether success depends on repeated attempts rather than stable behavior.
Tool calls Whether the agent uses shell, network, package managers, and repo APIs safely.
Policy violations Whether the agent touches forbidden files, secrets, credentials, or external services.

Do not collapse these into one score too early. A system that resolves 40% of tasks but violates network policy is not better than a system that resolves 32% cleanly in a sensitive repo. A system that passes tests but doubles review time may not be a productivity gain.

This scorecard should sit beside your broader coding agent evaluation metrics, not replace them with a single leaderboard number.

Build Tasks From Real Pull Requests

The best source material is completed work your team already understands: bug fixes, small features, migrations, dependency upgrades, test fixes, and infrastructure chores. Completed pull requests give you the intent, the base state, the final patch, and the tests that mattered.

For each task, capture these artifacts:

  • Base commit before the human fix.
  • Task prompt derived from the issue or pull request intent.
  • Clean repository snapshot with no answer-bearing history.
  • Dependency and build environment.
  • Held-out tests or validation script.
  • Reference solution for audit and diagnosis.
  • Metadata for task type, repo area, difficulty, owner, and expected runtime.

SWE-Bench++ describes a scalable recipe for turning pull requests into executable benchmark tasks: source the task, synthesize the environment, extract a test oracle, and run quality checks. That shape works for internal evals, but private repositories need stricter controls because the artifacts may include proprietary code, business logic, internal architecture, and reference solutions.

Task construction is not a clerical step. If the base commit is wrong, the dependency environment is stale, or the task prompt gives away the fix, your eval will measure benchmark defects rather than agent ability.

Prevent The Agent From Reading The Answer

Real pull requests are useful because they come from real work. They are dangerous because completed work leaves traces.

An agent should not be able to inspect the merged diff, branch names, commit messages, PR comments, review threads, linked issues with spoiler text, or Git history that reveals the solution. If it can, the eval becomes a retrieval task. The agent may look capable while simply recovering the human answer.

Use these controls when packaging each task:

  • Export a snapshot at the base commit instead of a full clone with useful history.
  • Remove remote branches, tags, and references that point to the solution.
  • Rewrite or redact task text that includes implementation details from reviews.
  • Keep reference solutions outside the agent-visible filesystem.
  • Block internet access unless a task explicitly needs an allowlisted package source.
  • Record all filesystem, shell, network, and repository API access during the run.

This is also where public benchmark contamination lessons apply. SWE-Bench+ reported solution leakage and weak tests at meaningful rates. A private benchmark can fail the same way if the answer is visible through local history or overly specific task metadata.

Use Hidden Tests, But Do Not Trust Them Alone

Held-out tests make private repo evals harder to game. They also create a false sense of safety if they are narrow.

A hidden test can confirm the expected bug fix. It may miss maintainability problems, brittle design, excessive rewrites, bad error handling, concurrency issues, or a change that passes the oracle while violating local conventions. That is why METR's finding about test-passing PRs matters to platform leaders: automated grading can overstate mergeability.

Pair executable tests with a maintainer rubric. Keep it short enough that busy reviewers will use it:

  1. Merge as is: correct, minimal, and consistent with repo norms.
  2. Merge after small edits: direction is sound, cleanup is limited.
  3. Needs revision: useful attempt, but material changes are required.
  4. Reject: wrong, risky, overbroad, or too costly to repair.
  5. Should have abstained: the task lacked enough evidence or belonged to a human.

That rubric turns an eval from "did it pass?" into "would we accept this workflow in production?"

Protect Eval Artifacts Like Source Code

A private eval export can contain repository snapshots, held-out tests, reference solutions, proprietary architecture, and sometimes secrets that were accidentally committed before cleanup. self-bench documentation explicitly warns that these exports should be kept private.

Treat the benchmark as a sensitive system:

  • Store task bundles in a private registry with access logs.
  • Separate eval artifacts from model training and fine-tuning data.
  • Scan snapshots, tests, logs, and agent transcripts for secrets.
  • Limit artifact retention, especially failed workspaces and command logs.
  • Use short-lived repo tokens with only the permissions required for the run.
  • Run tasks in disposable containers or VMs.
  • Deny outbound network by default, then add narrow allowlists when necessary.

Vendor privacy posture still matters. Business and API tiers may offer no-training-by-default commitments that consumer products do not. GitHub has stated that Copilot Business and Enterprise are excluded from its 2026 consumer training expansion. Anthropic's consumer policy changes apply to Claude Free, Pro, and Max, not commercial products. Those commitments help, but they do not replace local controls around code, credentials, logs, and benchmark storage.

For live repository access, pair the eval runner with a credential proxy for AI agents so tokens stay scoped, short-lived, and auditable.

Handle Untrusted Repo Text As An Attack Surface

Private evals often feed agents issue bodies, PR descriptions, comments, README files, and tool output. That text can contain prompt injection. The risk is higher when the agent has access to private code and broad repository credentials.

Recent GitHub MCP and Azure DevOps MCP security discussions show the pattern: malicious issues, comments, or tool results can steer agents into leaking private data or misusing credentials. In an eval runner, that can corrupt results and expose sensitive artifacts.

Set a clear boundary:

  • Task prompts and reviewed eval instructions can define goals.
  • Repository files, issue text, comments, dependency docs, and MCP tool output are context, not authority.
  • Shell commands, network access, credential reads, workflow edits, and repository writes require policy checks.
  • Agent logs should show which source text influenced sensitive actions.

This boundary is the same one used to manage coding agent prompt injection in README files, comments, and tool outputs.

Run Frozen And Fresh Eval Sets

Private repo evals need two tracks.

The first is a frozen baseline set. Use it to compare agents, prompts, harness changes, and vendor upgrades over time. Keep it stable so changes in score mean something.

The second is a rotating fresh set. Add recently completed tasks on a schedule so the organization cannot overfit prompts and harnesses to the baseline. Keep the fresh set private to the smallest group that needs access.

This creates a trade-off. Frozen tasks support longitudinal comparison. Fresh tasks protect against benchmark learning. You need both because platform teams make two different decisions: whether the workflow is improving, and whether the current workflow generalizes to new work.

Stratify both sets by task type and repo area. A single aggregate solve rate can hide a weak agent that succeeds on documentation and test chores but fails shared-library changes, migrations, or build-system work.

A Decision Checklist For Platform Teams

Use this checklist before approving coding agents for private repositories:

  1. Are tasks built from representative completed work, not only easy fixes?
  2. Does each task have a clean base commit and an executable validation path?
  3. Can the agent see Git history, merged diffs, branch names, or review text that reveals the answer?
  4. Are held-out tests stored outside the agent-visible workspace?
  5. Does maintainer review calibrate test pass against mergeability?
  6. Are results stratified by task type, repo area, language, and risk level?
  7. Are eval bundles, logs, snapshots, tests, and reference solutions stored privately?
  8. Are tokens scoped, short-lived, audited, and separated from human credentials?
  9. Is outbound network denied or allowlisted during runs?
  10. Are prompt-injection sources treated as untrusted context?
  11. Do you track cost per accepted task, review burden, retries, and policy violations?
  12. Do you maintain both frozen baseline tasks and rotating fresh tasks?

Private repo evals are the bridge between leaderboard confidence and production trust. Build them from real work, keep the answers hidden, protect the artifacts, and let maintainers judge the patches. Then your adoption decision is based on evidence from your codebase, not optimism from someone else's benchmark.

Get started

Deploy your fleet.

Put a fleet of sandboxed agents to work on your own infrastructure, provisioned in seconds and watched live from one console.

Get started →

Admin-provisioned · Self-host in one command · Your data never leaves your VM