← All posts

Prevent AI Coding Assistant Repeating Mistakes: Guide

To prevent AI coding assistant repeating mistakes, stop treating every correction as a better prompt. If the same failure comes back in a fresh session, you need a durable loop: capture the failure, classify it, promote it to the right artifact, verify it with examples or tests, and prune it when the repo changes.

Memory helps, but memory is soft. Repo instructions, path rules, tests, evals, hooks, policies, and code-review rules are the mechanisms that make a repeated lesson survive across sessions, teammates, tools, and machines.

Related reading: persistent memory for AI coding agents.

Why Agents Repeat Mistakes

Fresh sessions lose context. Local memories may not sync across machines or cloud agents. Long instruction files compete with code, tools, logs, and conversation history. Some rules are visible to the model but not enforced during an edit.

The pattern is familiar: the reviewer says "we do not use inline styles here," the agent fixes it, then the same issue appears in the next session. A Claude Code GitHub issue reported rules visible in context but repeatedly ignored during edit and write operations, with the user correcting the same behavior 5 to 10 times per session.

That maps to the vendor docs. Codex local memories are separate from checked-in AGENTS.md. Claude Code distinguishes memory and instructions from enforcement. Cursor states that large language models do not retain memory between completions and recommends adding focused rules after repeated mistakes.

Prevent AI Coding Assistant Repeating Mistakes With a Failure Loop

Use a simple loop after every repeated agent failure:

  1. Capture: record the failure with a link to the session, PR, CI run, or review comment.
  2. Classify: decide whether it is a convention, invariant, workflow, review judgment, or dangerous action.
  3. Promote: choose the right artifact, not the easiest one.
  4. Verify: add an example, test, eval, hook check, or review rule.
  5. Prune: revisit the rule when the repo or workflow changes.

This turns frustration into infrastructure. "I told it this already" becomes "we captured this once, scoped it, and made it harder to repeat."

Pick the Right Artifact

Use personal memory for helpful recall, such as a developer preference or recurring local setup detail. Do not put required team guidance only in memory. OpenAI's memory guidance frames memories as a recall layer, while team requirements belong in checked-in docs such as AGENTS.md.

Use repo instructions for team conventions: test commands, build order, known failures, timeouts, generated file rules, and directory orientation. Codex reads AGENTS.md before work and rebuilds the instruction chain each run. Nested files closer to the working directory can override broader guidance.

Use path-specific rules for local invariants. Cursor rules in .cursor/rules, GitHub Copilot path-specific instructions, nested AGENTS.md, and Claude project rules all support the same basic idea: put the lesson near the code it governs.

Use tests or evals when the behavior is checkable. OpenAI's eval workflow gives the pattern: define expected behavior, run test inputs, analyze results, and iterate. Do not depend blindly on a hosted eval product without checking current availability; keep the workflow.

Use hooks or policy for must-never-happen behavior. Codex PreToolUse can deny or rewrite supported calls. Claude Code hooks can run at lifecycle events such as PreToolUse, PostToolUse, UserPromptSubmit, Stop, task events, and compaction events. If the agent must not rewrite migrations or weaken auth, enforcement beats advice.

Related reading: local guardrails for Claude Code and Codex.

What Belongs in Team Rules

A good rule is consequential, scoped, and testable. It names the invariant, where it applies, what safe behavior looks like, and who owns it.

Example: "In services/billing/**, do not change invoice total calculation unless the task explicitly mentions billing semantics. If a test failure points here, explain the suspected behavior change and ask before editing calculation code."

Another example: "Do not edit generated files directly. Run the generator and include the command output in the final response."

A weak rule says "follow best practices." It will be ignored because it gives the model no local signal.

What Not to Put in Rules

Do not paste an entire style guide into every agent context. Do not add generic lint advice that CI already enforces. Do not encode rare edge cases that distract from common work. Do not duplicate long docs. Do not keep contradictory rules alive because nobody owns cleanup.

Instruction bloat is its own failure mode. If every session loads a long constitution, the important rule is easier to miss. Focused rules near the code are more useful than broad advice at the root.

Build a Failure Ledger

Create a small ledger for repeated failures. Each row should include failure title, source session or PR, root cause, chosen durable artifact, owner, scope, verification, and review or expiry date.

Example entries:

  • Skipped build before push, promoted to repo instruction plus stop hook that warns when final response lacks verification
  • Inline interfaces in frontend code, promoted to path-specific rule with safe example
  • Test rewritten to hide failure, promoted to hook/policy requiring explicit user request before test edits
  • Known flaky command timeout, promoted to Copilot or AGENTS instruction with correct command order and timeout

The ledger prevents random rule growth. It also lets platform teams measure whether repeated review comments are actually decreasing.

Add Regression Checks

Vendor benchmarks show why local checks matter. SWE-bench contains real GitHub issue tasks, and SWE-bench Verified has 500 human-filtered tasks, but UTBoost still found insufficient tests and erroneous patches incorrectly labeled as passed. Public benchmarks are useful signals. Your repo needs its own regressions.

For review-heavy behavior, OpenAI's Codex code-review rules pattern is strong: encode a non-obvious invariant near the code, state the safe path, and test the rule with a violation, a safe counterexample, and an unrelated change. OpenAI reported rule-guided variants recovering 98% of required custom findings versus 58.3% baseline in its primary eval suite.

A Practical Checklist

  • Collect the last 20 repeated agent review comments
  • Group them by convention, invariant, workflow, dangerous action, or review judgment
  • Promote each to memory, repo rule, path rule, test, eval, hook, policy, skill, or code-review rule
  • Keep each rule narrow and close to the code it affects
  • Add a violation example and a safe counterexample for judgment-heavy rules
  • Block must-never-happen actions with hooks or CI
  • Assign owners and expiry dates
  • Measure repeated comments, failed builds, rework, and false positives

The durable fix is not a longer prompt. It is a maintained failure ledger connected to rules, tests, hooks, and review checks. That is how a team teaches an AI coding assistant once and makes the lesson survive the next fresh session.

Get started

Deploy your fleet.

Put a fleet of sandboxed agents to work on your own infrastructure, provisioned in seconds and watched live from one console.

Get started →

Admin-provisioned · Self-host in one command · Your data never leaves your VM