Local Guardrails for Claude Code and Codex Projects
Local guardrails for Claude Code and Codex should stop harmful but plausible changes: weakening auth, rewriting old migrations, editing generated files, hiding a failing test, or changing billing behavior without owner approval. A sandbox can say where an agent may write. It cannot know which product invariant it just broke.
The practical model has four layers. Instructions orient the model. Permissions and sandboxing set coarse boundaries. Hooks, rules, policy scripts, and validation checks enforce deterministic tripwires. CI, CODEOWNERS, branch protection, pre-commit, SAST, and human review remain the organizational backstop.
Related reading: AI agent sandboxing.
Why Local Guardrails for Claude Code and Codex Need More Than Prompts
"Do not weaken authentication" in AGENTS.md or CLAUDE.md is guidance. A pre-tool hook that blocks edits to auth-sensitive files unless the task has approval is enforcement. Both are useful. They solve different problems.
Claude Code reads project instructions from CLAUDE.md and can also read AGENTS.md. Codex reads AGENTS.md and layers global, project, and nested directory instructions, with closer files taking precedence. These files help the agent choose a better path, but they are not enough for must-never-happen behavior.
Anthropic states that Claude Code permission rules are enforced by Claude Code, not by the model, with deny, ask, and allow precedence. OpenAI separates sandbox mode from approval policy: sandbox controls what Codex can technically do, while approval policy controls when Codex asks. That distinction is the foundation for sane local policy.
Use Instructions for Context
Use AGENTS.md and CLAUDE.md for norms the agent should see before work starts: repo layout, test commands, ownership boundaries, generated file rules, and examples of safe changes. Keep them short enough to earn their place in context.
Good instruction: "Do not modify tests to make failures pass unless the user explicitly asks for a test change. If a test appears wrong, explain why and ask before editing it."
Weak instruction: "Always write high-quality secure code." It sounds responsible, but it gives the agent no local constraint.
Use Permissions and Sandboxes for Boundaries
Codex's common local default, workspace-write with on-request, lets it edit inside the workspace and asks before internet or boundary-crossing actions. Claude Code supports fine-grained allow, ask, and deny rules, permission modes, persistent local approvals, hooks, and managed policies.
These controls answer coarse questions: can the agent write here, run this command, access the network, or call this tool? They should be narrow by default and easier to broaden for a specific session than to audit after broad access caused damage.
Use Hooks and Policy Scripts for Tripwires
Hooks are where tribal knowledge becomes testable. Claude Code hooks run deterministic lifecycle commands. Codex hooks can run scripts or MCP tools during lifecycle events, and PreToolUse can block or rewrite supported calls. PostToolUse cannot undo side effects, so high-risk checks belong before the action.
Start with tripwires that match real incidents:
- Block edits to old migrations unless an owner approval file or ticket label is present
- Ask before changing OAuth scopes, cookie flags, rate limits, or billing calculations
- Block generated file edits unless the generating command ran in the same session
- Deny test rewrites when the user asked for implementation changes
- Require verification evidence before the agent claims completion
- Stop commands that target production accounts from local sessions
Related reading: Claude Code hooks for permissions.
Claude Code Pattern
For Claude Code, put shared project settings in .claude/settings.json when teammates need the same permissions, hooks, telemetry, and plugins. Keep personal local overrides separate. Use CLAUDE.md or AGENTS.md for model-visible guidance, and hooks for deterministic enforcement.
Claude settings can be user, project, local, or managed. Managed settings override lower-precedence settings, which matters for enterprise rollout. If security owns a rule, do not bury it in a personal config.
Codex Pattern
For Codex, place required team guidance in checked-in AGENTS.md files. Use nested files when a rule only applies to a service or directory. Configure sandbox and approval policy intentionally. Use command rules and codex execpolicy check where command policy fits, while noting that command rules are documented as experimental.
Codex hooks are useful for lifecycle enforcement, especially PreToolUse checks that can deny or rewrite supported tool calls before the effect happens.
Make the Policy Portable
Claude and Codex expose different surfaces, so teams may need a neutral policy file compiled into each tool. Keep the neutral layer boring: path patterns, protected actions, approval sources, allowed commands, denied commands, and validation requirements.
Example policy idea: files under services/billing/** require billing-owner approval for calculation changes; files under migrations/** deny edits older than the current release window; files under auth/** require tests and owner review when cookie, token, or scope behavior changes.
Operate Guardrails Like Code
Guardrails can become stale or noisy. Give each rule an owner, scope, reason, test example, and review date. Track false positives and blocked incidents. Review changes through normal PR flow. Guardrail scripts themselves are part of your supply chain, so do not let agents edit them without review.
OWASP's Agent Control Standard calls for agents to be inspectable, traceable, and instrumentable. The local version of that idea is simple: when a guardrail blocks an action, the team should know which rule fired, what it saw, who owns it, and what the safe path is.
The goal is to let agents work while stopping them from "fixing" the wrong thing in a way that looks reasonable until review is already expensive. Put soft knowledge in instructions, hard constraints in hooks and policy, and final authority in CI and human review.