Claude Code Effort Levels for Engineering Leaders
Claude Code effort levels should be treated as a team policy knob for cost, latency, and thoroughness. The practical default is conservative: keep routine coding on Sonnet, reserve Opus for harder agentic work, use Fable only for the most difficult long-running tasks, and leave effort at high until your own evals prove that medium is good enough or that xhigh pays for itself.
The mistake is to manage effort as if it were only extra thinking time. Anthropic's docs frame effort as total work per turn: how many files Claude reads, how much it verifies, how many tools it calls, how far it pushes before asking, and how much explanation it writes. For an engineering organization, that makes effort part of the operating model, not a personal preference hidden inside one developer's terminal.
Claude Code effort levels are a workload policy
Claude Code supports low, medium, high, xhigh, and max on current effort-aware models. It also has a session-only ultracode mode that sends xhigh and coordinates more dynamic workflows for substantive work.
Current Claude Code docs list high as the default for every effort-supporting model except Opus 4.7, which defaults to xhigh. The API docs also say high is equivalent to leaving the effort parameter out. That matters because stale community advice about lower defaults still circulates. Teams should pin explicit settings and record the Claude Code version they assume in policy.
Effort is calibrated per model. The same word does not mean the same hidden budget across Sonnet, Opus, and Fable. A policy that says "all hard work uses xhigh" is incomplete unless it also says which model, which task class, which approval gate, and which budget applies.
The three controls: model, effort, and execution
A useful Claude Code governance policy has three layers:
| Layer | Decision | Why it matters |
|---|---|---|
| Model | Which model family is available and which one is the default | Sets capability, latency, and per-token price |
| Effort | Which effort level starts a session and which levels are capped | Controls how much work Claude attempts per turn |
| Execution | Which tools, permissions, budgets, and environments are allowed | Controls blast radius when the agent acts |
Do not confuse defaults with enforcement. Managed settings can set model, availableModels, enforceAvailableModels, effortLevel, permission rules, sandboxing, telemetry, and customization controls. But effortLevel in managed settings is a starting default. Where available, organization effort limits are the enforcement mechanism for effort caps.
For model restrictions, availableModels is the stronger control. It covers main sessions, aliases, fast mode, subagent and teammate models, skill or command model frontmatter, advisor model, and background agents. If Fable should not be generally available, do not rely on guidance in a README. Restrict it.
A default policy that will not surprise finance
Start with this baseline, then tune from real task data:
| Workload | Default model | Effort | Control |
|---|---|---|---|
| Routine edits, lint fixes, small refactors | Sonnet, or cheaper subagents where appropriate | high by default, test medium after evals |
File and diff budgets |
| Daily agentic coding | Sonnet 5 or Opus 5 | high |
Normal permission prompts and telemetry |
| Ambiguous debugging, architecture, large migrations | Opus 5 | high or xhigh |
Plan review before edits |
| High-stakes long-running work | Fable 5 by exception | high or xhigh |
Usage credits, spend caps, and human checkpoints |
| Broad exploration and log-heavy analysis | Haiku or Sonnet subagents when quality holds | low, medium, or high by eval |
Narrow tools and summary-only outputs |
The reason is simple economics. Anthropic pricing docs list Sonnet 5 at $2 per million input tokens and $10 per million output tokens, Opus 5 at $5 and $25, and Fable 5 at $10 and $50. Higher effort tends to produce more tool calls and more output. Output-heavy work on a more expensive model can become a direct spend multiplier.
That does not mean cheaper is always better. A weak model at high effort can still waste money if it retries, edits the wrong files, or needs repeated human rescue. The right comparison is task outcome per dollar, not tokens per turn.
Teams that need more detail on spend reports, token traces, and subagent fan-out should connect this policy to their coding agent cost monitoring process.
Set caps by role, not by enthusiasm
Enterprise org admins can cap maximum effort per model by custom role. Levels above the cap are hidden in /effort, and attempts to run above the cap through --effort or /effort run at the cap.
A practical role policy looks like this:
- Default developers: Sonnet available, Opus available for approved task classes, effort capped at
highorxhighdepending on team maturity. - Senior engineers and maintainers: Opus available,
xhighallowed for debugging, migrations, and architecture work. - Platform and incident owners: temporary access to Fable and higher effort for specific workspaces, with logging and spend review.
- CI, headless automation, and background agents: explicit model and effort settings, narrow tools, non-interactive billing rules understood before rollout.
Be careful with max. The docs describe it as a level that can show diminishing returns and overthinking. Treat it as an exception path for measured cases, not a badge for difficult work.
Do not let subagents hide spend
Subagents and agent teams are useful because they isolate context. They can keep the main session cleaner when one worker explores logs, another reads a subsystem, and another summarizes risk. They also create fresh contexts that consume tokens.
Anthropic's cost guidance highlights two cost traps. First, agent teams in plan mode use approximately seven times more tokens than standard sessions. Second, subagents can feel cheap because they save main-context space, while still charging for their own context, tools, and output.
Use subagents when isolation changes the quality of the work. Do not create a cast of overlapping assistants because it looks organized. Give each subagent a narrow description, narrow tools, and an explicit model. For broad exploration, a custom Explore subagent with model: haiku can force cheaper discovery when that quality level is acceptable.
For delegation patterns, pair this policy with Claude Code subagents best practices so teams limit agent fan-out deliberately.
Permission modes are part of the same decision
Effort controls how hard Claude works. Claude Code permissions control what that work can touch.
Manual prompts are safer for risky repositories, but they interrupt flow. acceptEdits accepts edits and common filesystem commands. plan explores without edits. auto uses a safety classifier. dontAsk denies tools unless they were already approved. bypassPermissions skips most prompts and belongs only in isolated environments.
For team rollout, the default should be boring: require plan mode for large or high-risk changes, keep manual approval around sensitive paths, and disable auto or bypassPermissions if your policy requires human approval. Managed settings can prevent developer override for those permission choices.
This is where cost and security meet. A high-effort agent with broad permissions can read more files, call more tools, and make more changes before a person looks closely. If the repository includes migrations, deployment config, dependency manifests, secrets-adjacent paths, or public APIs, permission gates should be stricter than the model default.
Teams should also define Claude Code hooks and permissions so effort policy and approval modes reinforce each other.
Vary effort by session, not every turn
It is tempting to change effort constantly: low for reading, high for editing, xhigh for debugging, medium for summaries. The platform docs warn that changing effort between cached requests can invalidate prompt caching. For long cached sessions, that creates avoidable cost and latency.
Set effort at the session or workload level. If the task changes enough to need a different effort policy, start a new session or make the change explicit. This also makes telemetry easier to interpret. A cost report that mixes five effort levels inside one task is harder to learn from than a report organized by workload class.
Map controls by deployment type
The governance surface depends on how Claude Code is deployed.
| Deployment | Primary controls | Policy note |
|---|---|---|
| Team and Enterprise seats | Org model restrictions, org default models, role-based effort limits, managed settings | Use caps where available and publish team defaults |
| Anthropic API and claude.ai enterprise orgs | Organization model and effort controls, API budget controls, telemetry | Track token economics directly |
| Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS | Managed files, MDM, provider budgets, gateway policy | Do not assume every Anthropic org control reaches the provider path |
| Gateway-managed environments | Central routing, spend controls, audit logs, model allowlists | Use the gateway to normalize policy across providers |
Provider divergence matters. If one team uses subscription seats and another uses a cloud provider route, the same written policy may need different enforcement mechanics. Write policy in terms of outcomes: allowed models, effort caps, permission gates, budget owners, and telemetry requirements.
Measure before lowering the default
Anthropic's cost docs give enterprise averages around $13 per developer per active day and $150 to $250 per developer per month, with 90 percent of users below $30 per active day. Those numbers are a baseline, not a substitute for your own pilot.
Before setting medium as the team default, run evals on real work:
- Small bug fix in a known area.
- Ambiguous bug with missing context.
- Refactor with tests and reviewable diff requirements.
- Large exploration task with logs or unfamiliar code.
- High-risk change touching migrations, auth, dependencies, or deployment config.
Track pass rate, reviewer intervention, files read, tool calls, tests run, changed files, reverted work, latency, and total cost. If medium matches high on routine tasks, use it there. If it skips files, tests, or verification on harder tasks, keep high or move the work to Opus rather than hoping prompt wording will fix the gap.
A rollout checklist for Claude Code governance
- Choose a default model by seat type and workload, not by individual preference.
- Set
availableModelsandenforceAvailableModelsfor hard model boundaries. - Use
effortLevelas a managed default, then apply org effort limits where available. - Keep
highas the baseline until evals justify lower effort for specific tasks. - Restrict
xhigh,max, and Fable to task classes with owners and budget visibility. - Give subagents explicit models, narrow tools, and clear reasons to exist.
- Require plan mode or approval gates for large diffs and high-risk paths.
- Disable risky permission modes if your organization requires human approval.
- Avoid changing effort repeatedly inside long cached sessions.
- Monitor usage through
/usage, analytics, OpenTelemetry, spend reports, and per-user or per-team alerts.
The decision is not "higher effort is better." The decision is which work deserves more model capability, which work deserves more agent persistence, and which work should stop for human review. Claude Code effort levels are useful when they are tied to workload classes, permission gates, and budget telemetry. Without those controls, effort becomes another hidden way for agent work to become expensive, slow, or hard to review.