BLOG

Field notes on running agents.

How we build, secure, and run a fleet of self-hosted AI agents.

Portable AI Coding Agent Sessions: Build or Buy Guide

Portable AI coding agent sessions are becoming a platform problem, not a convenience feature. Engineers now start work in one agent, hit a usage limit, switch…

Headless Claude Code Permissions Without Job Hangs

The answer: make unattended approval impossible to miss Headless Claude Code permissions should be explicit, narrow, and non-blocking. In practice, that means…

Visual QA Harness for AI Coding Agents Changing UI

For frontend and platform engineers trying to let coding agents change UI without shipping broken screens, the answer is a layered visual QA harness. Do not…

Context Engineering for Coding Agents in Large Repos

The Short Answer Context engineering for coding agents is the practice of deciding what the agent sees, when it sees it, how that context is refreshed, and…

Designing an MCP Permission Gateway for Claude Code

The short answer: put MCP policy behind a gateway An MCP permission gateway gives Claude Code one approved MCP endpoint, then enforces policy before the agent…

Claude Managed Agents vs Self-Hosted Coding Agents

If you are choosing between Claude Managed Agents, self-hosted execution, and owned open-source agent runtimes, the decision is not mainly about which agent…

Prevent AI Coding Agents From Deleting Files

To prevent AI coding agents from deleting files, do not rely on a prompt that says "never delete." Run agents in isolated workspaces, limit writes to the repo…

AI Coding Agent Ticket Triage Workflow for Team Handoffs

An AI coding agent ticket triage workflow should be read-heavy and review-first: the agent reads the GitHub issue or Jira work item, inspects repository…

Read-Only MCP Server for Kubernetes Logs and CI Failures

A read-only MCP server for Kubernetes logs and CI failures should enforce read-only access in Kubernetes RBAC, CI token scopes, MCP tool execution, redaction,…

Open Source Claude Code Alternatives for Self Hosting

For most platform teams evaluating open source Claude Code alternatives, the practical shortlist is OpenCode or Aider for terminal-first work, OpenHands for…

Best MCP Servers for Coding Agents

The best MCP servers for coding agents are the few that give agents high-signal engineering context under enforceable controls: source control, browser…

Engineering

Codex vs Claude Code cost and usage limits for teams

<article> <p>For engineering leaders and platform teams deciding policy, the practical answer on <strong>Codex vs Claude Code…

Engineering

Coding-Agent Observability: Harness Traces That Matter

<article> <p>You are a platform or security engineer deciding whether a coding agent can touch private repositories. The answer is yes only if…

Engineering

Spec-Driven Development for AI Coding Agent Teams

<article> <p>For an engineering manager or senior engineer trying to make agent-written code reviewable, <strong>spec-driven development for…

Engineering

MCP Tool Allowlist: Stop Tool Drift Before Agents Use It

<article> <p><strong>For platform engineers and security engineers responsible for MCP governance, the answer is direct: an <em>MCP…

Engineering

Claude Code auto mode vs skip permissions for teams

<article> <p>For staff engineers, platform-security leads, and team governance owners, the answer to <strong>Claude Code auto mode vs skip…

Coding Agent Session History Needs Clear UX Design

Coding agent session history should be designed as operational state, not as chat history. The user needs a readable transcript, the model needs a controlled…

Local MCP Server for Desktop App Control: Safe Design

A local MCP server is useful when it exposes a narrow desktop app boundary, not when it turns an agent into a general computer-control layer. Use it for typed…

Human Review Workflow for AI-Assisted Code at Scale

For platform and engineering leaders, a human review workflow for AI-assisted code should do one thing first: protect human judgment from being buried under…

Private Repo Evals for Coding Agents That Matter

Private repo evals answer the question public leaderboards cannot: will this coding agent make safe, useful, reviewable changes in our repository, under our…

Claude Code Effort Levels for Engineering Leaders

Claude Code effort levels should be treated as a team policy knob for cost, latency, and thoroughness. The practical default is conservative: keep routine…

Coding Agent Prompt Injection in READMEs and Comments

Coding agent prompt injection is a trust-boundary failure: untrusted repository text reaches the same reasoning path as trusted instructions, and the agent's…

Claude Code Custom Slash Commands for Team Workflows

Claude Code custom slash commands are now best treated as Skills. Legacy .claude/commands/*.md files still work, but the recommended pattern for reusable team…

Coding Agent Session Manager for Agent Fleets

A coding agent session manager is useful when your problem is no longer "resume the last chat" but "supervise a fleet of agents without losing control of…

Claude Code GitHub Actions Security Checklist

A safe Claude Code GitHub Actions workflow separates untrusted analysis from privileged execution. Treat it as an agentic CI workload with code access, prompt…

Prevent AI Coding Agents From Rewriting Tests in CI

To prevent AI coding agents from rewriting tests, treat existing tests as protected verifier assets. Let agents read tests and add new ones, but block or…

Engineering

One MCP Config for Claude Code, Cursor, and Codex

The short answer: one MCP config can be your source of truth, but it should not be one raw file copied into Claude Code, Cursor, and Codex. MCP standardizes…

Engineering

AI Coding Agent Over-Editing and Diff Size Limits

AI coding agent over-editing is a reviewability failure, not a style annoyance. The agent touches more files than the task needs, bundles behavior with…

Engineering

Coding Agent Action Bias and No Code Change Needed

The best coding-agent result is sometimes no diff. That sounds obvious until you measure it. FixedBench, a 2026 benchmark built from 200 already-fixed…

Engineering

How to Prevent AI Coding Agents Modifying Tests

If an agent can rewrite the verifier, a green build stops meaning what your team thinks it means. The practical way to prevent AI coding agents modifying tests…

Engineering

Claude Code LSP Setup for Production Repositories

The useful version of Claude Code LSP setup is not "turn on code intelligence and trust it." For a production repo, it is a small platform rollout: install the…

Engineering

AGENTS.md vs CLAUDE.md vs Cursor rules for agent teams

<article> <p>A platform team adds Codex to one repo. A frontend team keeps using Cursor. A backend team standardizes on Claude Code. Within a…

Engineering

Claude Code plan mode workflow for team approval gates

<article> <p>A senior engineer asks Claude Code to change a deployment workflow. The agent understands the request, finds the YAML file, edits a…

Engineering

Playwright CLI vs MCP for Claude Code in production

<article> <header> <p>Claude Code can drive a browser in more than one way. That is useful, but it also creates a decision that teams should…

Engineering

Claude Code checkpoints vs Git commits for team workflows

<article> <p>You ask Claude Code to clean up a tangled module. It touches five files, runs a formatter, tries a test command, changes direction…

Engineering

Claude Code monorepo setup for large engineering teams

<article> <p>A monorepo team usually notices the problem in a small, annoying way. Claude Code starts confidently, reads a few files, then reaches…

Engineering

AI coding agent database migration safety model

AI coding agent database migration safety starts with a simple question: what can the agent actually touch? For a platform or security engineer, the decision…

Engineering

Claude Code skills vs subagents vs MCP for teams

Choosing between Claude Code skills vs subagents vs MCP is not a question of which feature is most powerful. For a senior engineer designing a team workflow,…

Engineering

AI Coding Agent PR Reviewability Checklist for Teams

AI coding agents make pull requests cheaper to create. They do not make reviewer attention cheaper. For engineering managers and senior engineers, the…

Engineering

Claude Code subagents best practices for team governance

Claude Code subagents best practices matter most when a team moves from individual experimentation to shared automation. A subagent can keep noisy research,…

Engineering

Claude Code compaction strategy for long coding tasks

A good Claude Code compaction strategy treats compaction as a handoff boundary, not as durable memory. That distinction matters when a coding task spans hours,…

Engineering

Self-hosted Coding Agent Runtime Comparison Guide

The right self-hosted coding agent runtime is the one that matches your trust boundary. For most security-conscious teams, that means per-task isolated…

Engineering

Persistent Memory for AI Coding Agents: Governance

Persistent memory for AI coding agents should be treated as platform infrastructure, not as a bigger prompt. The useful design is layered: small repo…

Engineering

MCP Context Bloat: How to Govern Tool Growth

MCP context bloat is the hidden tax of giving coding agents too many tools at once. Each enabled MCP server can add tool names, descriptions, input schemas,…

Engineering

Claude Code Worktree Database Isolation Guide

Claude Code worktree database isolation means giving every agent its own runtime namespace as well as its own checkout. Git worktrees and Claude Code's…

Engineering

AI Coding Agent Secrets: Protect .env and Tokens

AI coding agent secrets should not rely on .env files, ignore rules, or tool permissions as the main boundary. The workable pattern is layered containment:…

Engineering

Multi-Agent Coding Orchestration Is Ops Now

Multi-agent coding orchestration works when agents run in parallel but review, isolation, approvals, and merge decisions stay controlled.

Engineering

Coding Agent Dependency Security: Privileged Installs

Coding agent dependency security starts by treating package installs as privileged operations, with policy approval, lockfiles, and registry controls.

Engineering

AGENTS.md Best Practices for Coding Agents

AGENTS.md best practices for coding agents: what to include, what to leave out, how to handle monorepos, and where enforcement belongs.

Engineering

Claude Code Hooks for Permissions: A Safer Pattern

Claude Code hooks for permissions can reduce approval fatigue when paired with narrow allowlists, sandboxing, managed settings, and audit logs.

Engineering

MCP Server Monitoring: What Reliability Requires

MCP server monitoring needs more than uptime checks. Track tool discovery, auth, latency, retries, timeouts, and per-tool safety controls.

Engineering

Cloud Coding Agent vs Local CLI Agent: Runtime Guide

Cloud coding agent vs local CLI agent is really a runtime decision: compare ownership, security, secrets, persistence, cost, and PR handoff before rollout.

Engineering

Coding Agent Evaluation Metrics for Real Repos

Coding agent evaluation metrics should measure repo tasks, review quality, cost, regressions, partial progress, and oversight beyond leaderboard resolve rate.

Engineering

Coding Agent Cost Monitoring: Control Spend Early

Coding agent cost monitoring helps platform teams control token, credit, and workflow spend by combining budgets, telemetry, attribution, and policy.

Engineering

Securing Coding Agents in CI/CD: Practical Baseline

Securing coding agents in CI/CD starts with one rule: treat untrusted GitHub content as hostile, then limit secrets, tokens, network, and write access.

Engineering

AI Code Review Agent: Evaluation Guide for PRs

An AI code review agent should be evaluated as a governed PR workflow, not a comment bot. Compare signal, cost, permissions, context, and human review impact.

Engineering

Credential Proxy for AI Agents Without Secret Exposure

A credential proxy for AI agents lets agents call private repos, APIs, CLIs, and MCP tools without exposing raw secrets to runtimes, prompts, or logs.

Engineering

Agent Harness for Coding Agents: Runtime Architecture

An agent harness for coding agents controls sandboxes, state, permissions, tool execution, review flow, lifecycle, and cleanup around safe AI coding work.

Engineering

AI agent observability for safe coding agent rollouts

AI agent observability for coding agents: trace runs, commands, diffs, tests, costs, approvals, and risky side effects before code reaches production.

Engineering

Long-running background AI agents need durable workers

Long-running background AI agents need durable workers, queues, checkpoints, approvals, sandboxing, cost caps, observability, PR review, and control.

Engineering

Persistent AI Agent Workspace Architecture Guide

Design a persistent AI agent workspace with durable files, sandbox snapshots, memory boundaries, rollback, tenant isolation, and clear retention policy.

Engineering

Self-hosted coding agent runtime: build, buy, operate

A self-hosted coding agent runtime gives policy control and data residency, but shifts sandboxing, secrets, audit, cleanup, and capacity to your team.

Engineering

Production Workflows for AI Coding Agents That Scale

Production workflows for AI coding agents need isolated workspaces, reviewable diffs, CI controls, protected merge gates, and accountable human ownership.

Engineering

MCP vs function calling: Practical architecture guide

MCP vs function calling explained for engineers choosing between direct tool calls, MCP servers, runtime discovery, auth boundaries, latency, and reuse.

Engineering

MCP Security for AI Agents: Production Controls

MCP security for AI agents needs token audience checks, sandboxed tools, schema pinning, approval UX, egress limits, and clear incident-ready audit trails.

Engineering

AI Agent Sandboxing: Secure Coding Agent Controls

AI agent sandboxing helps security leaders control source access, secrets, network egress, and coding agent execution risk before agents touch private code.