← All posts

Background AI Agent Approval Inbox: Team Design Guide

A background AI agent approval inbox starts to matter the first time your team has more than one unattended agent waiting on a human. One wants network access. One opened a PR that needs workflow approval. One needs a product decision. One is blocked on a file-write permission. A terminal prompt cannot be the control plane for that.

The useful design is a durable queue with state, scope, routing, audit, and resume semantics. It is not a notification list with approve and deny buttons. If the system cannot explain what is being approved, who owns it, how long the approval is valid, and how the agent resumes, the inbox will become another place where risky work hides.

Related reading: long-running background AI agents.

Why Terminal Prompts Do Not Scale

Terminal approvals work for one attended local session. They fail when agents run for minutes or hours, pause after the developer has left, create PRs, or resume from another client. They also fail as a team workflow because the person who sees the prompt may not be the owner of the path, service, environment, or risk.

Current products already expose the primitives. OpenAI Codex App Server supports stored threads, resume, and server-initiated approval requests keyed by thread and turn identifiers. GitHub Copilot cloud agent task states include waiting_for_user, alongside queued, in progress, completed, failed, idle, timed out, and cancelled. Claude Code supports allow, ask, and deny permissions, persistent local approvals, comments on prompts, hooks, and managed policies.

AWS Pizza Bot made the user experience concrete with an email-like inbox: All for thread history, Unread for completed work, and Action for paused work waiting on approval or an answer. The reaction to its launch showed the product question clearly: should this feel like email, Linear, a desktop app, a web app, or a ticket queue? For platform teams, the deeper question is authority.

What Belongs in a Background AI Agent Approval Inbox

Coding-agent approvals are not one object. A useful inbox should distinguish:

  • tool_call
  • permission_grant
  • network_egress
  • file_write
  • workflow_run
  • pr_review
  • merge_gate
  • human_question
  • policy_exception

OpenAI's separation of command and file approvals from permission and MCP elicitation flows points in the same direction. A command approval, a clarification question, and a PR workflow gate need different payloads, reviewers, scopes, and expiry rules.

The Approval Card Anatomy

Every card should answer what a reviewer would ask before clicking anything: which repo, which branch or worktree, which agent, which session, what payload, what command or diff preview, what reason, what risk level, what policy source, who requested it, who can approve it, when it expires, and where the agent will resume.

The minimum actions are approve, deny, edit-and-approve, request changes, delegate, snooze, and cancel run. The scope selector should default narrow: once, turn, session, repo, or org policy. Denial comments should become both agent context and audit evidence.

Related reading: MCP permission gateway.

State Model and Resume Semantics

The inbox needs durable state outside the live client. A practical state model includes created, notified, viewed, approved, denied, resumed, completed, failed, timed out, and cancelled. GitHub's cloud-agent states already show why waiting_for_user belongs in the lifecycle rather than in a side channel.

Resume is the sharp edge. Research on persistence layers warns about re-executed effects and concurrent resume races. The inbox should use stable request IDs, consume-once resolution, idempotency keys, and stored continuation state. Approval granted and effect committed should be separate events.

That distinction prevents a common audit lie. Clicking approve does not mean the command succeeded. It means a reviewer allowed the agent to attempt the next effect.

Routing Should Reuse Existing Authority

Do not invent a parallel chain of command if the organization already has CODEOWNERS, CI workflow approvals, deployment gates, GitHub environments, branch protection, and security exceptions. Route by path, repo, environment, permission type, tool, or risk class.

A file write under billing should go to the billing owner. A workflow approval on an agent-authored PR should follow repository policy. A network request to a vendor API should route differently from a local test command. High-risk mobile approvals should be constrained, not treated like clearing an informational notification.

Security and Audit Requirements

GitHub warns that auto-running workflows on Copilot-authored PR branches can expose write permissions or Actions secrets to unreviewed code. That is the type of risk an approval inbox has to surface, not hide. Claude Code's docs also make a useful distinction: permissions are enforced by Claude Code, not by the model.

A good audit log records requested, notified, viewed, decision, resumed, and effect result. It should distinguish policy auto-approval from human approval. It should show stale warnings when code, branch, policy, or the target resource changed after the request was created.

A Minimal Implementation Sketch

Start with five tabs: Needs action, Running, Completed unread, Failed or timed out, and All. Persist approval items in a service independent of the agent client. Store request payloads, risk labels, route decisions, reviewer actions, timestamps, expiry, and continuation references. Integrate deep links into GitHub or GitLab for PR review, checks, workflow approvals, and merge gates.

Then add policy. Low-risk repeat actions can be auto-approved by rule. Consequential actions should route to owners. Expired requests should force the agent to refresh context. Two reviewers racing on the same item should produce one final decision, not two partial resumes.

The value of a background AI agent approval inbox is control, not click reduction. It makes asynchronous agent work governable. The system should tell a clear story: what paused, why it needed a human, who decided, what resumed, and what actually happened next.

Get started

Deploy your fleet.

Put a fleet of sandboxed agents to work on your own infrastructure, provisioned in seconds and watched live from one console.

Get started →

Admin-provisioned · Self-host in one command · Your data never leaves your VM