← All posts Engineering

Coding-Agent Observability: Harness Traces That Matter

<article> <p>You are a platform or security engineer deciding whether a coding agent can touch private repositories. The answer is yes only if coding-agent observability covers the whole harness, including LLM request logs, tool calls, approvals, file changes, retries, costs, and outcomes. You need to reconstruct what the agent saw, which model calls it made, which tools it invoked, what changed on disk, who approved risky actions, what retries happened, what it cost, and why the run stopped.</p>

<p>The practical test is incident response. A developer reports that an agent opened a pull request with a suspicious dependency change. The final diff is not enough. You need the ordered trace: session root, user prompt, repository context, tool definitions, shell command, package output, approval decision, file writes, retry loop, compaction event, tests run, and final outcome. Without that chain, the investigation becomes guesswork.</p>

<p>This belongs with the same controls in a <a href="/guides/coding-agent-security-checklist">coding agent security checklist</a>, because traces are audit evidence, not only developer diagnostics.</p>

<h2>The Harness Is The System Of Record</h2>

<p>A coding-agent run is not a single API response. OpenAI's Codex agent-loop writeup describes the run as repeated cycles of model inference, tool calls, context updates, and code edits. The harness is the part that turns those cycles into work: it builds prompts, exposes tools, executes shell and file operations, handles approvals, appends observations, retries failures, compacts context, and eventually produces a patch or final response.</p>

<p>That makes the harness the system of record. If your observability stops at model latency, token totals, and final answer text, you can see that something happened but not why. Platform teams need traces to debug failed runs and compare vendors. Security teams need them as the audit trail for identity, repository, policy version, command, file path, data movement, approval source, and final diff.</p>

<p>The buying implication is direct: do not evaluate coding agents only on IDE experience or benchmark claims. Ask whether the product emits replayable, exportable, tool-complete traces across every runtime mode you plan to use: CLI, IDE, web, headless, subagent, MCP server, and CI-run agent.</p>

<h2>A Useful Trace Shape For Coding-Agent Observability</h2>

<p>The minimum useful trace starts with a stable session root. Every turn, model call, tool call, permission decision, file change, and outcome should attach to that root. OpenTelemetry GenAI conventions already cover model and provider metadata, request and response IDs, prompt and response metadata, tool definitions, tool-call IDs, tool arguments and results, token counts, and workflow or session identifiers. Content capture is opt-in.</p>

<p>A practical internal schema should preserve these fields even if vendor names differ:</p>

<table> <thead> <tr> <th>Trace layer</th> <th>Fields to capture</th> <th>Why it matters</th> </tr> </thead> <tbody> <tr> <td>Session</td> <td><code>session_id</code>, user, repo, branch, workspace, harness version, policy version</td> <td>Links a run to identity, source scope, and governance rules.</td> </tr> <tr> <td>Turn</td> <td><code>turn_id</code>, parent session, prompt metadata, context sources, compaction state</td> <td>Shows what the agent was responding to at each step.</td> </tr> <tr> <td>Model call</td> <td>provider, model, request ID, response ID, token counts, retries, latency, cost</td> <td>Supports debugging, vendor comparison, and cost attribution.</td> </tr> <tr> <td>Tool call</td> <td><code>tool_call_id</code>, tool name, arguments, result summary, exit code, duration</td> <td>Explains how the agent moved from intent to action.</td> </tr> <tr> <td>Approval</td> <td>decision, requester, approver or policy source, reason, timestamp</td> <td>Separates autonomous action from human or policy-authorized action.</td> </tr> <tr> <td>Filesystem</td> <td>path, operation, diff hash or artifact reference, accepted or reverted status</td> <td>Connects the trace to the actual code change under review.</td> </tr> <tr> <td>Outcome</td> <td>tests run, final status, stop reason, PR link or patch artifact, feedback</td> <td>Lets teams measure useful work instead of raw agent activity.</td> </tr> </tbody> </table>

<p>This is also where cost accounting becomes real. Per-seat quota views do not answer cost per run, per subagent, per model call, or per accepted pull request. The OpenAI Developer Community has asked for token usage per Codex conversation or task, which reflects the same platform need: aggregate usage is not operational attribution.</p>

<p>Use <a href="/guides/opentelemetry-for-ai-agents">OpenTelemetry for AI agents</a> as transport, then normalize vendor fields into a schema your team owns.</p>

<h2>OpenTelemetry Is Necessary, But Not Sufficient</h2>

<p>OpenTelemetry is becoming the common transport for agent telemetry. Codex has opt-in OTel export for runs, API requests, SSE events, prompts, tool approvals, and tool results, with prompts redacted unless explicitly enabled. Claude Code exports OTel metrics, logs or events, and optional traces for usage, costs, and tool activity, including session IDs, model and request IDs, tool-use IDs, API attempts, token and cost metrics, code-edit decisions, hooks, compaction, and redaction defaults.</p>

<p>The trade-off is maturity. The OpenTelemetry GenAI event documentation is still marked Development, and detailed inference events are opt-in. That should not stop adoption. It should change the architecture. Emit OTLP now, but put a versioned internal transform layer between vendor spans and dashboards. Your dashboards and audit queries should depend on your canonical schema, not directly on every current <code>gen_ai.*</code> attribute name.</p>

<p>Also verify coverage by entry point. A Codex GitHub issue reported that OTel coverage differed by runtime mode: interactive CLI had telemetry, while <code>codex exec</code> and <code>codex mcp-server</code> had gaps. Another OpenAI community report described missing tool-call spans in realtime traces. The operational lesson is simple: "tracing enabled" is not the same as tool-level completeness.</p>

<h2>Replay Without Leaking The Company</h2>

<p>Full replay wants prompts, file contents, tool outputs, model responses, and diffs. Security wants redaction, retention limits, encryption, and gated access to raw material. Treat this as a tiered telemetry design instead of a philosophical argument.</p>

<p>Keep metrics and structured events always on: session IDs, model IDs, tool names, token counts, costs, timings, approval decisions, exit codes, and file paths where policy allows. Sample detailed traces for normal debugging. Enable incident-triggered raw capture only under access policy, with retention limits and explicit handling for source code, secrets, and proprietary output.</p>

<p>SigNoz reports that content capture changed one single-tool DeepSeek Harness turn from 6,350 bytes to 73,078 bytes, mostly from repeated tool schemas. That is a storage and privacy warning at the same time. Raw content helps replay, but it can multiply telemetry volume and capture material you do not want in a general observability store.</p>

<p>A practical compromise is to store diffs and bulky tool outputs as access-controlled artifacts, then put hashes or artifact references in the trace. For sensitive repositories, the trace should prove what happened without making every dashboard user a source-code reader.</p>

<p>Raw replay policies should follow your <a href="/guides/ai-coding-agent-secrets">AI coding agent secrets</a> rules, since prompts and tool outputs can contain credentials or proprietary code.</p>

<h2>Control Points Belong Around Tools</h2>

<p>Model reasoning is not a reliable audit artifact. The useful evidence is observable behavior: selected tool, arguments, command output, file write, approval decision, tests run, diff accepted, or diff reverted. Hooks and pre-tool or post-tool events are natural places to capture that evidence and enforce policy.</p>

<p>Claude Code hooks fire around sessions, turns, and tool calls. DeepSeek Harness describes a plugin-based model where "Every capability is a plugin" and "Every run is traceable." Its product material says session logs record system prompts, reasoning, tool calls and results, subagent scheduling, context injection, resume, fork, search, and replay. SigNoz documents a LoongSuite plugin that turns DeepSeek Harness turns into OTel spans for entry, agent, reasoning step, LLM call, and tool execution, which suggests useful OTel support exists through third-party instrumentation rather than a stable native path.</p>

<p>For security engineering, the key is not whether the hook is elegant. The key is whether every tool boundary can produce an auditable event before and after execution. A shell command should record command text, working directory, policy decision, exit code, and output summary. A file write should record path, operation, diff reference, and whether the change survived final review. An MCP tool call should record server, tool name, arguments, result summary, and data egress classification where available.</p>

<h2>What An Incident Review Should Ask</h2>

<p>Suppose an agent modifies authentication code and adds a package. The incident review should not start with "what did the model think?" It should start with trace reconstruction.</p>

<p>First, identify the session: user, repository, branch, workspace, harness version, policy version, and runtime mode. Second, inspect context: user prompt, repository instructions, files read, tool results injected into context, and any compaction event. Third, inspect decisions: model calls, tool selections, shell commands, approvals, retries, and error recovery. Fourth, inspect artifacts: file changes, tests run, final diff, PR comments, and stop reason.</p>

<p>If any of those steps cannot be answered, the observability system has a gap. The gap may be acceptable for a low-risk pilot. It is not acceptable for production use against sensitive repositories.</p>

<h2>DeepSeek Harness Shows The Direction And The Risk</h2>

<p>DeepSeek Harness is a useful case study because it is positioned as a developer-preview, plugin-based coding-agent harness with source code included. Its session logs cover many of the fields platform teams ask for: prompts, tool calls, results, subagents, context injection, resume, fork, search, and replay.</p>

<p>The same case study shows why trace depth matters. The Agentic Harness Engineering preprint reports Terminal-Bench 2 pass@1 improving from 69.7% to 77.0% after ten harness-evolution iterations using component, experience, and decision observability. Use that as preprint evidence, not as a procurement guarantee. The important point is the mechanism: harness observability becomes an input to improving the harness itself.</p>

<p>Security review points in the same direction. A DeepSeek Harness security assessment used 14,560 controlled executions and trace-level source-to-sink analysis for indirect prompt injection. That reinforces the basic security claim: final answers and final diffs are not enough. You need traces to study how untrusted instructions moved through context, tools, and outputs.</p>

<h2>A Buying Checklist For Platform And Security Teams</h2>

<p>Use this checklist before standardizing on a coding-agent platform or harness:</p>

<ol> <li>Does it export OTLP traces, metrics, and events, or only vendor dashboards?</li> <li>Does it document the schema for sessions, turns, model calls, tool calls, approvals, file changes, retries, compaction, and outcomes?</li> <li>Are prompts, tool arguments, tool results, and model responses redacted by default, and can raw capture be enabled under policy?</li> <li>Can you verify tool-call completeness across CLI, IDE, web, headless, subagent, MCP-server, and CI modes?</li> <li>Can traces be joined to control-plane audit logs for web, admin, and policy activity?</li> <li>Can costs be attributed per run, model call, subagent, repository, branch, and accepted pull request?</li> <li>Can file diffs be stored as hashes or artifact references instead of inline source content?</li> <li>Can hooks or policies block, require approval for, or annotate risky shell, filesystem, network, and MCP actions?</li> <li>Can the vendor provide raw-event access, retention controls, encryption, and export paths for incident response?</li> <li>Can your team keep a canonical internal schema while OTel GenAI conventions continue to evolve?</li> </ol>

<p>The strategic decision is not logs versus traces. It is whether the agent harness can become accountable production infrastructure. Coding-agent observability should let you debug a failed run, investigate a suspicious change, compare vendors, measure cost per accepted outcome, and govern access without storing every secret-bearing byte by default.</p>

<p>If a product cannot show what the agent saw, what it called, what changed, who approved it, and why it stopped, it is not ready to be your system of record for private-code automation.</p> </article>

Get started

Deploy your fleet.

Put a fleet of sandboxed agents to work on your own infrastructure, provisioned in seconds and watched live from one console.

Get started →

Admin-provisioned · Self-host in one command · Your data never leaves your VM