← All posts

Local MCP Server for Desktop App Control: Safe Design

A local MCP server is useful when it exposes a narrow desktop app boundary, not when it turns an agent into a general computer-control layer. Use it for typed app context, stable object handles, and reviewable actions. Treat screenshots, mouse control, shell commands, AppleScript, and broad OS automation as a separate high-risk tier.

If your desktop product, internal tool, or developer workflow needs to connect to Claude Code, Codex, Cursor, VS Code, or another MCP-capable agent, the hard part is no longer proving MCP can call a tool. The hard part is deciding what the tool is allowed to touch on a user's machine.

What a local MCP server should actually do

A local MCP server should translate desktop app state into a small set of resources and tools that an agent can understand. Resources are best for read-only app state: current selection, document metadata, design context, file IDs, node IDs, window IDs, or workflow handles. Tools are for specific actions: export this asset, inspect this selection, generate code from this frame, apply this named operation, or write approved assets into a project.

That distinction matters because MCP tools are model-controlled. Once a host discovers a tool, the model may decide to call it. The client should still show inputs, ask for confirmation on sensitive operations, validate results, apply timeouts, and log usage. But the server author should not rely on the host UI alone. The desktop app still owns the local trust boundary.

The best pattern is semantic control. Give the agent get_design_context, export_selected_asset, or create_component_from_selection. Avoid starting with click, type, run_script, or move_mouse. Pixel coordinates are brittle. Generic script execution is hard to audit. Semantic app actions are testable, explainable, and easier to deny safely.

Figma shows the production shape

Figma is the clearest production example. Its desktop MCP server runs through the Figma desktop app, is enabled from Dev Mode, and exposes a local endpoint at http://127.0.0.1:3845/mcp. Figma documents setup for tools such as VS Code, Cursor, Claude Code, and Codex. For Codex, the setup uses Streamable HTTP with a server named figma-desktop.

The product lesson is more important than the URL. Figma does not frame the desktop server as the default answer for everyone. Its docs position the local server for specific organization and enterprise use cases, while recommending the remote MCP server for most users because the remote path can offer broader feature support and centralized account authorization.

That split is a useful decision rule. Use a desktop MCP server when the agent needs app-local state, selection context, enterprise routing constraints, or access to local project files that a remote endpoint cannot safely reach. Use a remote server when OAuth, rollout, policy, logging, and account-level governance matter more than local app state.

Choose transport as a security decision

The MCP 2026-07-28 transport model supports stdio and Streamable HTTP. For desktop control, both can be reasonable, but they create different operating models.

Transport Best fit Main risk Required controls
stdio Client-launched local subprocesses and developer tools Command execution through local configuration or inherited privileges Approved commands, restricted environment, sandboxing, timeouts, and logs
Streamable HTTP Desktop apps that host their own loopback endpoint Other local processes reaching the server 127.0.0.1 binding, Origin validation, authentication, and per-call authorization

For a desktop app that already runs continuously, Streamable HTTP is ergonomic. The agent connects to a loopback endpoint and sends JSON-RPC requests as HTTP POSTs. But loopback is not a security boundary by itself. MCP security guidance treats local MCP server compromise as a distinct threat because local servers may have direct access to the user's system and may be reachable by other local processes.

For Streamable HTTP local servers, the spec guidance is concrete: validate Origin, bind only to 127.0.0.1 where possible, and implement authentication. In practice, that usually means a user-enabled toggle in the desktop app, a generated local token, client registration or pairing, and authorization checks on every request.

Design the capability surface in risk tiers

Do not ship one flat set of agent tools. Classify capabilities by risk and make the UI, policy, and logs reflect that classification.

Tier Examples Default posture
Read app context Selection state, document metadata, node IDs, design context Allow after user enables the server, with audit logging
Generate or export Asset export, code generation, screenshot of selected frame Require visible parameters and target paths
Write app state Create, rename, move, delete, or mutate app objects Require confirmation and app-side policy checks
Write local files Download assets into a project or update generated files Constrain directories and show exact paths
Run scripts or control the OS UI AppleScript, JXA, shell, PyAutoGUI, mouse, keyboard, clipboard, OCR Deny by default, allow only for narrow admin-approved workflows

Community desktop-control MCP projects show real demand for AppleScript, Windows UI automation, screenshots, OCR, mouse control, keyboard input, clipboard access, and mobile device control. They also show why generic computer use should not be the default integration pattern for a product team. Those tools cross app boundaries. They can touch private data that the desktop vendor never intended to expose through its own API.

OS permission prompts help, especially on platforms such as macOS where Automation permissions can restrict cross-app control. They are not enough. The MCP server still needs allowlists, schema validation, URI validation, path sanitization, timeouts, output limits, and logs.

Stateful workflows need explicit handles

The 2026-07-28 MCP model is stateless and per-request oriented for Streamable HTTP. Desktop workflows often are not. A user selects a frame, an agent asks for context, exports an asset, then writes files into a local project. Do not hide that state in an implicit connection.

Return explicit handles. A tool can return a selection_id, file_id, node_id, export_job_id, or workflow_id. Later calls should validate that the handle is still authorized, still belongs to the current user, and still points to the same app object or a safe successor. If the selection changed, say so. If the file was closed, fail clearly.

This also improves reviewability. "Export node 12:45 from file abc to src/assets/button.svg" is a real approval prompt. "Use current selection" is weaker because the user and audit log may not agree about what the selection was when the call happened.

Approval belongs in both the host and the app

MCP clients should show tool inputs before calls and ask for confirmation on sensitive operations. That is necessary, but desktop app vendors should assume client approval UX will vary across Codex, Claude Code, Cursor, VS Code, and future hosts.

Put critical controls in the desktop app too. A practical model looks like this:

  • The user explicitly enables the MCP desktop server.
  • The app shows which local clients are connected where possible.
  • Read-only tools are separated from write tools.
  • Destructive actions require app-side confirmation or policy approval.
  • File writes are constrained to approved directories.
  • High-risk tools can be disabled by organization policy.

For enterprise users, this is where product policy catches up with SDK convenience. The TypeScript and Python MCP SDKs make server implementation accessible. They do not decide your permission model, your audit format, or your rollback story.

Log the full action trail

An audit log should let an operator reconstruct what happened without reading an agent transcript. Log the client identity if available, server version, app version, tool name, arguments after redaction, resource IDs, approval result, output summary, error state, and whether the operation changed app state or local files.

For write actions, include before and after identifiers where possible. For file writes, include the target path and whether the file was created, overwritten, or skipped. For exports, include source object IDs and output asset names. For denied actions, log the policy reason.

For operations, connect this trail to MCP server monitoring so tool calls, timeouts, and per-tool failures are visible outside the agent transcript.

Implementation checklist for desktop app control

  • Start with resources for read-only app state and narrow tools for named actions.
  • Use stable object handles instead of screen coordinates or implicit current state.
  • Keep Streamable HTTP endpoints on 127.0.0.1, validate Origin, and require authentication.
  • Validate all tool inputs with typed schemas and reject unknown fields where practical.
  • Validate resource URIs and sanitize file paths, especially for file:// resources.
  • Separate read, export, write, file-write, script, and OS-control capabilities.
  • Require confirmation for destructive actions and local file writes.
  • Set timeouts and output limits for every tool.
  • Log tool calls with redaction, resource IDs, approvals, results, errors, client identity, and app version.
  • Test with the official SDK examples and MCP Inspector before broad rollout.

The practical bottom line

A local MCP server is a good bridge when the desktop app has context an agent cannot get through a remote API. That bridge should be narrow, typed, authenticated, reviewable, and logged.

For platform teams, the decision is not "should agents control desktop apps?" The useful decision is which app-level capabilities deserve an MCP surface, which ones need human approval, and which ones should stay out of the agent path entirely. Start with semantic app actions. Add local file writes carefully. Treat raw OS automation as an exception, not the foundation.

Get started

Deploy your fleet.

Put a fleet of sandboxed agents to work on your own infrastructure, provisioned in seconds and watched live from one console.

Get started →

Admin-provisioned · Self-host in one command · Your data never leaves your VM