Read-Only MCP Server for Kubernetes Logs and CI Failures
A read-only MCP server for Kubernetes logs and CI failures should enforce read-only access in Kubernetes RBAC, CI token scopes, MCP tool execution, redaction, rate limits, and audit logs. Do not rely on prompts or hidden write tools. If the agent can call a mutating endpoint directly, the system is not read-only.
The platform goal is simple: let a coding agent inspect the evidence behind a failed build, failed deploy, or broken pull request, then propose a fix. The agent should read Kubernetes pod logs, warning events, resource status, GitHub Actions logs, or GitLab job traces. It should not run kubectl apply, exec into pods, rerun jobs, delete workflow logs, approve deployments, read secrets, or mutate repositories.
MCP fits this workflow because it standardizes how agents connect to external data and tools. The server becomes an authority boundary, so treat it like production infrastructure.
This design depends on the same controls used for MCP security, MCP tool allowlists, and securing coding agents in CI/CD.
What read-only must mean
Read-only has to hold at every layer. If one layer is missing, the others are compensating controls.
| Layer | Control | Block |
|---|---|---|
| Kubernetes identity | Namespace-scoped Role with only required read verbs and resources. | create, update, patch, delete, pods/exec, secrets. |
| CI identity | GitHub Actions: read or GitLab scoped read_api. |
Rerun, cancel, delete logs, write repo content, approve environments. |
| MCP server | Tool allowlist and execution-time authorization. | Direct invocation of hidden or mutating handlers. |
| Operations | Redaction, content caps, rate limits, and audit trails. | Secret leakage, API overload, and unaudited evidence access. |
The execution-time detail is not theoretical. The brief cites a 2026 Manifold Security report about a read-only bypass in a Kubernetes MCP server where restrictions were enforced at discovery but not execution. Test the direct call path, not only the visible tool list.
Start with the failure workflow
Do not expose a broad cluster assistant first. Start with the workflow your platform engineers actually need.
- The agent receives a failing PR, build, or deploy signal.
- It fetches failed CI jobs and relevant log excerpts.
- It maps commit, branch, image tag, or deployment metadata to a Kubernetes workload.
- It reads recent pod logs, previous container logs where needed, workload status, and warning events.
- It returns a cause analysis and patch recommendation.
That workflow needs evidence. It does not need cluster control.
Kubernetes: grant logs, events, and status only
For Kubernetes, the key permission is the pods/log subresource. Kubernetes documentation says pod logs are available through the Pod API and callers must have access to read the Pod log subresource. A minimal log-reader Role usually grants get and list on pods and pods/log in a namespace, plus read access to selected workload and event resources.
Keep the Role namespace-bound unless you have a clear multi-namespace operating model. A central gateway can still serve multiple teams, but each request should run under an identity scoped to the repo, service, namespace, or investigation.
Useful Kubernetes log tool parameters include:
namespacepodcontainerlabel_selectortail_linessince_secondspreviouslimit_bytes
Be careful with events. Kubernetes Event API documentation says events have limited retention and should be treated as informative, best-effort, supplemental data. Events are useful clues, not an authoritative timeline.
Watches need limits too. Kubernetes watches are first-class API verbs for streaming changes, but an MCP server should cap watch duration, namespace scope, concurrency, and result volume. Otherwise, a few agent sessions can create unnecessary API-server load.
CI: read logs, do not control jobs
For GitHub Actions, use fine-grained permissions that allow reading workflow runs, workflow jobs, and logs. GitHub's REST documentation says workflow run and log download endpoints require Actions: read, while deleting workflow logs requires Actions: write. Exclude write permissions from the MCP identity.
For GitLab, the Jobs API can list jobs by project or pipeline and retrieve a job log trace through GET /projects/:id/jobs/:job_id/trace. Prefer the narrowest useful token. GitLab documents that personal tokens can span all resources available to the user, group tokens span a group, project tokens are limited to one project, and read_api grants read API access.
Do not give the diagnostic server these capabilities:
- Retry or cancel jobs.
- Delete workflow logs.
- Approve environments or pending deployments.
- Write repository contents or pull requests.
- Read repository secrets.
- Publish packages or deployment artifacts.
GitHub's official MCP server is useful here because it includes an Actions toolset with tools such as actions_get, actions_list, and get_job_logs. It also supports toolsets and individual tool selection, which helps reduce context size and improve tool selection.
Treat logs as sensitive and untrusted
Logs are not harmless because they are read-only. They may contain tokens, customer data, internal hostnames, stack traces, request bodies, or test fixtures. They may also contain attacker-controlled text. A failed test can print instructions that try to steer the agent.
- Strip terminal control characters and unsafe formatting.
- Redact known secret patterns before content reaches the model.
- Cap bytes, lines, and time windows by default.
- Prefer failed spans over full logs.
- Label tool output as untrusted diagnostic data.
- Do not let log text choose new tools or widen permissions.
OpenTelemetry recommends data minimization and processors for removing, filtering, redacting, or transforming sensitive telemetry data. Its Logs Data Model also gives teams a stable shape for normalized logs before agent summarization.
Use short-lived identity
MCP authorization guidance requires OAuth 2.1 practices, secure token storage, HTTPS or localhost redirect URIs, PKCE, and short-lived access tokens where possible. Apply the same instinct to Kubernetes and CI identities.
For in-cluster deployments, Kubernetes bound service-account tokens are a better fit than long-lived static credentials. Kubernetes documentation says projected service-account tokens are time-bound and, by default, expire after one hour or when the Pod is deleted.
For CI providers, prefer GitHub Apps or fine-grained tokens over broad personal tokens. For GitLab, prefer project or group tokens with only the scope needed for job traces and pipeline metadata.
Design the tool surface narrowly
| Tool | Purpose | Guardrail |
|---|---|---|
ci_list_failed_jobs |
Find failed jobs for a workflow run or pipeline. | Repo and project allowlist. |
ci_get_job_log_excerpt |
Return failed log spans or a capped tail. | Byte cap, redaction, no delete or retry endpoint. |
k8s_get_pod_logs |
Read current or previous container logs. | Namespace RBAC, tail_lines, since_seconds. |
k8s_get_warning_events |
Fetch recent warning events for a workload. | Retention caveat in output. |
k8s_get_workload_status |
Inspect deployment, pod, and container state. | No mutation verbs, no secrets. |
AWS's EKS MCP Server runs in read-only mode by default and exposes tools such as get_pod_logs and get_k8s_events. The containers Kubernetes MCP server supports --read-only, --disable-destructive, toolsets, denied resources, telemetry, OAuth/OIDC for HTTP mode, and automatic credential redaction in MCP logging. Those features are useful, but still test execution directly.
Control API load and audit the trail
Agents are good at asking follow-up questions. That helps diagnosis and can stress shared APIs. GitHub documents common authenticated REST limits of 5,000 requests per hour, while GITHUB_TOKEN gets 1,000 requests per hour per repository. Secondary limits include concurrency and endpoint-per-minute constraints.
Set defaults: limit concurrent investigations, default logs to a small tail, cache immutable CI logs briefly, use extracted error frames first, cap Kubernetes watches, and return truncation metadata instead of hiding it.
At the MCP layer, log enough to reconstruct the investigation without storing full sensitive logs forever: agent session ID, human user or service principal, repository, pull request, workflow run, pipeline, job, cluster, namespace, workload, pod, container, tool name, filtered arguments, result size, truncation status, downstream API status, latency, denied calls, and policy reason.
Implementation checklist
- Create namespace-scoped Kubernetes Roles for pod status,
pods/log, selected workload reads, and events. - Exclude secrets, exec, port-forward, attach, and all mutating verbs.
- Use GitHub
Actions: reador equivalent read-only CI permissions. - Use GitLab project or group tokens with
read_apiwhere that is the narrowest useful option. - Allowlist specific MCP tools and enforce the allowlist at execution time.
- Redact sensitive log content before model exposure.
- Cap log lines, bytes, time windows, watch duration, and concurrency.
- Label Kubernetes events as supplemental because retention is limited.
- Store audit records that connect user, agent, repo, job, namespace, pod, and tool call.
- Test direct invocation of mutating methods, including paths hidden from discovery output.
The rule is simple: the agent may inspect evidence, but every write path belongs outside this server. If the platform keeps that boundary in identity, tool execution, and audit, a read-only MCP server becomes a useful triage surface instead of a quiet production control plane.