Our agents could already think. They couldn't reach a server.
That gap is why most "AI-powered" operations still need a person in the middle. Someone pastes credentials. Someone clicks deploy. The agent writes the plan, and a human runs it. That's a very expensive copy-paste job.
We closed the gap with thin MCPs. Here's the proof from day one: an SSH run (a secure remote login) at 10:16 found 29 plugins with 2 updates due. Chromium landed on our shared computer's disk at 186.8MiB. The build queue closed 4 of 7 parts by noon.
This is a 12-minute build log. By the end, you'll know:
- What a thin MCP is, and why it holds no judgement
- When an agent needs a gate and when it needs a whole computer
- The two bugs and one hosting limit a live WordPress run exposed
- What's still blocked on a person, and why we say so
- Five steps to build your first one
Only want the steps? Skip to the last section.
01 / what-is-a-thin-mcp
What is a thin MCP?
A thin MCP is a gate, not a brain. It's the smallest piece of code that gives an agent real server access and takes the laptop out of the path.
MCP stands for Model Context Protocol. It's an open standard for connecting AI apps to outside systems (think of it as a universal plug between an agent and your tools). The spec defines three things a server can expose: tools, resources, and prompts. It defines two transports (the ways messages travel): stdio for local processes and Streamable HTTP for remote ones. And it defines a discovery step. The client calls tools/list before it calls anything else.
The spec doesn't decide how much power you hand over. That's on purpose. Wiz puts it plainly: MCP standardises how models discover and call tools. It leaves authentication, authorisation, and transport security to whoever builds the server. The protocol is the wiring. The judgement is yours. Every MCP server architecture decision that matters sits in that gap.
So we made one: the gate holds the credential and the agent holds the judgement, and neither one ever holds both.
There are three stops in the signal path. Only one of them thinks.
- The agent reads the account and decides what should happen.
- The gate authenticates and refuses anything off its allowlist.
- The server runs the named call and returns the result.

02 / whats-inside-the-gate
What's inside the gate?
Three layers, in the order a request travels through them.
The key. One credential per client, stored as an environment variable on the gate. The agent never receives it and never logs it. It can't leak what it never holds. The rest of the industry landed here too. AWS ships OAuth for its own MCP server with short-lived tokens. AWS is explicit that authorising an agent grants it no extra permissions. Every request is still checked against your existing IAM policies (the access rules already on your AWS account).
The jacks. Named operations, spelled out. Ten tools on our WordPress gate. Seventy on the Vercel one. If there's no tool for a thing, the thing can't happen. That list is the entire attack surface (every way in for an attacker). It's the first time in this business anyone could say the attack surface out loud as a number.
The guard. Snapshot before writing, health check after, revert on failure. The server enforces this. A prompt doesn't request it. No amount of prompt engineering skips a gate that lives in the code path.

Nothing clever lives in the gate. That's the design, not a shortcut. Anything smart belongs in the agent. There, you change it by editing a sentence instead of redeploying code.
03 / when-does-an-agent-need-a-whole-computer
When does an agent need a whole computer?
Only when the job can't be one named call. Two shapes cover almost everything.
Use a thin MCP when the job is one named API call. A deploy, a rollback, an env var, a plugin update. No state, no shell, no filesystem. You can read the blast radius (how much breaks if it goes wrong) off the tool list.
Use a microVM when the job needs a real computer. A microVM is a tiny, isolated virtual machine. Ours is a Fly Sprite: a Firecracker microVM with a persistent disk. It sleeps when idle and wakes on request. The disk is the whole argument, and the real number makes it best. Chromium lands once at 186.8MiB instead of being paid for on every cold start. The browser session survives between calls. That's what makes login-gated pages workable.

Our rule: destructive writes stay behind named operations. The computer is for open-ended work, not routine execution. Hardware isolation protects us. It does nothing for the client whose live site the SSH key on that box can reach.
04 / what-did-we-build-and-what-broke
What did we build, and what broke?
We built three components. The plan changed twice once it hit a real site.
WordPress: 10 tools, verified live
wp-admin-mcp covers plugin and theme state, site health, update history, and response checks. It's multi-tenant off environment variables. Onboarding a client is three variables and no code change.
We verified it against our own WP Engine test install, not a fixture. The results: WordPress 7.0.4, PHP 8.4.25, homepage returning 200 in 34ms, and two plugin updates pending. One premium plugin was correctly flagged as blocked. It advertises a newer version but ships no download package. The morning run on the board showed the same kind of evidence: SSH at 10:16, 29 plugins listed, 2 updates due.
Then came the finding that rewrote the rollout plan.
On WP Engine, the web process can create a new file in wp-content/plugins but can't overwrite an existing one. A plugin upgrade is an overwrite. So every REST-driven update (one triggered through WordPress's web API) dies at the copy step with Could not copy file. The same upgrade through wp-cli over SSH, on the same site, succeeds.
Every check you'd normally run first gives a false positive here:
get_filesystem_method()reportsdirect.DISALLOW_FILE_MODSis false.is_writable()returns true.- Writing a brand new file works.
Only rewriting an existing file exposes it. That one behaviour killed the connector-plugin approach entirely. That covers ours and the off-the-shelf option whose source we read. Both call the same Plugin_Upgrader::upgrade() that the host refuses.
Two more bugs surfaced only because we ran a real update instead of trusting the happy path:
- The upgrader reported success for an upgrade that never landed. Both
get_plugins()and opcache (PHP's code cache) serve the old header inside the same request that did the upgrade. So the version read back unchanged and still looked fine. Now we prove success by re-reading the version from disk with opcache invalidated. A version that didn't move is reported as a failure. - The upgrader turned a plugin off and left it off. WordPress deactivates a plugin to swap its files. Our search plugin came back inactive after a "successful" update. Now we capture activation state before and reassert it after. A failed reactivation says so loudly, instead of leaving a client's plugin silently disabled.
Vercel: 70 tools, deployed, waiting on a token
We forked an MIT-licensed open-source server instead of writing one. Upstream is stdio only, and our agent platform can't connect to that. So we added a Streamable HTTP route over the same 70 tools. We left upstream's source untouched so the fork stays mergeable. tools/list returns all 70 with valid schemas. The endpoint returns 401 without a token and 200 with one.
It isn't doing real work yet. It needs a team-scoped API token. We won't reuse a full-account credential on an endpoint that anyone holding the MCP secret can reach. That's a deliberate stop, not a delay.
The shared computer: provisioned on paper, blocked on a human
The Sprite that replaces the one machine everything used to depend on is written, scripted, and idempotent (safe to run twice). It's blocked at step one on a browser login that only a person can complete. The platform mints its own tokens through a browser. There's no server-side path around it.
We're saying that out loud because a status board that only shows green isn't a status board.

05 / why-is-the-allowlist-the-security-model
Why is the allowlist the security model?
Because MCP adoption has outrun MCP hygiene, and the numbers aren't close.
- Research across 500+ scanned MCP servers found 38% had no authentication at all.
- A separate audit of 9,695 public MCP servers catalogued 4,982 security issues across 2,259 servers. That includes 2,054 instances of missing authentication, 880 cases of arbitrary file access, and 476 of command injection.
- Governance lags too. Only 21% of organisations have governance controls around non-human identities (service accounts, bots, and agents). Yet non-human identities outnumber people at 83% of organisations surveyed.
None of that argues against giving agents server access. It argues for the boring parts:
- Deny by default.
- One credential per client.
- Named operations only.
- Short-lived tokens where the platform supports them.
- A guard in the code path, not in a prompt.
Here's what doesn't work, and where this is weaker than a human. An allowlist can't stop a well-formed but wrong instruction. If update_plugin is on the list, an agent that picks the wrong plugin will succeed at doing the wrong thing. The allowlist bounds the blast radius. It doesn't supply judgement. That's what the snapshot is for. It's also why anything destructive still ends at a person in our operation.
06 / what-does-this-buy-a-team
What does this buy a team?
Every person on the team gets the same server access the founder has. Zero laptops sit in the deployment path.
Before this, the SSH keys, the deploy CLI, and the saved browser sessions all lived on one machine. If that machine was asleep, work queued behind it. That's fine for a side project. It's a terrible way to run client sites.
The gates move the capability off the hardware and onto an endpoint the whole team reaches. The permission model is tighter than the laptop ever had. It's the same layer we described in where MCP sits in the stack, now with a credential boundary in front of it.
The board sums up the win in three numbers: 4 parts done, 2 dashboard actions still on a person, and 0 laptops in the path. That's the whole point of the architecture.
07 / how-do-you-build-your-first-one
How do you build your first one?
Five steps, in order.
- Pick a job that's already one API call. Deploys, rollbacks, status reads. If you can't name the call, you don't want a thin MCP yet.
- Write the tool list before the code. That list is your MCP server architecture and your security review. It's short enough to read in a meeting.
- Put the credential on the gate. One per client, scoped as narrowly as the platform allows. Never in the agent, never in a prompt, never in a query string.
- Make the guard structural. Snapshot, verify, revert, all in the server. A rule in a prompt is a suggestion.
- Fork before you build. Our 70-tool Vercel gate is someone else's open-source server with a transport bolted on. The custom build was the one nobody had written yet.
Then run it against something real before you believe it. Every real finding in this post came from a live site. None came from a test fixture. Want the running cost of that choice? We published what agent access costs from our own books.