An OS for AI agents: not the one on your laptop

An OS for AI agents: not the one on your laptop. A sketch of a heavy stack of layers carrying one small agent, beside a slim glowing plate carrying many.

I keep picturing the day I stop running apps and start running agents. One plans my trips, one watches my investments, one reads my email, one plans my day. Each is a small workflow with a few sub-agents, talking to a language model and to a handful of services through the Model Context Protocol (MCP), with my credentials in hand and my instructions as its brief. I’ll want them kept apart from each other, for security and so that one can break without taking the rest down. My first thought was that they need an operating system (OS) of their own, a lean one, because the one on my laptop was built for a person. The more I turn it over, the more I think they need something else: not a kernel each, but one host that holds what each of them is allowed to do.

Why the one on your laptop doesn’t fit

A desktop operating system is shaped around one human. One identity carries all of your privileges, and permissions attach to files and applications, not to tasks. An agent inherits that shape badly: it either gets everything you have or can’t do anything useful.

Agents also read untrusted text all day. A hotel description, an email or a web page can carry instructions aimed at the agent, and a hijacked agent is an adversarial tenant on your machine. So the walls between agents, and between an agent and your account, have to be something the agent can’t talk its way out of.

The isolation we have was built for services, not for this. Put each agent in a container and you share the kernel, which isn’t a strong wall against a tenant that has turned on you, and you still pay for a Node or Python runtime per agent, tens of megabytes to run a few kilobytes of instructions. Put each in a microVM, a tiny virtual machine, and you add a guest kernel and a whole network stack on top, to call three endpoints on the web. Either way the agent sits idle most of the day, and the plumbing dwarfs the payload.

What an agent actually needs

Strip the trip planner down and it’s a loop with a table of what it may use. It needs:

  • A channel to a model, with a token budget.
  • Channels to a few tools, its MCP connections, each with a policy: what it may call, what needs my approval, how much it may spend.
  • Its state: instructions, skills, the conversation, a few notes. A key-value store is enough.
  • A channel to me, for approvals and reports.
  • Timers and triggers, so it wakes when something happens.

It doesn’t need a filesystem hierarchy, users and groups, the ability to spawn processes, its own network and encryption stack, a shell, an init system or any device. Nothing a kernel provides is on the list. That’s why a lean operating system per agent isn’t just overhead. It’s the wrong unit.

What it could look like

The serverless platforms solved exactly this shape years ago: millions of tiny tenants, mostly idle, each needing memory isolation and a handful of outbound calls. Their unit is the isolate, not the virtual machine. It starts in milliseconds and costs almost nothing while it waits. WebAssembly components take the same idea further: a component can only call what the host hands it, so there’s no ambient authority to abuse. The actor systems are the older form of the same thing, with supervision for independence and activation on demand for idleness.

So the picture I keep drawing is one host, many agents. Each agent is an isolate holding its instructions, its state and a table of capabilities. The host is the operating system, if you like, but there’s one per person, not one per agent.

One host, many agents: a sketched plate carrying a row of small agent cells, with threads to a glowing model core, three walled tool boxes and an approvals tray
One host per person: the agents live inside it as isolates, and everything they may use is handed to them by the host.

Four things fall out of that choice, and they matter more than the memory saved.

  • The agent never holds a credential. I do the OAuth login once, at setup, with the host. The agent’s manifest names its connections, and the host attaches the token when it forwards each call. An agent that has been hijacked can’t leak a token it has never seen.
  • Policy lives in the host, not in the prompt. The scopes a provider offers are coarse; a travel site may offer only full access. The host narrows them: search freely, book only after my approval, spend under a cap, log everything. If a hotel description hijacks the planner, all it can do is leave a booking request in my approval queue.
  • Isolate the tools, not the thinker. The agent loop is the host’s own code fed untrusted data. The dangerous code is elsewhere: a local MCP server is third-party code running on your machine, a headless browser is a large attack surface, and code execution is code execution. Those are the things that deserve a microVM each, run once as shared services that every agent reaches through the host. A remote MCP server needs no local isolation at all, only a rule about where the agent may send traffic and the token the host adds.
  • Sub-agents get a subset. Spawning an isolate is cheap, and a child can only receive capabilities its parent holds. The orchestrator and its workers become a trust hierarchy for free.
Isolate the tools, not the thinker: a small thin-walled agent cell beside three thick-walled boxes for the browser, code execution and a local MCP server
Thin walls around the agent, thick walls around the tools it calls.

The host itself is small. It keeps the capability table, holds the credentials in the system keychain, fronts the model with budgets and logging, wakes agents on their triggers, queues my approvals and writes the audit log. On top of a WebAssembly runtime, I’d guess that’s a few thousand lines, not a kernel. It should be hardened, inside a virtual machine if the computer is shared, but there’s one of it.

What I’m unsure about

  • The wall. An isolate is a weaker wall than a virtual machine against malicious code. The design holds only if the agent isolate never runs arbitrary code, so anything an agent writes runs inside the sandboxed tool. I think that’s the right line, and I’d want it tested.
  • Many agents or one. I may end up with one assistant with many skills rather than a fleet. The problem doesn’t go away, because each task still needs its own narrowed authority, but the shape of the host might change.
  • Scopes. Narrowing a provider’s permissions at the host works only as well as the host understands the calls. If providers start offering fine scopes themselves, part of this layer moves to them.
  • The vendors. The makers of the desktop operating systems could build this into the desktop and swallow the layer. That might be a good outcome.
  • Where it runs. On my own machine the thing to protect is my account from my agents; in a cloud tenancy it’s tenants from each other. The host shouldn’t care, and I haven’t worked through whether it can avoid caring.
  • The argument. The memory saved is real but it isn’t the point, because the model calls will dominate the bill. The point is that capabilities make the security story something I can explain, and the agents something I can update one at a time. Every piece here exists somewhere, in the serverless platforms, the actor systems and the capability-based kernels. What I haven’t seen is the whole thing composed for one person’s fleet, with my approvals treated as a resource to schedule.

Where to start

I’d start with the host interface, because it’s the whole point. Define what an agent sees: call the model, call a tool through the host, ask me, remember, and a trigger that wakes it. Build the host around a WebAssembly runtime, give it one MCP connection with one policy and an approval queue, and write one agent against it. Rust and Go are the comfortable languages for components today; a loop in C# will hit rough edges.

If you’d rather not build the host yet, there’s a shortcut. Cloudflare’s Durable Objects are stateful isolates with storage and alarms, woken on demand. With secrets kept out of the agent’s code and a gateway in front of the model, they map almost one to one onto the host above. It’s isolate-first rather than WebAssembly-first, and you take their isolation model and their lock-in, but a weekend gets you a fleet of one.

The operating system on your laptop will keep doing what it does. The agents need a smaller thing, with a sharper idea of permission.

If you’d like to build this, with me or without me, or to help fund it, I’d love to hear from you.

I’m Amir Pournasserian. I build AI and platform systems for a living, maintain FluentCMS and YeSvelte, and write here about what I find along the way.