❌

Vue normale

Reçu aujourd’hui — 28 septembre 2026Infra

The case for a cloud native agent harness

Coding agents became useful when they stopped being a chat box. Four things changed the shape of the problem: capable tools, a shared repository and filesystem, subagents, and skills that capture what the system learned from prior work. Together, those let an agent operate in an environment instead of merely answering questions about one.

These patterns grew up on developer laptops because developers already live in terminals, repos, and toolchains. They will not stay there. The same capabilities are heading toward people who will never open a terminal, and toward long-running workloads that no one is sitting in front of.

That is where the usual harness shape starts to creak. Most harnesses assume a desktop: one person, one machine, one local filesystem, one interactive process that is simultaneously the UI, the agent loop, the sandbox, the credential store, the tool host, and the session database. That design serves a single developer well. It fits poorly when an organization wants to run hundreds of sessions, govern which tools each one may call, survive a node dying mid-task, or let someone start a session from a laptop and pick it up from a phone.

The tempting fix is to put the desktop harness in a container and call it done. That moves where the process runs and leaves how the process is built untouched. Every concern is still coupled to one long-lived machine. Kubernetes taught this industry that a monolith in a container is still a monolith. The lesson applies to agents too.

A harness for this next phase should be designed as a distributed application from the start, with the agent loop separated from everything around it, so that hosting, exposure, and security fall out of the architecture rather than getting bolted on afterward. At Stacklok, we have been calling this a cloud-native harness. We recently open sourced our cloud-native harness, Mecatl, and I want to walk through what the shape looks like in practice.

The name, pronounced “meh-kah-tl,” is classical Nahuatl for cord, rope, or string. Ozz, the principal engineer who started the project, picked it. Rope, harness. It also comes with a mascot named Mecatito, which I mention only because I like him.

What a cloud-native harness is

A cloud-native harness is an agent runtime built as a distributed system. In Mecatl, the core engine owns the agent loop: reasoning, tool dispatch, permissions, hooks, and event emission. Everything else sits outside the engine behind explicit interfaces:

  • Clients. Terminal, API, and application clients that interact with the loop without owning it. Today that means a TUI (mecatui), gRPC and HTTP/SSE APIs, and a TypeScript SDK.
  • Execution environments. Assigned workspaces and command runners where the agent’s work happens.
  • Tool ecosystem. Built-in tools, streaming-HTTP MCP services, skills, and application-supplied integrations, exposed through a catalog.
  • Supporting services. Model providers, session state, event history, identity, and coordination.

The same loop runs in your terminal, as a service, or on Kubernetes without being replaced. That is the point of the split. The engine is one component, and every other component can be deployed, scaled, and secured on its own terms.

Here is what those boundaries let you do that a desktop harness cannot.

Run the loop like an application. The engine is a versioned, deployable component. You roll it out, log it, and observe it like any other service. In a Kubernetes deployment (via the mecak8s guide), workers are replaceable during normal operations. You can roll a new engine version and durable sessions stay put.

Keep sessions beyond a worker. Session state and event history live in durable storage under a single-writer coordination model. When a worker dies, a replacement process resumes from the last persisted turn boundary. It does not resume an in-flight operation, and work after the last successful save can be lost, so this is turn-level durability rather than a distributed transaction. It is enough that a pod eviction stops being a conversation-ending event.

Govern the tool ecosystem. A desktop harness is useful in large part because it has an unrestricted shell. A cloud-native harness does not need one. Mecatl exposes purpose-built tools, skills, and application integrations through an explicit catalog, with permission, audit, and execution-environment boundaries around each. A hosted agent starts from a catalog of things it is allowed to do.

Serve more than one kind of client. Because a client does not own the filesystem, credentials, or session state, the same loop can back a terminal, a remote service, an embedded application, and a Kubernetes deployment at the same time. A web UI, a desktop GUI, a Slack integration, and a collaborative document editor are all clients that can attach to the same runtime.

None of this means the desktop goes away. We started there, with a traditional TUI, so we could learn by using it every day. It has become my daily driver; since Mecatl is new, there are rough edges, but also plenty of delightful experiences. 

Where this is going, and where we need help

Mecatl is early, and the cloud-native harness is a direction as much as a current architecture. Several of the hardest problems are still design drafts, and they are exactly the kind of problems this community has worked on before.

Identity. When an agent calls an external system, who is calling? The user? The agent? The session? A subagent three levels down in a multitenant deployment? Our proposed answer is to make Mecatl its own SPIFFE trust domain and encode the full delegation chain in a JWT, so the receiving system has a “call stack” of every identity involved when it makes a policy decision. 

Tools beyond MCP. MCP is a good start and is supported out of the gate. But routing every tool’s input and output through the agent’s context window is expensive and often unnecessary. A PDF decoder should be able to operate directly on the workspace filesystem without the model ever seeing the bytes. We’re exploring scoped resource grants as short-lived, attenuable grants and direct tool-to-service data paths. 

Context attestation. Prompts are the new code and deserve the same treatment. If an agent or subagent runs with packaged context, that context should be versioned, signed, attributable, distributable, and subject to policy, like any other supply-chain artifact, and its provenance should show up in the identity chain above.

These are directions, not commitments about what ships today. The deployment and feature docs describe what works now; the linked design documents describe what we are arguing about.

If any of this is a problem you have, the project is at github.com/stacklok/mecatl, the docs are at mecatl.dev, and there is a Discord. Build a client against the SDK, tear apart the identity draft, or tell us where the TUI breaks. What would a cloud-native harness let you do that you cannot do today?

❌