❌

Vue normale

Reçu avant avant-hier

Is a Pod the right deployment unit for an AI agent?

Graphic: Is a Pod the Right Deployment Unit for an AI Agent?

When we first started building kagent, we didn’t run every agent in its own Kubernetes Pod, Service, and ServiceAccount. Instead, agents were simply executed inside the kagent runtime. It was the simplest architecture possible: one runtime hosting many agents.

It worked well for demos and proofs of concept.

As the number of agents grew, however, fundamental questions started to emerge.

  • How do we isolate one agent from another?
  • How does each agent get its own identity?
  • How do we enforce access and network policies?
  • How do we understand what an individual agent is doing?
  • Who owns an agent, and how do we support multi-tenancy?

These aren’t Kubernetes questions. They’re agent platform questions.

The Pod as the Deployment Unit

Our first answer was straightforward: run every agent in its own Pod, Service, and ServiceAccount.

That decision immediately solved many of our problems.

A Pod provides process and container isolation. A ServiceAccount gives every agent its own Kubernetes identity, allowing us to integrate naturally with authentication and authorization mechanisms. Existing network policies, admission policies, and security controls continue to work without modification. Observability systems can attribute logs, metrics, and traces to individual agents. Scheduling and resource management also became Kubernetes-native.

As the architecture evolved, we introduced stronger isolation mechanisms such as agent-sandbox in kagent, allowing agents to execute with tight security boundaries.

For a while, this felt like the right abstraction.

But Should Agents Be Best Represented as Pods?

The more we thought about agents, the more we realized they are quite different from traditional microservices.

Most services are expected to be continuously available.

Agents are not.

An agent may wake up only when assigned a task, execute for a few seconds or minutes, and then become completely idle. Keeping a dedicated Pod alive for every potential agent quickly becomes wasteful.

Agents also have execution patterns that don’t resemble long-running services:

  • An agent may dynamically create multiple subagents to perform certain subtasks in parallel.
  • An agent may impersonate a user or execute on behalf of a human.
  • An agent may pause while waiting for human approval before continuing.
  • An agent’s lifetime may be measured in seconds or minutes rather than days.

These characteristics naturally lead to a question:

Are Kubernetes Pods the right lifecycle abstraction for short-lived, bursty AI agents?

Pods are excellent execution environments. But that doesn’t necessarily mean they should also be the right abstraction for AI agents.

Enter Agent-substrate

Instead of treating every agent as a first-class Kubernetes workload, agent-substrate introduces an additional control plane above Kubernetes. Kubernetes continues to manage Pods, Services, networking, storage, and compute resources, while agent-substrate manages the lifecycle and placement of AI actors onto execution workers.

Agent-substrate introduces a set of abstractions that are similar to the Kubernetes concepts we are already familiar with. A WorkerPool is analogous to a NodePool, Workers are analogous to Nodes, and ActorTemplates correspond to the declarative specification of a Pod.

Comic Graphic: Workers analogous to Nodes: ActorTemplates correspond to the declarative specification of a Pod

Let’s look at what this abstraction looks like in practice. A WorkerPool defines a collection of execution workers that can host Actors. Example of the default WorkerPool in kagent:

apiVersion: ate.dev/v1alpha1
kind: WorkerPool
metadata:
  labels:
    app.kubernetes.io/instance: kagent
    app.kubernetes.io/name: kagent
  name: kagent-default
  namespace: kagent
spec:
  ateomImage: ghcr.io/kagent-dev/substrate/ateom-gvisor:v0.0.6
  replicas: 3

An ActorTemplate defines how an Actor should execute, much like a PodTemplate defines how a Pod should be created. Below is an example of a simple ActorTemplate in kagent. Note that it includes the runsc configuration, which serves as the execution entrypoint for gVisor. I omitted several kagent-specific fields, including the agent’s name and additional configuration details.

apiVersion: ate.dev/v1alpha1
kind: ActorTemplate
metadata:
  labels:
    app.kubernetes.io/managed-by: kagent
    kagent.dev/sandbox-agent: hello-substrate
  name: hello-substrate
  namespace: kagent
spec:
  containers:
  - command:
    - /app
    ...
    env:
    ...
    image: cr.kagent.dev/kagent-dev/kagent/golang-adk@sha256:e01479b52280b0eae9e2808cc68392ba98fd737782496ff256847257e6bb8ed1
    name: kagent
  pauseImage: gcr.io/gke-release/pause@sha256:bcbd57ba5653580ec647b16d8163cdd1112df3609129b01f912a8032e48265da
  runsc:
    amd64:
      sha256Hash: efd12935f6654c91a1389710eb8dfa4d12b6b9be00db87526dc2eb584ad00119
      url: gs://gvisor/releases/nightly/2026-06-02/x86_64/runsc
    arm64:
      ...
  snapshotsConfig:
    location: gs://ate-snapshots/kagent/hello-substrate
  workerPoolRef:
    name: kagent-default
    namespace: kagent

The Worker or Actor is not represented as a custom resource in Kubernetes. Kubernetes only sees WorkerPools and ActorTemplates. Agent-substrate, however, sees Workers and Actors. This separation allows the cluster to manage a fixed number of execution Pods while agent-substrate manages a much larger number of logical agents. You can use the substrate CLI or API to view them directly. Each Worker is mapped to a single unique Pod.

$ kubectl-ate get workers
NAMESPACE   POOL             POD                                         STATUS   ASSIGNED ACTOR
kagent      kagent-default   kagent-default-deployment-ddfcfbdd7-54pb7   FREE     <none>
kagent      kagent-default   kagent-default-deployment-ddfcfbdd7-jmjl5   FREE     <none>
kagent      kagent-default   kagent-default-deployment-ddfcfbdd7-z2mmh   FREE     <none>

$ kubectl-ate get actors
NAMESPACE   TEMPLATE                  ID                                                                STATUS             ATEOM POD   ATEOM IP   VERSION
kagent      hello-substrate           a786a0c4-c2c8-44e5-9ea5-67b64f41deb1                              STATUS_SUSPENDED   <none>                 5
kagent      hello-substrate           asr-kagent-hello-substrate-019efbb5-cc48-7601-8fc6-985e6239aa05   STATUS_SUSPENDED   <none>                 5
kagent      hello-substrate-linsun    0c82223d-cc14-40c8-a25c-5ee00fe153ae                              STATUS_SUSPENDED   <none>                 5
kagent      hello-substrate-linsun3   asr-ce96fc0ee592bf1e12336461                                      STATUS_SUSPENDED   <none>                 5

The important distinction is that an Actor, which represents (“acts as”)  an AI agent, is no longer itself a Kubernetes Pod.

Instead, an Actor is a logical entity that can be scheduled onto an agent-substrate Worker when work arrives and removed when execution completes. Workers remain long-running Pods managed by Kubernetes, while Actors are lightweight execution units that share those workers.

This abstraction allows us to continue leveraging Kubernetes for pod and service scheduling, networking, security, and resource management while supporting far more AI agents than the cluster could ever support as individual Pods.

In other words, Pods become the execution workers, not the deployment model for agents.

Challenging More Than Deployment Model

At first glance, agent-substrate may look like a more efficient scheduling layer.

In reality, it challenges a much deeper assumption: should a Pod be the primary representation of an AI agent at all?

Agent Identity

Should an agent’s identity really be tied to a Pod or its Service?

Or should identity belong to the ActorTemplate, namespace, tenant, and version, independent of whichever Worker happens to execute the Actor at a given moment? Christian Posta tried to explore this topic much deeper in his blog.

Security and Policy

Today, Kubernetes policies are attached to Pods, Services, or ServiceAccounts.

Should access control, network policy, and runtime permissions instead be expressed at the ActorTemplate level and selectively overridden for individual Actors? Can we use agentgateway to mediate the traffic and enforce policies?

Ownership and Multi-tenancy

Who owns an Actor?

Who owns an ActorTemplate?

How are quotas, billing, and lifecycle managed across teams and tenants when AI agent execution is no longer tied one-to-one with Pods?

Observability

When an Actor executes on different Workers over its lifetime, observability must follow the logical agent, not the underlying Pod.

Logs, traces, audit records, and execution history should all be associated with the Actor regardless of where it was scheduled.

Looking Ahead

Kubernetes remains an exceptional platform for running microservices and inference workloads at scale.

But AI agents introduce more unique characteristics than traditional cloud-native services. They are ephemeral, bursty, capable of spawning subagents on demand, and often act on behalf of users. The Pod may still be the right execution unit for AI agents, but it may no longer be the right deployment, identity, or lifecycle unit.

That is the question agent-substrate is exploring. Explore the agent-substrate project through kagent, join the agent-substrate community, and feel free to connect with me on LinkedIn.

Why sandboxing your agent is not enough

Why sandboxing your agent is not enough, comic graphic

The agentic AI space is moving incredibly fast. Not long ago, I learned about a cool project called agent-sandbox, which provides a sandboxed environment for AI agents by leveraging many of the building blocks we have already developed for Kubernetes pods, such as identities, storage and networking.

If you’ve ever read the horror stories about AI coding agents or followed projects like OpenClaw & NemoClaw, you know how important it is to provide a secure and isolated environment for your agents. Without proper isolation, agents can surprise you by doing things you never intended, such as deleting family photos or modifying critical files.

Just a few weeks ago at Open Source Summit North America in Minneapolis, while chatting with Bob Killen in the hallway track, I learned about a new project called agent-substrate. What immediately caught my attention was its ability to dynamically wake up agents based on invocation, thus allowing more agents on the same infrastructure resources while still providing the security benefits of sandboxed execution. 

Naturally, the first thing I did was discuss with our team how we could integrate it with kagent and agentgateway.

What Are the Differences Between the Two Projects?

Agent-sandbox

The agent-sandbox project provides a Sandbox Custom Resource Definition (CRD) and controller for Kubernetes under the umbrella of Kubernetes SIG Apps.

Its primary focus is on providing:

  • Strong identities for agents
  • Persistent storage that survives restarts
  • Lifecycle management of sandboxed pods
  • Security and isolation through the Sandbox controller

In short, agent-sandbox focuses on making agent execution secure, manageable, and Kubernetes-native.

Agent-substrate

The agent-substrate project is currently a standalone project and is not part of any Kubernetes SIG or other cloud native foundation project, though that may change in the future.

Built on top of Kubernetes, agent-substrate aims to go beyond sandboxing by focusing on:

  • Higher scale
  • Better resource efficiency
  • Lower latency execution
  • More dynamic lifecycle management for agents

My understanding is that agent-substrate provides the runtime building blocks needed to run AI agents securely at very high scale.

Instead of keeping agents running continuously as pods, agents execute in secure worker pods for short bursts, suspend when idle, and resume later on any available worker. The worker pod lifecycle is decoupled from the agent “actor,” which is managed by the agent-substrate control plane.

In this model, agents behave more like on-demand serverless workloads: they can be scheduled, paused, and resumed with minimal overhead, while still benefiting from Kubernetes-based sandbox isolation using lightweight runtimes such as gVisor or Kata Containers.

Do We Need Agent-substrate When We Already Have Agent -sandbox?

I believe the answer is yes.

Sandboxing your agents is necessary, but not sufficient.

In most Kubernetes environments, resources are constrained. You have to be selective about which agents run continuously. Many agents are only useful occasionally, and keeping them always on is inefficient.

This creates an awkward tradeoff:

  • Keep agents running idle and waste resources
  • Or constantly spin them up and down, adding overhead and latency

Neither option scales well.

This is where agent-substrate becomes interesting. While agent-sandbox focuses on security, isolation, and lifecycle management, agent-substrate focuses on density, efficiency, and operational scalability, while still preserving a secure execution model.

You can think of it as making large-scale agent fleets practical: not just safe, but economically viable.

Agent-substrate Integration with kagent

One of the things I appreciate about kagent is its simplicity and declarative YAML-based workflow. You always know what is running in your cluster, and you can recreate environments easily from source-controlled manifests.

With the efficiency introduced by agent-substrate, we can support many more agents using a shared pool of worker resources. Instead of assigning a dedicated pod per agent, we can use shared worker pools and templates that dynamically execute agents on demand.

Agents can now appear and disappear based on invocation, while still running on the same underlying infrastructure.

I’ve updated my AIRE agent to use agent-substrate as its runtime. This allows me to invoke agents only when needed, without keeping them running continuously.

It also enables multiple AIRE agents (each with different skills, echoing the idea of “Don’t put agents! build skills instead”) to share the same worker pool or even same pod.

The result is a more efficient system: agents remain available on demand, idle resource consumption drops significantly, and there is far less need to constantly scale pods up and down.

For example, my six AIRE agents map to six actor templates but only require a single worker pod to execute them, as long as they are not running concurrently. If concurrency increases, I can simply scale the worker pool (kagent-default) horizontally to increase the number of worker replicas.

kagent dashboard graphic

Final Thoughts

agent-sandbox and agent-substrate solve related but distinct problems.

Agent-sandbox asks: How do we run agents securely?
Agent-substrate asks: How do we run agents not only securely but also efficiently at scale?

As AI agents become more common in Kubernetes environments and AI costs remain one of the biggest concerns in adoption, we need to be more innovative and avoid tying an agent’s lifecycle too closely to Kubernetes pods.

Security, identity, isolation, and policy controls remain essential. At the same time, we need a runtime model that allows hundreds or thousands of agents to exist in a dormant state, waiting to be invoked, without requiring hundreds or thousands of pods for workloads that are idle most of the time.

The future is not just secure agents, it’s scalable, efficient, and ephemeral agents.

Additional Resources:

Agent Auth: A lawyer’s day in court

I’ve always thought about AI agents as microservices+.

They need everything a traditional microservice needs, and:

  • More authentication requirements because an agent may act on behalf of many different users.
  • More policy requirements because an agent’s behavior can be less predictable, requiring guardrails and policy enforcement.
  • More observability requirements, especially around context, prompts, tool calls, and the contents of requests and responses.
A cartoon example of a lawyers day in court, proving his authority to represent his client Alice.

When thinking about agent auth, I found myself reflecting on a traffic lawyer I hired years ago after receiving a traffic ticket for failing to stop for a school bus. It was my first, and so far only, traffic ticket.😅

The experience turned out to be a useful mental model for understanding agent auth.

Imagine a lawyer walking into court to represent Alice.

This is similar to an AI agent receiving a request from Alice and performing actions on her behalf.

The judge first asks the lawyer to prove who he is.

This is agent identity. Before the system can trust an agent, it needs to know exactly which agent is making the request.

Next, the judge asks, “Who are you representing today?”

This is principal identity. The system needs to know not only who the agent is, but also which user the agent is acting for.

The lawyer then presents documentation showing that he is authorized to represent Alice in this specific case.

In agent systems, this is often represented by an On-Behalf-Of (OBO) token or another delegation artifact. The token carries information about:

  • The identity of the principal (Alice)
  • The identity of the agent
  • The delegated permissions
  • The scope of the delegation

At this point, the judge knows three things:

  1. Who the lawyer is
  2. Who the lawyer represents
  3. What authority has been delegated to the lawyer

But that still isn’t enough.

The judge must also verify that the lawyer is allowed to represent Alice in this particular traffic case. This is where policy enforcement comes in.

Having a valid delegation does not automatically grant unlimited access. The requested action must still comply with the applicable policies and scopes.

In a real courtroom, the lawyer and the judge handle most of this complexity. They carry identities, verify credentials, validate representation rights, and enforce the rules of the court.

In an agentic system, we need similar infrastructure.

An agent platform must be able to:

  • Establish strong agent identities
  • Carry principal identities across requests
  • Issue and validate delegation tokens
  • Enforce authorization policies and scopes
  • Provide observability and audit trails for agent actions

This is where an AI native gateway can play an important role.

Rather than requiring every agent to independently implement identity propagation, delegation verification, policy enforcement, and auditing, the agent gateway and mesh can centralize these capabilities. The agent gateway and mesh become the equivalent of the court clerk, bailiff, and records office combined: ensuring identities are verified, delegations are valid, policies are enforced, and actions are auditable.

Combined with existing identity and service-mesh technologies such as SPIFFE, cert-manager, Istio, and agentgateway, we can build an agent platform where agents focus on business logic while the platform handles identity, delegation, policy enforcement, and observability.

The core idea is simple:

A lawyer is not the client.

An agent is not the user.

Both operate with their own identities while acting on behalf of someone else, under a specific delegation and within a defined scope. Agent auth is fundamentally about making that relationship explicit, verifiable, and enforceable.

❌