❌

Vue lecture

Manufacturing Trust for AI Agents | Docker’s WeAreDevelopers Keynote

Coding agents do more than write code. Agents can install dependencies, access networks, use credentials, and keep working after we’ve moved on to something else. Agents need access to get real work done. But the more access they have, the bigger the impact of a mistake. The challenge is building enough trust into the system to give agents more freedom without giving them access to everything. 

At WeAreDevelopers North America, Docker President Mark Cavage laid out Docker’s answer: give every agent a strong boundary, make its environment and authority reproducible, and let the work run where it makes sense, from a developer’s laptop to the cloud. 

In 30 minutes, Mark demonstrates what happens when an agent pushes beyond a container’s limits and how Docker Sandboxes create a stronger boundary. He then connects that foundation to Kits, cloud capacity for work that needs to continue beyond the laptop, and new partnerships helping developers put agents to work safely. 

agents trust factory transparent

Agents are actors, not workloads

Containers were built to isolate applications. Agents do more. They make decisions and act on the systems around them, so the boundary needs to extend to the environment they can reach. 

Mark’s onstage demo makes that clear. The container worked as designed. What the agent needed was containment around its full environment. Docker Sandboxes provide that containment. Each agent gets its own isolated microVM and kernel, separate from the developer’s host environment. 

Developers can choose the agent, model, and tools that fit the job, then define the sandbox’s access to files, networks, and secrets. Those policies stay outside the agent’s control. Together, the sandbox and its policies create a trusted environment for autonomous work. 

Docker Sandboxes are available today through the free, standalone CLI, so developers can bring stronger isolation to the agent workflows they already use. 

Make the agent and its authority reproducible

A strong boundary gives an agent a safe place to work. Developers also need a repeatable way to define what runs and what it can access. 

In the keynote, Mark announced the next generation of Docker Sandbox Kits, now built as standard OCI images, and invited ecosystem partners to help shape the specification as Docker works towards neutral governance. 

Docker Sandbox Kits package the agent, its tools, and the rules for what the sandbox can reach into one versioned, shareable artifact. Developers can share Kits through the same workflows they already use for container images, while keeping changes to an agent’s authority visible and reviewable. 

The result is a common foundation for the agent ecosystem, built on an open standard. 

Dockerfiles made software reproducible. Kits make authority reproducible

any model any harness any tool

Start on your laptop. Finish in the cloud.

Some agent work begins and ends on a developer laptop. Longer-running or parallel work needs capacity that stays available when the developer steps away. In the keynote, Mark announced Docker Cloud Sandboxes, bringing the same microVM-based isolation from the laptop to Docker-managed cloud compute. 

Using the familiar sbx workflow, developers can start locally, move their work to the cloud with one command, and let it continue after they close their laptop. They can also run tasks in parallel without provisioning or maintaining the infrastructure themselves. 

Docker Cloud Sandboxes are available today with pay-as-you-go pricing.

start local move to cloud move back

An ecosystem built around trust

An open agent ecosystem needs both: choice for the developers and shared standards the industry can build on.

Nous Research joined Mark onstage to show Hermes running as a first-class Kit in Docker Sandboxes. The demo showed what choice looks like in practice: a third-party agent packaged for a common environment, with Docker providing the underlying isolation and controls. Developers can choose the agent that best fits their work without having to rebuild the trusted execution layer around it.

Choice also depends on shared standards that the wider ecosystem can adopt. Docker has committed to submitting its Kits specification to the Cloud Native Computing Foundation (CNCF), the open source, vendor-neutral hub of cloud-native computing.

“Standards are what let an ecosystem move fast without fragmenting, and few companies understand that better than Docker through their involvement in efforts like the OCI and CNCF. By delivering Sandbox Kits as standard OCI images, Docker is giving the industry an open, repeatable way to package an AI agent, its tools, and its guardrails as one artifact. OCI is the foundation the cloud native ecosystem is built on, so a standard for agents that builds on OCI reaches the whole ecosystem at once. The cloud native community looks forward to working with Docker to bring this work under neutral governance.” 

Chris Aniszczyk

CTO at CNCF

Together, these partners show what an open agent ecosystem can look like: choice at the agent layer and a shared format for packaging an agent’s environment and authority. Docker provides the trusted foundation underneath, whether agents run locally or in the cloud.

Trust comes from the system around the agent

Taken together, the announcements in Mark’s keynote form one system. Sandboxes provide a deterministic boundary. Kits make the agent’s environment and authority reproducible. Cloud Sandboxes extend the same trust model to durable cloud capability. 

They give developers the freedom to choose their agents and more autonomy, while retaining control over what they can access and change. 

As agents take on bigger jobs, the infrastructure around them matters more. Docker provides the trusted foundation they need to work locally or scale in the cloud, while developers continue to stay in control of what ships.  

Ready to try Cloud Sandboxes?

For a limited time, new accounts can claim $250 in compute credit to get started.

Watch, Explore, Build

  •  

From Dockerfile to Kit: the Docker Sandboxes Kit Specification

Agents need containment, and a sandbox is only half of it. Something still has to say which agent runs there, what it gets, and what it may touch. That is a Kit: an ordinary OCI image, so the answer travels with the agent and means the same thing on any conforming runtime. Today we published the Docker Sandbox Kit Specification v3, open source under Apache 2.0 at docker/sandbox-kit-spec. Here is why I wrote it.

Everything that makes an agent useful is a grant

I run a lot of agents. They write code, run tests, install dependencies, call APIs, and work on infrastructure while I do something else. None of it happens without access, so I grant it one piece at a time: a bind mount, a token with broader scope than the task needs, a firewall rule that was quicker to open than to narrow. Each grant is reasonable on its own. Together they take back the isolation I was relying on, and none needed an exploit. The holes are configuration, added on purpose, usually by me.

I am worse at taking any of it back, and I could not reproduce the grants my setup depends on. No file records them. They live in shell history, dashboards, and my memory. I cannot hand that to a colleague or diff it against last week.

Containers package applications. Sandboxes contain agents.

A container packages applications. It shares the host kernel and uses namespaces and cgroups to give one fixed workload its own view of the filesystem, network, and processes. That is the right tool for software that runs, does its job, and touches only what it was handed.

An agent, however, is a probabilistic actor. It decides what to do next and then does it, to my filesystem, network, credentials, and cloud account. It will install a package that needs root, open a port nobody planned for, and try the next thing when the first is blocked. A container was not built for that: the boundary is the same kernel the actor is probing.

A Docker Sandbox is a microVM with its own kernel, so the boundary sits below anything the model can reach or rewrite. Inside one I can hand an agent root and let it loose, because the damage stops at the sandbox boundary. The sandbox is what lets me run an agent with the safeties off.

But an empty sandbox is not an environment. Something still has to say which agent runs, which tools and MCP servers it gets, which skills and instructions shape it, and exactly what it may touch.

What a Dockerfile cannot say

A Dockerfile answers everything about the software itself: how it is built, what gets packaged, how it starts. It was never standardized; OCI standardized the image it produces and how registries distribute it. What a Dockerfile does not describe is the outside: networks, credentials, volumes, tools, context. That half has lived in docker run flags, a Compose file, a CI config, and someone’s memory. Unversioned, unreviewable. A Kit writes it down with the content.

One image, one digest

If you have used sbx, you have used Kits. This is the third version of the format, and the change that matters is that a Kit is now an ordinary OCI image rather than its own artifact: no media type, no sidecar file, nothing for a registry to learn. The manifest carries the declarations in one annotation, vnd.docker.sandbox.kit.descriptor; the layers carry the content.

A Kit therefore builds with docker buildx build, pulls with docker pull, gets scanned and signed by the tooling you already run, and works in a FROM. Pinning the digest pins content, declarations, and metadata together. The tooling and distribution path are free; the format is something you learn: a grammar, a page per capability type, provides and requires, kind: set.

Two kinds of Kit exist. A workload runs and supplies the root filesystem. A mixin is an overlay: a CLI with its network rule, a credential binding, context for an agent. You launch one workload and any number of mixins.

Authority you can read

Part of the GitHub CLI mixin in the repository:

capabilities:
  - type: com.docker.sandbox/network-policy@2
    config:
      runtime:
        allow:
          - github.com
          - hosts: [api.github.com]
            methods: [GET, HEAD, POST, PATCH, PUT, DELETE]
        deny:
          - hosts: [api.github.com]
            methods: [DELETE]
            paths: [/repos/**]

  - type: com.docker.sandbox/credential@1
    optional: true
    config:
      service: github
      phase: runtime
      apiKey:
        name: GH_TOKEN
        proxyManaged: true
        inject:
          - {domain: api.github.com, header: Authorization, format: "Bearer %s"}

Read it as a permission slip. This Kit asks to reach GitHub and nowhere else, and for most of the API but not deletes under /repos/**, because deny wins. The token that can open a pull request cannot delete the repository. The credential is proxy-managed: a conforming runtime injects the real value into requests to the named domains, and inside the sandbox there is only a sentinel.

Two words carry weight: asks and conforming. A Kit grants itself nothing. Each entry is a request, and the host decides. A conforming runtime, one that implements the behaviour the specification describes, blocks hosts not on the list. Without one, the annotation is inert: an image and no enforcement. Docker Sandboxes is the first conforming runtime.

Everything a Kit needs goes through that one list, typed and versioned. Grants (network rules, credentials, volumes, ports, devices, skills paths) count toward “what may this Kit do”; entries that ask the runtime to act, like a lifecycle hook, do not. A required request the host cannot satisfy refuses the launch, rather than starting an agent with less authority than it declared, or more.

Composition is a function, not a sequence

Container images never solved multiple inheritance: a Dockerfile stage has one FROM. Mixins are overlays ordered by the dependency graph the Kits declare through provides and requires, never by the order you typed the flags, so the same set always composes to the same image.

The resolver is strict on purpose. Every requires is satisfied from inside the set or resolution fails; nothing is fetched to cover a gap. Exactly one workload is allowed. Two Kits providing the same name fail rather than one silently shadowing the other (composing the Claude workload with the Claude mixin is the canonical mistake). Where Kits overlap, declarations reconcile: network rules union, hooks run in dependency order, guidance becomes one document, licenses union. Incompatible requests are an error, not a coin flip.

A kind: set descriptor names other Kits; publishing it runs the same coherence rules at build time and merges them into one ordinary Kit. An incoherent set fails at your build, not at someone else’s launch.

The diff is the review

The Claude Code Kit in the repository declares the hosts it asks to reach, its credential, the volumes that persist between sessions, and its install and startup hooks. When the next version asks for another host or a second credential, that is a change in authority, not a software update, and it shows up in the pull request as added lines a human can refuse.

Review depends on somebody reading the diff, so the specification defines a second gate that does not. Every descriptor reduces to a normalized set of everything the host would have to grant; a runtime that gates updates records that set and compares the next version against it. A version inside what was granted may apply without asking. Any widening stops and asks, and removing a deny rule counts: if a later gh Kit dropped DELETE /repos/**, the runtime holds the upgrade. That is why the declarations had to live in the artifact, not beside it.

Why this is a specification and not a feature

A Kit that stopped meaning anything when run somewhere else would be lock-in, not a trust boundary. So the grammar is normative, every capability type has its own page describing what a conforming runtime must implement, and types version independently (network-policy@1 and @2 both exist today). Two conformance suites ship with it: one judges whether an artifact is a conforming Kit, the other whether a runtime behaves as the pages say. Every normative statement is covered by a check or a written waiver.

Docker maintains the specification today, and it should not stay under a single vendor: a format for deciding what an agent may do is worth less if it belongs to whoever sells you the runtime. Docker Sandboxes will be a first-class implementation, not the only one. If a Kit you want cannot be expressed, or a runtime duty cannot be implemented as stated, open an issue.

Try it

sbx is our sandbox CLI (brew install docker/tap/sbx). From a checkout of the repository:

cd examples
sbx run ./hello --kit ./gh .

Edit a descriptor and only that Kit rebuilds; docker buildx build publishes it to any registry. Docker Cloud Sandboxes runs the same Kits with the same trust model on elastic capacity. The specification, capability pages, and a worked tour are in docker/sandbox-kit-spec.

Not only agents

Agents forced this into the open because the authority they ask for is so large, but ordinary workloads have always arrived with unwritten expectations: the endpoints they call, the credentials they need, the volume that must survive a restart. That knowledge has lived in a Helm chart, a runbook, or a colleague. It is the same gap, less alarming when a web service gets it wrong. This specification is where any software writes down what it needs from the world around it; agents were the case urgent enough to have it built.

Dockerfiles made software reproducible. Kits make authority reproducible.

  •  

Docker and CNCF partner on an open spec for agent permissions

Docker and CNCF: Making what an agent may do as portable as the agent itself

Ten years ago, the software industry faced a choice. Every vendor could ship its own image format and its own runtime, and developers would have to pick a side. Or the industry could agree on one artifact. The second option won. Docker donated its image format and the Runc runtime to the Linux Foundation, and the Open Container Initiative (OCI) formed around them. Today, a Docker image built anywhere can run anywhere. The format is the backbone of the cloud native ecosystem and a de facto standard.

Now, we see a similar problem forming around agents. There is no shared format for what an agent is allowed to do. We are proposing the same kind of answer: one artifact, built on OCI, governed in the open.

The same problem, for agents

Containers were built for immutable software. The image is the application. If you want to change it, you rebuild it, and it behaves the same way everywhere. That’s why a container image describes how software is built and says nothing about what it may do once it runs. For a web service, that was fine. It got a network and a port, and that was enough.

Agents are mutable by definition. Claude Code and Codex install packages, call APIs, and use credentials on your behalf. They change the environment they run in, and they decide what to do next. So, every team writes its own rules for what an agent may reach: a network rule here, a token there, a volume mount to get a task done. Those rules live in shell history, in dashboards, and in someone’s memory. A few months in, nobody can answer a simple question: what is this agent allowed to do?

Every team answers that on its own today, and every runtime vendor could ship its own way of answering it. That’s the kind of fragmentation OCI was created to prevent.

What we are announcing

Kits are not new. Kits have been part of Docker Sandboxes as the way you package an agent, its tools, and what it may reach into something a team can share. What’s new is the artifact. A Kit is now an ordinary OCI image, and the format that describes it is open.

Today at WeAreDevelopers, we announced the Docker Sandbox Kit Spec, open source under Apache 2.0. A Kit carries three things in one image: the agent, its tools, and a typed list of everything it asks to reach, such as hosts, credentials, and volumes. Because the list is part of the image, pinning the image pins the agent and its requests together.

A Kit is not a new artifact type and not a fork of any OCI specification. It uses an extension point OCI already defines. It builds, pushes, pulls, signs, and scans like any image you run today, because it is one.

Today, we’re bringing the spec to CNCF, under their neutral governance, just like we did when the image format went to OCI.

Why it matters

Adoption is free. Because a Kit is an OCI image, every registry, scanner, and signing tool you already run handles it. There is nothing new to deploy.

The answer travels with the agent. Because the requests are in the image, “what may this agent do” has one answer. A teammate can pull it. A reviewer can diff it. A conforming runtime can enforce it. When a new version asks for more, the change shows up as added lines someone can refuse.

An ecosystem, not a format

Docker has always had an ecosystem-first mindset. The Dockerfile mattered because anyone could write one, any registry could store the result, and any runtime could run it. Kits follow the same approach, and we did not build them alone.

We have worked with AWS, Box, Datadog, Dynatrace, JFrog, NanoClaw, OpenClaw, Palo Alto Networks, Snyk, and more to build Kits for their tools. Cloud platforms, observability, security, artifact management, content, and agent frameworks are all represented. The Kits we unveiled during the opening keynote at WeAreDevelopers today are the product of all that work, and they are the first of many.

MCP gave agents a standard way to talk to a tool. Kits give the ecosystem a standard way to publish the whole arrangement: the agent, its tools, and what it asks to reach, in one image anyone can pull. That’s what turns a format into a supply chain. For example, a database vendor can publish a Kit that connects any agent to its service with the scope it recommends.An agent maker publishes its own Kit, so the request list comes from the people who built the agent. A platform team publishes one for the company, and every engineer starts from the same place.

None of that happens if the format belongs to one vendor. A standard for deciding what an agent may do is worth a good deal less if it belongs to whoever sells you the runtime. Docker Sandboxes is the first runtime that enforces it. It should not be the only one, and under CNCF governance, it will not be.

Containers made software portable. Kits make authority portable: the set of things you deliberately hand over to an agent travels with the agent, in the same image, and means the same thing wherever a conforming runtime opens it.

“Standards are what let an ecosystem move fast without fragmenting, and few companies understand that better than Docker. By delivering Sandbox Kits as standard OCI images, Docker is giving the industry an open, repeatable way to package an AI agent, its tools, and its guardrails as one artifact. OCI is the foundation the cloud native ecosystem is built on, so a standard for agents that builds on OCI reaches the whole ecosystem at once. The CNCF welcomes this, and we’re excited to work with Docker and the community on making it broadly adopted.“

Chris Aniszczyk

CTO at CNCF

Build a Kit

If you make a tool agents use, publish a Kit for it. If you run agents, start from one and share it with your team. The specification, the capability pages, and a worked tour of a real Kit are at docker/sandbox-kit-spec. If there is a Kit you cannot express, or a rule a runtime cannot implement, open an issue. Every Kit published and every issue filed is how a standard gets built.

  •  

Introducing Cloud Sandboxes: Start on Your Laptop, Finish in the Cloud

Run agents on your laptop, in the cloud, and move between them with one command, all safely. Earlier this year, we launched Docker Sandboxes: microVM environments where coding agents can work autonomously and safely. 

Since then, agents have started taking on long-horizon work: tasks that run for hours, not minutes. Agents need a new place to work. Today, we’re introducing Cloud Sandboxes: the same microVM-based sandbox, running on Docker-managed compute, with one command to move between them. 

The most important change in coding agents over the past year is that they can work for much longer. Tasks that used to need a developer checking in every few minutes (a large refactor, a dependency migration, a test suite that takes an hour) you can now hand off and review when they finish.

When agents worked in short bursts, the question was whether the model could hold a task together. Now that they work in hours, the question is where those hours happen. A laptop is built around a person. It sleeps when the lid closes, slows down on battery, and disconnects when you move. None of that matters for a task that takes thirty seconds. All of it matters for a task that takes all night.

We built Docker Sandboxes to answer the first question about agents: is it safe to run unattended? In Docker Sandboxes, agents run in a microVM with its own kernel and Docker daemon, isolated from your machine, so they can work autonomously while your files, network, and secrets are protected if an agent steps out of line.

Today, we’re introducing Cloud Sandboxes to answer the second question: how do I run a dozen agents at once, for five, ten, or 21 hours each, without watching any of them? Cloud Sandboxes is the same microVM, running on Docker-managed compute. The isolation model is identical. The CLI is identical. What changes is that the machines underneath are always on, and there are as many of them as you need.

Screenshot 2026 09 23 at 11.59.36 AM

Here’s what that means for you:

Close your laptop and keep the work going. Start an agent in the cloud before you leave for the day, disconnect, and review what they did in the morning.

Start local and move to the cloud when your task outgrows your laptop. Iterate with an agent on the code in front of you, then hand off long tasks and loops to a background agent.

$ sbx move my-project --to cloud

A move captures the sandbox’s filesystem and recreates it on the other side, so your work carries over. This works in both directions.

Run 100 tasks in parallel with nothing to provision. Docker runs the compute. You point your agents at it. Each one gets its own microVM, its own secrets, and its own network policy.

Try things cheaply. Run a pre-built agent for an hour to test an idea, then spin it down.

Everything an agent needs to run on its own

Agents that do long-horizon work without you need three things:

  1. A way to start with nothing to install
  2. Tools to do their jobs
  3. Limits on what it can reach

Cloud Sandboxes ship with all three. You can spin up and manage Cloud Sandboxes from the CLI or from the web console.

Kits. A kit is a pre-configured, pre-built sandbox for an agent. The leading coding agents are ready today, including Claude Code, Codex, Copilot, Antigravity, Open Code and Hermes. Or, you can easily add your own. It’s easy to spin up. sbx –cloud run codex starts one, for example. 

Screenshot 2026 09 23 at 12.01.19 PM

MCP. Connect the MCP servers your agents need (Jira, Linear, Grafana, incident.io, or any streamable HTTP endpoint) once, and all your agents can reach them through a single gateway whether they’re running in the cloud, locally, or in other clients like the ChatGPT desktop app.

Screenshot 2026 09 23 at 12.01.41 PM

Secrets. Store keys or tokens once. Cloud Sandboxes proxy injects it per request, so agents don’t see the actual secret. Prompt injections can’t touch secrets your agents never had in the first place.

Screenshot 2026 09 23 at 12.02.03 PM

Policies. Set network policies on what endpoints agents can access by defining them once. Centralized governance is coming soon for enterprises through Docker AI Governance.

You can access and manage all of this through the web console.

Why one sandbox in two places

We could have built Cloud Sandboxes as a separate product with its own commands. We didn’t, and the reason is simple: how much you can trust your agent shouldn’t depend on where it happens to be running.

As far as we know, no other agent sandbox works this way. We think the right answer to using agents is one strong isolation model that runs in both surfaces (locally and in the cloud), and an easy way (one command) to move between them.

These aren’t two tiers of one product. Interactive work belongs on your laptop. Work that takes hours belongs in the cloud. Most developers need both, and now you don’t have to pick.

Pricing

Cloud Sandboxes are pay-as-you-go. We meter compute by the second, and nothing else. A paused sandbox costs nothing. Volumes, egress, and hosting public images and Kits are free. You can bring your own model key and keep your existing provider for inference.

Size

vCPUs

Memory

Per hour

Micro

1

2 GiB

$0.07

Small (default)

2

4 GiB

$0.14

Medium

4

8 GiB

$0.28

Large

8

16 GiB

$0.56

XL

16

32 GiB

$1.12

Sandboxes run for one hour by default and up to 24 hours per session.

Get started

From the browser: sign in to the web console, choose a kit, and click Run. 

From the terminal:

$ brew install docker/tap/sbx
$ sbx login
$ sbx --cloud run claude

You’ll need sbx 0.45.1 or later and the pay-as-you-go plan, available on Docker Personal and Pro accounts. Docker Sandboxes on your laptop remain free and standalone, with no Docker Desktop required. Local and cloud sandboxes keep separate secrets, templates, and network policies, so read the differences in the docs before you move a local workflow.

For a limited time, new accounts get $250 in free Cloud Sandboxes credit. Claim it here.

We think the way software gets built is changing. Developers will spend less time typing and more time directing: handing agents real tasks, stepping back, and reviewing what comes back. For that to work, agents need environments that are safe enough to run unsupervised and durable enough to run for hours. That’s what Docker Sandboxes is for, and as of today, it runs on your laptop and in the cloud.

Start one before you close your laptop tonight.

  •  

Meet the Ecosystem: Partners and Customers at WeAreDevelopers with Docker

As teams put AI agents to work, they need to move quickly without losing control of what they deploy. They’re combining models, tools, and infrastructure from across a fast-changing ecosystem. Making those pieces work together and keeping them accountable as the stack evolves is becoming a core part of building AI applications.

Docker’s approach to this challenge is providing a trusted, common foundation for containment, curation, and control of agent workloads at its core, while pairing those capabilities with an open ecosystem of partners and tools.

That ecosystem spans model providers, MCP tools and gateways, enterprise applications, data and memory platforms, identity, security, observability, and code quality. It also includes the cloud providers, systems integrators, and channel partners that help organizations bring these capabilities into production. 

Integrating this ecosystem gives teams the freedom to choose the models, platforms, and clouds that fit their needs while maintaining a consistent foundation for governance. Developers remain in the lead: choosing what agents can access, directing their work, and verifying the outcomes. The goal is to give them the tools and guardrails to build with confidence as models, frameworks, and requirements change.

At WeAreDevelopers World Congress North America, September 23–25 in San Jose, partners and customers are bringing that ecosystem to life at the Docker Pavilion. Customer sessions will show how these technologies come together in practice, from repeatable AI deployments at the edge to simpler development with payment APIs. Lightning talks and demos will explore enterprise knowledge and agent memory, collaboration between agents, security and incident response, and verification of generated code. 

Here’s who you can meet and what they’ll be sharing.

Customer talks — September 24

Customers bring another essential perspective: how these technologies come together in the systems they build.

  • Spectro Cloud: In “Repeatable Agentic Workloads on Palette,” Colton Shaw will demonstrate how a versioned cluster profile brings together hardened images, local inference, and agent workloads for repeatable edge deployments, including environments without a cloud connection. 12:15–12:30 PM.
  • Joint panel “From TokenMaxxing to True AI Ownership,” hosted by Per Krogslund from Docker and executives from Spectro Cloud and J.P. Morgan Payments, for a conversation about moving beyond token consumption toward ownership of how AI is deployed, governed, and put to work. September 24, 3:45 PM.
  • J.P. Morgan Payments: In “Insert Coin: docker compose up with J.P. Morgan Payments,” Alan Torrance will show how developers can run Unicorn Finance with one command and no API keys. The open source example brings a client, mock server, and the real OpenAPI specifications behind J.P. Morgan’s Payments APIs together in two containers. 4:30–4:45 PM.

Partner talks — Sep 24, 2026

  • Palo Alto Networks: Investigate agent activity through searchable audit records and live detections in Cortex XSIAM, with Cameron Hyde showing the integration in action. 11:15–11:30 AM.
  • Datadog: Follow an agent security incident from detection to investigation and response, with Amrita Lakhanpal connecting AI Guard, service context, and incident management. 12:45–1:00 PM.
  • ClickHouse: Reduce unnecessary components in your database’s base image. Zoe Steinkamp will walk through running ClickHouse on Docker Hardened Images. 1:15–1:30 PM.
  • Prediction Guard: Explore how execution isolation and controls over model calls work together, with Sharan Shirodkar testing both against a poisoned tool output. 3:15–3:30 PM.
  • Snyk: See the prompts, file activity, and generated code behind an agent’s work, with Javier Garza demonstrating the Evo Agentic Development Security Sandbox Kit. 5:00–5:15 PM.

Partner talks — Sep 25, 2026

  • GitGuardian: Put controls around the moments an agent reads files, edits code, or runs commands, with Dwayne McDaniel showing how hooks can help protect secrets. 9:00–9:15 AM.
  • Mend.io: Add runtime guardrails to detect malicious inputs, prevent unsafe actions, and record agent activity, with Gary M Segal demonstrating the approach. 9:30–9:45 AM.
  • Merge: Give agents access to an integration catalog while keeping third-party credentials outside the sandbox, with Gil Feig explaining how the pieces connect. 9:45–10:00 AM.
  • BAND: Explore how separately sandboxed coding agents can exchange tasks, messages, and artifacts, with Vlad Luzin demonstrating collaboration through Jam. 12:15–12:30 PM.
  • Chainloop: Give reviewers evidence of what an agent actually did. Daniel Liszka will demonstrate signed session records and policy checks on a pull request. 1:15–1:30 PM.
  • Box: Turn enterprise documents into deliverables that people can review, with Carter Rabasa demonstrating governed document access, evidence checks, and isolated code execution. 2:30–2:45 PM.
  • SurrealDB: Build agents with memory you can inspect over time, with Chiru Boggavarapu showing how to trace what an agent knew and when. 2:45–3:00 PM.
  • Cognee: Give agents temporary access to company knowledge and remove it when the task is finished, with Vasilije Markovic demonstrating a practical architecture. 3:45–4:00 PM.
  • Sonar: Guide and verify agent-generated changes using Sonar Vortex and the SonarQube CLI, with Manish Kapur demonstrating the workflow inside a sandbox. 4:45–5:00 PM.

These sessions bring together the people building the tools and the teams putting them to work. It’s an opportunity to compare approaches, ask questions, and see how the ecosystem can help you tackle your next engineering challenge.

Come visit us at WeAreDevelopers. Meet our partners, customers, and speakers, catch a lightning talk, and see their technologies in action. Plan your visit to San Jose.

  •  

6 Benefits of Sandbox Environments (and How Docker Sandboxes Delivers Them)

In our State of Agentic AI report, 60% of organizations reported having AI agents running in production. Those agents install packages, run scripts, and call external services on their own, and much of that work now happens on developer laptops, with developer credentials. Running untrusted or experimental code directly on your machine has always carried risk, and handing that same machine to an autonomous agent raises the stakes.

A sandbox environment gives code a separate, controlled space to run in, with limited access to the machine underneath and external systems. How strictly it holds that line depends on how the sandbox is built, which is where the differences between them start to matter.

The benefits of sandbox environments are worth understanding on their own, and they compound when the thing running inside is an agent working unattended with permissions auto-approved. Below are six, from isolation and credential handling to the policy you enforce at runtime, and how Docker Sandboxes delivers each one.

Key takeaways

  • A sandbox gives you a hard isolation boundary, so untrusted code or autonomous agents run without access to the host machine.
  • Docker’s sandbox environments offer benefits like isolation, policy you control, safe credentials, disposability, a real Linux dev environment, and the same sandbox technology for every agent.
  • A sandbox enforces the network and filesystem policy you define at runtime, which is what makes it the enforcement point for governance.
  • For AI agents, these benefits combine into full autonomy inside a boundary that allows them to get work done, safely.
docker 6 Benefits of Sandbox Environments

1. Isolation

Everything in this list builds on isolation, and the strength of that boundary is what makes a sandbox trustworthy. For Docker Sandboxes, each sandbox runs in its own microVM: a lightweight virtual machine with its own Linux kernel, isolated from the host by a hardware-backed hypervisor boundary. 

That boundary is the same kind of isolation a full virtual machine gives you, and it’s what lets you hand an agent real freedom. Because a Docker sandbox runs its own kernel, a compromised or runaway agent can’t reach the host, other sandboxes, or anything outside its environment. If it tries to escape, it hits a wall. So an agent can install packages, pull untrusted dependencies, and run code unattended. But when something inside goes wrong, the damage stays in the sandbox and disappears when you discard it. That containment is what makes it safe to let an agent run at full speed.

ⓘ MicroVM vs. container isolation: A (Linux) container shares the host’s kernel, so its isolation depends on kernel-level controls. Note that when using Docker Desktop, in order to provide an environment for running Linux containers, you’re already using a VM for hosting containers, so they are isolated from the host OS. However, all containers still share the same kernel (the one of the Linux VM). Hence, you won’t have strong isolation between containers.

2. Network and filesystem controls you define

Isolation sets the outer wall. The controls you define decide what the workload can reach while inside it. Most sandboxes let you scope network and filesystem access to some degree: which domains and IP ranges the workload can reach, and which paths on the host, if any, it can read or write. How precisely you can express that policy varies between tools, and it’s worth checking before you commit, because broad-strokes rules leave gaps that an agent will eventually find.

Docker Sandboxes lets you set that policy per sandbox and enforces it at the boundary at runtime, so the rules hold even when the code inside tries something you didn’t anticipate. The same controls that keep an experiment from making unauthorized outbound connections also shut down data exfiltration and block access to untrusted or malicious services. Restricting the filesystem keeps sensitive host paths, like SSH keys and cloud credentials, out of reach.

3. Secure credential handling

Agents need credentials to do useful work: a token to push to a repo, an API key to call a service. The risk is that a credential sitting inside the environment can be read, logged, or leaked by whatever runs there. Most sandboxes pass secrets in as environment variables or mounted files, which puts the value inside the boundary where the workload can read it, and so can anything the workload runs.

Docker Sandboxes keeps credentials out of the environment entirely. They stay in the host keychain, and the sandbox injects them into outbound network requests at the boundary, so the workload gets the benefit of the credential while the value itself stays on the host. An agent that can’t read a secret also can’t exfiltrate it, write it to a log, or hand it off to a prompt-injected instruction. The credential does its job on the request path while the sensitive material stays under your control.

4. Ephemeral, disposable environments you can recreate fast

A sandbox is quick to create and easy to throw away, so you can treat every one as disposable. When a task finishes, or when an agent goes off the rails, you can delete the environment and everything inside goes with it, from installed packages to running processes to any changes the agent made to the system. But if your working directory is mounted from the host, the files the agent creates or edits there stay on your machine even after the environment is gone.

The recreation side is just as valuable. Because a sandbox is defined in code, you can spin up an identical environment on demand, configured the same way every time, down to the packages and settings. This is the infrastructure-as-code approach applied to your workspace: reproducible, versionable, and consistent across a team. For agents, disposability also unlocks parallelism. You can run several agents at once, each in its own fresh environment, and tear them all down when the work is done.

5. A real Linux dev environment with a full Docker daemon

Isolation doesn’t have to mean a stripped-down box. A sandbox worth using gives the workload a real Linux environment with the tools a developer or an agent actually needs, so you can install packages, run services, start databases, and compile code inside the boundary. Environments vary widely in how complete they are, and a thin one pushes work back onto the host, which defeats the point of having a boundary at all.

Docker Sandboxes includes a full Docker daemon, isolated within the sandbox, so an agent can build and run containers as part of its work with no path back to the host daemon. That’s a meaningful capability for agentic workflows, where a single task might involve building an image, running a test suite in a container, and tearing it all down. The environment behaves like a genuine machine, which is what makes it a viable place to do real work.

6. The same sandbox technology for every agent

Developers will often move between agents. One task suits Claude Code, another suits Gemini CLI, Copilot CLI, Codex, Kiro, or OpenCode. If each agent brought its own isolation model, you’d be securing a different environment for every tool, and each vendor’s model could shift with a version bump.

A single sandbox technology solves this by running every agent the same way, inside the same kind of isolated environment with the same policy engine. You define network, filesystem, and credential policy once, and it applies no matter which agent is doing the work. For a platform or security team, that consistency is what makes governance enforceable at scale: one boundary to reason about, one set of controls to audit, across every agent your developers adopt.

Who gets the most from sandbox environments

The same six benefits pay off differently depending on your role.

  • Individual developers
    • You get freedom to experiment. You can try a risky dependency, run an unfamiliar tool, or let an agent work unattended, knowing the environment is contained and disposable. When something breaks, you delete it and start clean, and your machine is never in the blast radius.
  • Platform teams
    • You get consistency and control. A sandbox defined once gives every developer the same environment and the same policy, across whichever agents they use. That means less setup for your developers to think about and a single standard you can maintain centrally.
  • Security teams
    • You get containment and oversight. A sandbox limits what an agent can reach and gives you one boundary to monitor across every tool. You can approve agent adoption because the environment enforces your policy at runtime, which is the heart of securing AI agents in production. Every environment is disposable, so there’s nothing persistent to compromise.

Why this matters for AI agents

Put the six together and you get the reason why sandboxes might become the standard way to run agents. An agent needs autonomy to be useful. It has to install things, run code, and call services without a human approving each step. Autonomy on your host machine is dangerous, but put it inside a sandbox and it’s safe.

Isolation contains what the agent can do, and the controls you define scope what it can reach. Credentials stay out of its hands, so a compromised agent has nothing to leak. When a run goes sideways, disposability lets you throw the environment out and start over in seconds. And a real Linux dev environment means the agent can do genuine work, and running every agent on one sandbox technology keeps all of this consistent no matter which tool your team reaches for. Together, these benefits let an agent operate at full speed while keeping the blast radius of any mistake close to zero.

Run agents safely with Docker Sandboxes

These benefits depend on each other, and a gap in any one becomes the weak point a runaway agent finds first. Isolation without credential handling still leaks your secrets, and a dev environment you can’t tear down cleanly turns into a liability the first time an agent misbehaves.

Running agents safely means delivering all six together, and that’s what Docker Sandboxes is built to do. Containment comes from microVM isolation, the controls are the network and filesystem policy you set, and credentials stay in the host keychain, injecting at the boundary so the agent never sees them. Environments are disposable and defined in code, the workspace is a real Linux system with a full Docker daemon, and the same sandbox technology runs every major coding agent the same way.

And when you’re ready to run agents safely across a team, Docker AI Governance extends the same boundary into org-wide policy. You define network, filesystem, and tool-access rules once, govern which credentials a session can use, and apply it on every developer’s machine, with an audit trail security can defend.

Get started with Docker Sandboxes → 

Explore Docker AI Governance →

Frequently asked questions

What is a sandbox environment used for?

Sandboxes give coding agents and the code they run an isolated, disposable place to execute, fully separated from the host. The main use is running AI coding agents like Claude Code, Codex, or Gemini CLI unattended, letting them install packages, run services, and even run Docker inside the sandbox, and trying risky changes you’d rather keep off your machine.

What is the main benefit of a sandbox environment?

Isolation. A sandbox keeps whatever runs inside from reaching the host, so a mistake, a malicious package, or a misbehaving agent stays contained.

Are sandbox environments only for security?

No. Security is a major benefit, but sandboxes also improve reproducibility, speed up onboarding, and let developers and agents experiment freely, because the environment is disposable and defined in code.

Do sandbox environments slow developers down?

They don’t have to. MicroVM-based sandboxes like Docker Sandboxes start in seconds and give you a full Linux environment right away, so isolation adds safety at very little cost to speed.

How do sandboxes help with AI agents?

They let an agent run with full autonomy while containing what it can reach. Isolation limits the blast radius, the policy you define scopes access, and credential handling keeps secrets out of the agent’s hands.

  •  

YOLO Mode: Agent Autonomy Without the Guardrails

AI agents have come a long way in both capability and everyday use since generative AI went mainstream in late 2022. In Stack Overflow’s 2025 Developer Survey, 84% of developers said they use or plan to use AI tools in their workflow, up from 76% a year earlier. As those tools shift from suggesting code to writing files and running commands on their own, one practical question follows. How much should an agent be allowed to do without stopping to ask? Turn that dial all the way up and you reach what developers call YOLO mode.

It’s worth understanding YOLO mode before you enable it, because its main risk is easy to misread. The risk comes down to where an agent runs.  On your own machine, one mistaken command can delete  files, expose your credentials, and make network requests you may not want. Inside a proper boundary, however, developers can use agents in YOLO mode to unlock a new level of productivity, without jeopardizing security.

Key takeaways

  • YOLO mode is when an AI agent auto-approves every action, with no confirmation prompts.
  • It’s popular because it’s fast, and risky for the same reason. The danger isn’t the autonomy, it’s where the autonomy runs.
  • On your host, a bad command or prompt injection reaches real files and credentials. Inside an isolated sandbox, the blast radius is contained.
  • Run YOLO mode where it can’t do real damage, in an isolated, disposable environment with scoped access and no real secrets.

What is YOLO mode?

YOLO mode is the community nickname for running an AI agent with every action auto-approved. When turned on, agents can read files, write code, run shell commands, and call tools without stopping for user approval. While in Claude Code it’s the –dangerously-skip-permissions flag, other common agents each have their own version of the same switch.

  • Codex CLI has `–full-auto`, plus `–dangerously-bypass-approvals-and-sandbox` when you drop the sandbox too.
  • Gemini CLI uses `–yolo`, or the Ctrl+Y toggle mid-session.
  • GitHub Copilot CLI has `–allow-all`, also aliased as `–yolo`.
  • Cursor exposes it as auto-run in settings rather than a flag.

The names differ, but the behavior is the same: remove the prompts and let the agent go. 

YOLO mode showed up in Cursor first, then Claude Code, and by 2026 it’s a standard toggle in most coding agents. But when people ask what YOLO mode is, they’re usually asking whether they should use it, and the answer is that it depends entirely on where the agent is running.

Why developers turn it on

On a regular task, a careful agent asks for permission constantly. “Can I edit this file, run this test, install this package, call this tool?” 

Dozens of prompts for one feature. While these constant permission requests can help prevent agents from going rogue, each approval forces you to context switch and breaks the flow that made the agent worth using. A few reasons why developers are leveraging YOLO mode include:

  • Context switching: Every approval pulls a developer out of their flow, taxing mental focus and overall productivity. 
  • Prompt fatigue: Excessive querying, refinement, and approvals force creative coding to take a back seat to tedious prompt wrangling and debugging.  
  • Low-risk, routine work: Agents can often handle repetitive tasks that would otherwise take developers away from creative coding and innovation. 
  • Momentum: An agent is most useful when it has the freedom to keep moving, but a steady stream of prompts breaks that.

If you turn approvals off, these friction points disappear for the most part, and the agent can deliver the speed it promised. But what’s the cost of giving agents the autonomy of YOLO mode?

Why is YOLO mode risky?

When you remove the prompts, you remove the last human check before an action runs, which amplifies the security risks agents already carry. If the agent is working directly on your host, that action has the full run of your machine, including your files, environment variables, credentials, and network. A confused or compromised agent can do a significant amount of damage when nothing stands between an agent’s decision and your system.

On an unprotected host, YOLO mode introduces risks such as:

  • Destructive commands: A vague or mistaken instruction runs something like rm -rf against the wrong directory, and nothing pauses to catch it.
  • Secret and credential exposure: The agent can read environment variables, .ssh keys, tokens, and .env files, then use or leak them.
  • Prompt injection: The agent acts on whatever it reads, so a hidden instruction in a web page, an issue, a code comment, or a document can redirect it, and the attacker never needs access to your machine.
  • Data exfiltration: A mistaken or hijacked agent sends sensitive data out over the network.
  • Unintended broad changes: Edits and config changes reach past the task at hand into your other projects.
  • Network and lateral reach: The agent can hit internal endpoints and outside services, or act with your credentials to push code and call APIs.

And unfortunately, keeping manual approvals on doesn’t remove all risk. Once permission fatigue kicks in, it can be all too easy to accidentally approve the wrong request. So the safeguard belongs in the environment the agent runs in, where a bad command or a tired click has a greatly reduced scope of impact.

The fix isn’t fewer permissions, it’s a boundary

If prompts aren’t the answer, what is? A boundary the agent can’t cross. Guardrails only work when something outside the agent enforces them. The agent needs a bounding box, with constraints set before it runs and clear limits on what it can touch. Inside that box, it should be free to move as fast as it wants. The goal is to shape the environment so that a mistake can’t damage your systems or leak your secrets.

Comparing YOLO mode with and without a sandboxed environment.

In practice, that means running the agent in an isolated, ephemeral environment instead of on your host. Done well, the agent gets a real place to work. It can install packages, run services, and edit files, but it can’t see your credentials, reach your other projects, or touch the host.

Unlike a container that shares the host kernel, a microVM puts a hardware-level boundary around the agent, so the isolation holds even if the agent tries to break out, and it does that without the speed penalty people expect. If a run goes sideways, you destroy the environment and start clean. This is the core idea behind sandbox security and why agents need isolation in the first place.

What does YOLO mode look like at scale?

For one developer on a sandboxed laptop, YOLO mode is a personal choice. Across a team, it becomes a policy question. A hundred developers each deciding on their own when to skip permissions is the ungoverned-autonomy problem that keeps security leaders up at night. The picture that works at scale is one where the safe path is the default. Every agent runs inside an isolated, disposable environment, configured once at the organization level so it holds for everyone.

This is the problem AI Governance is built to solve. You define the rules once across the surfaces that matter, network access, the filesystem, and the tools an agent can reach, then enforce them automatically at every developer’s machine. Governance turns a per-developer judgment call into a consistent, repeatable capability. Clear boundaries are what let an organization extend autonomy to its agents while keeping the risk contained. Once the boundary is standard, YOLO mode is fast and safe for everyone.

What it unlocks for developers

Once the boundary is in place, the developer can stop supervising every step, and the payoff kicks in:

  • Deep focus: Give direction, step away, and come back to a cloned repo, passing tests, and an open pull request. No interruptions pulling you off your own work.
  • Long, autonomous runs: The agent edits, runs the tests, reads the failures, and retries until the task is done, the kind of run a wall of prompts would stall.
  • Agents in parallel: Point several at different tasks, each in its own disposable environment, and let them run at once.
  • You review the outcome: Your job moves up to the pull request, the tests, and the diff, where your judgment matters most.

That’s the real appeal, and the sandbox is what makes it safe to lean on.

Unlock agent autonomy, safely

YOLO mode is really a question in disguise. How much autonomy can you give an agent before the risk outweighs the speed? Framed that way, the answer stops being about the agent and starts being about its environment. Give an agent the run of your laptop and even a small mistake is expensive. But give it a boundary it can’t cross and you get the speed with almost none of the exposure.

That’s exactly what Docker Sandboxes is built for. Each agent runs in its own disposable microVM with control over networking, filesystem access, and resource limits, so you can run agents in YOLO mode safely from day one. For teams that want those boundaries applied consistently rather than agent by agent, Docker AI Governance sets and enforces the rules everywhere developers work. Define the box. Then let the agent go as fast as it likes.

Get started with Docker Sandboxes → 

Explore Docker AI Governance →

Frequently asked questions

Is YOLO mode safe?

It depends entirely on where the agent runs. On your host machine, YOLO mode is risky, because a mistake or a prompt injection can reach your files and credentials. Inside an isolated, disposable environment with scoped access and no real secrets, the blast radius is contained and YOLO mode is reasonable to use.

What does –dangerously-skip-permissions do in Claude Code?

It turns off the confirmation prompts, so Claude Code reads, writes, runs commands, and calls tools without asking for approval at each step. It trades the safety of human review for speed. It’s the most common way people run Claude Code in YOLO mode.

How do I use YOLO mode safely?

Run the agent inside an isolated sandbox rather than on your main machine, give it scoped network access and throwaway credentials instead of your real ones, work against a cloned or disposable copy of your project, and keep a way to inspect what it did. The goal is a boundary the agent can’t cross, not a more careful set of prompts.

Is auto mode the same as YOLO mode?

Not exactly. Full YOLO mode approves everything. Some tools now offer a classifier-gated auto mode that runs safe actions automatically while still blocking or flagging dangerous ones. That’s a useful middle ground, but it’s a filter on top of the agent, not a boundary around it. Isolation still matters.

  •  

Building Reproducible AI Evaluation Workflows with Docker Sandboxes

AI evaluation has never been easier to start. Reproducing it reliably is another story. Developers now have access to more benchmarks, evaluation libraries, model APIs, and agent frameworks than ever before. But keeping the prompt, model, and scoring method fixed doesn’t necessarily make a run reproducible. The execution environment matters too.

Python dependencies change. Local tools drift. Setup steps go undocumented. A workflow that succeeds on one machine may behave differently on another. Most discussions about evaluation focus on what should be measured: benchmarks, scoring methods, or judge models. Much less attention is given to how those evaluations are executed. Yet that execution layer often determines whether someone else can reproduce the same workflow weeks or months later.

When I started exploring Docker Sandboxes, I wasn’t trying to build another evaluation framework. I had a much smaller question.

Could Docker Sandboxes and an SBX Kit make evaluation workflows easier to rerun, inspect, and compare?

That question eventually became the SBX AI Evaluation Kit, an open-source Docker Sandboxes Mixin Kit focused on repeatable execution, structured evaluation records, and runtime evidence. The current implementation does not execute AI models or automatically derive evaluation judgments. Instead, it executes configured commands consistently and preserves evidence of what actually ran.

In Practice

In practice, the workflow starts by choosing where the evaluation command should run through the execution block:

execution:
  executor: sbx
  command:
    - python3
    - -c
    - print("hello from sbx")

With executor: sbx, the runner delegates command execution to Docker Sandboxes and writes the runtime evidence into the resulting artifact.

The repository is also packaged as an SBX Mixin Kit, so it can be applied when starting a Claude sandbox:

sbx run claude --kit .

The runner reads the configured executor and delegates the command to SBX, which executes it inside the sandbox:

python run_evaluation.py

From Documentation to an Executable Workflow

Each evaluation is defined in a YAML file that describes the evaluation and the command to run. The repository validates that definition, executes it, and produces a structured JSON record of the result. The difference is in what gets recorded. A written evaluation captures what someone intended to do. An execution-backed evaluation captures what actually happened.

Separating Evaluation from Execution

I wanted the evaluation definition to stay independent of where it ran. A workflow written during local development shouldn’t need to change simply because it later executes inside Docker Sandboxes.

To keep those concerns separate, I introduced an executor abstraction. The evaluation describes what should run; the executor determines where it runs.

With the local executor, the configured command runs on the host. With the SBX executor, command execution is delegated to Docker Sandboxes. Switching between the two only requires changing the executor configuration, not rewriting the surrounding evaluation workflow.

image1

Figure 1. Evaluation definitions remain independent of the execution environment. The same workflow can use either the local or SBX executor while producing runtime evidence in the same structure.

Capturing Evidence Instead of Assumptions

For each execution, the runner records enough information to inspect what actually happened:

  • the selected executor,
  • the command that was executed,
  • standard output (stdout) and standard error (stderr),
  • the exit code,
  • and the execution time.

These details are stored in the evaluation artifact. The repository also generates a digest of the evaluation configuration. This creates a deterministic link between the evaluation configuration and the artifact it produced, without trying to replace full experiment-tracking systems.

{
  "executor": "sbx",
  "command": ["python3", "-c", "print(\"hello from sbx\")"],
  "stdout": "hello from sbx\n",
  "stderr": "",
  "exit_code": 0,
  "duration_ms": 120.0
}

Scaling from One Evaluation to Many

Real-world evaluation rarely consists of one isolated run. Teams compare prompts, validate behavior, measure regressions between releases, and test multiple scenarios. That led to evaluation suites.

Rather than changing how an individual evaluation works, a suite groups multiple evaluation definitions into a single repeatable workflow. Each evaluation still produces its own structured artifact, while the suite also generates an aggregated summary of the overall run.

Reusable SBX Kits Beyond Evaluation

The same pattern isn’t limited to evaluation. An SBX Kit can package more than a development environment; it can also package the setup an engineering workflow depends on. The same model could support regression testing, policy checks, security analysis, code-generation experiments, and other workflows that depend on consistent execution and inspectable results.

Conclusion

The SBX AI Evaluation Kit doesn’t replace evaluation frameworks, benchmarks, or scoring systems. Its job is narrower: execute configured evaluation workflows in a way that is easier to rerun and inspect.

The question I came away with is simple: before comparing benchmark scores or choosing a judge model, can someone else reliably run the same workflow under comparable conditions?

You can explore the code, experiment with custom evaluation YAMLs, and run the workflow yourself in the sbx-ai-eval-kit repository on GitHub.

Resources

  •  

Below the Harness: Governing a Multi-Model, Multi-Harness World

We believe the future is a multi-model, multi-harness world. And we think it needs a new trust model.

In 1988, Norm Hardy described a problem that had been quietly breaking systems for years: the confused deputy. A program that takes action using its permissions instead of yours.

Today, every AI agent is that deputy. It inherits your authority: Your credentials, your repo access, your ability to call APIs. But its behavior is probabilistic. It might be acting on an instruction found in its environment, on a step it invented, or on a confident wrong answer. 

The industry didn’t fix the confused deputy problem by making the deputy itself more careful. They fixed it by moving its authority a layer away. Forty years on, that’s still the answer.

Everyone is converging on the same future

Three facts are pushing the industry toward the same conclusion.

  1. Agents are expensive loops. An agent takes many steps, and you pay for every token of every one. We all can agree that it makes no economic sense to call the latest frontier model for simple tasks. 
  2. The leader of frontier capability changes often. We’re all aware that the top model of the day (and its vendor) changes every couple of months.
  3. Your workflows may need custom models. Many teams are recognizing that intelligence is commodifying and the differentiator is custom models, derived from custom context.

As a result, all of us are quickly ending up with a portfolio of multiple models across multiple harnesses. 

A similar convergence is happening one layer up. Developers pick certain tools for the right task, the way they always have. For example, perhaps Claude Code for long refactors, Codex for daily work, Hermes for quick scripts. 

It’s reasonable to expect the future of work to be multi-model and multi-harness.

Which makes trust the defining question

A lot of agents work the same way. 

They read material that is often out of our control: support tickets, web pages, documentation, and code written by strangers. But they act with authority you granted: your credentials, repo access, production APIs, and the open internet. And they usually do both from a developer’s laptop, outside typical security guardrails like VPCs and IAM.

Private data and the ability to act autonomously, together, is what makes an agent worth deploying. Your deputy needs the ability to execute in order to be useful. Which means the interesting question is no longer which model is best. It’s what happens when one of these deputies is wrong, or manipulated.

Per-harness guardrails break down

The obvious answer is that each harness ships its own guardrails. Many do. But relied on as your security boundary, they fail in three ways.

  1. The agent talks past them. Guardrails inside the harness are enforced in the same loop the agent is running. Deny it a git push and it reaches for the API. Deny the API and it opens a gist. Deny the gist and it tucks the data into a channel you trust and never inspect. Researchers showed last year that a single malicious issue filed in a public GitHub repo could steer a coding agent into reading a company’s private repositories and publishing the contents in a pull request the agent opened itself. Nothing was hacked since every step used the agent’s own legitimate access, through a channel everyone trusts. A boundary the agent can negotiate with is not a boundary.
  2. The rails move without you. Most harness’s isolation models are closed source and ship on their vendor’s schedule. The major coding agents have each revised their default sandbox and approval behavior several times in the past year alone. Updates to sandboxing models should be treated as a security event. Multiply this by ten harnesses and your security posture is, at any moment, whatever is the patchwork of your half dozen vendors’ measures.
  3. The rails don’t cover the fleet. The custom agent your platform team built has exactly the guardrails your platform team wrote. The agent inside your support SaaS has whatever its vendor chose, and most expose no isolation controls to you at all. Every new harness means building or auditing governance again, from scratch, differently. You end up with a dozen implementations that drift apart, each blind to the others’ traffic, with no single place to set a rule and no single record to understand why something went wrong.

Safety cannot depend on the agent making the right decision, or on someone else’s release schedule.

A layer below

So here is what we believe. The future is multi-model and multi-agent. And given that future, we believe every organization will need a layer below: a runtime layer, below the harness, that all of them run on top of.

The reasoning is straightforward. Strip away the model, the vendor, and the framework, and an agent has two ways to affect anything. It runs code, which touches files and opens network connections. Or it calls a tool, which acts on a system. Everything an agent does travels one of those paths. And both paths cross the same surface: the runtime, where processes execute, credentials get used, and requests leave the machine. Every agent passes through it, no matter which model powers it, which vendor shipped it, or whether you built it yourself. That makes it the one place where rules you define can be enforced across all your agents. It is also the same fix as 1988, applied to today’s deputy: the authority sits a layer away.

Put enforcement there and each of the three failure scenarios we spoke about flips around.

Your agents can’t talk past themselves. The boundary for an agent sits outside the loop the agent is running, so it holds steady no matter if the model is with you, hallucinating, or compromised. A hard neutral boundary at the runtime is more effective than a prompt-level boundary the agent creates for itself.

The rails stop moving randomly. Policy is yours, written once, covering execution, tool calls, credentials, and spend. Now, a model or agent vendor making an update won’t randomly change your security posture.

The rails cover your whole fleet. A policy you write up will apply to every harness. And every action, by all your agents, lands in one record: what ran, what it touched, which rule decided. 

This is what lets you be nuanced about agents. Without a boundary below your harnesses, you have three bad options: block agents completely, allow all of them and hope for the best, or wedge a manual approval into every step and give up the productivity you wanted.

A boundary at the runtime gives you a fourth option. When consequences are bounded even if an agent goes off the rails, you can start granting it true autonomy, which is the goal.

We expect models to keep changing and new harnesses to land in all of our toolkits. That part is healthy. The boundary underneath them is the part that should hold steady.

At We Are Developers in San Jose, Tushar Jain, Docker’s CTO, will talk more about this world: multiple models, multiple harnesses, and a single runtime under it all.

  •  

Secure by default is your only way forward

Every worker a company employs, be it a person or a program, builds on a foundation someone else assembled, and that includes the newest hire on your team. This new hire got to work the moment they arrived, building with what your company already has in place and they’re shipping code at a pace your reviews can’t keep up with. Also, everything they make is going out under your name. If it were a human, they’d spend the first week asking where things live and who maintains what. This one never asks. It treats everything it finds as trustworthy, so everything it builds carries that unexamined trust forward. And because this new hire is an agent that’s working all night at machine-class throughput, the foundational problems that used to surface slowly now surface all at once.

The foundation that nobody audited

The line between a supply chain attack and an AI attack no longer exists. Take a look at what the average foundation holds, because most of it comes from outside the company. For a long time now, public base images have carried hundreds of packages that your application never uses. Every one of those packages adds to the attack surface. Almost none of them ever get reviewed because no team has time to read code it didn’t choose and doesn’t use. In most stacks, something like a ten-year-old Java service is keeping the business running on software whose maintainers stopped patching years ago. Platform teams have been coping in their own ways, usually with a golden-image program somebody built years ago and a scanner pointed at it all. Because the images underneath are so bloated, that scanner cries wolf about four hundred times a week. All of this together is why audit season now eats up most of a quarter.

Attackers know all of this, and they’ve been working on the foundation layer all year. They’ve poisoned packages and developer tools, and they’ve had real success harvesting coding-assistant credentials at scale. Most foundations were built for a world that no longer exists.

What a good foundation takes

The good news is that none of this is unsolvable. A foundation can be strengthened to carry what’s now being built on top of it. It has to meet a few requirements, and each one depends on who does the security work, because when the vendor doesn’t, your team picks up the slack. A foundation holds when every part of it is built from source by someone who signs the work and stands behind it. Nothing should ship that your application doesn’t need, because anything extra adds surface area to defend later. Patching needs the same treatment because new vulnerabilities keep landing no matter how clean an image starts. A fix should come with contractual backing and a date. You should know exactly what’s inside every image the day it ships. And none of this should force you to move your stack onto a different distribution just to get safer images. A migration like that becomes a quarter-long project in its own right, and the foundation can’t protect anything until the move is complete.

This is exactly what Docker Hardened Images were built for. They stay compatible with the Alpine and Debian images teams already run, so adoption amounts to a one-line change to the FROM line in your Dockerfile, with no migration project attached. The images are also minimal by design, carrying only what your application needs, which reduces the attack surface by up to 95% and leaves near-zero critical and high CVEs from day one. The difference is immediately visible in scanning. Scans complete much faster with low noise, and the few findings that do remain are worth directing the team’s attention to. When a CVE does get disclosed, the remediated image is available within seven days of the upstream fix, and what once consumed a sprint of engineering time closes as a pull request. The same evidence carries through to audits, which most organizations will eventually face. Every hardened image ships with a signed SBOM (Software Bill of Materials) and build provenance, a verifiable record of the image’s contents and build process. You present auditors with proof that already exists, and no one needs to spend weeks reconstructing it.

Furthermore, a hardened base image by itself may not be enough, because minimal images almost always need customization before they fit production workflows. Teams add their own CA certificates and init scripts, install additional system packages through apt and apk, or adopt separate products entirely to cover what the base image cannot, fragmenting their foundation across vendors. That’s usually where a hardened foundation breaks down, because customizing an image invalidates the provenance and the SBOM, and with them the assurances you paid for. Not with Docker.

Hardened system packages give everything you add the same built-from-source treatment, ensure your customizations run through the same hardened pipeline, and keep the guarantees intact, with the SLA still behind them. With Docker, the entire foundation stays within a single ecosystem.

One thing stays inevitable no matter how well you do all of this. The software you depend on will eventually go unsupported upstream, and without coverage, the security patches stop, and the compliance answers get harder every quarter. Extended Lifecycle Support closes that gap with commercially backed patches for up to five years past end of life, so the move to whatever comes next happens on your timeline and your terms, instead of upstream’s. That is what a solid foundation looks like, and it has never mattered more, because your newest employee, the agent, is stress-testing what everyone before it built.

The new layer

Agents build on this foundation the same way every human before them has, and the trust it carries passes into what they build. But there’s a new reality now. Agents have created a new layer on top, and it matters almost as much as the foundation itself. They pull packages from the foundation and wire tools together, running what they build as soon as it exists. They’re also non-deterministic and ephemeral. The same task can go differently every run, and the agent session that did the work no longer exists by the time anyone comes back with questions.

Every control in the standard stack was built for a human worker, one with a permanent identity and a predictable pace, whose work can be reviewed before it ships. Agents have none of those traits. The market’s first response was to ask for human permission before every agent action, and when the prompts got too cumbersome, teams moved to isolating agents. That created its own gap because the endpoint tools meant to watch the work sit on the host, and the more you isolate the agent, the less those tools see. There has never been a control surface built for a workflow like this, and retrofitting the old parts leaves teams stuck between prompt fatigue and blind spots.

So Docker built the missing layer, one that adds to your defense in depth without replacing anything you already run. At Docker, every agent session runs in its own disposable, MicroVM-based Docker Sandbox. The sandbox walls the agent off from the host at the operating-system level. Credentials get proxied in for the task at hand and never stored inside, and you decide what gets piped in and out of the box. Our own security team has blocked coding agents on the host outright and runs them in sandboxes with full autonomy, several at a time. An infostealer that lands in one of those boxes finds nothing to grab. Call it YOLO mode with guardrails.

The tools agents reach for are the next layer, built on the same foundation. Agents interact with the outside world through MCP (Model Context Protocol) servers, connectors that let them call external tools and access data. An agent grabbing connectors off the open internet is the package problem all over again. So Docker ships hardened MCP servers through the same catalog as the hardened images, built and signed the same way. The MCP Catalog and Toolkit give your teams one trusted place to find and run them. Every tool call routes through the MCP Gateway, where it is authenticated, authorized, and logged before reaching the external system. That turns enforcement from advisory to strict. 

Docker Scout enforces the policy at build time, so the secure path remains the default without anyone having to police it by hand. And where the box sits stops mattering, whether it’s a laptop or the cloud, because the boundary travels with the work, as Docker containers always have.

The winning playbook already exists

Docker wrote this playbook the first time. In the 2010s, software pulled in parts its builders didn’t control, and shipping outpaced review. Slowing down was never on the table, so Docker packaged the application and its dependencies into one portable, isolated unit, and speed and safety started pulling in the same direction. That bet is a large part of how the modern software supply chain took shape, and now we’re making it again for agents. One foundation and one boundary serve people and agents on the same supply chain, under the same policy. Security gets quieter, and development gets faster. There’s no separate AI security program to buy. Docker has been making the case that security is a developer experience problem from the start.

See it live in San Jose

We’re bringing all of it to WeAreDevelopers World Congress in San Jose, September 23 to 25. Docker’s CISO Mark Lechner will take the stage with One boundary for the agentic era, the boundary his own team lives inside, and the Docker Zone will run live demos all three days.

The newest hire starts Monday either way. What will you have ready for them to build on?

  •  

Moving from Minimus to Docker Hardened Images

The hardened-images space gets better when more people are working on the problem, and Minimus has been a valuable part of that work. That changed this week, when they announced they are ending operations. Though we were competitors, we both believed strongly in the importance of reducing vulnerabilities at the foundation of the software supply chain. Their efforts to bring needed awareness to this challenge will be missed, and our thoughts go out to Minimus employees who are impacted by this decision.

While the human side of this story deserves the most attention, there’s also a practical side: if you’re a customer running Minimus images in production, you’re now facing a migration you didn’t plan for. Their notice commits to a 60-day maintenance window, with images receiving upstream updates until the registry goes offline on October 22, 2026. Images already pulled will keep running after that date, but no further updates will ship to them, and any new CVE stays unpatched from that point on.

If you need a hand, Docker is offering free migration assistance to Minimus customers. Write to minimus@docker.com to walk through your specific image list, your compliance requirements, or questions around your migration plans, and a technical migration expert will get back to you. You don’t need a sales call to start migrating to DHI today.

Docker’s free, open source catalog is available to everyone under Apache 2.0, allows production use, and has no user caps. The migration is about as easy as these things get, a drop-in with minimal workflow changes. It’s more of a swap than a rebuild. It’s easy to find your images’ equivalents in the DHI catalog, and for most of your services, the whole change is updating the FROM line. Use the migration guide for the step-by-step process and the checklist to track each image through the swap and verification. The worked examples show full migrations end to end, and Gordon, Docker’s AI assistant, runs the first pass with you.

Whether you decide to migrate to Docker or somewhere else, we recommend you start that process now, while the maintenance window keeps your current images patched. You can browse the full DHI catalog on Docker Hub, make the first swap, and, of course, reach out to us if you need help.

Docker Hardened Images

Docker Hardened Images are minimal, hardened images built from source and continuously maintained by Docker. The catalog covers 4,000+ images, compatible with Alpine and Debian, so your Dockerfiles and CI keep working as they are. Every image ships near-zero CVEs with full, unsuppressed CVE visibility, and each carries a complete SBOM, SLSA Build Level 3 provenance, and cryptographic signatures. Docker manages the full lifecycle of your image, and teams moving from standard public images see up to 95% CVE reduction and up to 90% attack-surface reduction. Paid tiers add SLA-backed remediation, FIPS and STIG variants, customizations, and up to five years of coverage for versions past end of life.

  •  

MinIO End of Life: How to Stay Patched and Audit-Ready with Docker ELS

MinIO reached end of life in February 2026. Docker Extended Lifecycle Support (ELS) keeps end-of-life software like it patched, compliant, and audit-ready for up to five years, covering versions upstream no longer supports all the way up to entire projects.

On February 13, 2026, the MinIO open-source project was archived upstream. A project with more than a billion Docker pulls stopped shipping releases, bug fixes, and security patches overnight. From that day forward, every environment running MinIO is exposed. New CVEs in MinIO and its Go dependency tree now arrive with no upstream patch behind them, and an audit reads that as unsupported software in production.

And MinIO is only the newest instance of a wider problem. Black Duck’s 2026 Open Source Security and Risk Analysis report found that 93% of commercial codebases carry components with no development activity in at least two years. The same pattern runs across the stack. Node 18, Python 3.8, and older Airflow releases still run in production long after upstream support ended, and frameworks like FedRAMP, DORA, and the Cyber Resilience Act treat unpatched end-of-life software as an audit finding. The migration deadline ends up set by the audit calendar instead of the roadmap.

Docker Hardened Images Extended Lifecycle Support exists to hand that schedule back to you. The model is simple. Request an ELS image, and Docker builds and maintains it for up to five years past upstream end of life. The maintained MinIO image is the newest proof of that model.

MinIO lives on as the newest ELS update

The archive lands on the storage layer, where migrations are measured in petabytes. Moving a production object store to a different system is slow, expensive work, and the CVE exposure keeps growing while that work runs.

Teams running MinIO have three options

  1. Move to a commercial replacement and take on new licensing and lock-in.
  2. Carry the patches yourself, which means staffing sustained Go security engineering for a project that no longer ships fixes.
  3. Keep what you run and put a vendor on the hook for it. 

Doing nothing is not a fourth option. 

Docker identified the archive as a live exposure across its customers’ software supply chains and built the answer into the catalog, where MinIO lives on as a maintained, hardened image. Docker tracks new CVEs across MinIO and its full Go dependency graph, transitive dependencies included at no extra cost, then backports the fixes, rebuilds, and ships. Your object store stays supported and your audits stay clean.

Extended Lifecycle Support for your whole fleet

What ELS does for MinIO, it does for any end-of-life component you need to keep. An EOL finding forces a choice between two bad projects. Rush the migration and risk breaking production, or file the exception and watch the list grow every quarter. ELS removes that deadline. Patches and audit evidence keep flowing on the images already in production while the migration happens on the roadmap’s schedule.

The entitlement is built for how end of life actually arrives, on staggered dates across a fleet. Applied to a repository, it covers every available ELS version there. When one migration completes, you re-point it at the next repository, and the coverage moves with the risk.

Coverage is not limited to a fixed list either. Docker watches the end-of-life calendar and builds ahead of it, and anything you don’t see in the catalog, you can request. The span runs from end-of-life versions of supported software all the way up to entire archived projects. Nginx, Node, and Python ELS images are already there.

ELS is a paid add-on to a Docker Hardened Images subscription, and it runs on the same rails as the rest of DHI:

  • Name it, get it. Tell Docker the end-of-life line your production depends on. Docker builds it hardened and maintains it at the line’s newest patch version.
  • Adopt without a migration. ELS-tagged images appear in the standard DHI catalog alongside LTS tags. Same registry, same workflow, a FROM-line change.
  • Stay patched for years. Critical and high-severity CVEs are patched on a 14-day SLA, for up to five years past end of life.
  • Evidence included. Every ELS image holds the same standard as the rest of the catalog. Built from source and signed, with SBOMs, VEX statements, and SLSA Build Level 3 provenance maintained for the life of the image.

Those attestations are the difference between extended support and an extended liability. A legacy app with a giant SBOM and no exploitability data just lights up your scanners. ELS ships the evidence with the image, so auditors see signed proof of what’s patched and what’s not exploitable.

If there’s a version in your fleet you can’t migrate off and can’t leave unpatched, that’s an ELS conversation. Browse the DHI catalog to see what’s already covered, and talk to us about the versions you need to keep alive. 

  •  

Running AI agents in GitHub Actions with Docker Sandboxes

In July 2026, GitHub Agentic Workflows added Docker Sandboxes as a supported agent runtime. It means that in your CI an AI coding agent can have broad control of its environment, including being able to run Docker containers, while the environment itself is isolated in a microVM with a network policy and secrets injection like the current best practices for AI isolation advice. 

Agentic isolation matters because useful coding agents do more than read a repository and suggest a patch. They install tools, run arbitrary shell commands, execute project code, start databases, and occasionally discover surprising new meanings for the word “cleanup.” Those capabilities make the agent useful, and direct access to a CI runner gives every mistake a larger blast radius.

Now, with sbx integrated, the boundary for the Agent is a disposable environment with substantial freedom inside and narrow access to everything outside it.

I put together a small example to see what that looks like in practice. The agent runs on a GitHub-hosted Ubuntu runner, enters a Docker Sandbox (sbx), runs a Java integration test suite with PostgreSQL using Testcontainers, finds an intentionally seeded bug, fixes it, and opens a draft pull request. The Github Agentic Workflows offers the integration out-of-the-box, so the setup requires zero custom configuration for actions.

What are GitHub Agentic Workflows?

GitHub Actions remains the CI system. It schedules the job, provides the Ubuntu runner, manages permissions and secrets, and records the result.

GitHub Agentic Workflows, usually shortened to gh-aw, is an open-source GitHub CLI extension and compiler. You describe an agentic workflow in a Markdown file that combines execution configuration in YAML frontmatter with the agent’s task in the body. Running gh aw compile turns that source into a conventional GitHub Actions workflow with a .lock.yml suffix.

The relationship looks like this:

Markdown workflow
    |
    | gh aw compile
    v
Generated GitHub Actions .lock.yml
    |
    | runs on ubuntu-24.04
    v
Docker Sandbox microVM
    |
    v
Copilot agent and its tools

docker-sbx belongs to gh-aw‘s agent runtime configuration. The runs-on field still selects ubuntu-24.04, and the compiled file is a standard GitHub Actions workflow. It installs the sandbox tooling, authenticates it, checks the runner, starts the agent in the sandbox, and cleans everything up afterward.

That integration landed in gh-aw and shipped in version 0.82.9.

Configuring sbx in GitHub Actions

Here is the configuration from the sample’s sandbox-explorer.md:

---
name: "Docker Sandboxes sample: exploratory test"

on:
  workflow_dispatch:

runs-on: ubuntu-24.04

permissions:
  contents: read
  copilot-requests: write

engine: copilot

network:
  allowed:
    - defaults
    - github
    - containers
    - java

sandbox:
  agent:
    id: awf
    runtime: docker-sbx
    sudo: true

tools:
  edit:
  bash: [":*"]

safe-outputs:
  create-pull-request:
    title-prefix: "[docker-sbx sample] "
    draft: true
    protected-files: blocked
    allowed-files:
      - "src/**"
---

The three lines under sandbox.agent select the Docker Sandbox runtime. Inside it, the agent has the sudo and unrestricted shell access needed to build the application and start its test infrastructure.

Outside the sandbox, the workflow keeps a much smaller surface. Its network block allowlists the destinations this job needs, while the agent’s GitHub token can read repository contents and send requests to Copilot. Pull request creation happens in a separate safe-output job whose patch may contain files only under src/**.

How much autonomy a CI agent should receive depends on the job. For this one, the split is useful: broad shell access inside the sandbox, small network and repository surfaces outside it, and a draft PR that still expects human review.

The isolation boundary is a micro VM

While it’s common to assume that “Docker” implies a single application container, this setup actually uses a microVM as the primary isolation boundary.

With sbx, every sandbox is a dedicated environment with its own kernel, filesystem, and network stack. Most importantly, it runs its own private Docker daemon. This means the agent gets full root privileges inside the VM without ever gaining control over the host’s Docker daemon. The only bridge between them is the explicit shared workspace of the repository.

Having a private daemon is a game-changer for integration testing. In this demo, the app runs Testcontainers exactly as a developer would on their local machine. The resulting structure looks like this:

GitHub Actions runner
└── Docker Sandbox microVM
    ├── GitHub Agentic Workflows agent
    └── Private Docker daemon
        ├── Maven / Java 21 container
        └── PostgreSQL Testcontainers container

To keep the environment clean, the test launcher runs Maven inside a pinned container, passing the sandbox’s Docker socket through so it can talk to the private daemon:

docker run --rm \
  --add-host=host.testcontainers.internal:host-gateway \
  -e TESTCONTAINERS_HOST_OVERRIDE=host.testcontainers.internal \
  -v "$PWD:/workspace" \
  -w /workspace \
  -v /var/run/docker.sock:/var/run/docker.sock \
  maven:3.9.9-eclipse-temurin-21@sha256:3a4ab3276a087bf276f79cae96b1af04f53731bec53fb2e651aca79e4b10211e \
  mvn --batch-mode "$@" test

Testcontainers then uses that socket to spin up the PostgreSQL database. It sounds like a lot of layers—a container running a build that starts another container, all inside a microVM on a CI runner but each layer serves a specific purpose in ensuring the agent remains isolated yet fully capable.

Giving the agent a defect worth finding

The sample is a small Java 21 registration service. Its requirements say that email addresses are case-insensitive. The seeded implementation stores them as provided and relies on PostgreSQL’s case-sensitive unique constraint. An existing Testcontainers integration test catches exact duplicates but says nothing about the latter case.

The Markdown portion of the workflow asks the agent to inspect the requirement and code, run the baseline suite, and add a test for two addresses that differ only in case. If the invariant fails, the agent should make the smallest source correction. Before touching the application, it records uname, Docker version, Docker information, and a tiny Alpine container run, leaving specific evidence in the workflow log about where the work executed.

The task itself is plain Markdown beneath the frontmatter in the yaml file. The important part for us (after some commands for recording the environment for debugging) is:

Act as a bounded exploratory tester for this repository.
... 

Then:
1. Read `REQUIREMENTS.md` and the relevant source and test files.
2. Run `./scripts/test-in-docker.sh` without changing anything.
3. Add a PostgreSQL Testcontainers test that checks registration of two
   addresses that differ only in letter case.
4. Run the focused test and explain the observed behavior.
5. If the implementation violates the documented invariant, make the
   smallest fix under `src/`.
6. Run the complete test suite again.
7. Create one draft pull request containing the regression test and fix.

And the prompt level guardrails to suggest the correct behavior: 

Do not modify dependency manifests, workflow files, scripts, documentation,
or generated files. Do not weaken or delete existing tests. Include the
commands run and their results in the pull request description.

The real run of course followed that path: its baseline passed, then the new case-variation test failed with:

expected: <false> but was: <true>

The agent normalized the email before inserting it, reran the complete suite, and got two passing integration tests.

The log reported Docker client and server version 29.7.1 with the default context. It is the correct Docker version currently in the sbx default sandbox template. This is the sandbox’s private daemon, the one Testcontainers library used to launch PostgreSQL for the integration tests. 

image2 1

The complete workflow passed on GitHub’s hosted ubuntu-24.04 runner. The run took 11 minutes and 16 seconds.

The safe-output job then opened a draft PR containing exactly two files under src/**: the regression test and the one-line normalization fix. Workflow configuration, scripts, dependencies, and documentation were outside its allowed patch surface.

image1 2

The generated draft pull request stayed inside the declared source-only boundary.

Running the workflow yourself

Start by installing the gh-aw:

gh extension install github/gh-aw

The compiled Docker Sandbox runtime needs Docker credentials to authenticate and pull its sandbox template. Add DOCKER_USERNAME and DOCKER_PAT under the sample repository’s Settings > Secrets and variables > Actions, or let the GitHub CLI prompt for both values:

gh secret set DOCKER_USERNAME
gh secret set DOCKER_PAT

The repository’s Copilot entitlement and copilot-requests: write were sufficient for the successful sample. Repositories without that entitlement can use a supported COPILOT_GITHUB_TOKEN secret as documented by gh-aw.

Also enable Allow GitHub Actions to create and approve pull requests in the repository’s Actions settings. Then compile the Markdown source and commit both the source and generated workflow:

gh aw compile sandbox-explorer

git add .github/workflows/sandbox-explorer.md \
  .github/workflows/sandbox-explorer.lock.yml
git commit -m "Compile Docker Sandboxes sample workflow"
git push

The .lock.yml is generated code. Changes belong in the Markdown source, followed by another compile.

Finally, start the workflow and watch it:

gh aw run sandbox-explorer
gh run watch

The sample works on GitHub’s hosted ubuntu-24.04 runner as committed. A self-hosted Linux runner needs an appropriate KVM-capable setup, plus the Docker and system access required by Docker Sandboxes.

Try sbx on your laptop

Support for isolating your agents in CI is fantastic, but the easiest way to understand Docker Sandboxes is to put one around an agent on a local project. Follow the Docker Sandboxes setup for your platform, sign in, move to a repository, and run an installed agent:

sbx login
cd ~/my-project

sbx run <claude|codex|opencode>

Give it a task that needs real tools, such as running tests, building an image, or starting a Testcontainers dependency. sbx is much easier to evaluate and understand when the workload is your actual development loop.

And if your experiment grows into an organization-wide agent rollout, Docker AI Governance is the next thing to explore. It applies organization and team policies for sandbox network, filesystem, and MCP access, and records policy decisions in audit logs. Those records help to identify the source client, including sbx, and the machine hostname, so the same policy and audit model can easily cover your  team’s laptops and your CI runners.

  •  

Docker Verified Publisher Applications Are Now Self-Serve

Curating trusted content for the agentic software era

While AI made it easier for organizations to keep up with the latest innovations, it also made it harder to know what to trust. When software is selected at machine speed, the question is no longer “is this popular?” It’s “do we know who published this?”

Docker Hub has always been where developers go to answer that question. Starting today, software vendors looking to make their trusted content discoverable to developers by becoming Docker Verified Publishers will enjoy a faster application process, with less friction, and plans that fit their specific growth needs.

With the Docker Verified Publisher (DVP) program, Docker Hub turns into a trusted, discoverable, and measurable distribution channel. Organizations accepted into the program earn verified status and prioritized ranking. DVP publishers also gain access to analytics reports that show which versions are getting the most traction and which companies are pulling them, turning open-source reach into a commercial pipeline.

What’s new in Docker Verified Publisher Applications

Applying to become a Docker Verified Publisher (DVP) is now self-serve. You can now apply directly in Docker Hub, our team reviews your application, and if you’re approved, you become part of our trusted ecosystem on Docker Hub.

This marks a significant improvement in how we onboard and evaluate publishers. Previously, companies interested in becoming a Docker Verified Publisher needed to contact our sales team to be considered for the program. While the Docker team still evaluates every single application manually, this change makes it significantly easier to apply to the program.

Within our new self-serve process, you can choose between two different plans that suit your needs as you grow. 

Turn pulls into reach: One badge for all your content

The verified publisher program helps you grow your impact on Docker’s ecosystem. With the badge and priority search ranking due to trusted status, developers evaluating options on Hub see your verified content first.

DVP analytics also help close the gaps you have in understanding your users and product offerings. Summary and trends reports show which repositories are gaining ground and where adoption is shifting across versions and releases. Domain-level reports on Growth turn anonymous traffic into named domains, so the teams already running your software show up in your sales and partner pipeline.

In addition, DVP is designed to mean the same thing across every content type on Hub. Docker Hub isn’t just images anymore. Developers come to Hub for MCP servers, models, sandboxes, agents, and more; everything that’s needed for an agentic stack. DVP offers one review, one badge, one answer to “who published this” no matter what you’re publishing. Whatever you distribute next, your verification comes with you.

What DVP means for developers

The Verified Publisher badge means Docker has manually reviewed the publisher behind that content and confirmed they are who they claim to be. Publishers such as Google, Microsoft, AWS, Datadog, Grafana Labs, n8n, and many more rely on DVP to build trust, increase visibility, and grow adoption of their content on Docker Hub.

Pulling your images from Docker Verified Publishers is a good step towards improving your security posture, but also needs to be paired with other good consumption practices. This means, for example, reviewing the specific artifact you pull, pinning to digests rather than mutable tags, verifying provenance and any signatures at the image level, and checking for CVEs. 

And while publisher verification is an important link in the trust chain, we continue building towards stronger, more secure publishing flows across Docker Hub. Stay tuned for more in this space.

Get started

You can apply to the Docker Verified Publisher Program from the Explore page in Docker Hub. Verification is done by the Docker team, and you’ll get a checkout link as soon as you’re approved.

  •  

17,600 Actions: Agent Security Is a Systems Problem

Everyone has been talking about the OpenAI/Hugging Face incident, and I was initially skeptical that Docker had much to add. After several weeks of customer conversations, I think we do. The useful lesson is not that an AI agent escaped a sandbox. It is what 17,600 actions expose about security systems designed for human tempo.

Hugging Face reconstructed approximately 17,600 attacker actions across a four-and-a-half-day campaign in July, including roughly two and a half days inside its infrastructure.

Put thirty seconds of human review against each action and you get 147 hours of work. Hugging Face grouped those actions into roughly 6,280 clusters; review only those and you still get more than 52 hours. Either way, manual approval and ordinary alert triage were never plausible primary controls for this workload.

This was not a fundamentally new category of attack. If you have any background in incident response, the sequence is familiar: exploit a reachable service, obtain code execution, collect credentials, escalate privilege, and cross the next trust boundary.

What was different was the rate, persistence, and ability to carry authority and state across short-lived environments. The agent could fail, rebuild its tooling, recover its communications channels, and continue from another environment without starting the campaign over. The useful mental model is a capable attacker fused with a fuzzer: it can reason about each result and keep probing without fatigue. That is the AI part of the story.

This is not a commentary on model alignment, and I will offer no opinion here on which model or agent framework you or your organization should use. The security question is familiar: what authority should be given to a workload that can execute code, hold credentials, reach the network, and continue testing paths long after a human would have stopped?

What happened

OpenAI was running internal cyber-capability evaluations against a benchmark called ExploitGym, which tasks agents with finding and exploiting vulnerabilities. The evaluation used reduced refusal behavior and none of the production-grade classifiers a deployed model would normally have.

The agents did not have direct internet access. They were allowed to install packages through an internally hosted package-registry cache. The agents found a previously unknown vulnerability in that permitted service and used it to reach the open internet.

The permitted service was the escape path. That’s the important point in the incident.

From there, the agents escalated privilege inside OpenAI’s research environment, reached a user-hosted CyberGym-style code-execution harness running on an external sandbox, and obtained root inside that harness. The sandbox platform itself was not compromised. Having apparently worked out that Hugging Face might host the benchmark’s reference answers, the agents went after Hugging Face directly, assessing that cheating was the fastest path to a passing score.

Hugging Face’s reconstruction documents local-file disclosure, code execution through a template engine, cloud metadata access, Kubernetes privilege escalation, exposure of a production secret containing 136 keys, access to the corporate network through a stolen credential, and limited write access to source control. It found no evidence that public models, datasets, packages, or released software were modified, and the customer content accessed was limited to a handful of benchmark-related datasets. That is a good outcome relative to what was possible. It is not evidence that the attack path was harmless.

The individual weaknesses were familiar: an unsafe processing path, accessible cloud metadata, broad privileges, long-lived credentials, and a shared administrative identity. A capable human attacker could have chained them, too. The difference was the volume of paths tested, the speed with which a failed path was replaced, and the amount of evidence defenders had to interpret while the campaign continued.

If you want the full blow-by-blow, Hugging Face published a detailed technical writeup and interactive replay. Both are worth reading.

The controls this requires

That arithmetic, not the novelty of the attack, is what changes the operating model. A human cannot sit in the control loop for thousands of actions at machine speed.

There are also two related threat models to account for.

In the first, the agent discovers and exploits a path on its own, as the agents did in this incident. In the second, an external party redirects the agent through a poisoned tool, compromised dependency, malicious document, prompt injection, or attacker-controlled data. In that case, the agent becomes the attack vector.

The same authority is dangerous in either case. The controls therefore have to govern what enters the environment, which capabilities the agent receives, and what happens when its behavior leaves the expected boundary.

What can go wrongControl required
Untrusted content attempts to change the agent’s instructionsTreat external content as data, never authority; preserve provenance and trust boundaries; do not allow retrieved content to modify policy or grant capabilities
A tool, package, image, or MCP server is poisonedStrict allowlisting; pin versions and digests; verify signatures and provenance; require SBOMs; scan and review updates before promotion
External data becomes executable inputUse structured and typed tool interfaces; validate and escape arguments; prohibit unsafe shell interpolation and other paths from data to command execution; run commands inside a sandbox and constrain which commands can be invoked
A permitted service becomes an escape pathExpose the narrowest operation-specific interface; scope access by task and destination; monitor permitted paths for behavior that diverges from their purpose
The workload reaches ambient credentials or control-plane servicesRemove ambient credentials; use short-lived, task-scoped workload identity; block metadata and control-plane endpoints the task does not require
The workload exceeds its intended execution boundaryUse a hardened isolation boundary; reject privileged workloads, host filesystem mounts, raw device access, and unrestricted host sockets unless explicitly required
State or authority persists beyond replacement of the workerGovern durable write paths; separate read from write authority; correlate identity and activity across runs
The agent, or an attacker acting through it, crosses into another systemRequire explicit authorization for external and destructive actions; scope credentials to the target operation; make write paths observable and revocable
Thousands of individually plausible actions hide behavior that has diverged from the taskEvaluate sequences and bursts against policy and expected behavior; correlate activity across tools and systems; contain at machine speed

The point is not to put human approval in front of every action. That defeats much of the value of using agents and eventually produces approval fatigue. Routine and reversible actions should have very little friction. More sensitive, external, destructive, or unusual actions should receive stronger scrutiny.

Review and sanitization still matter, particularly for packages, tools, and paths where external data could become a command. But sanitization alone is not a reliable answer to prompt injection. Natural-language content is too ambiguous for us to assume we can always identify and remove the malicious part. The stronger boundary is architectural: untrusted content must not be able to grant itself authority, change policy, or create capabilities the agent did not already have.

Done well, governance is not what limits agent autonomy. It is what makes it possible to safely give agents more of it.

Where Docker fits today, and where we do not

We are proud to be founding authors of the Agent Baseline. We worked with other industry experts to distill the problem into six outcomes: Discover, Constrain, Authorize, Observe, Validate, and Respond.

If Docker Sandboxes sit in one specific bucket, it’s “Constrain,” but really, we believe they’re foundational, and where you would instrument or implement all six. They give each agent a dedicated microVM and enforceable boundaries around local compute, filesystem access, and network reach, as well as providing the base (and thus ground truth) layer to observe. That is a real and useful layer.

Docker AI Governance addresses parts of Authorize and Observe by giving organizations a centralized way to define and enforce controls around agent environments, including network and filesystem policies and access to MCP servers and tools.

Together, Sandboxes and AI Governance provide a meaningful part of the answer today: a hardened execution environment and centralized policy enforcement around it. They do not repair a vulnerable service the agent is authorized to contact, narrow a credential issued by another system, or replace the customer’s own security architecture. No vendor, Docker included, can claim its technology would have made this particular incident a non-event.

But a deterministic enforcement boundary is still necessary. It gives an organization one place to apply least capability and least privilege, and one place to observe what the agent was actually allowed to do. If an agent is using a package registry as an egress proxy rather than a package registry, that’s the kind of divergence the telemetry needs to help surface, especially when viewed across a sequence of requests rather than one request at a time.

The broader problem remains difficult. The useful unit of observation is not always one tool call. It may be a burst of activity, a target, a protocol, a credential, or a pattern visible only across systems. A package request can be normal. Repeatedly probing the service behind it, discovering credentials, and using them to reach another system should change the assessment.

That’s the agent-security challenge beyond basic containment. We need to constrain authority, but also observe activity at the right granularity, recognize when it deserves more scrutiny, and respond at the same tempo as the agent. For all of us, Docker included, there is still substantial work ahead across observation, validation, and response.

The operational tradeoff

Security, capability, and autonomy all matter, and they will always be in tension. Said differently, none of this is free.

Short-lived credentials expire during long-running tasks. Narrow egress policies break legitimate package installation. Admission controls reject tools developers assumed they could run. Cross-system detection costs money and produces false positives. A write approval inserted at the wrong point can eliminate most of the productivity the agent was supposed to provide.

Teams will be tempted to loosen each control until the agent works again. That is understandable. The failure mode created by a strict policy is immediate and visible; the failure mode created by excessive authority remains invisible until an incident.

The answer is not to remove the controls or ask a human to approve everything. It is to make friction proportional to consequence, test the failure modes, measure the operational cost, and weigh it against the risk and potential blast radius.

How I work

I use agents every day, and I assume that a sufficiently capable agent will eventually try something I did not anticipate (perhaps on a daily basis…).

For the most part, I do not run one general-purpose agent with access to everything. I use task-focused agents, each packaged as a separate kit, built on free Docker Hardened Images and run in Docker Sandboxes.

Each kit starts with a specific job, then receives only the software, network access, files, credentials, and external capabilities required for that job.

In most cases, the agent has very few restrictions inside its sandbox. That is intentional. What matters is that god mode inside the sandbox does not become god mode over my laptop, my credentials, or every service I can reach.

I do a lot of desk research. Those agents can access the open internet. They’re not useful if they can’t. But their image has no compilers, package manager, general-purpose network debugging tools, or development toolchain, and it runs with deliberately limited system permissions. They can retrieve and analyze public information, but have very little machinery with which to turn something they encounter into an exploit or act on another system. They have no reason to hold my source code or production credentials.

My production coding agent has a much richer environment. It runs pi, can use multiple models, compile code, run tests, and use the tools required for real engineering work. Its network access is restricted to an explicit allow list of services I use, including Docker, GitHub, Snowflake, and Cloudflare. It does not receive arbitrary internet access or arbitrary tools simply because a coding task occasionally needs the network.

My home kit can interact with an Arduino, but it does not receive direct access to the host or the device. A host-side MCP server brokers the allowed operations. The agent can request a defined Arduino capability through that interface; it cannot turn that permission into general access to every device connected to the machine.

My development kit is where I experiment. It runs with balanced network access, but no ambient host secrets and no unrestricted access to host files. When it needs Google Workspace, Snowflake, or another host service, host-side daemons broker those calls. The agent sees the capability I have chosen to expose, not the underlying credential or the rest of the service. Those brokers can enforce which operations are allowed and which are blocked.

These are deliberately different environments. The research agent would be poor at production coding. The coding agent cannot reach every site the research agent can. The home agent cannot turn an Arduino operation into arbitrary host access. The development agent can query a service without possessing the credential that authorizes the query.

That constraint is the feature.

Conclusion: Security at agent speed

The OpenAI/Hugging Face incident was not the failure of a single boundary. It was a chain of reasonable-seeming permissions and familiar weaknesses that became something very different when an agent could test thousands of paths, preserve state across runs, and carry authority from one system into the next.

We will not anticipate every vulnerability an agent might find or every way it might combine the access we give it. The architecture cannot depend on perfect agent behavior, perfect software, or a human noticing every dangerous action in time.

So, the starting point is still least capability and least privilege: give an agent the narrowest interface, credentials, tools, and network access its task requires. Put those controls at a deterministic enforcement boundary. Make the resulting activity observable, not only as isolated requests, but as sequences and patterns across systems. When the behavior leaves the expected envelope, containment has to happen at agent speed.

Docker Sandboxes and Docker AI Governance provide important parts of that architecture today: hardened execution boundaries and centrally enforced policy around them. They do not secure every service an agent is permitted to contact, and they do not eliminate the need for an organization to decide what authority each agent should have. The broader work across Discover, Constrain, Authorize, Observe, Validate, and Respond is why we helped create the Agent Baseline in the first place.

The goal is not to build an agent that never tries the wrong thing. The goal is to build a system where trying the wrong thing does not give it the keys to everything else.

  •  

Coding Agent Horror Stories: The Command You Already Approved

This is Part 5 of our AI Coding Agent Horror Stories series, a look at real security incidents involving AI coding agents, and how Docker Sandboxes contain agent execution at the boundary rather than at the command line.

In Part 1, we walked through six categories of AI coding agent failures and why they keep happening. The agent runs as you, with your filesystem permissions and your credentials, and nothing sits between the model’s decision and the shell’s execution. Part 2 went deep on the rm -rf ~/ incident. Part 3 moved the same problem into a production cloud environment. Part 4 followed the credentials themselves through a supply chain attack. 

This one is about the safety net. Most teams running a coding agent today have some version of a list of commands the agent may run without asking, and the assumption underneath it is that anything dangerous will show up as a prompt you can refuse. In January, researchers at Pillar Security showed that the assumption doesn’t hold.

Today’s Horror Story: The Approval That Ran Something Else

On January 14, 2026, researchers at Pillar Security disclosed CVE-2026-22708, a flaw in Cursor. When the agent ran in Auto-Run Mode with an allowlist enabled, a handful of shell built-ins executed without appearing in that allowlist and without asking for approval. Anything that could get text in front of the agent, a README or a dependency or an issue comment, could use them to change environment variables silently. A command the developer then approved, something as ordinary as git branch, would run the attacker’s code instead. Cursor rated it High and patched it in version 2.3.

No memory corruption was involved here and no permission was escalated. The developer was shown an accurate prompt, approved a command that was genuinely harmless, and got arbitrary code execution anyway, because the meaning of that command had been changed a minute earlier by something they were never shown.

In this issue, you’ll learn:

  • How shell built-in slipped past an allowlist that was working exactly as designed
  • Why the attack still worked when the allowlist was completely empty
  • What Docker Sandboxes contain here, and the two things they do not
  • How kits, organisation policy and audit logs cover what a per-laptop allowlist misses
image1 1

Caption: Comic illustrating how an injected instruction changes environment settings without triggering an approval prompt, so that a command the developer legitimately approves runs the attacker’s payload instead.

The Problem

Typically, programs read settings from their environment when they start up. Git checks one called PAGER to work out which program displays its output, and Python checks one called PYTHONWARNINGS. Nobody thinks about these, which is rather the point. The commands that change them are shell built-ins, and Pillar’s research names export, typeset and declare specifically, a detail reported independently at disclosure. Built-ins are not programs sitting on disk, and the checker was looking for programs on disk, so they went through without ever being surfaced.

Which means the whole attack is two lines.

# This one runs silently. You are never asked.
export PAGER="open -a Calculator"

# This one you are asked about, and you say yes, because obviously.
git branch

Git looked up PAGER to work out how to show the branch list, found the attacker’s command sitting in it, and ran that instead. Pillar notes this worked even with a completely empty allowlist, which is the most restrictive setting on offer.

An allowlist checks whether the command in front of it is on the list, which is fine for cutting down interruptions, and nobody wants to approve ls for the ninetieth time in a morning. But the name of a command does not tell you what that command will do. The check reads the name, waves it through, and the setting that decides what actually happens was changed a minute earlier by something the check was never shown.

Cursor’s documentation now describes the allowlist as best-effort and warns that bypasses are possible. Pillar went further and argued that agents should be handed full command execution inside an isolated environment, and that the industry ought to deprecate allowlists altogether.

The Scale of the Problem

None of the underlying trick is new. Pillar’s write-up points back to Elttam’s 2020 research on environment variables, which showed how these settings could be turned into code execution.

It sat there for six years without troubling anybody very much. Pulling it off meant already being on someone’s machine, setting several things in the right order, running each step yourself, and anyone with that much access had faster ways to cause damage.

Then coding agents arrived and removed every one of those obstacles at once. They act on instructions found in files they were told to read, they run several steps in a row without stopping to check, and they run as you. A technique that used to need somebody sitting at your keyboard now arrives in a repository you cloned this morning.

It is the same shape as the s1ngularity attack from Part 4. There, a poisoned package borrowed an agent that was already logged in. Here, poisoned text borrows a command that was already approved. Neither one breaks anything. Both of them use permission that was handed over deliberately, for something nobody intended.

Technical Breakdown: How the Attack Works

image2

Caption: Diagram showing how an injected instruction changes the shell environment out of sight, so that an allowlisted command carries the attacker’s payload when the developer approves it.

The attack has two halves, and the split between them is the entire trick.

1. The half you never see

The agent reads a file it was asked to read, and that file contains an instruction meant for the agent rather than for you. Built-ins then quietly set the environment. Nothing appears on your screen.

Pillar demonstrated a longer version of this, chaining several settings together, PYTHONWARNINGS, BROWSER, and PERL5OPT among them, so that every later python3 command on that machine would run attacker code. The details differ, but the principle is the same: change what a program reads at startup, and you change what it does.

2. The half you approve

Then you run git branch or python3 script.py, or the agent runs it for you under your allowlist. These are the commands people add to allowlists to stop the constant interrupting, so the better tuned your list is, the more reliably the trigger fires. The payload runs with your permissions.

Some variants skip the approval altogether. One writes extra lines into ~/.zshrc, so the code runs again every time you open a terminal. You could finish the project, delete the repository, and still be running it next month.

The Impact

The full chain in Pillar’s research ends with the victim’s SSH private keys leaving the machine.

Work backwards and the whole thing started with a piece of text in a file, read by an agent doing exactly what it was asked to do. No memory bug. No privilege escalation. Nothing in any log that looks the slightest bit out of place.

Pillar reported it in August 2025 and the fix shipped that January. Cursor engaged with the report and made a real change, so anything the parser cannot classify now requires approval, which closes the paths that were demonstrated. Five months is a fair measure of how awkward this is to fix at the layer where it was found rather than a complaint about the vendor.

The wider problem has not gone anywhere, because it was never really about shell built-ins. It is about a check that studies the command while somebody rearranges the furniture around it.

image4

Caption: Diagram showing the same payload running inside the microVM, and what it can and cannot reach from there.

How Docker Sandboxes Contain This at the Execution Layer

Docker Sandboxes run AI coding agents in isolated microVMs, each with its own kernel, filesystem, and deny-by-default network, so a compromised dependency an agent pulls cannot reach the host, its credentials, or other workloads. Inside that box the agent can run anything, including with sudo, which is exactly what Pillar recommends. There is no allowlist to slip past. We made the longer argument for why a shared kernel is the wrong shape for this in The Untrusted Autonomous Workload.

So run the same attack again, this time in a sandbox, and watch where it gets to.

The injection still lands. The environment gets changed, git branch still triggers it, and the payload runs. Nothing about a sandbox stops that. Then the payload goes looking for your SSH key and does not find one. Your home directory sits on the other side of the boundary, so there is no ~/.ssh/id_rsa inside the box to copy.

It can still use the key. Sandboxes forwards an SSH agent socket into the box so that ordinary work like git push keeps working, which means code inside can ask that agent to authenticate on its behalf. It cannot take the key anywhere, but it can borrow it for as long as the sandbox runs. Your network policy is what limits that, since SSH needs a rule naming the exact destination address and port before it connects to anything.

The ~/.zshrc trick fails outright, because that file lives on your host and a poisoned copy written inside the box disappears along with the box.

Getting data out is harder than people expect. HTTP and HTTPS leave only through a proxy on your host that checks every request against your rules, anything else over TCP needs a rule naming the address and port, and UDP and ICMP are blocked outright.

Two caveats, both stated plainly in Docker’s security documentation. The first is your workspace, which is live on your host by default, so Git hooks and Makefile targets are still within reach and a poisoned hook will not turn up in git diff. Running with --clone hands the agent its own copy.

The second is the shared agent skills store. Supported agents mount the same host-side store read-write unless you opt out at creation time, which is what lets an agent refine a skill and keep it. Every sandbox sharing that store sits inside one trust boundary, so a skill modified inside one becomes an input to the next that loads it. The store is sandbox state though, and a modified skill does not by itself execute on your host, so the risk runs sandbox to sandbox rather than sandbox to host.

Isolation has its own seams. In July, Pillar published a series of sandbox escapes across four coding agents, and the mechanism was never a broken sandbox but a file written inside one that a tool outside later trusted. Both caveats above are that shape.

None of this stops the injection. It changes what the injection can get to, which is the only part of this problem with a dependable answer.

Codify the Boundary with Kits

image3

Caption: Diagram showing how a kit declares an agent’s tools, files and network rules, while real credentials stay on the host and are injected by the forward proxy on the way out.

The allowlist failed here partly because it is a list, edited on each laptop, that an injection can reach around. Kits are Docker’s answer to the editing-on-each-laptop half of that.

A kit is a declarative YAML artifact that extends a sandbox agent with credentials, network policies, environment variables, startup commands and files. Rather than every developer maintaining a personal allowlist, you write the boundary once, deny-by-default network plus only the destinations a task genuinely needs, and hand the same kit to everybody. It gets reviewed, versioned and diffed like any other file in the repository. The kit spec reference covers the fields, and docker/sbx-kits-contrib has working examples.

This lands directly on the SSH question above. A forwarded SSH agent is a live credential limited only by network policy, so leaving that policy to whoever remembers to run sbx policy deny is the same per-laptop weak point this whole post has been complaining about. A kit can bake the network rule in, so untrusted work has no SSH egress unless the destination was declared up front.

What This Looks Like in Practice

The vulnerability is in the editor, so what you want is the setup that puts the editor’s terminal inside the box. Cursor is built on VS Code and connects the same way, over Remote – SSH, with the editor staying on your machine while files, terminals and extensions run in the sandbox. You will need Docker Sandboxes 0.37.0 or later, SSH access configured, and Cursor’s Remote – SSH support installed. The Cursor integration guide has the full walkthrough.

# One-time setup: configure your SSH client for sandboxes.
sbx setup ssh
# Check the sandbox is reachable, then open the Command Palette,
# run Remote-SSH: Connect to Host, and enter &lt;name&gt;.sbx
ssh demo.sbx
# See what this sandbox is currently allowed to reach.
sbx policy ls
# Shut egress down and open only what the task needs.
sbx policy deny network "**"
sbx policy allow network "github.com,registry.npmjs.org"

Those last two commands come with a catch. If your organisation has governance switched on, the org policy replaces local policy and sbx policy allow and sbx policy deny will have no effect on your machine. You can spot it in the output of sbx policy ls, which begins with a Governance: Managed by <org> line when that is the case. Depending on how admins scope things, some rule types may be delegated back to local control, but a local allow will never beat an organisation-level deny.

Same editor, same agent, same allowlist, same payload. All that changed is which machine the terminal is on.

What happensOn your laptopInside a sandbox
The payload runsYesYes
Where it runsYour machine, as youA microVM with its own kernel
Your SSH key fileCan be read and copiedNot there
SSH authenticationAvailable, key includedAvailable, key stays outside
The ~/.zshrc trickPersists indefinitelyGone with the sandbox
Sending data outOpen by defaultOnly where policy allows
Who sets the rulesEach developerThe organisation
Evidence afterwardsNoneA logged policy decision

Making This Hold Across a Team

A kit gets the boundary out of one developer’s head and into a file the team shares, but a file can still be ignored or edited on the machine that matters. Docker AI Governance moves the settings up one more level. Network and filesystem rules are defined once by your admins and reach developers through the login they already use, so there is nothing to configure per machine and nobody quietly reopening what security closed. A shared kit is the boundary as a suggestion. Governance is the boundary as a ceiling.

The part that matters most for this story is the record it keeps. What made CVE-2026-22708 work was that the first half was invisible, with no prompt and nothing written down anywhere you would think to look. Under governance every policy decision produces an event carrying the user, the timestamp and the rule that fired, and those events stream into whatever SIEM your security team already uses.

So an attack that succeeds inside the sandbox and then reaches for somewhere it should not leaves a trail behind it. That is a good deal better than a check that finds nothing wrong and mentions it to nobody.

Best Practices

1. Treat export like any other command. Anything that changes environment settings can change what your next command does, even when that next command is on your allowlist.

2. Do not mistake an allowlist for a boundary. It reduces interruptions. The vendor documentation now says outright that it is best-effort and not a security guarantee.

3. Isolate before the first command, not after something looks wrong. Untrusted means anything you did not write and have not read, which covers most of a dependency tree.

4. Use --clone for code you have not vetted, and opt out of the shared skills store. Otherwise Git hooks and build scripts stay live on your host, a poisoned hook will not appear in git diff, and a skill modified inside one sandbox is waiting for the next sandbox that loads it.

5. Remember a forwarded SSH agent is a live credential. The key file staying on your machine is not the same as the key being unusable, so restrict egress for untrusted work.

6. Read your own policy. Run sbx policy ls. Deny-by-default with a long allow list is closer to allow-by-default than it looks.

Take Action

  • Install Docker Sandboxes. Visit the Docker Sandboxes documentation to install sbx and run your first agent inside a microVM.
  • Connect your editor. The Remote – SSH integration puts your terminals inside the boundary while the editor stays where it is, so your workflow does not really change.
  • Codify the boundary with a kit. Define the network and credential rules your team needs once, and hand the same artifact to everybody instead of a personal allowlist.
  • Read the security model. The documentation is straight about what is isolated and what is not, including the workspace and shared skills store behaviour that --clone and the opt-out flag change.
  • Turn on audit logging. Docker AI Governance streams policy decisions into your SIEM, which turns a silent bypass into something somebody can actually investigate.

Conclusion

The uncomfortable thing about CVE-2026-22708 is that nobody in the story did anything wrong.

Cursor built an allowlist, which is what everyone asked for. The developer approved git branch, which any of us would have approved. The check inspected the command and found it acceptable, which is exactly its job. The attack worked anyway.

Getting an agent to correctly judge every instruction it reads is a problem that gets harder as agents get more capable, and it has no clean ending. Limiting what an agent can reach is a problem we solved a long time ago. The more useful move is to stop needing the first one, and to write down what the agent may reach as an artifact you can review, rather than a list each laptop keeps for itself.

Coming up in our series: Issue 6 looks at the ClawHub infostealer campaign, where malicious skills reached developer machines through a marketplace ranking exploit, and at what sandboxed skill execution, and that shared skills store, change about a registry you cannot personally audit.

Learn more

  •  

Make zero CVEs your new default

Supply-chain attacks have stopped being isolated incidents somewhere in the past year. The compromises now reach the tools the industry trusts to defend itself, with Trivy and KICS among this year’s targets. Mark Lechner, Docker’s Chief Information Security Officer, called the latest wave ‘a permanent shift in the threat landscape’, and the months since have borne that out. The volume is growing at the same time. Over a quarter of production code is now AI-authored, and agents pull in dependencies at machine speed. If you run a platform team or a security program, this is the math you are already living with. More code and more images arrive every week, almost none of it written by your own engineers, and all of it has become your responsibility the moment it ships.

None of this is news to us. Securing the software supply chain is the problem we’re here to solve. The latest round of updates widens the trusted foundation Docker is building under your supply chain, and tightens how it’s enforced. More of the software inside your images is now built and patched by Docker itself, and security coverage continues after software reaches end of life. Images can be tailored to your environment without losing their guarantees, and policy enforcement now reaches every developer machine.

A trusted foundation for the whole supply chain

Screenshot 2026 08 05 at 14 49 37 Hardened Images catalog Docker Hub

It all starts from one principle, and Docker Hardened Images was built on it. Security that doesn’t get adopted doesn’t secure anything. The entire catalog is free for every developer, because a secure baseline shouldn’t be a premium feature. Every image is compatible with Alpine and Debian, the distributions your teams already run, and Docker builds every one of them itself, from source. Adoption is a FROM-line change, not a migration project. And every image is independently verifiable, with signed SBOMs (software bills of materials) and SLSA Build Level 3 provenance, so your auditors work from evidence instead of vendor claims.

A year in, the numbers make the case. The catalog has grown past 4,000 hardened images, plus MCP servers, Helm charts, and ELS images. It draws more than 3.5 million pulls a week, with over a million builds running regularly to keep all of it patched, and open source projects like n8n run production on DHI. The catalog grows the way it always has, driven by what customers request. But the goal was never just a catalog. The goal is one trusted foundation under your whole software supply chain, where the images you run, the packages inside them, the charts that deploy them, and the tools your agents call all carry the same provenance. Security becomes the default from day one, and it holds, without asking your teams to change how they work.

Built from source, down to every package

The hardening keeps reaching deeper into the stack. Docker Hardened System Packages take hardening below the image, to the packages inside it, across both Alpine and Debian, with every package built from upstream source, patched, and maintained by Docker in the same SLSA Build Level 3 pipeline that builds the images themselves. And the repository behind them is open to more than the catalog. DHI Enterprise customers can point apt or apk directly at Docker’s hardened package repository and bring the same packages into images they build themselves, extending the hardened supply chain beyond the images Docker ships to every image your organization builds.

The coverage keeps widening. What began with Alpine now spans Debian, with Python, the catalog’s most pulled image, among the first to ship fully hardened. The work compounds every week, and the Debian and Alpine package lists are public, so you can watch the catalog harden in real time.

If you’ve spent time chasing base-image CVEs, you know why this matters. System packages are notorious for slow fixes; a patch can sit waiting on the distribution’s next release for months or years. Docker doesn’t wait. We patch at the package level, ahead of upstream when it counts, and the fix lands in every image that uses that package, in one build wave instead of image by image. Entire businesses have been built on delivering community-distribution security updates faster than the community. With DHI, that speed is included.

The guarantees hold up under inspection, too. Packages you add through DHI customization, tailoring an image to your workloads, come from that same hardened repository, not an unverified public mirror, so they are hardened system packages in their own right and the SLA that covers the base image extends through everything you add. And because one vendor stands behind the image, the packages inside it, the CVE investigation, and the patch, your auditors get a single chain of signed provenance instead of a stack of vendor assurances.

Your distribution, meanwhile, stays your distribution. Building a hardened package ecosystem from source is a serious engineering commitment, and Docker made it twice, for Alpine and for Debian, so keeping your house standard never costs you your security posture.

Patch past end of life

Production software has a habit of outliving its maintainers. Migrations wait on budgets, dependencies, and test cycles, and CVEs don’t wait with them. That’s the problem DHI Extended Lifecycle Support (ELS) exists for. It keeps end-of-life software patched, with SBOMs and provenance maintained, for up to five more years.

ELS isn’t limited to a set catalog, either. Docker watches the end-of-life calendar and builds coverage ahead of it, and anything you don’t see, you can request. MinIO is the newest addition. Upstream archived the project in February 2026, yet in the DHI catalog it lives on, patched and hardened, and your migration runs on your schedule instead of upstream’s.

Customize at scale, manage as code

Nobody runs stock images in production. You add CA certificates, agents, and the packages your applications demand. The trouble is that in most of this market, the first change you make is where the vendor’s guarantees end, and everything after it is yours to carry. DHI customization works the other way around. You define what your images need, and Docker manages the full lifecycle of your customized images, rebuilding them through the same hardened pipeline on every upstream patch. The SBOM, the attestations, and the SLA travel with the customization instead of dying at it.

Customization operates at scale, too. Bulk customizations run through the UI, CLI, and API, with YAML configuration and GitHub Actions support, so you can tailor hundreds of repositories in one pass and let the rebuilds take care of themselves. And if your platform runs on Terraform, customization is code as well. The DHI Terraform provider mirrors and customizes hardened images with the same pull requests and reviews as the rest of your infrastructure.

The savings are real infrastructure, not a rounding error. Customers tell us they’ve shut off the CI pipelines that existed only to rebuild images, because Docker rebuilds for them. The blind redeploy cadence goes with those pipelines. You ship an update when a fix actually needs to go out, knowing exactly what changed, instead of rebuilding everything on a schedule and hoping QA catches what moved.

For organizations whose data-residency requirements keep images inside the EU, EU-hosted customizations arrive in September. Your customized images will live in Docker Hub’s EU region with the same SBOMs, attestations, and SLA as everywhere else. Residency stops being the reason your hardening program waits.

Harden beyond base images

The same standard keeps moving up the stack. The catalog now carries fully supported Helm charts, so your Kubernetes deployments start hardened too. And it carries a growing set of hardened MCP servers, because the tools your agents call deserve the same scrutiny as the images they run on.

Govern it all with Docker Scout policy

Scanning tells you what’s wrong. Policy is how you keep it from shipping. And enforcement is where most supply-chain programs quietly fail, because hardened artifacts only protect you when your teams actually use them. Developers move fast and default to what works, and the developer machine is exactly where the current wave of attacks aims.

Docker Scout policy closes that gap. It evaluates flexible, customizable policies from the CLI and inside CI, and it ships with the same policies Docker uses to verify every hardened image in the catalog. The policies are written in Rego, the industry standard, and they’re portable, so the same rules that gate a build in your CI travel with your teams to every developer machine in your organization. Gating at the registry matters, but it stops at the registry; developers can route around it all day. Policy that travels to the machine is how you hold every image you run, and every image your teams build, to the bar Docker holds itself to.

It’s an additive control. It works alongside the scanners you already run, and it’s already in the Docker subscription you have.

The foundation is already in your stack

The supply-chain problem is not going to shrink. More code is coming, agents are becoming contributors, and the patch windows regulators expect keep getting shorter. Point tools won’t carry that weight. A foundation that’s secure by default will, backed by an ecosystem that keeps it that way. That is exactly what Docker’s security portfolio delivers. Hardened content on the distributions you already run, customization that keeps its guarantees, support that outlasts upstream, and policy you control, from one vendor accountable for all of it.

And none of it asks you to adopt something new. It’s all in the Docker you already run. Your builds, tools, and pipelines stay the same. Your CVE count doesn’t.

Browse the DHI catalog and pull your first hardened image today. And if you want the full story, how all of this works together, with your questions answered live, join our live webinar in early September. We’d love to see you there.

  •  

Reproducible ESP32 Firmware Development with Docker and Docker Sandboxes

Firmware development has always been challenging: mismatched toolchains, “it works on my machine” builds, and the tension between maintaining legacy products and shipping new features. In this article we explore how you can use Docker and Docker sandboxes to ease firmware development, especially for ESP32 projects. Nowadays, teams end up supporting multiple hardware revisions, several ESP-IDF releases, and long-term customer deployments, all while iterating on new capabilities like Wi-Fi 6, Matter, or power optimizations.

The official espressif/idf Docker image solves the reproducibility problem. Docker Sandboxes (the sbx CLI) solve a newer one: letting AI coding agents work on your firmware at full speed without giving them the keys to your laptop. This article walks through a practical workflow that combines both: clean builds, parallel environments for new and legacy firmware, and safe unsupervised AI sessions.

Part 1: The Baseline – Building with the Official Image

The espressif/idf image ships a complete, pinned ESP-IDF installation: the framework itself, the Xtensa/RISC-V toolchains, Python environment, CMake, ninja, everything. A build needs one command:

docker run --rm -v $PWD:/project -w /project \
  -u $UID -e HOME=/tmp \
  espressif/idf:release-v5.4 idf.py build

A few details worth understanding rather than cargo-culting:

  • -u $UID -e HOME=/tmp makes the container run as your user, so build artifacts in build/ aren’t owned by root. HOME=/tmp gives the IDF tools a writable home for their caches.
  • Pin your tag. latest tracks the master branch and will break you eventually. vX.Y tags are fixed releases; release-vX.Y tags track the release branch and receive bugfixes. For products in maintenance, exact vX.Y.Z tags are the safest; for active development, release-vX.Y is a good balance.
  • If your mounted project is owned by a different user than the one in the container, Git will complain about “dubious ownership”. The image supports -e IDF_GIT_SAFE_DIR='/project' to whitelist the path (use : to separate multiple paths).
  • Enable the compiler cache with -e IDF_CCACHE_ENABLE=1 and persist it across runs by mounting a volume for it. Full rebuilds of a mid-size project drop from minutes to seconds.

Flashing and monitoring

On Linux, pass the serial device through:

docker run --rm -it \
  --device=/dev/ttyUSB0 \
  --group-add $(getent group dialout | cut -d: -f3) \
  -v $PWD:/project -w /project \
  -u $UID -e HOME=/tmp \
  espressif/idf:release-v5.4 idf.py flash monitor

The --group-add is needed because you’re running as $UID, not root, and the device node belongs to dialout.

On macOS and Windows, Docker Desktop cannot pass USB devices into containers. The clean workaround is a network serial bridge using RFC2217, which esptool supports natively. On the host:

pip install esptool
esp_rfc2217_server -p 4000 /dev/cu.usbserial-1420

Inside the container, point idf.py at the network port:

idf.py --port 'rfc2217://host.docker.internal:4000?ign_set_control' flash monitor

This looks like a hack but it’s actually a feature: once the serial port is a network endpoint, anything can reach it. Containers, CI runners, and (as we’ll see) sandboxed AI agents. Keep this trick in mind; it’s the linchpin of Part 3.

Hide it behind a Makefile

Nobody should type these commands twice. A small Makefile keeps the interface stable even if the plumbing changes:

IDF_IMAGE ?= espressif/idf:release-v5.4
PORT      ?= /dev/ttyUSB0

DOCKER_RUN = docker run --rm -it \
  --device=$(PORT) \
  --group-add $(shell getent group dialout | cut -d: -f3) \
  -v $(PWD):/project -w /project \
  -v idf-ccache:/ccache -e CCACHE_DIR=/ccache -e IDF_CCACHE_ENABLE=1 \
  -u $(shell id -u) -e HOME=/tmp -e IDF_GIT_SAFE_DIR=/project \
  $(IDF_IMAGE)

build:
    $(DOCKER_RUN) idf.py build

flash:
    $(DOCKER_RUN) idf.py flash

monitor:
    $(DOCKER_RUN) idf.py monitor

menuconfig:
    $(DOCKER_RUN) idf.py menuconfig

shell:
    $(DOCKER_RUN) bash

Now make build works identically for every developer and in CI, and switching IDF versions is make build IDF_IMAGE=espressif/idf:release-v5.3.

Part 2: Parallel Environments – New Features and Legacy, Side by Side

This is where the container approach stops being merely convenient and starts changing how you work. Because each container is fully isolated, you can run two different IDF versions against two different boards at the same time, on the same machine.

# Terminal 1 - new feature branch, IDF 5.4, experimental board
docker run --rm -it --device=/dev/esp32-experimental \
  -v $PWD/new-feature:/project -w /project \
  -u $UID -e HOME=/tmp \
  espressif/idf:release-v5.4

# Terminal 2 - legacy firmware, IDF 5.3, production board
docker run --rm -it --device=/dev/esp32-production \
  -v $PWD/legacy:/project -w /project \
  -u $UID -e HOME=/tmp \
  espressif/idf:release-v5.3

Typical uses: flashing experimental code on one board while a long-running soak test or customer demo stays untouched on the other; A/B-comparing power consumption between firmware versions; reproducing a field bug on the exact legacy toolchain while the fix is developed on the current one.

Stable device names with udev

/dev/ttyUSB0 and /dev/ttyUSB1 swap depending on plug order, which will eventually make you flash the wrong board. On Linux, pin them with udev rules keyed on the adapter’s serial number:

# find the serial numbers
udevadm info -a /dev/ttyUSB0 | grep '{serial}'
# /etc/udev/rules.d/99-esp32.rules
SUBSYSTEM=="tty", ATTRS{serial}=="A50285BI", SYMLINK+="esp32-experimental"
SUBSYSTEM=="tty", ATTRS{serial}=="B7743NM0", SYMLINK+="esp32-production"

After udevadm control --reload, the symlinks survive reboots and re-plugs, and your Makefile targets can reference boards by role instead of by enumeration accident.

Or codify it with Compose

If the two-environment setup is permanent, a compose.yaml documents it better than shell history:

services:
  new-feature:
    image: espressif/idf:release-v5.4
    volumes: ["./new-feature:/project"]
    working_dir: /project
    devices: ["/dev/esp32-experimental:/dev/ttyUSB0"]
    stdin_open: true
    tty: true

  legacy:
    image: espressif/idf:release-v5.3
    volumes: ["./legacy:/project"]
    working_dir: /project
    devices: ["/dev/esp32-production:/dev/ttyUSB0"]
    stdin_open: true
    tty: true

docker compose run new-feature idf.py flash monitor and the mapping from role to physical board is version-controlled.

Part 3: Docker Sandboxes – Letting AI Agents Work Unsupervised

Coding agents like Claude Code are genuinely useful for firmware work: porting components between IDF versions, writing unit tests, chasing config drift in sdkconfig. But to be useful they need to run things: builds, flashes, pip install, sometimes Docker itself. Giving an agent that freedom directly on your host, in bypass-permissions mode, is uncomfortable for good reasons.

Docker Sandboxes solve this with a stronger primitive than a container: each sandbox is a microVM with its own kernel, filesystem, network stack, and its own private Docker daemon. The agent can install packages, modify system config, build and run containers, and none of it touches your host. Your workspace directory syncs into the sandbox at the same path, so file paths in error messages match between the two worlds.

The CLI is small and clear:

# start Claude Code in a sandbox for the current project
sbx run claude

# work on a specific directory
sbx run claude ~/firmware/new-feature

# see what's running, resource usage, network requests
sbx

# list and clean up
sbx ls
sbx rm new-feature

Three properties matter for firmware work in particular:

  1. Disposability. The agent can trash its environment experimenting with esptool versions, partition tables, or custom toolchains. sbx rm and it never happened. Your host IDF setup, if you even have one, is untouched.
  2. Network policy. Sandboxes route traffic through a host-side proxy with three modes: open, balanced (default-deny with pre-approved developer and package-manager domains), and locked down. An agent that decides to curl your firmware to somewhere unexpected simply can’t.
  3. Credential isolation. API keys and tokens are injected by the host-side proxy into outgoing requests; the sandbox itself never sees them. A prompt-injected agent can’t exfiltrate what it doesn’t have.

But how does the agent flash a board?

Here’s where the RFC2217 trick from Part 1 pays off. The sandbox is a VM; there is no USB passthrough. But there is a network path to the host. So expose the serial port as a network service on the host:

esp_rfc2217_server -p 4000 /dev/esp32-experimental

and tell the agent (in your project’s CLAUDE.md or equivalent) to flash with:

idf.py --port 'rfc2217://host.docker.internal:4000?ign_set_control' flash monitor

Now the agent’s whole loop runs end-to-end inside the sandbox: edit, build in a container it spawned itself, flash real hardware, read the monitor output, fix the bug. The only thing it can reach on your machine is one serial port you explicitly published. That’s a remarkably good trade: full hardware-in-the-loop autonomy, minimal blast radius.

Run one sandbox per board and you get the parallel-environment pattern from Part 2, agent edition: an agent iterating on the experimental board via port 4000 while you, or a second locked-down agent, watch the production board via port 4001.

Honest caveats

Sandboxes are newer technology than containers, and it shows in places. MicroVM isolation is available on macOS (Apple Silicon), Windows 11, and Linux with KVM. Build performance inside the microVM is noticeably slower than native containers: fine for agent sessions, annoying for your own tight inner loop. And the agent runs in bypass-permissions mode by design; the isolation is the permission system, so review the diff before merging, same as you would for any contributor.

Part 4: Putting It Together – A Daily Workflow

  • Regular development: VS Code Dev Containers with the espressif/idf image (plus the Espressif IDF extension inside the container). Same image as CI, full IntelliSense, native-container speed.
  • AI-assisted experimentation: sbx run claude --branch <feature>. The branch flag keeps the agent’s commits on a worktree, so your checkout stays clean; review and merge when it’s done.
  • Multi-board testing: parallel containers (you) or parallel sandboxes (agents), one per device, with udev-stable names and one esp_rfc2217_server per board.
  • CI: GitHub Actions with the official espressif/esp-idf-ci-action, pinned to the same IDF version as your dev image. If a build passes locally, it passes in CI. It’s the same bits.
# .github/workflows/build.yml
jobs:
  build:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with: { submodules: recursive }
      - uses: espressif/esp-idf-ci-action@v1
        with:
          esp_idf_version: v5.4
          target: esp32s3

Pro Tips

  • Pin exact image tags (release-v5.4, not latest), and record the tag in the repo (Makefile or compose file) so the toolchain version is part of the code review.
  • One project folder per product line (new-feature/, legacy/) with its own pinned image. Never share a build/ directory between IDF versions.
  • IDF_GIT_SAFE_DIR=/project kills the Git ownership warnings; IDF_CCACHE_ENABLE=1 plus a ccache volume kills the rebuild times.
  • Add --group-add for the dialout GID when combining --device with -u $UID.
  • On macOS/Windows, and always with sandboxes, RFC2217 is your serial transport. One server per board, one port per server.
  • Put the flash/monitor commands and port mapping in CLAUDE.md so agents discover the hardware setup without being told each session.
  • If your team standardizes on extra tools (clang-tidy, cppcheck, a particular esptool), bake a thin custom image FROM espressif/idf:release-v5.4 rather than installing them in every session.

Conclusion

Docker turned ESP32 builds from a fragile, machine-specific ritual into something reproducible enough to trust. Parallel containers turn one desk into a small hardware lab, with legacy and next-gen firmware coexisting without friction. And Docker Sandboxes close the last gap: they make it reasonable, not reckless, to hand an AI agent a real board and let it work.

If you’re still installing ESP-IDF directly on your host machine in 2026, you’re working harder than necessary. Try the two-board setup this week: new firmware iterating on one device, stable firmware soaking on the other. Then hand one of them to an agent in a sandbox and see how far it gets.

Happy hacking!

Learn more

  •  

Docker VMM Public Beta: A Complete Overhaul, Built for Performance

Today we’re announcing the public beta of a fully rebuilt Docker VMM: a new first-party virtualization layer underneath Docker Desktop, optimized for containers, and now available on both Mac and Windows starting with Docker Desktop v4.86. 

What’s Changed, and Why It Matters

Part of the magic of Docker Desktop is how it provides a seamless deployment of the Linux-native Docker engine on other platforms, like macOS and Windows. To support that, Desktop automatically creates and manages a VM and all the complicated integration of your local network and filesystem, in a safe and performant way. 

Creating that VM is the job of a virtual machine monitor, the layer that sits between your hardware and the containers Docker runs. Most developers never think about it. But when it’s slow, unstable, or holding onto your machine’s memory it should have released, you notice it constantly. 

Docker Desktop has always relied on a third-party VMM for this. Now it runs on Docker VMM, built by us from the ground up. That means we own the full stack, and we can tune every part of the engine for container workloads specifically. That translates directly to you: an engine that improves continuously, responds to developer feedback, and ships on our own schedule.

This matters for everyone running Docker Desktop today. Performance, stability, and governance improvements at the virtualization layer enhance the experience across the board, for every workflow, on every team.

Isometric diagram of the Docker stack: Host, DockerVMM, and Docker Engine layers supporting apps and containers today, with agent support is coming next.

Image 1: Isometric diagram of the Docker stack: Host, DockerVMM, and Docker Engine layers.

The Performance Improvements Are Real

Here’s what you’ll notice when you start using the beta release of Docker VMM:

Faster startup. Container startup is measurably faster across the board, from first launch to project switches to restart recovery. 

Better file I/O. File sharing between container and host is significantly faster. When you’re in an edit-compile-test loop, you’ll see improvements every single build. 

Smarter memory management. Docker VMM returns memory to the host when containers are idle, so Docker Desktop isn’t holding onto RAM you’re not using. 

Improved stability on Windows. For the first time, Windows developers get a VMM built and maintained by Docker, with performance and stability work coming straight from us.

Stronger isolation, better performance. DockerVMM still runs in a fully isolated VM, optimized for performance. On Windows, that means the isolation you’d expect from Hyper-V with the speed you’d expect from WSL2. 

One Engine, Everywhere You Run Docker

The virtualization engine powering Docker VMM also powers Docker Sandboxes (SBX). That’s not a coincidence; it’s intentional. Every improvement lands in both products, so you get them wherever you choose to run Docker. 

This matters beyond performance. As we build deeper capabilities into the engine, including enterprise admin controls and tighter governance for dev environments, they surface across both products. Longer term, we’re building toward a unified runtime that spans laptop, cloud, and on-prem, where containers, Compose apps, and agents are all first-class on one foundation. Docker VMM is how Docker Desktop gets there, and this is step one. 

How To Enable It

On Mac: If you are already using Docker VMM in Settings, you will be automatically updated to the new engine when you upgrade to v4.86. 

On Windows: Open Settings > General and you will see a new “Docker VMM” option. Switch it to opt in. 

Feature Flag for Docker VMM in Settings 

Image 2: Feature Flag for Docker VMM in Settings 

No feature flag, no waitlist. Any Docker Desktop user on v4.86 or later can switch today. Note: Linux support will be available at GA. 

What’s Next

Beta runs through fall, focused on real developer workflows: builds, file syncs, and the container startup patterns you hit every day. 

GA is targeted for the end of October 2026, when Docker VMM becomes the default engine for new Docker Desktop installs across Mac, Windows, and Linux. GA is the baseline, and from there, the pace picks up. Everything we build next sits on this foundation. 

Try It Today

Update to Docker Desktop v4.86 to get started. 

Noticing a difference? Have ideas for where you’d want us to go next? We’re collecting feedback through in-product responses, our community Slack, and support channels. 

This is the best Docker Desktop has ever run, and it only gets better from here. 

Learn more

Docker VMM is available today in public beta in Docker Desktop v4.86 for Mac and Windows. Follow the Docker blog to stay up to date on GA and what comes next. 

  •  

A new security baseline for enterprise agentic adoption

Agent Baseline is a blueprint for AI adoption that defines six security outcomes for putting enterprise agents to work without giving them unchecked authority.

Consider this scenario: a customer-support agent receives a ticket with an attachment. Hidden inside the attachment is an instruction: query the customer database and send the results to an external address.

The agent has everything it needs to comply. It can read tickets, query internal systems, call tools, and connect to the internet. The instruction is malicious, but it looks like part of the work.

What stops the agent before customer data leaves the company?

That is the practical security problem enterprises face as agents move from experiments into daily operations. The problem is not only whether a model can recognize a malicious instruction. It is whether the systems around the model limit what the agent can reach, what authority it can use, and what actions it can take when the model gets the decision wrong.

Agents turn familiar controls into a new systems problem

Enterprises already know how to manage identities, isolate workloads, restrict networks, test software, collect logs, and respond to incidents. Those controls remain necessary.

Agents change how the controls must work together. An agent can be reprogrammed at runtime through natural-language instructions. It can choose how to pursue a goal, call tools, use delegated credentials, and spawn other agents. Its effective capabilities may change as models, prompts, tools, MCP servers, and permissions change.

A coding agent illustrates the problem. Give it a bug to fix and it may read source code and internal documentation, install packages, call an external API, delegate tasks to sub-agents, and commit a change. Each step may be reasonable on its own. The risk emerges from the combination: one runtime-programmable actor moving across systems under delegated authority, faster than a person can review every decision.

Security teams therefore need to answer three questions about every agent:

  1. What is operating, and what can it do?
  2. Is it staying inside approved boundaries?
  3. If something goes wrong, can we prove what happened and stop it?

Most organizations can answer parts of these questions. Far fewer can answer them for one agent, one task, and one run across every model, tool, credential, policy decision, and downstream action

Enter the Agent Baseline: an open blueprint for building, operating and governing enterprise agents.

Agent Baseline was created by Docker, Snyk and Keycard to define the minimum security outcomes an enterprise agent deployment should meet.

The current v1.0 draft contains 35 controls across six outcomes:

  • Discover: Maintain an accurate record of every agent, its owner, purpose, components, dependencies, and effective access.
  • Constrain: Limit the agent’s runtime, data, tools, network reach, compute, and duration to what its approved purpose requires.
  • Authorize: Bind consequential actions to a distinct identity, task, target, scope, and period of validity.
  • Observe: Connect intent, identity, policy, tool use, actions, and outcomes with a stable run or trace ID.
  • Validate: Test the agent in the configuration and environment in which it will operate, then verify its outputs and outcomes.
  • Respond: Stop the agent, revoke its authority, quarantine affected components, preserve evidence, and determine impact.

We officially launched the Agent Baseline  at Black Hat 2026, to a full house during the event Securing your AI Agent: The Road to Software Factory.  If you’re curious to hear how it went, check the video below:

Eli Aleyner, VP of Strategy, Docker

The Agent Baseline in Practice

Here is how the baseline contains the support-ticket incident:

“Discover” establishes what is at risk. The agent registry identifies the agent’s owner and purpose, the model and tools it is actually running, the database it can query, the credentials it may use, and any downstream agents it can call. This is current runtime evidence, not the configuration approved six months ago.

“Constrain” blocks the path out. The agent runs inside an isolated environment with a capability profile built for customer support. Its filesystem access is limited. Its network policy denies unapproved destinations by default. When it attempts to reach the external address, the request fails and generates evidence instead of quietly succeeding.

“Authorize” limits the value of compromised access. The agent does not carry a standing credential with broad database rights. It receives short-lived authority tied to the customer-support task, the permitted records, and the allowed action. If it delegates work, the downstream agent cannot receive more authority than the original agent held.

Together, “Constrain” and “Authorize” make the blast radius measurable, which is far better done before an incident than during one. A compromised run reaches in three directions: what it can execute and touch on the host, what identity it can prove and use, and what it can connect to outside. Each direction has a control that shrinks it.

image

The blocked request and the odd query land under one run ID. That is “Observe”: correlated evidence, so the story does not have to be pieced together from five logs a week later. And none of it was a surprise, because “Validate” had already tested this agent against prompt injection in the configuration it actually runs in.

“Respond” contains it. The run is stopped and its active grants revoked, the evidence is preserved, and the affected customer records are scoped so the team knows exactly what the run reached. Essential tickets keep moving through an approved manual fallback while the investigation runs.

None of this depends on the model behaving. Most teams already run three or four of these controls; the usual gap is that they do not connect, so one fires in one place and the evidence lands somewhere else.

Securing organizations in the decade of agents

Agents can accomplish a wide range of tasks. A single agent can navigate seamlessly through the inner and outer loops of development, go through PRDs, write code, commit it, and ultimately push changes to production, much like a human engineer. It also has the ability to do that incredibly fast, using different tools, and creating sub-agents that work in parallel, leveraging the same tools and authentication of the original agent.

Agent governance has become a recurring requirement in our work with customers. They want the productivity of coding agents without giving those agents unchecked access to developer machines, credentials, source code, and external services. 

This led to the development of Docker Sandboxes, microVM sandboxes that run AI agents securely, and Docker AI Governance, a centralized control layer for managing what AI agents can access and do across an organization. These new products, along with the existing Docker MCP Gateway and Docker Hardened Images now give organizations of all sizes an underlying infrastructure with which to manage agentic risk.

Read more about Agent Baseline:

We published Agent Baseline v1.0-draft on July 30, 2026, and presented it at Securing Your AI Agent: The Road to the Software Factory during Black Hat USA 2026. You can watch the session on demand on the link below.

The draft is open for community review until September 30, 2026. We are looking for implementation feedback, missing controls, evidence that a control is ineffective, and cases where a requirement creates disproportionate operational burden.

  • Download the white paper here 
  • Visit agentbaseline.org and contribute your comment to the architecture

Agents will keep gaining access and autonomy. The standard cannot be that they behave perfectly. The standard must be that we know what they can do, enforce where they can go, trace what they did, and stop them when something goes wrong.

  •