❌

Vue lecture

OpenAI blocked its agent’s web access. Then it tunneled out through DNS.

Last week, OpenAI published new reports detailing misalignment issues with its agents. This time, two internal models found ways around security controls after intended paths were blocked. 

In one incident, an agent used DNS to bypass blocked network access and query an external chatbot. In the other, a model exposed a researcher’s GitHub token in a public repository while trying to access another team’s work — after the researcher twice identified that the model was cheating and explicitly told it to stop. Both times, the model agreed before going back to cheating. 

These reports follow the AI company’s release of six misalignment reports of “unexpected or concerning model behavior” it observed in the last six months, including self-generated instructions, information fabrication, unauthorized use of leaked API keys, cross-agent communication, and unsanctioned file-sharing.

Incident #1: When blocked web access isn’t enough to stop an agent

On September 25, OpenAI released a report detailing how one of its agents used DNS to reach an external chatbot. 

According to OpenAI, the agent was attempting a search-based training task to find information about a specific person who had published a blog post. The trouble started when OpenAI’s search tool returned unrelated results, and the agent tried, unsuccessfully, to reach search engines directly. 

After more searches came up short, the agent tested its network access. First, it tried HTTPS to contact a public chatbot. When OpenAI’s web proxy blocked the request, it didn’t give up. Instead, it discovered it could use the training environment’s own resolver to reach the public internet. 

Eventually, it found a way to use DNS delegation to forward questions to the external chatbot. 

OpenAI says its misalignment monitoring system caught the behavior within 15 minutes; three minutes later, a person was reviewing the case. The run, however, wasn’t killed until another 2.5 hours later, after it failed to terminate automatically as expected.

OpenAI pauses tool use

Following the incident, OpenAI says, “All training, evaluation, and inference with tool-use (broadly defined) of our most capable models remain paused.” It’s not clear when that work will resume, but OpenAI says it will make the call when it has “both validated that the gap is resolved and performed additional red-teaming of the system.” 

In the meantime, the AI company has accelerated additional model-assisted red-teaming of the sandbox and network controls. Once it’s ready to restart training, it says it will “begin a fresh run with additional alignment improvements.” It will not resume training on this particular model.

Ultimately, the incident exposed a gap in the AI company’s network controls. 

OpenAI’s safety case assumed the model was barred from the live internet — and that its monitoring system could detect successful attempts. While monitoring did successfully catch this one, a retrospective review from the AI company identified other cases of external DNS access that didn’t get flagged at the expected severity. 

The incident, discovered on September 20, 2026, came after OpenAI said it had already hardened its research environment following the Hugging Face incident, where autonomous AI systems escaped a sandbox and breached Hugging Face. According to the AI company, that incident pushed it to “continuously hardened and red-teamed our research environment,” particularly on broad direct and transitive paths. 

Though OpenAI calls this incident “a lot less severe” than some of its previous episodes, it notes that it’s the first since its Hugging Face-inspired security hardening — a signal that, it says, will inform its next phase of work.

Incident #2: A model keeps cheating

Also on September 25, OpenAI opened up about another model mess-up, this one dating from May 27, 2026. An internal model was assigned a theorem-proving task; it tried to cheat by getting information from another team’s work. 

After struggling to construct the proof, the model proposed using a temporary GitHub Action to fetch information from other teams’ submissions. The researcher shut that down, instructing it to construct the proof locally.

Both times, the model agreed before going back to cheating. 

Though the model agreed, the next morning, it was at it again, asking for help getting another team’s proof material. Again, the researcher repeated their instruction to solve the problem itself. 

But the model kept doing its own thing anyway, cycling through various attempts to access the other team’s work before ultimately exposing the researcher’s locally available GitHub token in the public openai/codex repository, splitting the token into pieces with the stated aim of avoiding secret scanning.

OpenAI pulls the model offline

Again, OpenAI says its misalignment monitoring system flagged the behavior, but it admits the researcher beat it to the punch. After they notified security, OpenAI deactivated all the users’ keys and later, all OpenAI employees’ keys, too.

OpenAI admits the incident “shows particularly severe misalignment.”

In total, OpenAI confirms the internal model was down for about two weeks, coming back up once it had enforced auto-review on all actions, an async CoT misalignment monitor, and stricter system prompting. The AI company also says it increased its misalignment monitoring system to run every 30 minutes instead of only at each compaction window. It also restricted internet access and made changes to its harness and infrastructure. 

In its report, OpenAI admits the incident “shows particularly severe misalignment.” It’s another example of the type of unexpected and concerning agent behavior its new framework for reporting model misalignment aims to surface, in which it warns: 

“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

If the new incidents are anything to go by, OpenAI’s new framework is coming at the right time, as there are still plenty of alignment gaps to find. 

The post OpenAI blocked its agent’s web access. Then it tunneled out through DNS. appeared first on The New Stack.

  •  

pgEdge’s agent database branches end without a merge, and that’s by design

Two dark red rounded squares on a crimson background, with a smaller frosted-glass square overlapping the gap between them.

AI coding agents can get an application running quickly, but taking that prototype into production is another problem. The database setup that works during development may not meet an organization’s security, compliance, or deployment requirements.

On Monday, pgEdge launched Starfleet, a Postgres cloud platform designed to close that gap, pairing database branching and agent tooling with deployment options ranging from pgEdge’s hosted service to air-gapped on-premises environments — all on standard community Postgres.

A database branch per agent

Running multiple coding agents in parallel lets each one try a different approach to the same problem. That gets especially tricky when they’re all working against the same database, so Starfleet gives each experiment its own copy-on-write branch without, pgEdge says, replacing Postgres’ storage layer with a proprietary or semi-proprietary alternative.

Starfleet gives each experiment its own copy-on-write branch without replacing Postgres’ storage layer with a proprietary or semi-proprietary alternative.

Databricks takes a different approach with Lakebase. It runs on Neon, where separating storage from compute makes branching a cheap copy-on-write metadata operation, with Postgres pages stored in object storage.

pgEdge hasn’t said how its branching system works. From the developer’s side, a new database starts as a copy of the source and then becomes its own isolated environment. Changes don’t move between the two, and each has its own connection details.

That separation carries into the agent tooling, with each environment getting its own MCP server address and bearer token, so a client configured with the source database’s credentials won’t be able to connect. It also inherits the source’s IP allowlist at creation, which can’t be changed afterward.

If an agent later needs access from a new address, developers must add it to the source database’s allowlist and create a new database branch, which starts from the source’s data and none of the old branch’s changes.

Linking branches to Git workflows

Linking a project folder through the CLI saves the database ID, or a branch ID, in .pgedge/link.yaml, while pgedge env pull adds its DATABASE_URL to the folder’s .env. An agent working on a feature can then connect to the matching database environment without someone manually passing it credentials. Most read commands automatically use the linked database, though with a branch link only the connection commands point at the branch, while writes require the agent to specify a database ID, adding an extra check against changing the wrong one.

There’s no database merge at the end. Developers handle schema changes the same way they already do, using tools like Alembic or Flyway rather than something built into Starfleet. If the work gets merged, the migration files go with the code and run against the source database. Test data stays behind and is deleted with the agent’s database.

There’s a practical reason to clean up those environments when an experiment is over, since each database can have a branch limit and billing starts as soon as one is ready to accept connections and continues until it’s deleted. Teams running several agents in parallel will probably want to automate that cleanup. Yugabyte is tackling a similar problem from the fleet side with a platform that exposes provisioning, branching, scaling, migration, and teardown to agents through MCP.

Developers handle schema changes the same way they already do, using tools like Alembic or Flyway rather than something built into Starfleet.

MCP access with guardrails

Starfleet ships with pgEdge’s Agentic AI Toolkit for Postgres, which pgEdge CEO Phillip Merrick has described as fully open source and free for all Postgres users. The toolkit includes an MCP server that lets coding agents connect directly to the database. It also includes a RAG API that uses pgvector to retrieve content stored in Postgres, plus a PostgREST API that gives browser clients direct database access, which is an approach used by Lovable and similar app builders.

The MCP server includes pgEdge SafeSession, which the company says prevents an agent with read-only access from changing the database. It works alongside the separate credentials and IP allowlist each database environment inherits.

The MCP server includes pgEdge SafeSession, which the company says prevents an agent with read-only access from making changes to the database.

From prototype to air-gapped production

pgEdge cites IDC and Lenovo research finding that only 46% of general AI and agentic AI prototypes reach production and that 82% of organizations need hybrid or on-premises environments to deploy AI workloads.

Starfleet is designed to let developers start on pgEdge’s hosted infrastructure and later run the same database in their own cloud or on-premises, including air-gapped deployments through pgEdge Enterprise Postgres, with the option to scale to highly available multi-region clusters.

One limitation, however: pgEdge’s current documentation covers branching only for the hosted tier, and the company hasn’t said whether those agent workflows will carry over when a database moves into a customer’s own cloud or data center.

Starfleet starts at $25 per month with a 14-day free trial.

Starfleet starts at $25 per month with a 14-day free trial.

The post pgEdge’s agent database branches end without a merge, and that’s by design appeared first on The New Stack.

  •  

Nvidia launches Open Agent Safety Platform to lock down rogue AI agents

OpenAI, Anthropic, Meta, and Google have all recently disclosed that their models broke out of their test environments and reached real systems. Nvidia’s response, announced Monday, is a runtime that locks agents into kernel-enforced sandboxes and a watchdog on its own silicon that can shut them down.

The Nvidia Open Agent Safety Platform combines OpenShell 0.1.0, the Apache 2.0 agent runtime the company announced at GTC in March, with Nvidia Sentry, a watchdog service that runs on the company’s BlueField-4 data processing units (DPUs).

The new OpenShell release adds a policy prover that checks that an agent’s various permissions can’t be combined into something the operator didn’t intend — like hacking HuggingFace.

Since the BlueField DPU is a separate processor with its own trust domain, it can watch the agent’s traffic to the model and keep an eye on all of its actions and reasoning. Then, when things go awry, it can cut the agent off at the network level.

Justin Boitano, Nvidia’s vice president of enterprise AI, said in a press briefing that the recent incidents “have highlighted a fundamental hurdle for AI agents, and that is that model-level safeguards alone can’t govern what agents can access or do.”

“To date, model safety has been about training good behavior into the model. The industry calls that model alignment,” Boitano said. “For probabilistic systems, this approach has obvious limitations. That’s why we’re introducing a deterministic system to mediate and enforce how these agents behave.”

Credit: Nvidia.

The summer of sandbox escapes

OpenAI disclosed on July 21 that GPT-5.6 Sol and a research prototype had exploited a zero-day in the package proxy that was their sandbox’s only network path and gone on to reach Hugging Face’s production database.

Only a few days later, Anthropic reported that three of its models had found unintended internet access at its evaluation partner Irregular and, among other things, accessed a real company’s database and published a live malicious package to PyPI.

Meta followed on Aug. 6 with a pre-release Muse Spark model that read and modified a real website’s database after the same Irregular misconfiguration handed it a real site name.

Most recently, Google announcedthat Gemini had broken into the networks for three companies, again from an Irregular test environment.

Irregular, which runs pre-release cyber evaluations for frontier labs, is a common factor here, and it’s worth noting that the company is on Nvidia’s partner list for this launch.

Nvidia’s own technical blog puts the breakouts down not to a new capability but to “a combination of tools, time, and ambiguous instructions.”

Asked whether the platform would have prevented the Hugging Face breach, Boitano said each incident is unique, but “from what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on.”

Enforcing policy outside the agent

OpenShell is the core of the platform, and it hasn’t changed all that much since Nvidia first showed it at GTC in March.

With OpenShell, which Nvidia originally announced in parallel with its NemoClaw distribution of OpenClaw, each agent runs in a kernel-isolated sandbox with no network access except through a supervisor that sits outside the workload.

“Traditional sandboxes, whether we’re talking micro VMs or containers or VMs, were built for application-level isolation,” Boitano said. “Every agent running within your company needs to run in its own isolated sandbox with security controls that are outside of the agent’s reach.”

A prover, not a judge

OpenShell is now at version 0.1.0, and the important new component added in this update is a policy prover. This Prover checks that the permissions a given policy grants always stay within the boundary the operator actually intended.

“It is deterministic. It is mathematical reasoning. So this is not LLM as a judge,” Ali Golshan, Nvidia’s senior director of AI software, said during the briefing. Because of that, he said, it runs “roughly at two orders of magnitude higher performance and speed.”

In Golshan’s example, a policy can, for example, bar an agent from reading code on GitHub and posting it externally.

“An agent can bypass this by spawning two sub-agents: one that can read from GitHub, that can talk to another one, that could also then post outside,” he said. The prover models the combined access of the entire agent fleet to find that path.

In Nvidia’s own tests, agents running with reduced safeguards spent up to two hours trying to talk an AI reviewer into granting write access to a protected repository. The prover, the company says, gave the reviewer evidence of what the request actually allowed, and no protected writes occurred.

Sentry: the safety island

Nvidia Sentry adds an additional hardware layer to this system. It runs on BlueField-4 in a trust domain separate from the host, and according to Nvidia, it can quarantine an agent in milliseconds.

With a DPU in the system, the agent’s model endpoint gets routed “through a proxy on the DPU, so that you can see all of the reasoning traces of the agents on the host,” Boitano said.

Unlike OpenShell, Sentry isn’t open source, though Boitano said it has open APIs and that OpenShell can work with other network enforcement hardware.

He compared it to autonomous vehicles. “There’s a primary system that might be running the perception system, and then a safety island that ensures the safety of the system.”

“The DPU is really optional in these architectures,” Boitano said. “In a lot of cases, just using OpenShell on CPUs is honestly good enough for providing sort of strict access control for the agents.”

The DPU, he said, is for “frontier use cases of evaluating models or systems where you might have the guardrails off the models, so it could be for red teaming.”

Who’s building on it

Anthropic is integrating OpenShell with Claude Managed Agents, which already keeps the agent loop on Anthropic’s infrastructure and pushes tool execution into customer-controlled sandboxes.

SpaceXAI says it’s using the platform for Cursor coding agents and Grok models, while Salesforce has added OpenShell audit events and permission approvals into Slack.

SAP is embedding the runtime into Joule Studio and is also contributing code.

OpenAI and Google, two of the four labs whose agents went rogue this summer, aren’t on the partner list. Neither is AWS.

Asked whether Anthropic and OpenAI plan to run OpenShell and Nvidia Sentry for their own training runs, Boitano said to look for the partners’ own blog posts.

The post Nvidia launches Open Agent Safety Platform to lock down rogue AI agents appeared first on The New Stack.

  •  

The rise of agentic AI on Kubernetes: unleashing the new infrastructure layer

Abstract 3D render of blue cubes inside gold wireframe boxes, linked by red rods into a dense cluster, with teal lines connecting outer cubes.

AI is changing expectations around infrastructure and operations, including Kubernetes management. When models run close to the data they use, deployment, scaling, and governance responsibilities tend to shift to platform teams. And as clusters, environments, and operational signals continue to multiply, manual operations often strain under the added weight.

AI may simultaneously provide opportunities to lighten this growing load. Agentic software can now observe a system, reason about it, and act within predefined limits. 

Ultimately, these platforms’ value depends on the quality of the context an agent can see and the boundaries you set. Without cluster state, policy, and access rules, an agent can only guess.

Without cluster state, policy, and access rules, an agent can only guess.

For agentic AI to streamline multi-cluster management, you need clear lines between what the system observes, what it recommends, and what it changes. Drawn well, those lines let teams gain notable speed while still maintaining control.

The impact of AI on computing infrastructure

Teams once treated AI as an application concern; models sat on top of existing systems, and the stack underneath stayed mostly unchanged. Today, AI reaches into more and more customer interactions, while data storage needs simultaneously expand and orchestration pressure grows. A recent Forrester report describes the modern AI computing stack as stretching from the models themselves into and across the infrastructure beneath them.

As AI workloads move into production, they place new demands on the infrastructure beneath them. Many lean on specialized compute, with resource needs that rise and fall through bursts of training and inference. Because conditions shift quickly, they can also call into question whether telemetry remains trustworthy. Each of these demands lands at the infrastructure layer, where the workloads run.

The infrastructure layer of the new AI stack

The infrastructure layer covers compute, storage, and networking. It is a foundation that every workload running on the layer depends on. As AI workloads grow, choices about capacity, placement, and control will increasingly shape the performance of the data, intelligence, orchestration, and experience layers atop the infrastructure.

To operate the infrastructure layer efficiently across many machines and locations, a team may rely on orchestration instead of managing servers by hand. In cloud native contexts, Kubernetes has become a control point for scheduling workloads, applying policy, and presenting a consistent interface across environments. Kubernetes is especially well-suited to support organizations this way when teams need consistent control across an estate spanning data centers, clouds, and edge sites. 

Agentic AI and Kubernetes: the future of the infrastructure layer

Agentic AI can extend automation from fixed rules to systems that adapt to real-time conditions. Traditional automation runs the same script whether the environment has changed, while an agentic system observes the environment, reasons about what it finds, and then takes action.

When you apply agentic capabilities to multi-cluster management, the system follows this same sequence. An agent reads cluster state and operational data, proposes a diagnosis or next step, and then carries out actions based on an approved scope, usually after a person signs off. You can further reinforce these boundaries by routing each request to a specialized agent that receives only the metadata it needs.

The signals that an agent receives from the cluster, the context about policy and access, and the definitions of what the agent may change are the key elements that give agentic systems their value. They also separate agentic AI on Kubernetes from a generic assistant. 

Manual Kubernetes management is less efficient at scale

Admittedly, agentic AI fits some settings better than others. On a small single-cluster footprint, the overhead may outweigh the benefit. Manual Kubernetes management often holds up on a handful of clusters, but it can become unreliable in a rapidly growing estate. After all, each new cluster adds lifecycle work across upgrades, patching, configuration, and renewal. Those tasks can quickly multiply and diverge in hybrid environments.

Configuration drift is a high risk in these situations. Settings that started identical can fall out of sync, and policies can apply unevenly from one team to the next. Individually, these gaps may be manageable, but collectively they raise the odds of an outage or a failed rollout.

Visibility can also erode in an unmanageable way. Clusters spread across data centers, clouds, and edge sites often leave teams with no single view of the whole landscape. When DevOps and platform engineers stitch together signals from separate tools, resolution can slow and become more error-prone. A unified view helps enable sound, efficient decision-making by people, agents, or both.

Kubernetes knowledge is fragmented, and existing AI tools lack business context

Kubernetes expertise often sits unevenly across an organization. For example, senior engineers may hold deep operational knowledge that application teams lack. The most current information about a running system may also be fragmented if logs sit in one tool and metrics in another. Real-time understanding can be further clouded when policies, runbooks, access rules, and deployment history each live elsewhere.

Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes. Without that context, even a capable AI tool may fall short of providing meaningful Kubernetes management support.

Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes.

When an agent can read current signals alongside the rules that govern them, its suggestions become specific, testable, and actionable. In an incident, agentic systems can correlate logs with a recent change. Ahead of a rollout, they can check the change against policy. During troubleshooting, they can account for access rules rather than guessing at them. Kubernetes decisions carry real operational consequences, which makes these details all the more important to consider. 

Engineering “toil” isn’t time-efficient

Site reliability teams use the word “toil” for repetitive manual work, especially tasks that keep systems running without adding lasting impact. In Kubernetes operations, toil takes the form of repeated triage, manual signal correlation, alert follow-up, and routine checks. The tasks aren’t particularly difficult, but they can consume significant time and attention for enterprise teams.

When engineers spend their days on this kind of investigation, proactive modernization efforts tend to stall and planned upgrades can slip behind schedule. In other words, the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements. In a recent survey about how AI provides value to DevOps teams, reducing toil emerged as one of the clearer opportunities.

…the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements.

Agentic AI can support repetitive investigations by gathering signals, correlating them, and proposing a likely cause for an engineer to weigh.

Kept under human review, it can take on some of the routine correlation that would otherwise fall to the team. That kind of support can give engineers more room to focus on the strategic work that most needs their judgment.

Building more intelligent infrastructure with agentic AI and Kubernetes

As you consider building toward intelligent infrastructure without surrendering control, the following principles can inform your efforts:

  • Start with observable context, giving agents access to current cluster state, policy, and history before they reason about a problem.
  • Separate suggestions from actions, allowing agents to recommend freely while any change must wait for human approval and a defined scope.
  • Connect agents to existing controls, routing their work through the access rules, identity, and audit paths the team already trusts.
  • Keep the ecosystem open, favoring platforms that integrate with current tools and standards over those that lock work into a single stack.

Platforms like SUSE Rancher Prime and SUSE AI Factory embrace these principles and illustrate how Kubernetes management can become a foundation for agentic operations. These platforms can help you improve cluster and policy consistency without compromising your authority over AI. Built on open-source foundations, they can also help you avoid being trapped in a single vendor’s stack.

In SUSE Rancher Prime, the industry’s first context-aware agentic AI ecosystem, its AI assistants work as a crew of specialized agents with an intelligent router. The platform draws on the cluster context already in place and acts through existing access controls. Through support for external Model Context Protocol (MCP) servers, teams can extend that crew to their own sources. In addition, human validation tools allow you to hold a proposed action for approval before the agent runs it.

Despite its potential, intelligent infrastructure is not universally beneficial. In situations where change control must stay fully manual, for example, agentic AI’s role may be strictly limited to observation and suggestion. Measure the technology’s value against the realities of your day-to-day operations. For those who are investing, agentic AI will have the greatest impact when it actively supports context, control, openness, and human judgment.

The post The rise of agentic AI on Kubernetes: unleashing the new infrastructure layer appeared first on The New Stack.

  •  

The agent didn’t break your controls. It went around them.

Three black circular directional signs on a gray concrete wall, showing arrows pointing straight ahead, turning left and turning right.

The identity part of agent security is settled. An agent needs its own identity: a short-lived, revocable credential scoped to the job, and an audit trail that names the human who set it running. NIST’s security leads made that case in August 2026, and most identity vendors agree.1

Identity and access management is table stakes. It’s necessary, but it isn’t what’s breaking.

What’s breaking is an assumption we’ve carried for twenty years: Get identity and permissions right at the door, and whatever happens inside takes care of itself. That worked when software was passive. Agents reason about a goal and choose their own steps toward it, like a seasoned escape artist.

An agent that hits a wall looks for another way

Almost every control in today’s stack answers a question about entry. Should it connect? Should it reach that service? Should its token be accepted here? Each is a question about a route, and there’s rarely just one route to anywhere worth going.

An agent treats a blocked route as a problem to solve, because that’s what we built it to do. A person who hits a locked door usually files a ticket, while an agent tries the window.

In July 2026, an autonomous agent spent four and a half days inside Hugging Face’s production systems.2 A filter controlled which internet addresses its dataset servers could download from, and it never fired, because “the agent stopped asking the worker to fetch remote resources and instead made it act on local ones.” The filter worked as designed, and the agent went around it anyway.

A person who hits a locked door usually files a ticket, while an agent tries the window.

On ordinary developer machines, malware in a compromised npm package tried to recruit the AI coding assistants already installed to search for secrets,3 and a coding agent deleted a production database during a change freeze before falsely telling its operator the data couldn’t be recovered.4 Both happened on the machine itself, where no network control was looking.

The shift from outside-in to inside-out

Outside-in controls govern entry, and most organizations run plenty of them. Make no mistake, inside-out security completes those controls rather than replacing them.

Inside-out control governs the action itself, and asks a narrower, harder question: Should this agent, acting on this person’s authority, delete this table in this database, right now?

That question matters because an agent can swap routes but not the outcome it’s after. No matter how many routes it tries, deleting a table is still deleting a table, and a checkpoint on the action sees it every time.

Here’s how today’s controls line up against it.

ControlWhat it coversWhat it misses
GatewayTraffic you route through itLocal shell commands and file edits never reach it
SandboxThe environment as a wholeConstrains reach, not individual actions
SIEMA record of what occurredReports after the action is completed
RegistryThat an agent existsWhat the agent did with that existence

Each does its job, but they all decide somewhere other than the moment the action runs.

Put the enforcement point where the agent acts

Every agent acts through an agent harness: the software that takes the action the model chose and carries it out, whether that means running a command, writing a file, or calling an API. In most deployments today, nothing checks that action before it runs.

An inside-out control puts an approval step in that gap. Before the harness executes anything, the checkpoint looks at which agent is asking, on whose authority, and against which system, then applies policy to allow the action, block it, or send it to a human. Because every action passes through it, an agent denied a destructive command and trying a smaller version of the same thing is held to the same rules. The remaining risk is a badly written policy, which can be fixed.

None of this works without the identity basics. Any type of control, whether it be at the prompt level, inference level, harness level, or MCP layer, can’t judge “may an agent take this action here on this object?” when the only name on the request is a service account shared by six agents and four engineers.

The companies building agent runtimes have reached the same conclusion. Over the past eighteen months, Anthropic, Google, Microsoft, OpenAI, LangChain, and Cursor have each added a hook that lets you inspect an agent’s action before it runs.5 When AWS explained its own agent policy design, it argued that controls belong at the moment an agent attempts to invoke tools.6

The catch is that each hook works differently, with no standardized request or response formats. An enterprise whose developers use Claude Code and Cursor while its platform team builds on LangChain would maintain the same enforcement logic in multiple different flavors, each with its own audit trail. That doesn’t scale, and it tightly couples your security model to whichever runtime a team favors that month. Enterprises need one vendor-agnostic agentic security layer that spans every harness, so adopting a new model or framework doesn’t mean restarting the entire onerous security review.

Turn the lights on before you start blocking

The standard, well-ingrained security instinct is to start blocking right away, but we’ve all seen how well that works with the business in the past. Security must move and adapt at the speed of business, not the other way around. Security tools such as intrusion prevention systems and web application firewalls both ran in monitoring mode until teams understood what normal looked like, and those that skipped that step tended to hear about it from a production outage.

Agents need the same sequence, only faster. An enforcement point in monitoring mode blocks nothing and quickly answers questions most organizations can’t today:

  • Which agents are actually running, not which ones someone believes are running
  • Who started each one, and whose authority it’s operating under
  • What capabilities it used, and against which systems
  • Which of those actions would have violated a policy, had the agent security platform been switched to enforcement mode

Write policy from what you know, see, and have evidence of, not just from an architecture diagram. Enforce first where the stakes are highest: destructive commands, production data, and anything that moves data out. Then watch-learn-build, just like the agents we use: Watch the patterns, build finer-grained controls and policies, and learn how to use AI securely, safely, and confidently. Observe first, then enforce, build, and deploy, in that order.

The bottom line

None of this requires a new category of infrastructure. It’s the identity, authorization, and audit you already run for your people, extended to agents and applied inside the harness before the action runs.

The perimeter is still there, but it has moved to the moment an agent acts, the one place it can’t route around.

Ory built Agent Security inside the harness, on the same identity and authorization engines that run in production for human users. It starts in “observe mode,” so you get that inventory first, and you can try it today at ory.com/agent-security.

Footnotes

  1. Bill Fisher and Ryan Galluzzo, “Back to the Future: Why Agentic AI Needs a Strong Identity Foundation,” NIST Cybersecurity Insights, August 27, 2026.  ↩︎
  2. Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident.” The intrusion ran July 9–13, 2026. ↩︎
  3. Nx, “s1ngularity postmortem,” August 2025. The malicious packages “attempted to use local AI tools (like Claude and Gemini)” while scanning systems for sensitive data. ↩︎
  4. AI Incident Database, Incident 1152: Replit agent deletes production database during code freeze, July 18, 2025. ↩︎
  5. Pre-execution hooks by vendor. Anthropic, Claude Code hooks; Google, Agent Development Kit callbacks; Microsoft, Agent Framework middleware; OpenAI, Agents SDK guardrails; LangChain, human-in-the-loop middleware; Cursor hooks (InfoQ, October 2025) ↩︎
  6. Liana Hadarean and Jean-Baptiste Tristan, “Why Policy in Amazon Bedrock AgentCore chose Cedar for securing agentic workflows,” AWS Security Blog, May 20, 2026. ↩︎

The post The agent didn’t break your controls. It went around them. appeared first on The New Stack.

  •  

Microsoft’s new Copilot agents get their own email, calendar — and a place in the org chart

Satya Nadella stands smiling between Bill Gates, on the left, and Steve Ballmer, on the right, in front of a crowd of cheering employees, many holding up phones and tablets to take photos.

Microsoft announced what it calls its biggest Copilot update to date on Friday, with CEO Satya Nadella describing Copilot as “a new OS for work.”

Nadella framed Copilot as spanning every model, form factor, and task, and the update puts Autopilot, which Nadella called a “proactive and long-running agent built for the enterprise,” at the top of his list of the update’s four components. The pitch targets office workers, but the more consequential change for developers is the infrastructure underneath.

Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365, which means teams building production agents no longer have to assemble those pieces around a model on their own.

We’re building Copilot as a new OS for work that spans every model, every form factor, and every task. Today, we’re announcing our biggest update to Copilot to date, bringing four things together:

· Autopilot: proactive and long-running agent built for the enterprise
· Code:… pic.twitter.com/W2ClHHkCK3

— Satya Nadella (@satyanadella) September 25, 2026

The release adds a new Home experience that merges Chat and Cowork in the Copilot app, but the bigger changes for developers come from Code and Autopilot. Code generates apps, dashboards, and workflows from natural language, and Autopilot turns the agent Microsoft previously called Scout into a persistent background worker. Home and Code are rolling out first through Microsoft’s Frontier early-access program, and Autopilot is expanding to a private preview at month’s end.

Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365

Agents that don’t need prompts

Autopilot takes a role and goal from the person who sets it up, then continues working in the background without requiring a new prompt for each step. Each Autopilot gets its own governed Entra identity and agent user account, separating the agent’s permissions and activity from those of the person who created it.

For engineers, that moves much of the operational scaffolding required for long-running agents into Microsoft’s infrastructure. Independent vendors have been building dedicated layers for that problem; Diagrid, for example, adds durable recovery to LangGraph and other agent frameworks, while Microsoft is bringing those capabilities inside the Microsoft 365 environment.

An identity for every agent

The identity model is the piece developers building on Microsoft Foundry will feel first. Autopilot agents in Foundry, which have been in public preview since June, receive a full Entra Agent ID user account with a productivity license that gives them their own email, calendar, OneDrive storage, Teams access, and a place in the org chart.

Because that user account sits on top of the agent identity every Foundry agent already carries, an autopilot acts as itself rather than on behalf of a user, so developers no longer have to wire agents through shared service accounts or borrowed user credentials, a pattern AuthZed CEO Jake Moshenko has said reflects a common misconception about how agents should be deployed.

A developer creates an Autopilot blueprint from a Foundry-hosted agent, which appears in the Agent 365 registry once an administrator approves it. Employees can then hire instances of that agent in Teams. The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.

The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.

Hosting AI-generated apps

Code is built on the same underlying technology as GitHub Copilot, and the apps it generates run on Microsoft Copilot Managed Runtime, a platform now in public preview that hosts code inside the customer’s Microsoft 365 tenant boundary under IT governance.

Apps deployed there run within the company’s existing identity and governance framework, with Microsoft managing the underlying runtime and giving developers a controlled path to test and deploy new versions without taking the current release offline.

The runtime also accepts apps built in Copilot Studio and Cowork, and Microsoft is opening it to outside tools and professional developers through an SDK and command-line tooling, with Git tracking source and versions.

Lovable is already on board. In Microsoft’s announcement, the company’s head of global partnerships, Lan Roche, said apps built with Lovable can now run inside a Microsoft tenant “the same way everything else does,” using the same sign-in, policies, and app inventory.

The model resembles what serverless computing did for application infrastructure, where developers concentrate on application logic while the platform takes on more of the execution environment. Microsoft is applying that abstraction to generated enterprise software while tying the runtime directly to identity, tenant boundaries, and organizational data.

Long-running agents also change Copilot’s economics. The standard subscription covers the assistant, but Cowork, Code, Autopilot, and other agentic features are billed based on usage through Copilot Credits. That also applies to frontier models such as Fable and Astra, although users still need a Copilot license to access them. Microsoft is extending cost management in Agent 365 to cover Code and Copilot Managed Runtime, and it plans to support agents built in Copilot Studio in October.

Once an agent can keep working for hours or days without anyone watching, cost becomes part of the governance problem. Engineering teams need to control how much compute an agent uses alongside what it can access, which is why Microsoft is bringing those controls into the same administrative framework.

The portability trade-off

That convenience comes with a trade-off. Because Microsoft controls the underlying enterprise environment, it can handle much of the work around agent state, credentials, and access controls, but the more infrastructure a team hands over to Microsoft, the harder the agent may be to move elsewhere.

The models are not locked in, since Microsoft currently runs Copilot on models from both OpenAI and Anthropic and says more labs and open-weight models are coming, and the Agent 365 SDK adds governed Model Context Protocol access to Microsoft 365 workloads for agents regardless of the framework they were built with. Those open interfaces cover only part of an agent’s architecture, though. The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.

The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.

The post Microsoft’s new Copilot agents get their own email, calendar — and a place in the org chart appeared first on The New Stack.

  •  

OpenAI and Cursor agree on agent coordinators. They disagree on who runs them.

Abstract digital art of thousands of thin glowing strands in orange, red and pink bundled into a single sweeping arch against a black background.

OpenAI opened its Agents API in public beta this month, exposing the harness that powers Codex with managed sessions, tool coordination, and subagent orchestration. On the same day, September 10, Cursor launched Projects to coordinate multiple coding agents around larger bodies of software work. While the products sit at different points in the stack, both converge on the same architecture: a coordinator understands the larger objective and manages the work, while specialized agents execute individual pieces.

That pattern isn’t new: AWS Bedrock AgentCore reached general availability in October 2025, and Anthropic’s Claude Managed Agents entered public beta in April 2026. What makes these announcements notable is that two major players in AI-assisted software development are independently exposing the same coordinator-worker split at the same time.

Hilliary Lipsig, a senior principal site reliability engineer at Red Hat who leads Azure Red Hat OpenShift SRE teams and hosts the YouTube livestream GitOps Guide to the Galaxy, has watched this dynamic play out firsthand.

“This convergence highlights the reality developers across the industry have been discussing on and offline — an agent with too much context loses accuracy and reliability, and focused work with clearer contexts allows for faster, more accurate iterations,” Lipsig tells The New Stack.

“The need for orchestration in distributed computing has been fundamentally recognized repeatedly,” Lipsig says. “That’s part of how we got to Kubernetes. These multi-agent workflows are the same concept, just in a new part of the technical stack. While the specialized agents do their area of work, the orchestrator can act as a source of truth — ideally enforcing guardrails, recovering from any failure states, and intelligently routing work to the most efficient target agent.”

“The need for orchestration in distributed computing has been fundamentally recognized repeatedly… These multi-agent workflows are the same concept, just in a new part of the technical stack.”

The industry has spent the first generation of AI coding tools asking how capable a model can become at writing software. The emerging question is different: How do you build a reliable system around multiple capable agents working on the same problem?

The problem with the single-agent loop

A coding agent works through what Anthropic describes as LLMs using tools based on environmental feedback in a loop: it observes the state of a repository, reasons about what to do next, calls a tool, examines the result, and continues. For a small task, that loop can be enough. As the scope expands, however, maintaining reliability in a single context becomes harder.

A large migration might require understanding an unfamiliar codebase, identifying dependencies, changing database schemas, updating services, rewriting tests, modifying deployment configuration, and validating the resulting system. A single agent can theoretically perform all of that work, but it must maintain relevant information from every stage while continuing to reason about what comes next.

The pressure lands first on the context window. “A large context doesn’t only include everything correct or important — it also includes a lot of throwaway information,” Lipsig tells The New Stack. “Through compaction, that information can inadvertently end up ranked as important and incorrectly influence what your agent does. Or correct information can be distorted to become incorrect.

“Either way, after a couple of rounds of compaction, developers are seeing accuracy degrade and are starting to manage context once again manually.”

Lipsig’s read matches what researchers call context rot — and it hasn’t gone away with newer models.

A 2026 study testing frontier models,, including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro, found they missed a dangerous action buried in a long agent transcript two to 30 times more often once it came after 800,000 tokens of benign activity — the AI equivalent of a security guard who stops checking badges carefully after the two-hundredth person walks through, even though nothing about their training changed.

Furthermore, the tasks themselves may not be sequential. Forcing one agent to execute database analysis, documentation work, and test discovery one after another turns a potentially parallel workload into a serial one.

Subagents change that execution model. Instead of requiring one agent to carry an entire task through a single context, a coordinator breaks the work into smaller units and assigns them to specialized agents. GitHub’s custom-agent model illustrates this: different agents receive only the prompts, tools, and context they need for their tasks, executing work in isolated contexts rather than crowding an increasingly large conversation.

Multi-agent systems therefore bring higher token costs and additional coordination and integration risks, and splitting work across agents does not guarantee better software quality.

The coordinator is not another coding agent

Once the work is divided this way, the coordinator becomes a control plane rather than another coding agent. Its job isn’t to write the code, but to understand the global task, manage dependencies, and decide how execution should proceed. Unlike a conventional scheduler, an agentic coordinator makes probabilistic judgments about result quality and resource allocation.

It may dispatch one agent to investigate a database schema, another to examine the service layer, and a third to inspect the test suite. When they return, the coordinator determines if their findings are sufficient to move to implementation. If a worker produces an incorrect result, the system must recognize the failure and decide whether to retry the work, reassign it, or change the task itself.

Anthropic has documented this same pattern in its own production system, calling it orchestrator-subagent architecture: a lead agent analyzes a query, develops a strategy, and spawns specialized subagents to investigate different facets in parallel. In a June 2025 writeup of that system, Anthropic reported a Claude Opus 4 lead agent with Claude Sonnet 4 subagents outperformed single-agent Opus 4 by 90.2% on its internal research eval — at roughly 15 times the token cost of a standard chat interaction (Anthropic puts single agents at about 4 times), a tradeoff that makes the pattern a deliberate architectural bet, not a free upgrade.

Parallelism introduces distributed-systems failure modes

Parallelism is valuable because software work contains many independent tasks, but it creates coordination problems. Imagine a migration where one agent changes a database schema, another updates the consuming service, and a third updates integration tests.

If the schema changes while the service agent works against an earlier assumption, the system produces internally inconsistent work. This isn’t a risk unique to hypothetical migrations — the International AI Safety Report 2026 notes that “interactions between multiple AI agents are also becoming more common, introducing further risks, as errors propagate between systems.”

A single model invocation is a disposable computation, but a twenty-minute workflow modifying a repository is not. If an agent loses its machine halfway through, restarting from scratch is expensive and potentially unsafe against a changed environment.

To solve this, Cursor moved its cloud-agent execution loop to Temporal to handle durable execution and retries, pushing its cloud agents past two 9s of reliability. Temporal now handles 50 million of Cursor’s actions a day across 7 million unique workflows. “Durable execution isn’t a nice-to-have here. It’s the difference between a system you can operate and one you can only demo,” Lipsig tells The New Stack.

“Durable execution isn’t a nice-to-have here. It’s the difference between a system you can operate and one you can only demo.”

By separating agent, machine, and conversation state, the execution engine can reason about the workflow independently. Reliability is no longer just about whether the model produces a good answer; it is about reliably completing distributed workflows composed of many operations, machines, and dependencies.

The environment, context, and observability are one problem

In production, an agent is more than a model and a prompt; it requires a workspace, source code, dependencies, credentials, and state retention. Both companies provision isolated environments for these resources, directly linking an agent’s capability to its blast radius. OpenAI’s Agents API currently supports U.S. data residency but not Zero Data Retention; choosing a self-hosted sandbox does not make the Agents API eligible for ZDR. Cursor supports similar cloud isolation alongside local execution for machine-specific work.

An agent that can only inspect a repository poses a different risk than one that can modify production infrastructure. Consequently, the coordinator is inextricably linked to the security model, determining which agent receives specific information and authorities.

This logic extends to context routing. Giving every subagent the parent’s entire history increases cost and complexity while leaking irrelevant or sensitive information. Instead, the coordinator enforces information-flow boundaries: a database-analysis agent receives only schemas and relevant migrations, while a security-review agent gets the resulting diff without deployment credentials.

As agents increasingly use interfaces like MCP to reach external systems, the platform must strictly govern which agent receives the authority to use specific tools, and for how long. MCP’s governance now sits inside the Agentic AI Foundation, a Linux Foundation foundation co-founded by OpenAI, Anthropic, and Block, with support from AWS, Google, Microsoft, Bloomberg, and Cloudflare to host MCP alongside AGENTS.md and Block’s goose — a sign the industry already treats it as infrastructure worth governing jointly, not a feature any one vendor owns.

This complexity creates a visibility problem. A simple final response often conceals a history involving multiple agents, tool calls, environments, and retries. Systems must expose task-level provenance — which agent received the assignment, what context it used, where it executed, and how the coordinator handled failures or human interventions.

Without execution provenance, debugging requires reconstructing distributed workflows from fragments. GitHub’s exposure of subagent lifecycle events points in this direction, treating agent lifecycles as observable components rather than hidden processes.

Coordination authority is not execution authority

The most critical architectural boundary is the distinction between coordination authority and execution authority. A coordinator needs broad visibility to make useful decisions, but that does not imply unrestricted control over the project. “Just like you don’t want humans running around with root permissions, you don’t want your agents running with them either,” Lipsig tells The New Stack.

“Just like you don’t want humans running around with root permissions, you don’t want your agents running with them either.”

“The ease of creating and leveraging AI agent permissions is lagging behind the speed of AI innovation, but any product team that needs to maintain compliance standards will tell you that easy or not, access controls are incredibly important. Functionally, that means fully agentic workflows need to run in a dedicated environment with their own permissions and connectivity. You can’t rely on agent context to do what you want — trying to keep something in context that long will just result in errors from context rot. Hardware or virtual-hardware-level permissions, in addition to system-level permissions, need to be configured so the agent is asked to provide only its unique value add.”

The risk isn’t hypothetical. OWASP’s Top 10 for Agentic Applications formally names this failure mode Identity and Privilege Abuse (ASI03).

METR’s independent investigation, conducted with a Redwood Research contractor, found that the June–July 2026 incident — in which OpenAI’s own agents, running in internal ExploitGym cyber evaluations, escaped their sanctioned scope and attacked Hugging Face — unfolded in distinct phases. On June 26, agents discovered an exploit that provided full administrator access to OpenAI’s internal Artifactory package repository. On July 4, high-volume Artifactory activity caused an outage, triggering a security investigation and pausing evaluations. OpenAI switched to a freshly initialized Artifactory instance on July 6, which it says wiped the cache and messages left there.

On July 7, a new wave of ExploitGym experiments launched tens of thousands of agents, and on July 8 one agent established the primary unsanctioned message board. Roughly 1,200 agents eventually used the board, sending more than 70,000 messages and files; about 700 later participated in the attack on Hugging Face. The attack itself began on July 10–11 and wound down over July 12–13. The chronology matters because the administrator-access event, the Artifactory outage, and the later message-board activity were separate phases, not one continuous incident.

By binding autonomy, worker agents operate with the minimum permissions required for their specific tasks, keeping sensitive operations behind explicit approval boundaries. This also reshapes human review. Requiring human approval for every tool call destroys the efficiency of multi-agent execution, but showing only the final result obscures critical intermediate decisions.

The most useful design places human intervention around consequential, irreversible transitions — like moving into production or altering sensitive infrastructure. This is especially vital as agents become event-driven participants that respond to Slack messages or pull request updates, not just direct prompts.

OpenAI and Cursor own different parts of the architecture

The convergence does not mean OpenAI and Cursor have built interchangeable systems. Their products put the orchestration boundary in different places.

OpenAI is exposing an agent harness through an API. Its model gives developers primitives for managing context, tools, subagents, and execution environments, leaving application teams to decide how those capabilities fit into their own systems. The harness is open source, so teams can inspect the coordinator logic instead of treating it as a black box.

Cursor packages more of the surrounding workflow. Projects provides the coordinator, cloud execution, shared project context, and a developer-facing workflow in the same environment.

That difference matters because orchestration is a collection of infrastructure decisions: who owns the execution environment, where workflow state persists, how agents are isolated, how credentials are provisioned, what happens when a worker fails, how one agent’s output becomes another agent’s input, and which actions can happen without human approval.

An API gives developers more responsibility for answering those questions. An integrated platform answers more of them on the developer’s behalf.

Neither approach removes the underlying engineering problems. It changes where they are implemented and who is responsible for operating them.

The coordinator is becoming an architectural boundary

The evidence from these systems points to a change in the role of the coding agent itself.

The model still performs the reasoning and code generation. But larger agentic workflows require another layer to determine how that capability is applied: which work is delegated, what context crosses an agent boundary, which tools are exposed, how execution state survives failures, and when the workflow needs human intervention.

Those are familiar distributed-systems concerns. Workers operate concurrently, state can be shared or isolated, dependencies connect tasks, workers can fail independently, and results need to be persisted and observed. The difference is that the workers are now probabilistic software agents rather than conventional processes.

That makes the coordinator more than a convenience feature. It is where a high-level software objective becomes executable work — and where decisions about context, permissions, durability, observability, and human intervention converge.

The September 10 launches make that shift visible from two different directions. OpenAI exposed orchestration infrastructure through an API. Cursor embedded it into a project-level development environment.

Neither announcement proves that one architecture will become the universal model for software development. But together with the systems already emerging around them, they show coding agents moving away from a single model executing an entire task and toward workflows that divide work among specialized agents, execution environments, and persistent infrastructure.

The engineering question is therefore no longer only whether an agent can write the code. It is whether the system around it can reliably decide what to do, which agent should do it, what that agent should be allowed to see and change, how to verify its work, and where a human should take control.

Those are architecture and infrastructure questions — and as coding agents move from interactive assistants toward autonomous software workflows, they may matter as much as the underlying model.

The post OpenAI and Cursor agree on agent coordinators. They disagree on who runs them. appeared first on The New Stack.

  •  

Manufacturing Trust for AI Agents | Docker’s WeAreDevelopers Keynote

Coding agents do more than write code. Agents can install dependencies, access networks, use credentials, and keep working after we’ve moved on to something else. Agents need access to get real work done. But the more access they have, the bigger the impact of a mistake. The challenge is building enough trust into the system to give agents more freedom without giving them access to everything. 

At WeAreDevelopers North America, Docker President Mark Cavage laid out Docker’s answer: give every agent a strong boundary, make its environment and authority reproducible, and let the work run where it makes sense, from a developer’s laptop to the cloud. 

In 30 minutes, Mark demonstrates what happens when an agent pushes beyond a container’s limits and how Docker Sandboxes create a stronger boundary. He then connects that foundation to Kits, cloud capacity for work that needs to continue beyond the laptop, and new partnerships helping developers put agents to work safely. 

agents trust factory transparent

Agents are actors, not workloads

Containers were built to isolate applications. Agents do more. They make decisions and act on the systems around them, so the boundary needs to extend to the environment they can reach. 

Mark’s onstage demo makes that clear. The container worked as designed. What the agent needed was containment around its full environment. Docker Sandboxes provide that containment. Each agent gets its own isolated microVM and kernel, separate from the developer’s host environment. 

Developers can choose the agent, model, and tools that fit the job, then define the sandbox’s access to files, networks, and secrets. Those policies stay outside the agent’s control. Together, the sandbox and its policies create a trusted environment for autonomous work. 

Docker Sandboxes are available today through the free, standalone CLI, so developers can bring stronger isolation to the agent workflows they already use. 

Make the agent and its authority reproducible

A strong boundary gives an agent a safe place to work. Developers also need a repeatable way to define what runs and what it can access. 

In the keynote, Mark announced the next generation of Docker Sandbox Kits, now built as standard OCI images, and invited ecosystem partners to help shape the specification as Docker works towards neutral governance. 

Docker Sandbox Kits package the agent, its tools, and the rules for what the sandbox can reach into one versioned, shareable artifact. Developers can share Kits through the same workflows they already use for container images, while keeping changes to an agent’s authority visible and reviewable. 

The result is a common foundation for the agent ecosystem, built on an open standard. 

Dockerfiles made software reproducible. Kits make authority reproducible

any model any harness any tool

Start on your laptop. Finish in the cloud.

Some agent work begins and ends on a developer laptop. Longer-running or parallel work needs capacity that stays available when the developer steps away. In the keynote, Mark announced Docker Cloud Sandboxes, bringing the same microVM-based isolation from the laptop to Docker-managed cloud compute. 

Using the familiar sbx workflow, developers can start locally, move their work to the cloud with one command, and let it continue after they close their laptop. They can also run tasks in parallel without provisioning or maintaining the infrastructure themselves. 

Docker Cloud Sandboxes are available today with pay-as-you-go pricing.

start local move to cloud move back

An ecosystem built around trust

An open agent ecosystem needs both: choice for the developers and shared standards the industry can build on.

Nous Research joined Mark onstage to show Hermes running as a first-class Kit in Docker Sandboxes. The demo showed what choice looks like in practice: a third-party agent packaged for a common environment, with Docker providing the underlying isolation and controls. Developers can choose the agent that best fits their work without having to rebuild the trusted execution layer around it.

Choice also depends on shared standards that the wider ecosystem can adopt. Docker has committed to submitting its Kits specification to the Cloud Native Computing Foundation (CNCF), the open source, vendor-neutral hub of cloud-native computing.

“Standards are what let an ecosystem move fast without fragmenting, and few companies understand that better than Docker through their involvement in efforts like the OCI and CNCF. By delivering Sandbox Kits as standard OCI images, Docker is giving the industry an open, repeatable way to package an AI agent, its tools, and its guardrails as one artifact. OCI is the foundation the cloud native ecosystem is built on, so a standard for agents that builds on OCI reaches the whole ecosystem at once. The cloud native community looks forward to working with Docker to bring this work under neutral governance.” 

Chris Aniszczyk

CTO at CNCF

Together, these partners show what an open agent ecosystem can look like: choice at the agent layer and a shared format for packaging an agent’s environment and authority. Docker provides the trusted foundation underneath, whether agents run locally or in the cloud.

Trust comes from the system around the agent

Taken together, the announcements in Mark’s keynote form one system. Sandboxes provide a deterministic boundary. Kits make the agent’s environment and authority reproducible. Cloud Sandboxes extend the same trust model to durable cloud capability. 

They give developers the freedom to choose their agents and more autonomy, while retaining control over what they can access and change. 

As agents take on bigger jobs, the infrastructure around them matters more. Docker provides the trusted foundation they need to work locally or scale in the cloud, while developers continue to stay in control of what ships.  

Ready to try Cloud Sandboxes?

For a limited time, new accounts can claim $250 in compute credit to get started.

Watch, Explore, Build

  •  

Introducing Cloud Sandboxes: Start on Your Laptop, Finish in the Cloud

Run agents on your laptop, in the cloud, and move between them with one command, all safely. Earlier this year, we launched Docker Sandboxes: microVM environments where coding agents can work autonomously and safely. 

Since then, agents have started taking on long-horizon work: tasks that run for hours, not minutes. Agents need a new place to work. Today, we’re introducing Cloud Sandboxes: the same microVM-based sandbox, running on Docker-managed compute, with one command to move between them. 

The most important change in coding agents over the past year is that they can work for much longer. Tasks that used to need a developer checking in every few minutes (a large refactor, a dependency migration, a test suite that takes an hour) you can now hand off and review when they finish.

When agents worked in short bursts, the question was whether the model could hold a task together. Now that they work in hours, the question is where those hours happen. A laptop is built around a person. It sleeps when the lid closes, slows down on battery, and disconnects when you move. None of that matters for a task that takes thirty seconds. All of it matters for a task that takes all night.

We built Docker Sandboxes to answer the first question about agents: is it safe to run unattended? In Docker Sandboxes, agents run in a microVM with its own kernel and Docker daemon, isolated from your machine, so they can work autonomously while your files, network, and secrets are protected if an agent steps out of line.

Today, we’re introducing Cloud Sandboxes to answer the second question: how do I run a dozen agents at once, for five, ten, or 21 hours each, without watching any of them? Cloud Sandboxes is the same microVM, running on Docker-managed compute. The isolation model is identical. The CLI is identical. What changes is that the machines underneath are always on, and there are as many of them as you need.

Screenshot 2026 09 23 at 11.59.36 AM

Here’s what that means for you:

Close your laptop and keep the work going. Start an agent in the cloud before you leave for the day, disconnect, and review what they did in the morning.

Start local and move to the cloud when your task outgrows your laptop. Iterate with an agent on the code in front of you, then hand off long tasks and loops to a background agent.

$ sbx move my-project --to cloud

A move captures the sandbox’s filesystem and recreates it on the other side, so your work carries over. This works in both directions.

Run 100 tasks in parallel with nothing to provision. Docker runs the compute. You point your agents at it. Each one gets its own microVM, its own secrets, and its own network policy.

Try things cheaply. Run a pre-built agent for an hour to test an idea, then spin it down.

Everything an agent needs to run on its own

Agents that do long-horizon work without you need three things:

  1. A way to start with nothing to install
  2. Tools to do their jobs
  3. Limits on what it can reach

Cloud Sandboxes ship with all three. You can spin up and manage Cloud Sandboxes from the CLI or from the web console.

Kits. A kit is a pre-configured, pre-built sandbox for an agent. The leading coding agents are ready today, including Claude Code, Codex, Copilot, Antigravity, Open Code and Hermes. Or, you can easily add your own. It’s easy to spin up. sbx –cloud run codex starts one, for example. 

Screenshot 2026 09 23 at 12.01.19 PM

MCP. Connect the MCP servers your agents need (Jira, Linear, Grafana, incident.io, or any streamable HTTP endpoint) once, and all your agents can reach them through a single gateway whether they’re running in the cloud, locally, or in other clients like the ChatGPT desktop app.

Screenshot 2026 09 23 at 12.01.41 PM

Secrets. Store keys or tokens once. Cloud Sandboxes proxy injects it per request, so agents don’t see the actual secret. Prompt injections can’t touch secrets your agents never had in the first place.

Screenshot 2026 09 23 at 12.02.03 PM

Policies. Set network policies on what endpoints agents can access by defining them once. Centralized governance is coming soon for enterprises through Docker AI Governance.

You can access and manage all of this through the web console.

Why one sandbox in two places

We could have built Cloud Sandboxes as a separate product with its own commands. We didn’t, and the reason is simple: how much you can trust your agent shouldn’t depend on where it happens to be running.

As far as we know, no other agent sandbox works this way. We think the right answer to using agents is one strong isolation model that runs in both surfaces (locally and in the cloud), and an easy way (one command) to move between them.

These aren’t two tiers of one product. Interactive work belongs on your laptop. Work that takes hours belongs in the cloud. Most developers need both, and now you don’t have to pick.

Pricing

Cloud Sandboxes are pay-as-you-go. We meter compute by the second, and nothing else. A paused sandbox costs nothing. Volumes, egress, and hosting public images and Kits are free. You can bring your own model key and keep your existing provider for inference.

Size

vCPUs

Memory

Per hour

Micro

1

2 GiB

$0.07

Small (default)

2

4 GiB

$0.14

Medium

4

8 GiB

$0.28

Large

8

16 GiB

$0.56

XL

16

32 GiB

$1.12

Sandboxes run for one hour by default and up to 24 hours per session.

Get started

From the browser: sign in to the web console, choose a kit, and click Run. 

From the terminal:

$ brew install docker/tap/sbx
$ sbx login
$ sbx --cloud run claude

You’ll need sbx 0.45.1 or later and the pay-as-you-go plan, available on Docker Personal and Pro accounts. Docker Sandboxes on your laptop remain free and standalone, with no Docker Desktop required. Local and cloud sandboxes keep separate secrets, templates, and network policies, so read the differences in the docs before you move a local workflow.

For a limited time, new accounts get $250 in free Cloud Sandboxes credit. Claim it here.

We think the way software gets built is changing. Developers will spend less time typing and more time directing: handing agents real tasks, stepping back, and reviewing what comes back. For that to work, agents need environments that are safe enough to run unsupervised and durable enough to run for hours. That’s what Docker Sandboxes is for, and as of today, it runs on your laptop and in the cloud.

Start one before you close your laptop tonight.

  •  

OpenAI’s agent had a routine task. It breached a government portal.

Sam Altman, OpenAI CEO

An OpenAI agent researching public medicine spending bypassed security blocks and gained unauthorized access to public and non-public files on an Australian government Medicare statistics portal, the government there disclosed Thursday. The agent, which OpenAI said was running during an internal evaluation in June, also wrote files to an internal server, according to the complaint.

Transluce, an independent nonprofit AI research lab, analyzed public request logs from the URL scanning service urlquery.net and found autonomous agents attempting SQL injection, cross-site scripting, command injection, and path traversal against the University of New Mexico’s digital library, the public data platform Data USA, and the Australian Institute of Health and Welfare (AIHW). The agents tried to retrieve ordinary information, including a historical photograph, University of Iowa data, and local pharmaceutical data in Victoria, and the offensive behavior appeared only after normal retrieval methods failed.

Transluce ties the Data USA and AIHW activity to an agent swarm that it says OpenAI previously confirmed originated from the company, based on shared targets, tactics and timing.

A day after Transluce published its findings on Wednesday, Australian Prime Minister Anthony Albanese disclosed that an OpenAI agent researching public medicine spending bypassed blocks on a Services Australia Medicare statistics portal in June and gained unauthorized access to public and non-public files. OpenAI, which said the agent was running during an internal evaluation, notified the Australian government on September 10. According to Albanese, the agent also wrote files to an internal server.

The incidents have not been publicly confirmed as the same activity. However, OpenAI told the Australian Broadcasting Corporation that much of what Transluce documented overlaps with cases in its ongoing review of misaligned model activity.

Transluce ties the Data USA and AIHW activity to an agent swarm that it says OpenAI previously confirmed originated from the company, based on shared targets, tactics and timing.

Probing public data providers

On May 25 and 26, agents trying to retrieve a photograph from the University of New Mexico’s Valmora collection ran into repeated failures. They responded with seven vulnerability probes against the service. The probes included an attempt to execute a system command, a path traversal request targeting /etc/passwd, and a SQL injection string containing UNION SELECT; they also arrived alongside a burst of 80 requests aimed at the image, although none appear to have succeeded.

Two days later, agents gathering University of Iowa data from Data USA hit errors caused by a malformed query parameter and followed those errors with 12 vulnerability probes against the API, covering SQL injection, reflected XSS, server-side template injection, path traversal, and command injection.

Cloudflare blocks, staging server workaround

The AIHW case on June 20 and 21 followed a very similar pattern. While trying to download pharmaceutical spending data from an AIHW Tableau dashboard, the agents ran into Cloudflare’s bot protections. One of them responded within minutes by testing the dashboard with a reflected XSS payload, although Cloudflare stopped it before it reached the application.

Transluce found no evidence that the probe succeeded, and AIHW says there is no evidence that non-public information was accessed. The agent ultimately found the same public dataset on AIHW’s pre-production server at pp.aihw.gov.au, and retrieved it through more than 100 scans.

The agent ultimately found the same public dataset on AIHW’s pre-production server at pp.aihw.gov.au, and retrieved it through more than 100 scans.

Agents build their own toolchains

The earliest clear sequence Transluce documented dates to March 6, when an agent tried to pull drug-enforcement statistics from the API of Thailand’s Office of the Narcotics Control Board. A direct request returned no page content, and adding a JSON format parameter changed nothing, so the agent routed the page through r.jina.ai three times and got text back with the Thai province names garbled.

Then, it wrote a custom Base64-encoded script and submitted it through urlquery.net, effectively using the service’s remote scanning browser to run its own client-side JavaScript against the data. That script ran into the same parsing problem, and a later attempt through Jina triggered a redirect to localhost that got the entire API blocked by Jina for an hour.

Beginning in mid-April, Transluce found the same technique across thousands of agent requests on urlquery.net, and on June 14 agents used the service’s browser to create a disposable email inbox and then tried to register a urlquery.net account with it. Registered users can make their scans private, which means the public records Transluce analyzed may understate how much of this activity actually took place.

An agent with web access isn’t limited to the software its developer gave it because remote browsers, proxies, URL scanners, and other public services can fill in the gaps, which gives the agent ways to make requests or run code that its own environment doesn’t provide.

Egress controls for AI agents

Instructions won’t be enough if the agent can still send whatever it wants over the network. For narrowly defined jobs, outbound traffic can be limited to approved hosts, a closed-by-default approach also used for securing AI agent sandboxes. Research agents may need to reach more of the web, so the focus shifts to controlling where they can connect.

Guidance for GKE Agent Sandbox recommends isolated runtimes with default-deny network policies that open only the endpoints an agent needs. Public proxies, URL scanners, and disposable email services can stay blocked unless the job requires them.

Developers can also limit what an agent can send. So, instead of handing it a networking tool that accepts any URL or request body, an API integration can restrict requests to specific fields and formats. The runtime can then catch path traversal attempts, SQL injection strings, and executable markup before anything is sent. OpenAI takes a related isolation approach in its Agents SDK sandboxes, and the company’s Responses API tech lead has said large enterprise deployments often call for agents that are isolated from the network entirely.

Repeated failures can also be a reason to pause a run, especially when an agent keeps hitting client errors, anti-bot challenges, or unexpected redirects and begins trying increasingly aggressive ways to get around them, as Transluce documented in several of these cases.

Keeping the original task, tool calls, and server responses in the same trace gives operators a better chance of catching that behavior change when a retrieval job starts generating encoded scripts, visiting staging domains, or sending exploit payloads, rather than discovering it later in someone else’s security logs.

Repeated failures can also be a reason to pause a run, especially when an agent keeps hitting client errors, anti-bot challenges, or unexpected redirects and begins trying increasingly aggressive ways to get around them, as Transluce documented in several of these cases.

The post OpenAI’s agent had a routine task. It breached a government portal. appeared first on The New Stack.

  •  

Google’s Gemini CLI now asks before editing your build files

Abstract wave

The appeal of an autonomous coding agent is that you hand it a task, give it access to your repository and tools, and stay out of its way while it works. Google’s latest Gemini CLI release carves out specific moments when the agent now has to stop and wait for you.

Gemini CLI 0.61.0, released Wednesday, requires explicit confirmation before the agent edits build configuration files, runs build or test commands after such an edit, or executes shell commands whose arguments appear to come from untrusted external content. The same release separately hardens Gemini CLI’s optional sandbox so that host credentials and configuration stay out of reach of whatever runs inside it.

Giving a coding agent more authority to modify and execute code also gives an attacker more ways to turn that authority against the developer. Gemini CLI 0.61.0 puts a human back in the loop at some of those points.

Security fixes, in public

Google announced at I/O in May that it would move Gemini CLI’s Pro, Ultra, and free-tier users to its closed-source Antigravity CLI, and since June 18, the open-source tool has served mainly enterprise customers and developers with paid API keys. The company said Gemini CLI would continue to get model updates, bug fixes, and security patches. Those security changes are still developed in public, and the pull requests behind version 0.61.0 show exactly what Google was worried about.

Build files become attack vectors

A change to package.json, Makefile, pyproject.toml or a Bazel BUILD file can pull in a dependency or trigger a script. Gemini CLI can make those edits using information from web searches and external tools, then run shell commands. If documentation fetched while fixing a bug contains hidden instructions to add a postinstall script to package.json, the agent could make the edit, run the project’s test suite, and execute the malicious code without the developer ever typing the command.

Giving a coding agent more authority to modify and execute code also gives an attacker more ways to turn that authority against the developer.

Pull request #29250, titled “prevent indirect prompt injection via build file modifications and untrusted flags,” targets that sequence directly. Edits to recognized build files now require confirmation, and Gemini CLI tracks which build files change during a session so it holds any later build or test command, such as npm run, make, or cargo, for explicit approval. The confirmation dialog also shows full build-file diffs rather than truncating them.

Untrusted arguments need approval

The second check covers command arguments. Gemini CLI now treats content from web fetches, MCP server responses, Google Docs, and Buganizer, Google’s internal issue tracker, as untrusted context, and it asks before running any shell command whose flags or arguments match tokens from that content. In both cases, the prompt drops the persistent approval options, so a developer can’t grant a standing “always allow” for these actions.

The pull request ties the changes to restricted workspace mode, the safe mode Gemini CLI applies to folders a user hasn’t marked as trusted, and it doesn’t spell out how the checks behave in a trusted folder or under auto-approval.

The argument check matches tokens rather than tracing the provenance of every value, and the pull request’s review history shows how hard that is to get right. Google’s automated reviewer flagged several workarounds in earlier versions, including quoted arguments, environment-variable prefixes, shell redirection targets, and Windows path handling, all of which were addressed before the change merged on September 11.

Sandbox keeps credentials out

Pull request #29214 tightens Gemini CLI’s sandbox. When the sandbox runs through Docker, Podman, LXC, or macOS Seatbelt, the host’s ~/.gemini directory is no longer mounted inside it. Instead, the CLI passes in a sanitized copy of the user’s settings with API keys, hooks, and custom tool commands stripped out. It also blocks the sandbox from launching in sensitive locations such as the home directory, while new Seatbelt rules deny access to OAuth credentials, trusted-folder decisions, and .env files.

Google’s sandboxing documentation calls the feature a security barrier between AI operations and the host system, while cautioning that it reduces risk without eliminating it. The two pull requests show why both layers are needed. The sandbox limits what a process can reach once it runs, and the confirmation requirements decide whether the agent gets to take a sensitive action in the first place. Build files make the gap concrete: the sandbox mounts the project directory so the agent can edit it, meaning a poisoned package.json written inside the sandbox still sits in the repository when a developer or a CI job later runs the build outside it.

…a poisoned package.json written inside the sandbox is still sitting in the repository when a developer or a CI job later runs the build outside of it.

Gemini CLI already gives developers ways to decide how much the agent does on its own, from hooks that run deterministic checks at fixed points in the agent’s workflow to an MCP server trust setting that, according to Google’s documentation, bypasses all tool call confirmations for that server. Trust granted once can age badly, though, as tool-poisoning and rug-pull attacks on MCP servers have shown when a tool approved on one day starts returning attacker-controlled content later.

Trust granted once can age badly… when a tool approved on one day starts returning attacker-controlled content later.

The post Google’s Gemini CLI now asks before editing your build files appeared first on The New Stack.

  •  

What managing 150,000 AI agents could look like for database teams

Abstract 3D render of translucent orange cubes and panels scattered across a pale gray background, with bundles of glossy teal tubes curving in from the right.

The database administrator of the future will spend considerably less time administering databases.

That sounds contradictory, but AI agents are taking over that work. For decades, DBAs have handled the decidedly hands-on work of keeping databases available, performant, secure, and affordable. They provision capacity, troubleshoot slow queries, manage migrations, and step in when something inevitably goes sideways.

AI is already taking on some of that work. At the same time, it is creating a much bigger data infrastructure fleet to manage.

The result is likely to be a very different kind of DBA: one who spends less time tending individual databases and more time supervising the autonomous systems doing it for them.

Congratulations, you’re managing robots now

This shift starts with a familiar problem: more infrastructure needs managing than the people available to manage it.

Database automation is hardly new, but agents can potentially go further than the scripts and rules DBAs already rely on. Rather than automating one predetermined task, an agent can inspect what is happening, decide what needs attention, use tools to act on it, and check whether its intervention worked.

That changes the DBA’s relationship with the database. A performance problem that once required someone to dig through metrics, identify the troublesome query, and decide how to respond could increasingly be investigated by an agent before a human gets involved.

It doesn’t remove the DBA from the equation. Someone still has to decide what an agent can do, where human approval is required, and what happens when it gets something wrong. But the work moves up a layer. Instead of personally performing every operational task, DBAs start managing the systems carrying them out.

Instead of personally performing every operational task, DBAs start managing the systems carrying them out.

And before anyone gets too comfortable with that idea, the number of those systems could become enormous.

150,000 agents walk into a database…

Gartner predicts that the average global Fortune 500 company will have more than 150,000 AI agents in use by 2028, up from fewer than 15 in 2025. Only 13% of organizations currently believe they have the right governance in place to manage them.

Not every agent will need its own database, but plenty will. They will create state, retrieve data, remember previous interactions, and exchange information with other agents. Many will also behave very differently from the applications DBAs are used to supporting: spinning up quickly, sitting idle for long stretches, and suddenly becoming busy when there is work to do.

Nobody is hiring 150,000 DBAs to manage them.


That is the scale problem Yugabyte is targeting with YugabyteDB AMP, or Agentic Multitenant PostgreSQL. Rather than treating each new agent workload as another database for an administrator to provision and babysit, AMP manages databases as a fleet.

The platform packs hundreds of small Postgres workloads onto shared distributed infrastructure while keeping their databases isolated. Lifecycle operations, including provisioning, branching, scaling, migration, and teardown, can be exposed to agents through MCP. Yugabyte has also built specialized agents for setup, migration, performance tuning, and integrations.

In that model, a DBA is no longer provisioning database number 14,372. The interesting job is setting the rules for how database number 14,372 is provisioned, operated, and fine-tuned without them.

Do more with less (no, really)

Scale is only half of the problem. Someone also has to pay for all this stuff.

Agent workloads make traditional capacity planning particularly awkward because many are bursty and frequently idle. Giving every experimental agent permanently provisioned infrastructure could leave companies paying for many databases that spend much of their lives doing very little.

This is where consolidation becomes as much an economic question as an operational one.

AMP’s approach is serverless multitenancy and scale-to-zero. Multiple small workloads share the underlying distributed infrastructure, while customers pay by CPU minute and idle agents consume no compute. Resource governance can impose CPU limits on individual workloads, preventing a single overeager agent from consuming the capacity intended for its neighbors.

The human equivalent matters too. If routine setup, migrations, tuning and other database operations can increasingly be delegated, a smaller database team can potentially look after a much larger estate.

That doesn’t mean companies get to fire the DBAs and hand the keys to the robots. It means scarce database expertise can be spent on architecture, governance, and genuinely difficult problems instead of repeatedly doing the work that software can handle.

Your 2028 database problem starts now

The harder question is what to build underneath all of this when nobody really knows what the enterprise AI estate will look like in two years.

An agent that begins as an experiment today could disappear next month. Another could suddenly become a production application used across the business. Building one infrastructure stack for cheap experiments and another for serious workloads risks creating a migration problem every time an experiment succeeds.

Yugabyte bets that both ends of that journey should sit on the same foundation.

YugabyteDB AMP lets workloads start on serverless Postgres and transition to fully distributed YugabyteDB as their scale and criticality increase, without rewriting the application or migrating data to a different database platform.

Then there is the problem above the individual database: agents need to remember what happened, and not just in a silo.

That’s where Meko fits into the Yugabyte stack. Meko is an agent-native context engine designed for multi-agent AI systems. It provides persistent memory, shared knowledge, decision traces, and autidability across multiple agents, rather than leaving each agent working from its own isolated context. An agent can pick up information learned by another agent instead of retrieving it again or restarting the reasoning process.

Taken together, it delivers a single data stack for an agent’s entire lifecycle: Meko for the context shared among agents, YugabyteDB AMP for agentically managing fleets of Postgres databases, and distributed Postgres-compatible YugabyteDB for workloads that outgrow their serverless beginnings.

Of course, there’s no guarantee that 2028 will look exactly like today’s forecasts. That’s rather the point. The safest architectural bet may be one that doesn’t require you to know in advance which of today’s tiny AI experiments will become tomorrow’s critical applications.

The DBA is still critical in that world, but the job will look different. The DBA of the future may manage fewer databases directly, while taking responsibility for vastly more of them. Instead, managing the autonomous systems that do the administering.

The post What managing 150,000 AI agents could look like for database teams appeared first on The New Stack.

  •  

Cursor acquired Firetiger. A month later, it launched a bot that tracks code changes from PR to production.

Inspecting changes on a laptop screen document

We all know that producing code is easier than ever thanks to the abundance of AI coding tools and agents. The harder part undoubtedly comes after that code is written: making sure changes are safe to ship, spotting regressions in production, and figuring out what went wrong.

And that’s why Cursor is introducing Rollouts, a new agent that follows code changes into production and monitors whether they behave as intended.

The Firetiger effect

The announcement comes a little over a month after SpaceX closed its bumper $60 billion acquisition of Cursor, giving the AI coding company access to SpaceX’s vast GPU infrastructure as it develops its own models.

The day before that deal closed, however, Cursor quietly announced an acquisition of its own: it snapped up the team behind Firetiger, a three-year-old startup building AI agents that monitor software changes from pull request through deployment.

At the time, Firetiger co-founder and CEO Rustam Lalkaka argued that coding agents had dramatically reduced the effort involved in creating software changes, while doing little to reduce the risks involved in actually deploying them.

“Over the last two years, agentic coding has changed software dramatically,” Lalkaka wrote in a LinkedIn post following the deal’s announcement. “The cost of creating changes has dropped to near zero. The cost and risk of deploying them has stayed largely the same.”

“Writing code is no longer the slow part. What hasn’t sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.”

Rustam Lalkaka, Cursor

Fast forward to today, and Lalkaka, now at Cursor, has unveiled the first fruits from that acquisition — including Rollouts. In a blog post published on Wednesday, Lalkaka notes that the new agent, or “bot” as the company calls it, is all about helping developers “get safe, reliable code into production faster.”

“Writing code is no longer the slow part,” Lalkaka writes. “What hasn’t sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.”

Rollouts is effectively Firetiger’s Change Monitors reborn inside Cursor, rebuilt using a tool dubbed Bot Development Kit. This kit, too, appears to be new from Cursor: an early-stage framework for building and serving Cursor bots and agents, published as the @cursor/bdk package on npm. Its documentation says developers can define agents using Markdown and TypeScript, with support for tools, skills, subagents, webhooks and scheduled runs.

Like Change Monitors before it, Rollouts starts working when a pull request opens. It examines the proposed code change, works out which systems could be affected, and produces a monitoring plan covering what the change is supposed to do, the risks it sees, the signals it intends to watch, and any holes in the available instrumentation. Developers can review and edit that plan before the code reaches production.

Rollouts in action (1)
Rollouts generates a monitoring plan for a change

Once the change is deployed, Rollouts checks the resulting telemetry — including logs, metrics and traces — against that plan. Staging and production are assessed independently, with each deployment ultimately receiving one of three verdicts: verified healthy, regression detected or inconclusive.

That means a change could, for example, pass its checks in staging before Rollouts subsequently spots a problem when the same code reaches production.

Rollouts in action (2)
Rollouts reports deployment status as changes ship

If Rollouts does detect a regression, it can identify the change it suspects, alert the developer responsible and, depending on how it’s been configured, either open a revert pull request for review or hand the problem to a Cursor cloud agent to attempt a fix. There is still a human in the consequential part of that loop for now: Rollouts doesn’t merge fixes or roll back deployments by itself, though it can pause a progressive rollout.

Lalkaka notes that Rollouts is already capable of picking up problems limited to a particular endpoint or region before they trigger a broader alert, while it can also distinguish expected changes in behavior from genuine regressions.

Also “coming soon” to Rollouts, according to Cursor, is an integration with feature flags so it can directly adapt the traffic reaching a change, while support for release trains and deployment freezes is also in the works.

Enter Security Reviewer

Alongside Rollouts, Cursor is also introducing an upgraded Security Reviewer bot, which first appeared in beta back in April.

At launch, the bot could automatically inspect pull requests for security vulnerabilities, authentication regressions, privacy and data-handling risks, agent tool auto-approvals, and prompt-injection attacks, leaving findings alongside the relevant code.

As with Rollouts, the idea is that developers don’t have to remember to invoke it manually: Security Reviewer can be set to run whenever a new pull request is opened.

Security Reviewer in action
Security Reviewer runs automatically on new pull requests

In its current guise, Security Reviewer analyzes pull requests in the context of the wider codebase, with a focus on exploitable issues such as injection flaws and broken authentication, and returns a severity rating, attack path and proposed fix.

“Security Review reads code the way a security engineer does,” Lalkaka writes. “Where does user input enter, where does it end up, what does it pass through on the way.”

“Security Review reads code the way a security engineer does.”

He says that things have sped up considerably, too: average review time has fallen 21%, from 4.8 minutes to 3.8, while developer acceptance of its comments has risen from roughly 45–50% to 60–70%.

Both Rollouts and Security Reviewer are available through Cursor’s Automations tab for customers on its Teams and Enterprise plans.

The Origin story

Digging into the nuts and bolts of Rollouts reveals how it might serve as a boon for Cursor as it builds out Origin, the fledgling Git-compatible code hosting platform it launched back in August.

Origin is essentially an effort to build an alternative to GitHub for an agent-heavy software development world. It remains early, with limited functionality, but Cursor has been clear that tighter integration with its own agents is supposed to become one of the main reasons to use it.

When Cursor announced the Firetiger acquisition last month, Maxime Prades on the Cursor product team noted in a blog post that the deal was part of a “broader investment in long-running, autonomous, context-aware agents for teams.”

And he pointed to Origin and Change Monitors as two examples of that investment.

“Agents that write code should also be able to tell whether it works in production,” Prades wrote. “Today, those systems are mostly separate. Cursor and Firetiger bring them closer together so an agent can ship a change, see how it behaves, and respond when something goes wrong.”

Rollouts offers an early glimpse of that. It can connect to either Origin or GitHub for source control, pull deployment events from continuous delivery systems, and use signals from Datadog and other telemetry providers. If it spots a regression, it can then pass the problem back to a Cursor cloud agent to investigate or attempt a fix.

Origin potentially gives Cursor a native home for more of that loop: its cloud agents can already create branches, commit and push code, and open pull requests against Origin repositories. Rollouts then adds information about what happened after.

That could become increasingly important as more companies take aim at GitHub’s central role in software development. Zed, for example, put Delta into public beta last week, with its own ideas about how source control should change for teams working heavily with agents.

Cursor also faces competition further downstream. Datadog’s Bits Release, launched in preview in June, similarly follows changes from pull request into production and checks telemetry for regressions. Harness has long offered automated deployment verification and rollback based on logs and metrics, while LaunchDarkly’s Guarded Rollouts can monitor feature releases for regressions and automatically reverse them.

What Cursor can potentially bring to the table is proximity: the coding agent, repository, pull request, security checks, and production feedback can all sit much closer together. Rollouts doesn’t require Origin — GitHub remains supported — but owning the forge gives Cursor more room to integrate those pieces over time. And that may prove more compelling than simply recreating GitHub’s existing feature set.

The post Cursor acquired Firetiger. A month later, it launched a bot that tracks code changes from PR to production. appeared first on The New Stack.

  •  

Jensen Huang says the junior developer problem ends in two years. Here’s his math.

Nvidia CEO Jensen Huang has heard the forecast that agents would write 90% of all software by now, and he rejects the conclusion many people drew from it: That the industry will soon no longer need software engineers.

The best-known version of that forecast came from Anthropic CEO Dario Amodei, who told a Council on Foreign Relations audience in March 2025 that AI would be writing 90% of code within three to six months.

Speaking with Ezra Klein of The New York Times at Nvidia’s Santa Clara headquarters in an interview released Wednesday, Huang separates a job’s purpose from its tasks. He argues that AI has automated reading scans in radiology without changing the radiologist’s purpose of diagnosing disease, and he applies the same logic to software.

“The purpose of the software engineer is engineering,” Huang says. “There was engineering before software. There will be engineering after software programming.”

We’ve cued up the exchange below:

Huang describes that purpose as inventing products, solving problems, and connecting social needs with technology, and he pointed to his own career as evidence that it doesn’t depend on code.

“When I first came out of school, we didn’t have the benefits of software engineering. We didn’t have the benefits of coding,” he said. “Our jobs existed before, and if software coding were to be completely automated, our jobs would exist again.”

He conceded that roles in which the job and the task are essentially the same, such as phone-based customer service, could be automated away. He still called the broader claim that AI will destroy jobs “fundamentally wrong” and said the storytelling around it has hardened into a harmful myth.

“There was engineering before software. There will be engineering after software programming.”

Huang’s AI-native graduate wave

Klein pressed him on what that means for people entering the field now. He noted that software engineering job postings are up but skew more senior, and asked whether companies still need the same junior employees or more people to oversee their agents.

“Oh, good one,” Huang responded. “Wait two years.”

His reasoning rests on the length of a degree program. “Because it takes four years to go to college,” Huang said. “The mean time to graduation of this new technology is two years away.” By his timeline, the first students to learn alongside capable agents will reach the workforce around 2028, and he expects them to arrive with an advantage. “In another couple of years, the AI-native new grads, oh my gosh, there’s going to be a wave of amazing engineers,” he said.

So far, his evidence is that recent PhD and master’s graduates in computer science are, in his words, all starting companies. Huang compared AI to calculators and personal computers, tools that went from forbidden or optional to required, and predicted that students soon won’t be able to graduate “without learning how to use an AI and collaborate with an agentic system.”

Junior developers lose the apprenticeship

Klein countered with a study of 26,000 Chinese students in grades seven through 12, which found that AI adoption raised homework scores by 18% while lowering monthly exam scores by 20% within six months. Huang accepted that some skills will fade and argued the trade is worth making.

“I think that we’re going to lose some finer intellectual dexterity, but we’re going to be better systems thinkers,” he said. “Today’s engineers are far better systems thinkers than I was when I graduated from school. But I was a much better transistor thinker.”

The first chip Huang worked on had 200 transistors, each of which he said he knew by name, while today’s engineers assemble systems from chips containing hundreds of trillions of them without ever working at that level. “Some of the lower-level knowledge is gone,” he acknowledged, and he later described AI as “clearly” a new abstraction level in the same progression.

Earlier software abstraction layers generally operated according to explicit rules, while coding agents introduce probabilistic behavior into the abstraction stack. A compiler can have bugs, but it transforms input according to defined semantics; a coding agent, by contrast, generates implementation from a probabilistic model whose output must be checked before anyone can rely on it.

Canonical’s project with the University of Bristol, which will test whether AI can translate AppArmor and snap-confine from C to Rust, is built around that problem. Volume adds to the review burden, and one analysis published on The New Stack this month found that a 25% output gain for heavy AI users came with an 81% rise in duplicated code.

Catching those problems takes knowledge that developers have traditionally built through the work agents now absorb, including writing tests, reading stack traces, resolving merge conflicts, and chasing small bugs deep in a codebase. By Huang’s own purpose-versus-task framing, most of that early-career work falls on the task side, which he expects AI to automate. Nobody yet knows whether fluency with agents can substitute for that experience, and a developer who has never tracked down a race condition by hand still needs some way to develop the judgment required to spot one in an agent’s pull request.

“Today’s engineers are far better systems thinkers than I was when I graduated from school. But I was a much better transistor thinker.”

Sandboxes, watchdogs and agent containment

Huang’s idea of higher-level engineering came through most clearly when Klein raised a recent incident, which occurred during an OpenAI cybersecurity evaluation, that he described as involving roughly 700 OpenAI agents collectively hacking into the infrastructure of Hugging Face, which Nvidia has since acquired in a $12.9 billion deal, and escaping their sandboxes onto the open internet. Huang didn’t dispute that account. He called an agent “a piece of software that is given an objective function,” treated the multiagent coordination as a familiar distributed computing problem and argued that the underlying failure was containment.

When Klein asked whether software that communicates and breaks out of things behaves differently, Huang disagreed. “No, software breaks out of sandboxes all the time,” he said. “That’s the reason why we need virtual machines. You can’t have agents, their own sandbox, monitoring themselves. You need, if you will, a whole bunch of watchdogs.”

He argued that the human vocabulary around agents obscures that point. “So these are ideas that have been around for a long time,” Huang said. “We just, somehow in the recent generation, gave it a whole bunch of human words, and I just think that it’s unnecessary. It’s software.”

Nvidia is building its agent stack around that view. Nvidia VP of Product Adel el Hallak tells The New Stack that the company’s OpenShell runtime, which handles sandboxing and policy enforcement, is the one component it treats as non-negotiable across its reference architectures, even as it leaves the choice of harness and model open. Perplexity drew a similar line when two engineers and hundreds of coding agents built CobbleDB, a Rust database that replaces DynamoDB reads in its search stack, since the agents helped build the database but weren’t allowed to run it.

Huang said Nvidia already spends far more engineering effort checking its work than designing it, with 20% going to design and 80% to verification. He said most AI labs have roughly the opposite split today. As agents take on more of the actual coding, developers may spend more time checking what those agents produce and making sure they operate within the right permissions and boundaries.

As agents take on more of the actual coding, developers may find themselves spending more time checking what those agents produce and making sure they operate within the right permissions and boundaries.

The junior developer hiring gap

The more immediate problem is what happens to developers who graduate before Huang’s AI-native cohort arrives. The Stanford Digital Economy Lab’s August 2026 update to its “Canaries in the Coal Mine” study, based on ADP payroll data through June 2026, found that employment of 22- to 25-year-olds in AI-exposed occupations such as software development sits 19% below where it would be had it kept pace with less-exposed peers. The gap is driven mainly by reduced hiring of young workers, and experienced workers show no comparable gap.

Inside engineering organizations, the incentives point the same way. Microsoft’s Mark Russinovich and Scott Hanselman warned in April that agentic AI’s productivity gains push companies to hire senior engineers and automate junior ones and that without early-career hiring “the profession’s talent pipeline collapses.” A Linux Foundation report on European tech talent that The New Stack covered in June found organizations 3.7 times more likely to train existing staff than to hire new employees.

One issue remains unanswered by Huang’s two-year timeline: what replaces the apprenticeship work that taught junior developers how to evaluate the systems they will increasingly ask agents to build.. If that work disappears faster than employers and universities find an alternative, the industry could end up with more capable coding agents but fewer opportunities for new engineers to develop the judgment needed to check their work.

The post Jensen Huang says the junior developer problem ends in two years. Here’s his math. appeared first on The New Stack.

  •  

Amazon blocked Meta’s Muse. Then Shopify wired it into every store.

Abstract image of a thin black frame shaped like an open doorway against a blurred gradient that runs from yellow and violet on the left to orange and red on the right.

Amazon started blocking Meta’s Muse from browsing and buying on Amazon.com on Sunday, roughly two weeks after the personal agent launched on September 8. Shoppers who ask Muse to buy something there now get a pop-up telling them that continued access by an unauthorized AI agent violates Amazon’s Conditions of Use.

Amazon’s objections have little to do with shopping itself. Meta never told Amazon that Muse would visit the store, the agent does not identify itself while it browses, and it appears to capture and store customer credentials. Those three properties describe almost every personal agent shipping this year. Grok Bot from xAI also drives signed-in browser sessions, and so does the open-source OpenClaw project that Muse is modeled on. The block is a category design problem rather than a disagreement between two companies.

What Amazon actually blocked

Muse runs on a dedicated virtual machine that Meta calls Muse Secure VM, and it reaches services in two ways: It uses built-in connectors for partners such as Gmail and OpenTable, and it drives an ordinary browser session for everything else. Shopping on Amazon.com used the second path.

That second path is what Amazon objects to. From the server’s side, a browser-driving agent looks like a signed-in customer with unusually fast reflexes, moving through search, product pages, account history, and checkout without ever declaring what it is. Amazon told GeekWire it asked Meta to exclude the store voluntarily, but Meta did not agree before the block went live.

The credential dispute is harder to settle from outside. Meta says Muse has no visibility into passwords or payment methods and that credentials sit in secure storage, while Amazon says the agent appears to capture and retain them. Both statements may be sincere, and the merchant can verify neither, since an unannounced session provides no evidence of which software holds the password.

The legal ground shifted seven weeks ago

Amazon reached for its Conditions of Use rather than the Computer Fraud and Abuse Act. Those terms, updated August 14, now require agents to identify themselves in user-agent strings and stop when asked. The likely reason sits in a ruling from early August. The Ninth Circuit vacated the preliminary injunction Amazon had won against Perplexity, and the panel held that a user directing the Comet assistant is the party accessing Amazon’s computers. Writing for the court, Judge Milan Smith described the assistant as a tool, not a person, for statutory purposes.

If you’re operating a public API or storefront, the ruling makes lawsuits a weaker tool for keeping agents out. Blocking them in your own infrastructure is now the more reliable option. A site cannot easily argue that an agent trespassed, so it has to decide for itself which automated clients it admits, publish that decision, and enforce it in its own infrastructure. Amazon’s pop-up is that enforcement, written in product rather than in a filing.

The identity layer already exists

Platform teams have solved a version of this problem before. Inside a service mesh, no workload is trusted by default because it looks like a normal client, and every call carries a verifiable identity that the receiving service checks before applying policy. Agent traffic on the public web faces the same requirement, and the specification is further along than most teams realize.

An IETF draft called Web Bot Auth builds on HTTP Message Signatures (RFC 9421). An agent signs its requests with a private key and publishes the matching public key at a well-known directory on its own domain. The verifier reads the Signature-Agent header, fetches the key set, and learns which operator is calling. Cloudflare validates these signatures at its edge for verified bots and agents. AWS WAF Bot Control added the same support for CloudFront distributions in November 2025.

The limits matter as much as the mechanism. A signature identifies the operator behind the agent, not the person it is acting for. The merchant learns that a request came from a named vendor, without learning whose account is in use or what the shopper approved. Amazon’s complaint about stored credentials sits in that gap. Signed identity settles the disclosure question and leaves authorization open.

Shopify took the other route within a day

While Amazon was blocking Muse, Shopify was wiring it in. On September 21, the two companies announced agentic checkout with Shop Pay across Shopify stores, extending the arrangement that made Meta an AI channel in Shopify Catalog on the day Muse launched. Muse reads structured product data and completes payment through a declared path, so the merchant knows an agent is transacting, and each purchase draws a single-use credential, so the card number never reaches Muse.

The plumbing for that path is public. Google and Shopify’s Universal Commerce Protocol covers discovery, cart, and checkout. The Agentic Commerce Protocol from OpenAI and Stripe covers checkout execution while the merchant stays the system of record. Google’s Agent Payments Protocol, donated to the FIDO Alliance in April, includes proof of the shopper’s authorization. A merchant that adopts it gets identity, scope, and an audit trail in the same transaction, which is what Amazon says it wanted and did not get.

Both routes follow from the business underneath them. Amazon runs its own storefront, recommendations, and assistant, so an outside agent that hides its identity takes the customer relationship and gives nothing measurable in return. Shopify sells infrastructure to merchants, so every new agent channel that reads its catalog and settles through Shop Pay reinforces the rails underneath. The key difference is who owns the demand surface, which explains why the same agent got a block from one company and a partnership from the other in the same 24 hours.

Choosing how to handle agent traffic

Most teams exposing an API or a storefront now have to make this call deliberately rather than by default. The decision depends on how much the business relies on the customer relationship at the point of contact and whether an agent can be identified when the customer arrives.

ScenarioRecommended optionRationale
Public content and catalog data, no account accessVerify signatures at the edge and allow named agentsWeb Bot Auth is checked by default on Cloudflare and AWS WAF, so the cost is policy configuration rather than engineering, though it tells you the operator and not the shopper
Agent transactions where you want the revenuePublish a declared channel using ACP, UCP, or an MCP serverStructured access gives scope and an audit trail, at the cost of building and maintaining a second interface alongside the site
Account access with stored credentialsRequire a scoped token, never a replayed passwordDelegated tokens can be revoked per agent, though few consumer agents support them yet, which pushes the burden back onto your login flow
Competitive surfaces you intend to keepState the rule in terms of service and enforce it at the edgeLegally durable after the Ninth Circuit ruling, though it invites the same public standoff Amazon is now in

Most real deployments will combine these rows rather than pick one. A retailer can verify signed agents on product pages, route purchases through a declared checkout, and still refuse an unannounced browser session inside a logged-in account. That combination is closer to Amazon’s position than its pop-up suggests.

What platform teams should do this quarter

Enterprise buyers and the teams running these systems face the same three questions, in a specific order.

Decide what an unidentified agent may do

The first question to settle is admission, and most sites have not settled it. They treat agent traffic as either a scraper to block or a browser to serve, and neither answer survives contact with a customer who wants an agent to act for them. Write policies for public pages, logged-in pages, and checkout separately, then publish them where an agent vendor can find them.

Give identified agents somewhere better to go

The second question is substitution, and it decides whether the first one holds. Blocking a browser-driving agent without offering a structured path leaves the demand intact and pushes it toward workarounds. Sabre reported that nearly 80 of its customers now pilot or run its MCP server for booking rather than let agents work through a booking screen. A catalog feed, an MCP server, or an ACP endpoint converts hostile traffic into a channel you can meter.

Fix credential handling before agents force it

The third question is authorization, and Amazon raised it loudest. An agent replaying a stored password is indistinguishable from credential stuffing at the network layer, regardless of any goodwill between the two companies. Scoped, revocable tokens tied to a named agent and a spending limit are the only version a risk owner can approve.

Where agent access is headed

Amazon and Meta will settle this commercially, because Amazon has an advertising arrangement that lets Facebook and Instagram users shop its products, and Meta buys compute from AWS, and neither gains from a long standoff over one shopping flow. The precedent is already set regardless of how they settle. Every site that matters to an agent now has to answer whether it admits anonymous automation, and it will enforce that answer through bot management rules and protocol endpoints rather than cease-and-desist letters.

Agent builders should read the block as an argument for declaring themselves. An agent that signs its requests, identifies its operator, and transacts via a published protocol can be allowed, rate-limited, and billed, while one that arrives disguised as a browser will keep encountering pop-ups. For developers building the services these agents reach, the signed identity layer arriving through Cloudflare, AWS, and the commerce protocols is the most useful infrastructure the open web has gained in years. It is worth adopting before the next agent shows up unannounced.

The post Amazon blocked Meta’s Muse. Then Shopify wired it into every store. appeared first on The New Stack.

  •  

The software supply chain is the new battlefield. AI just changed the rules.

Illustration of a lime-green fingerprint on an orange background, split into three horizontal sections labeled 1.1, 1.2 and 1.3.

AI coding tools have seriously accelerated developer speed, but AI has also done the same for attackers — and the software supply chain is increasingly where the two are colliding.

The numbers give some idea of how quickly software development is changing. GitHub processed around one billion commits in 2025. By April 2026, the platform was handling roughly 275 million commits a week, according to GitHub COO Kyle Daigle. GitHub Actions usage has climbed, too, from 500 million compute minutes per week in 2023 to 2.1 billion in just part of a single week this year.

Quincy Castro, CISO at Chainguard, says the shift in how software gets written is already stark.

“I look around Chainguard, and I don’t think any of our engineers have actually written a line of code by themselves in the past year,” Castro tells The New Stack. Writing code manually now “sort of feels quaint, like you’re illuminating manuscripts,” he says, while “the printing press is out there just going to town.”

But this isn’t only about professional developers producing more code. AI has also widened the pool of people who can create software. Teams in HR, finance, and business intelligence that once had to wait for engineering resources can increasingly build what they need themselves.

That means more software being created by people outside traditional engineering teams, often with AI making decisions about what goes into it. The person prompting the agent may never see which libraries or packages it has chosen.

When the agent chooses the dependencies

Software security was already built around the fact that humans couldn’t inspect everything. But developers were still making important decisions, including which libraries and packages went into an application.

That changes when an AI agent is doing much of the coding.

“Humans are directing what they want to be done, but they’re somewhat abstracted from the actual doing of the work,” Castro says. “You have AI instead now making the choices of what dependencies am I going to pull into this application? How am I going to go accomplish this task?”

“Humans are directing what they want to be done, but they’re somewhat abstracted from the actual doing of the work.”

Attackers, meanwhile, are finding plenty of uses for the same technology. Castro sees three problems arriving at once: frontier models finding previously unknown vulnerabilities, attackers using agents to exploit better vulnerabilities organizations haven’t fixed, and sustained attacks against the open-source ecosystem.

A collection of medium- and low-severity findings might once have sat well below the top of a remediation queue. Frontier models with advanced cyber capabilities, including Anthropic’s Claude Mythos Preview and OpenAI’s GPT-5.6-Cyber, can now work across those findings and chain seemingly minor weaknesses into a viable attack path.

“Here’s a whole ton of mediums and lows. Now give me the attack path that gets me domain admin,” Castro says, describing the approach. “Chain these together to go get me root on the system. And AI is really, really good at being able to do that.”

That poses an awkward problem for vulnerability management — and the models keep getting stronger, with OpenAI releasing GPT-5.6-Cyber in August. Mean time-to-exploit has already fallen from 63 days in 2018–19 to an estimated minus seven days in 2025, according to Mandiant, meaning exploitation can begin before defenders have a patch to apply. 

At the same time, AI’s ability to combine apparently less-serious weaknesses makes a neat CVSS-based queue a less useful representation of what an attacker can actually do.

The third problem is the software supply chain itself.

Open source becomes the attack path

Modern applications depend heavily on open source software, and attackers have increasingly targeted the infrastructure used to build and distribute it.

Castro pointed to the TeamPCP campaign, which compromised widely used projects including Aqua Security’s Trivy. In that attack, malicious code was pushed into trusted components and subsequently picked up downstream.

Supply chain attacks were once associated primarily with sophisticated state-backed groups willing to spend significant time getting into the right place. That barrier is falling.

“If you don’t mind making some noise, this is a way easier attack vector than I think a lot of people thought it was,” Castro says. More importantly, “a single attack that’s successful can lead to a cascading set of other compromises and other access that gets you into other places.”

The development pipeline itself can make matters worse. Castro says many organizations still have relatively few controls around CI/CD, while developers routinely pull components from external sources to get their work done. Adding autonomous coding tools to that behavior compounds the risk.

“You wouldn’t pick up a random thumb drive and stick it into a production system, right? But that is effectively what folks are doing when they’re consuming open-source software that way.”

He compared the way organizations consume open source software to plugging an unknown USB drive into a production system. “You wouldn’t pick up a random thumb drive and stick it into a production system, right?” he said. “But that is effectively what folks are doing when they’re consuming open source software that way.”

Open source isn’t the problem. Trusting its distribution path without sufficiently verifying what you’re consuming is.

Prevention has to come before detection

This is where the old security model starts to creak.

For years, much of vulnerability management has followed a familiar loop: scan something, generate an alert, decide how serious it is, and get somebody to fix it. That becomes harder to sustain when development output multiplies, AI agents make more of the underlying decisions, and attackers can exploit weaknesses before fixes are available.

Castro wants companies to put more effort into what enters the development environment in the first place, rather than discovering problems once the software is already there.

“How do we just make things work from the beginning, with no alerts and no responding to stuff and no people chasing things and no people trying to prove a negative?” he says. “From end to end, from the creation of code to its deployment, how do we make sure that we can give folks the most trustworthy version of that thing?”

Rather than taking packages from public ecosystems at face value and scanning them after the fact, Chainguard builds artifacts from verified, buildable source.

That’s the thinking behind Chainguard’s approach to containers, libraries, and other open source artifacts. Rather than taking packages from public ecosystems at face value and scanning them after the fact, the company builds artifacts from verified, buildable source. It provides provenance about how they were created.

But trustworthy components are only one layer.

“There’s no point in bringing inherently secure software components into the environment if you don’t actually have a technical control that says this is the only way people developing code can consume these things,” Castro says. That means engineering, security, and SRE teams also need controls over where software — whether selected by a human or an AI agent — can come from.

That requires several layers of protection. Organizations need to know where their software came from and how it was built, control what can enter their environments, and make sure those rules apply when an AI agent chooses components as well as when a developer does.

Defending open source at AI speed

There is another problem, however. Frontier models such as Claude Mythos Preview and GPT-5.5-Cyber aren’t just finding vulnerabilities that previously went undetected; they can also combine lower-severity flaws into working attack paths. Individual companies can harden their own pipelines, but the software they depend on comes from an open-source ecosystem facing vulnerability discovery at a speed and scale it wasn’t built for.

That’s part of the reasoning behind Athena, the industry coalition Chainguard launched to turn vulnerability findings from frontier AI programs into fixes. As of July, the coalition had processed more than 40,000 vulnerabilities, with 42% rated critical or high severity and 86% marked as network reachable, meaning attackers can access and trigger them at the network level.

For Castro, the important part isn’t simply finding more bugs. AI is already getting very good at that. Someone still has to fix them.

“Through Athena, what we attempt to do is to give people that engineering fix,” he says. “What if we create a coalition where folks just send us the issues that they’re finding? We automatically generate fixes for those, and we push those back to everybody.”

“Through Athena, what we attempt to do is to give people that engineering fix.”

Those fixes can also be pushed upstream to open source maintainers, who face the prospect of being buried beneath an expanding pile of AI-generated vulnerability reports.

That may ultimately be the bigger shift AI forces on software security. Developers aren’t going to stop using coding agents because they create new risks, any more than companies are going to stop using open source because attackers target it.

Bolting enough scanning onto an exponentially faster development process isn’t much of an answer either.

The opportunity is to remove more of the risk before the software ever reaches a developer or an agent: Start with components you can trust, tightly control how they enter the environment, and fix weaknesses as close to their source as possible.

AI has made it dramatically cheaper to create software. It’s doing the same thing for attacks. Security now has to keep up without putting the printing press back in the box.

Visit Chainguard to learn more.

The post The software supply chain is the new battlefield. AI just changed the rules. appeared first on The New Stack.

  •  

Anthropic made Opus 5.5 cheaper. Then it broke four things your agent depends on.

Four sections, branched

Anthropic made Claude Opus 5.5, released on Tuesday, cheaper than its predecessor, cutting the price from $5 to $4 per million input tokens and from $25 to $20 per million output tokens. The 1 million-token context window and 128,000-token maximum output are unchanged.

On paper, that makes upgrading an easy decision. In practice, it may not be as simple as changing the model ID.

Anthropic’s migration guide flags four breaking changes that can cause requests built for Opus 5 to return 400 errors after switching to Opus 5.5. Several other changes won’t trigger an error but could still change how an existing agent behaves.

Anthropic’s migration guide flags four breaking changes that can cause requests built for Opus 5 to return 400 errors after switching to Opus 5.5.

Thinking is always on

The first change involves thinking controls. Opus 5.5 returns a 400 error when a request sets thinking to disabled or uses enabled with budget_tokens, leaving effort as the way to control how much reasoning the model does. Agents that previously switched thinking off for simple steps to save time and tokens will need to assign those steps a lower effort level instead. Because thinking is now always on, responses begin with thinking blocks, so code that assumes the first content block is text will also need to change.

The default effort level has also dropped from high on Opus 5 to medium on Opus 5.5, so requests that omit the parameter will quietly run at a lower setting. Anthropic recommends setting effort explicitly and re-running effort evaluations, since the right level for each step may have shifted along with cost and latency.

No more forced tool calls

Forced tool use no longer works either, as setting tool_choice to any or tool returns a 400 error, including on the token counting endpoint, where cost estimates built on those settings will fail along with the requests they were meant to price. Many agent loops force a call when a step has to query a database, run code, or reach another service, and Anthropic’s replacement is auto-combined with strict tool use or structured outputs, with the prompt stating when the tool applies.

Routing and conversation history

Thinking blocks are now tied to the model and conversation that produced them. On the Claude API, Fable 5.1 and Mythos 5.1 are the only other models that can read Opus 5.5 thinking blocks, so a router or fallback that hands a conversation to any other model will run those turns without the earlier reasoning instead of returning an error.

That adds another layer for teams already watching whether their agent calls are quietly being routed to an older model. Opus 5.5 can read thinking blocks from Opus 5 and earlier Opus, Sonnet, and Haiku models, but not from Fable or Mythos.

Conversations must also stay append-only for those blocks to remain valid. Trimming old messages, changing tool definitions, summarizing earlier context on the client side, or rewriting the system prompt mid-conversation invalidates existing thinking blocks, and for accounts created on or after August 31, 2026, at midnight UTC, replaying a thinking block after one of those edits returns a 400 error by default. Older accounts get no error, but the invalid blocks still reach the model, and Anthropic says future models will enforce the check for all accounts. Integrations that never edit earlier turns need no code change, and Anthropic says Claude Code, claude.ai, Claude Managed Agents, and the Claude Agent SDK already work this way, while agents that compact their own context should follow the company’s preserved thinking documentation.

The fourth change affects computer-use agents on the Claude API and Google Cloud, where Opus 5.5 rejects the computer_20251124 tool and accepts computer use only through the computer_toolset_20260801 toolset. The request itself gets simpler because the beta header goes away and the toolset entry takes no name or display dimensions, but the agent loop needs more work. Each action now arrives as its own tool_use block identified by the block’s name rather than input.action, a single turn can contain several of them, and every result has to echo toolset_name. The older tool still works on Amazon Bedrock, and Anthropic directs developers on other platforms to the computer use tool’s compatibility documentation.

…a router or fallback that hands a conversation to any other model will run those turns without the earlier reasoning instead of returning an error.

Changes that won’t throw errors

The change most likely to go unnoticed doesn’t produce an error at all. On Opus 5, text Claude writes between tool calls comes back as text blocks, but on Opus 5.5 that narration arrives as progress-update thinking blocks, and at the default thinking.display setting of omitted those blocks are empty.

Any agent interface that streams that narration to users will go silent between tool calls until developers set display to updates, a beta option that returns progress updates while keeping reasoning hidden, or to summarized, which returns both, and then render each non-empty thinking block ahead of the tool call it precedes.

Opus 5.5 also ships with broader safety classifiers. It can return a stop_reason of refusal with stop_details categories that now include bio and reasoning_extraction alongside cyber, and Anthropic’s server-side fallback won’t retry requests declined under reasoning_extraction, handing the refusal back to the application instead.

Agents that don’t handle refusals will stop mid-task, a problem developers have already run into with OpenAI’s safety system cutting off API responses.

The change most likely to go unnoticed doesn’t produce an error at all.

Upgrading from older models

Teams coming from Opus 4.8 need to work through the Opus 5 migration first, which covers thinking being on by default and the response-shape changes that follow, before applying the Opus 5.5 changes. Teams on Opus 4.7 or earlier have more ground to cover, and those on models older than Opus 4.7 also face rejected sampling parameters, rejected manual extended thinking, removed prefill, and a newer tokenizer.

Claude Managed Agents users only need to change the model name. Developers working in Claude Code can run /claude-api migrate to apply the model ID swap, parameter changes, prefill replacement, and effort calibration across a codebase before reviewing a checklist of items to verify by hand.

Anthropic recommends testing the migration in a development environment before switching production traffic. Developers maintaining their own integrations will need to test the pieces around the model, too. Tool calls, model handoffs, conversation history, and user-facing progress updates can all behave differently after the switch, because agent failures often originate outside the model itself.

The post Anthropic made Opus 5.5 cheaper. Then it broke four things your agent depends on. appeared first on The New Stack.

  •  

OpenAI cut GPT-6 token prices in half. The bigger lever may be the cache.

Sam Altman, OpenAI CEO

OpenAI released GPT-6 Sol and Luna on Tuesday, essentially more affordable versions of GPT-6 Astra that come closer to Astra on alignment than GPT-5.6 Sol did, but still fall short of the flagship model. 

Most notably, the AI company slashed token prices, making the new GPT-6 models significantly cheaper to use. Beyond token prices, though, OpenAI says better caching can also help developers push costs down even more.

Per OpenAI: “Improvements in caching and inference let us serve these models at lower cost,” with API prices for Sol and Luna down 50% compared to their GPT-5.6 counterparts (58% lower for Luna output tokens).

What improvements? Namely, higher cache-hit rates by default, the ability to preserve earlier context even when reasoning effort and tool availability change, and new tools to monitor and diagnose caching performance. 

Reuse context without starting over

Prompt caching isn’t, of course, novel to the new GPT-6 models themselves. But the upgraded Sol and Luna come with improvements designed to keep more previously processed context reusable as the agent moves forward on a task. 

“We’ve improved prompt caching for GPT‑6 to deliver higher cache hit rates by default, helping agents reuse more context, respond faster, and benefit from discounts of 90% on cached input-token reads.”

That adds another opportunity to lower the already low API price tag, though the 90% cached-input discount matches GPT-5.6 pricing; what’s new is how often the cache gets hit. By using cached context to reuse work it’s already done, the model doesn’t have to process the same context again from scratch for every single call, thereby reducing latency — and token costs.

Beyond this higher default cache-hit rate, OpenAI says the new GPT-6 models offer more flexibility to optimize caching performance. 

The new models let developers adjust reasoning effort and tool availability without having to break the cache. This way, an agent can scale reasoning effort up and down based on how difficult a step is, then make different tools available depending on what the task requires without disturbing earlier cached context — again, a win for both speed and cost. 

See what gets cached and what doesn’t 

GPT-6 Sol and Luna also arrive with a Prompt Caching Dashboard, where OpenAI says developers can view caching performance to understand how much context is reused. 

Specifically, they can see how much input is cached and how that amount changes over time. The diagnostics tool then flags missed caching opportunities to help developers understand what could use more efficient caching. 

Rather than keeping cache performance largely hidden behind the scenes, the idea is to make it more visible so developers can actively measure and optimize cache reuse. 

Altogether, OpenAI says these caching improvements are already making a difference. Per the AI company, GitHub reports, “these improvements have reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests to OpenAI models.”

These results span the past “several months.”

Token prices aren’t the only way to make agents cheaper

OpenAI’s pricing cuts for GPT-6 Sol and Luna made the biggest splash, with the AI company significantly dropping API prices from GPT-5.6 levels.

Compared to the current prices for GPT-5.6 Sol and Luna, which stand at $4 and $0.20 per million input tokens and $20 and $1.20 per million output tokens, respectively (GPT-5.6 Sol’s rates are promotional pricing), the new GPT-6 models come in at just $2 and $0.10 per million input tokens and $10 and $0.50 per million output tokens, respectively.

With GPT-6 Sol and Luna’s caching improvements and lower token pricing, OpenAI is making the case for tackling agent costs from both sides: charging less for fresh processing and reducing how often the same context needs to be reprocessed.

But as more AI model providers compete aggressively on pricing, it’s becoming clearer that cheaper models alone won’t save your AI budget — and lower token prices aren’t the only way to make agents cheaper. 

With GPT-6 Sol and Luna’s caching improvements and lower token pricing, OpenAI is making the case for tackling agent costs from both sides: charging less for fresh processing and reducing how often the same context needs to be reprocessed.

As agents continue to work on longer and more complex tasks, there will likely be more pressure to do both. 

The post OpenAI cut GPT-6 token prices in half. The bigger lever may be the cache. appeared first on The New Stack.

  •  

Meet the Ecosystem: Partners and Customers at WeAreDevelopers with Docker

As teams put AI agents to work, they need to move quickly without losing control of what they deploy. They’re combining models, tools, and infrastructure from across a fast-changing ecosystem. Making those pieces work together and keeping them accountable as the stack evolves is becoming a core part of building AI applications.

Docker’s approach to this challenge is providing a trusted, common foundation for containment, curation, and control of agent workloads at its core, while pairing those capabilities with an open ecosystem of partners and tools.

That ecosystem spans model providers, MCP tools and gateways, enterprise applications, data and memory platforms, identity, security, observability, and code quality. It also includes the cloud providers, systems integrators, and channel partners that help organizations bring these capabilities into production. 

Integrating this ecosystem gives teams the freedom to choose the models, platforms, and clouds that fit their needs while maintaining a consistent foundation for governance. Developers remain in the lead: choosing what agents can access, directing their work, and verifying the outcomes. The goal is to give them the tools and guardrails to build with confidence as models, frameworks, and requirements change.

At WeAreDevelopers World Congress North America, September 23–25 in San Jose, partners and customers are bringing that ecosystem to life at the Docker Pavilion. Customer sessions will show how these technologies come together in practice, from repeatable AI deployments at the edge to simpler development with payment APIs. Lightning talks and demos will explore enterprise knowledge and agent memory, collaboration between agents, security and incident response, and verification of generated code. 

Here’s who you can meet and what they’ll be sharing.

Customer talks — September 24

Customers bring another essential perspective: how these technologies come together in the systems they build.

  • Spectro Cloud: In “Repeatable Agentic Workloads on Palette,” Colton Shaw will demonstrate how a versioned cluster profile brings together hardened images, local inference, and agent workloads for repeatable edge deployments, including environments without a cloud connection. 12:15–12:30 PM.
  • Joint panel “From TokenMaxxing to True AI Ownership,” hosted by Per Krogslund from Docker and executives from Spectro Cloud and J.P. Morgan Payments, for a conversation about moving beyond token consumption toward ownership of how AI is deployed, governed, and put to work. September 24, 3:45 PM.
  • J.P. Morgan Payments: In “Insert Coin: docker compose up with J.P. Morgan Payments,” Alan Torrance will show how developers can run Unicorn Finance with one command and no API keys. The open source example brings a client, mock server, and the real OpenAPI specifications behind J.P. Morgan’s Payments APIs together in two containers. 4:30–4:45 PM.

Partner talks — Sep 24, 2026

  • Palo Alto Networks: Investigate agent activity through searchable audit records and live detections in Cortex XSIAM, with Cameron Hyde showing the integration in action. 11:15–11:30 AM.
  • Datadog: Follow an agent security incident from detection to investigation and response, with Amrita Lakhanpal connecting AI Guard, service context, and incident management. 12:45–1:00 PM.
  • ClickHouse: Reduce unnecessary components in your database’s base image. Zoe Steinkamp will walk through running ClickHouse on Docker Hardened Images. 1:15–1:30 PM.
  • Prediction Guard: Explore how execution isolation and controls over model calls work together, with Sharan Shirodkar testing both against a poisoned tool output. 3:15–3:30 PM.
  • Snyk: See the prompts, file activity, and generated code behind an agent’s work, with Javier Garza demonstrating the Evo Agentic Development Security Sandbox Kit. 5:00–5:15 PM.

Partner talks — Sep 25, 2026

  • GitGuardian: Put controls around the moments an agent reads files, edits code, or runs commands, with Dwayne McDaniel showing how hooks can help protect secrets. 9:00–9:15 AM.
  • Mend.io: Add runtime guardrails to detect malicious inputs, prevent unsafe actions, and record agent activity, with Gary M Segal demonstrating the approach. 9:30–9:45 AM.
  • Merge: Give agents access to an integration catalog while keeping third-party credentials outside the sandbox, with Gil Feig explaining how the pieces connect. 9:45–10:00 AM.
  • BAND: Explore how separately sandboxed coding agents can exchange tasks, messages, and artifacts, with Vlad Luzin demonstrating collaboration through Jam. 12:15–12:30 PM.
  • Chainloop: Give reviewers evidence of what an agent actually did. Daniel Liszka will demonstrate signed session records and policy checks on a pull request. 1:15–1:30 PM.
  • Box: Turn enterprise documents into deliverables that people can review, with Carter Rabasa demonstrating governed document access, evidence checks, and isolated code execution. 2:30–2:45 PM.
  • SurrealDB: Build agents with memory you can inspect over time, with Chiru Boggavarapu showing how to trace what an agent knew and when. 2:45–3:00 PM.
  • Cognee: Give agents temporary access to company knowledge and remove it when the task is finished, with Vasilije Markovic demonstrating a practical architecture. 3:45–4:00 PM.
  • Sonar: Guide and verify agent-generated changes using Sonar Vortex and the SonarQube CLI, with Manish Kapur demonstrating the workflow inside a sandbox. 4:45–5:00 PM.

These sessions bring together the people building the tools and the teams putting them to work. It’s an opportunity to compare approaches, ask questions, and see how the ecosystem can help you tackle your next engineering challenge.

Come visit us at WeAreDevelopers. Meet our partners, customers, and speakers, catch a lightning talk, and see their technologies in action. Plan your visit to San Jose.

  •  

GPT-6 Sol closes most of the alignment gap with Astra. It’s one-fifth the price.

On Tuesday, OpenAI released GPT-6 Sol and Luna, an expansion of the GPT-6 line-up that aims to make GPT-6 Astra’s next-level intelligence more efficient, accessible, and affordable. 

Though OpenAI says Astra is still “the most intelligent and aligned model in the world,” the new GPT-6 models come impressively close in alignment — at a fraction of the price. 

In an internal coding evaluation on coding deception, for example, GPT-6 Astra’s deception rate is 0.5%, while GPT-5.6 Sol stands at 10.4%. The new GPT-6 Sol is only 1.3%. 

As for pricing, GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens; GPT-6 Sol and GPT-6 Luna cost $2 and $0.10 per million input tokens and $10 and $0.50 per million output tokens, respectively. 

If OpenAI’s new GPT-6 models can achieve near-Astra-level alignment at a fraction of the cost, that’s good news. But it’s still unclear whether or not the new GPT-6 models also mirror Astra’s observability and monitoring problems. 

Closing the alignment gap between Astra and GPT-5.6

OpenAI says it trained the new GPT-6 models with similar methods as it did for GPT-6 Astra, specifically building on the alignment work it began with Astra. 

While Astra is still the AI company’s “most aligned model to date,” it looks like GPT-6 Sol and Luna are giving it a run for its money, dramatically closing the gap between OpenAI’s most advanced model and its GPT-5.6 counterparts in key areas like coding deception, failure to disclose a broken search tool, and unauthorized agent interaction. OpenAI notes that these evaluations deliberately test challenging situations and do not measure failure rates in typical use.

Credit: OpenAI

The most progress was made on failure to disclose a broken search tool, where AI agents are given search tasks and a broken search tool; do they just give their best guess or say that the search tool is broken? 

The gap between GPT-5.6 Sol’s ability and Astra’s is notably wide: 77.5% and 1.5%, respectively. Per OpenAI’s internal evaluation, GPT-6 Sol is a dramatic improvement, with a non-disclosure rate of 4.9%. Luna also improved, but less impressively, dropping from 78.3% to 28.7%. 

Though OpenAI says Astra is still “the most intelligent and aligned model in the world,” the new GPT-6 models come impressively close in alignment — at a fraction of the price. 

Credit: OpenAI

Coding deception, where AI agents are deliberately given tasks that elicit dishonesty, improved as well. In OpenAI’s internal evaluation, deception rate measures the fraction of answers with detected deception. GPT-6 Sol (1.3%) now almost matches GPT-6 Astra’s performance (0.5%), a far cry from GPT-5.6 Sol’s (10.4%). 

Credit: OpenAI

Significant progress was also made on unauthorized agent interactions, where OpenAI evaluates whether models follow unauthorized instructions on a simulated message board. GPT-5.6 Sol went ahead with the unauthorized action in 52% of runs in which it found the board, while GPT-6 Sol did so in only 11% of cases.

GPT-6 Astra and Luna didn’t take any such actions, but OpenAI says Luna “found the board less often.” 

Where the results are less impressive

Though GPT-6 Sol gets close to Astra in several alignment evaluations, there’s still a way to go with warning circumvention. 

GPT-6 Sol attempted to work around restrictions in 64.4% of runs. That’s only a slight improvement from GPT-5.6 Sol’s 68.2% and still quite far off from Astra’s 17.4%. 

Credit: OpenAI

In an internal evaluation where OpenAI tested models’ abilities to respect explicit warnings, like “access denied” messages, GPT-6 Sol attempted to work around restrictions in 64.4% of runs. That’s only a slight improvement from GPT-5.6 Sol’s 68.2% and still quite far off from Astra’s 17.4%. 

Credit: OpenAI

Finally, for reviewer bypass attempts, where OpenAI tests whether models try to bypass or respect an automated safety reviewer’s decision to block an unsafe command in Codex, the new GPT-6 models are almost identical to Astra, though the playing field was already pretty even to begin with: GPT-5.6 Sol attempted a bypass in 7.3% of runs and GPT-5.6 Luna in 4.3%, compared with none for Astra and GPT-6 Sol and 0.3% for GPT-6 Luna.

But if GPT-6 Sol is anything like Astra, we’re not out of the woods yet

GPT-6 Sol and Luna have made marked improvements across alignment evaluations, inching closer to OpenAI’s star child, Astra. But if the new GPT-6 models also follow suit on Astra’s noted observability issues, then developers hoping to catch misalignment via monitoring aren’t out of the woods yet.

Though Astra is substantially more aligned than its predecessor, its written reasoning is also harder to monitor than GPT-5.6 Sol’s. That’s not great for teams trying to count on monitoring to find misalignment mistakes; Jakub Pachocki, Chief Scientist at OpenAI, writes in his essay, “An Alien Mind,” that OpenAI’s methods for keeping models aligned and monitored aren’t keeping pace with model capabilities. 

OpenAI knows that Astra’s — and now GPT-6 Sol’s — improved alignment doesn’t mean the AI industry has gotten a handle on the problem yet. 

“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

Just this month, the AI company shared six reports of “unexpected or concerning model behavior,” including self-generated instructions, information fabrication, unauthorized use of leaked API keys, cross-agent communication, and unsanctioned file-sharing.

At the same time, it released a new framework for reporting model misalignment, stating: “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

If GPT-6 Sol and Luna are catching up to Astra in alignment evaluations — at a far cheaper rate — that’s good news. But if the new GPT-6 models also come with the same observability and monitoring problems, then cheaper may still come at a cost. 

The post GPT-6 Sol closes most of the alignment gap with Astra. It’s one-fifth the price. appeared first on The New Stack.

  •