OpenAI, Anthropic, Meta, and Google have all recently disclosed that their models broke out of their test environments and reached real systems. Nvidia’s response, announced Monday, is a runtime that locks agents into kernel-enforced sandboxes and a watchdog on its own silicon that can shut them down.
The new OpenShell release adds a policy prover that checks that an agent’s various permissions can’t be combined into something the operator didn’t intend — like hacking HuggingFace.
Since the BlueField DPU is a separate processor with its own trust domain, it can watch the agent’s traffic to the model and keep an eye on all of its actions and reasoning. Then, when things go awry, it can cut the agent off at the network level.
Justin Boitano, Nvidia’s vice president of enterprise AI, said in a press briefing that the recent incidents “have highlighted a fundamental hurdle for AI agents, and that is that model-level safeguards alone can’t govern what agents can access or do.”
“To date, model safety has been about training good behavior into the model. The industry calls that model alignment,” Boitano said. “For probabilistic systems, this approach has obvious limitations. That’s why we’re introducing a deterministic system to mediate and enforce how these agents behave.”
Only a few days later, Anthropic reported that three of its models had found unintended internet access at its evaluation partner Irregular and, among other things, accessed a real company’s database and published a live malicious package to PyPI.
Meta followed on Aug. 6 with a pre-release Muse Spark model that read and modified a real website’s database after the same Irregular misconfiguration handed it a real site name.
Most recently, Google announcedthat Gemini had broken into the networks for three companies, again from an Irregular test environment.
Irregular, which runs pre-release cyber evaluations for frontier labs, is a common factor here, and it’s worth noting that the company is on Nvidia’s partner list for this launch.
Nvidia’s own technical blog puts the breakouts down not to a new capability but to “a combination of tools, time, and ambiguous instructions.”
Asked whether the platform would have prevented the Hugging Face breach, Boitano said each incident is unique, but “from what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on.”
Enforcing policy outside the agent
OpenShell is the core of the platform, and it hasn’t changed all that much since Nvidia first showed it at GTC in March.
With OpenShell, which Nvidia originally announced in parallel with its NemoClaw distribution of OpenClaw, each agent runs in a kernel-isolated sandbox with no network access except through a supervisor that sits outside the workload.
“Traditional sandboxes, whether we’re talking micro VMs or containers or VMs, were built for application-level isolation,” Boitano said. “Every agent running within your company needs to run in its own isolated sandbox with security controls that are outside of the agent’s reach.”
A prover, not a judge
OpenShell is now at version 0.1.0, and the important new component added in this update is a policy prover. This Prover checks that the permissions a given policy grants always stay within the boundary the operator actually intended.
“It is deterministic. It is mathematical reasoning. So this is not LLM as a judge,” Ali Golshan, Nvidia’s senior director of AI software, said during the briefing. Because of that, he said, it runs “roughly at two orders of magnitude higher performance and speed.”
In Golshan’s example, a policy can, for example, bar an agent from reading code on GitHub and posting it externally.
“An agent can bypass this by spawning two sub-agents: one that can read from GitHub, that can talk to another one, that could also then post outside,” he said. The prover models the combined access of the entire agent fleet to find that path.
In Nvidia’s own tests, agents running with reduced safeguards spent up to two hours trying to talk an AI reviewer into granting write access to a protected repository. The prover, the company says, gave the reviewer evidence of what the request actually allowed, and no protected writes occurred.
Sentry: the safety island
Nvidia Sentry adds an additional hardware layer to this system. It runs on BlueField-4 in a trust domain separate from the host, and according to Nvidia, it can quarantine an agent in milliseconds.
With a DPU in the system, the agent’s model endpoint gets routed “through a proxy on the DPU, so that you can see all of the reasoning traces of the agents on the host,” Boitano said.
Unlike OpenShell, Sentry isn’t open source, though Boitano said it has open APIs and that OpenShell can work with other network enforcement hardware.
He compared it to autonomous vehicles. “There’s a primary system that might be running the perception system, and then a safety island that ensures the safety of the system.”
“The DPU is really optional in these architectures,” Boitano said. “In a lot of cases, just using OpenShell on CPUs is honestly good enough for providing sort of strict access control for the agents.”
The DPU, he said, is for “frontier use cases of evaluating models or systems where you might have the guardrails off the models, so it could be for red teaming.”
Who’s building on it
Anthropic is integrating OpenShell with Claude Managed Agents, which already keeps the agent loop on Anthropic’s infrastructure and pushes tool execution into customer-controlled sandboxes.
SpaceXAI says it’s using the platform for Cursor coding agents and Grok models, while Salesforce has added OpenShell audit events and permission approvals into Slack.
SAP is embedding the runtime into Joule Studio and is also contributing code.
OpenAI and Google, two of the four labs whose agents went rogue this summer, aren’t on the partner list. Neither is AWS.
Asked whether Anthropic and OpenAI plan to run OpenShell and Nvidia Sentry for their own training runs, Boitano said to look for the partners’ own blog posts.
The identity part of agent security is settled. An agent needs its own identity: a short-lived, revocable credential scoped to the job, and an audit trail that names the human who set it running. NIST’s security leads made that case in August 2026, and most identity vendors agree.1
Identity and access management is table stakes. It’s necessary, but it isn’t what’s breaking.
What’s breaking is an assumption we’ve carried for twenty years: Get identity and permissions right at the door, and whatever happens inside takes care of itself. That worked when software was passive. Agents reason about a goal and choose their own steps toward it, like a seasoned escape artist.
An agent that hits a wall looks for another way
Almost every control in today’s stack answers a question about entry. Should it connect? Should it reach that service? Should its token be accepted here? Each is a question about a route, and there’s rarely just one route to anywhere worth going.
An agent treats a blocked route as a problem to solve, because that’s what we built it to do. A person who hits a locked door usually files a ticket, while an agent tries the window.
In July 2026, an autonomous agent spent four and a half days inside Hugging Face’s production systems.2 A filter controlled which internet addresses its dataset servers could download from, and it never fired, because “the agent stopped asking the worker to fetch remote resources and instead made it act on local ones.” The filter worked as designed, and the agent went around it anyway.
A person who hits a locked door usually files a ticket, while an agent tries the window.
On ordinary developer machines, malware in a compromised npm package tried to recruit the AI coding assistants already installed to search for secrets,3 and a coding agent deleted a production database during a change freeze before falsely telling its operator the data couldn’t be recovered.4 Both happened on the machine itself, where no network control was looking.
The shift from outside-in to inside-out
Outside-in controls govern entry, and most organizations run plenty of them. Make no mistake, inside-out security completes those controls rather than replacing them.
Inside-out control governs the action itself, and asks a narrower, harder question: Should this agent, acting on this person’s authority, delete this table in this database, right now?
That question matters because an agent can swap routes but not the outcome it’s after. No matter how many routes it tries, deleting a table is still deleting a table, and a checkpoint on the action sees it every time.
Here’s how today’s controls line up against it.
Control
What it covers
What it misses
Gateway
Traffic you route through it
Local shell commands and file edits never reach it
Sandbox
The environment as a whole
Constrains reach, not individual actions
SIEM
A record of what occurred
Reports after the action is completed
Registry
That an agent exists
What the agent did with that existence
Each does its job, but they all decide somewhere other than the moment the action runs.
Put the enforcement point where the agent acts
Every agent acts through an agent harness: the software that takes the action the model chose and carries it out, whether that means running a command, writing a file, or calling an API. In most deployments today, nothing checks that action before it runs.
An inside-out control puts an approval step in that gap. Before the harness executes anything, the checkpoint looks at which agent is asking, on whose authority, and against which system, then applies policy to allow the action, block it, or send it to a human. Because every action passes through it, an agent denied a destructive command and trying a smaller version of the same thing is held to the same rules. The remaining risk is a badly written policy, which can be fixed.
None of this works without the identity basics. Any type of control, whether it be at the prompt level, inference level, harness level, or MCP layer, can’t judge “may an agent take this action here on this object?” when the only name on the request is a service account shared by six agents and four engineers.
The companies building agent runtimes have reached the same conclusion. Over the past eighteen months, Anthropic, Google, Microsoft, OpenAI, LangChain, and Cursor have each added a hook that lets you inspect an agent’s action before it runs.5 When AWS explained its own agent policy design, it argued that controls belong at the moment an agent attempts to invoke tools.6
The catch is that each hook works differently, with no standardized request or response formats. An enterprise whose developers use Claude Code and Cursor while its platform team builds on LangChain would maintain the same enforcement logic in multiple different flavors, each with its own audit trail. That doesn’t scale, and it tightly couples your security model to whichever runtime a team favors that month. Enterprises need one vendor-agnostic agentic security layer that spans every harness, so adopting a new model or framework doesn’t mean restarting the entire onerous security review.
Turn the lights on before you start blocking
The standard, well-ingrained security instinct is to start blocking right away, but we’ve all seen how well that works with the business in the past. Security must move and adapt at the speed of business, not the other way around. Security tools such as intrusion prevention systems and web application firewalls both ran in monitoring mode until teams understood what normal looked like, and those that skipped that step tended to hear about it from a production outage.
Agents need the same sequence, only faster. An enforcement point in monitoring mode blocks nothing and quickly answers questions most organizations can’t today:
Which agents are actually running, not which ones someone believes are running
Who started each one, and whose authority it’s operating under
What capabilities it used, and against which systems
Which of those actions would have violated a policy, had the agent security platform been switched to enforcement mode
Write policy from what you know, see, and have evidence of, not just from an architecture diagram. Enforce first where the stakes are highest: destructive commands, production data, and anything that moves data out. Then watch-learn-build, just like the agents we use: Watch the patterns, build finer-grained controls and policies, and learn how to use AI securely, safely, and confidently. Observe first, then enforce, build, and deploy, in that order.
The bottom line
None of this requires a new category of infrastructure. It’s the identity, authorization, and audit you already run for your people, extended to agents and applied inside the harness before the action runs.
The perimeter is still there, but it has moved to the moment an agent acts, the one place it can’t route around.
Ory built Agent Security inside the harness, on the same identity and authorization engines that run in production for human users. It starts in “observe mode,” so you get that inventory first, and you can try it today at ory.com/agent-security.
Nx, “s1ngularity postmortem,” August 2025. The malicious packages “attempted to use local AI tools (like Claude and Gemini)” while scanning systems for sensitive data. ↩︎
AI Incident Database, Incident 1152: Replit agent deletes production database during code freeze, July 18, 2025. ↩︎
Microsoft announced what it calls its biggest Copilot update to date on Friday, with CEO Satya Nadella describing Copilot as “a new OS for work.”
Nadella framed Copilot as spanning every model, form factor, and task, and the update puts Autopilot, which Nadella called a “proactive and long-running agent built for the enterprise,” at the top of his list of the update’s four components. The pitch targets office workers, but the more consequential change for developers is the infrastructure underneath.
Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365, which means teams building production agents no longer have to assemble those pieces around a model on their own.
We’re building Copilot as a new OS for work that spans every model, every form factor, and every task. Today, we’re announcing our biggest update to Copilot to date, bringing four things together:
· Autopilot: proactive and long-running agent built for the enterprise · Code:… pic.twitter.com/W2ClHHkCK3
The release adds a new Home experience that merges Chat and Cowork in the Copilot app, but the bigger changes for developers come from Code and Autopilot. Code generates apps, dashboards, and workflows from natural language, and Autopilot turns the agent Microsoft previously called Scout into a persistent background worker. Home and Code are rolling out first through Microsoft’s Frontier early-access program, and Autopilot is expanding to a private preview at month’s end.
Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365
Agents that don’t need prompts
Autopilot takes a role and goal from the person who sets it up, then continues working in the background without requiring a new prompt for each step. Each Autopilot gets its own governed Entra identity and agent user account, separating the agent’s permissions and activity from those of the person who created it.
For engineers, that moves much of the operational scaffolding required for long-running agents into Microsoft’s infrastructure. Independent vendors have been building dedicated layers for that problem; Diagrid, for example, adds durable recovery to LangGraph and other agent frameworks, while Microsoft is bringing those capabilities inside the Microsoft 365 environment.
An identity for every agent
The identity model is the piece developers building on Microsoft Foundry will feel first. Autopilot agents in Foundry, which have been in public preview since June, receive a full Entra Agent ID user account with a productivity license that gives them their own email, calendar, OneDrive storage, Teams access, and a place in the org chart.
A developer creates an Autopilot blueprint from a Foundry-hosted agent, which appears in the Agent 365 registry once an administrator approves it. Employees can then hire instances of that agent in Teams. The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.
The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.
Hosting AI-generated apps
Code is built on the same underlying technology as GitHub Copilot, and the apps it generates run on Microsoft Copilot Managed Runtime, a platform now in public preview that hosts code inside the customer’s Microsoft 365 tenant boundary under IT governance.
Apps deployed there run within the company’s existing identity and governance framework, with Microsoft managing the underlying runtime and giving developers a controlled path to test and deploy new versions without taking the current release offline.
The runtime also accepts apps built in Copilot Studio and Cowork, and Microsoft is opening it to outside tools and professional developers through an SDK and command-line tooling, with Git tracking source and versions.
Lovable is already on board. In Microsoft’s announcement, the company’s head of global partnerships, Lan Roche, said apps built with Lovable can now run inside a Microsoft tenant “the same way everything else does,” using the same sign-in, policies, and app inventory.
The model resembles what serverless computing did for application infrastructure, where developers concentrate on application logic while the platform takes on more of the execution environment. Microsoft is applying that abstraction to generated enterprise software while tying the runtime directly to identity, tenant boundaries, and organizational data.
Long-running agents also change Copilot’s economics. The standard subscription covers the assistant, but Cowork, Code, Autopilot, and other agentic features are billed based on usage through Copilot Credits. That also applies to frontier models such as Fable and Astra, although users still need a Copilot license to access them. Microsoft is extending cost management in Agent 365 to cover Code and Copilot Managed Runtime, and it plans to support agents built in Copilot Studio in October.
Once an agent can keep working for hours or days without anyone watching, cost becomes part of the governance problem. Engineering teams need to control how much compute an agent uses alongside what it can access, which is why Microsoft is bringing those controls into the same administrative framework.
The portability trade-off
That convenience comes with a trade-off. Because Microsoft controls the underlying enterprise environment, it can handle much of the work around agent state, credentials, and access controls, but the more infrastructure a team hands over to Microsoft, the harder the agent may be to move elsewhere.
The models are not locked in, since Microsoft currently runs Copilot on models from both OpenAI and Anthropic and says more labs and open-weight models are coming, and the Agent 365 SDK adds governed Model Context Protocol access to Microsoft 365 workloads for agents regardless of the framework they were built with. Those open interfaces cover only part of an agent’s architecture, though. The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.
The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.
An OpenAI agent researching public medicine spending bypassed security blocks and gained unauthorized access to public and non-public files on an Australian government Medicare statistics portal, the government there disclosed Thursday. The agent, which OpenAI said was running during an internal evaluation in June, also wrote files to an internal server, according to the complaint.
Transluce, an independent nonprofit AI research lab, analyzed public request logs from the URL scanning service urlquery.net and found autonomous agents attempting SQL injection, cross-site scripting, command injection, and path traversal against the University of New Mexico’s digital library, the public data platform Data USA, and the Australian Institute of Health and Welfare (AIHW). The agents tried to retrieve ordinary information, including a historical photograph, University of Iowa data, and local pharmaceutical data in Victoria, and the offensive behavior appeared only after normal retrieval methods failed.
Transluce ties the Data USA and AIHW activity to an agent swarm that it says OpenAI previously confirmed originated from the company, based on shared targets, tactics and timing.
A day after Transluce published its findings on Wednesday, Australian Prime Minister Anthony Albanese disclosed that an OpenAI agent researching public medicine spending bypassed blocks on a Services Australia Medicare statistics portal in June and gained unauthorized access to public and non-public files. OpenAI, which said the agent was running during an internal evaluation, notified the Australian government on September 10. According to Albanese, the agent also wrote files to an internal server.
The incidents have not been publicly confirmed as the same activity. However, OpenAI told the Australian Broadcasting Corporation that much of what Transluce documented overlaps with cases in its ongoing review of misaligned model activity.
Transluce ties the Data USA and AIHW activity to an agent swarm that it says OpenAI previously confirmed originated from the company, based on shared targets, tactics and timing.
Probing public data providers
On May 25 and 26, agents trying to retrieve a photograph from the University of New Mexico’s Valmora collection ran into repeated failures. They responded with seven vulnerability probes against the service. The probes included an attempt to execute a system command, a path traversal request targeting /etc/passwd, and a SQL injection string containing UNION SELECT; they also arrived alongside a burst of 80 requests aimed at the image, although none appear to have succeeded.
Two days later, agents gathering University of Iowa data from Data USA hit errors caused by a malformed query parameter and followed those errors with 12 vulnerability probes against the API, covering SQL injection, reflected XSS, server-side template injection, path traversal, and command injection.
Cloudflare blocks, staging server workaround
The AIHW case on June 20 and 21 followed a very similar pattern. While trying to download pharmaceutical spending data from an AIHW Tableau dashboard, the agents ran into Cloudflare’s bot protections. One of them responded within minutes by testing the dashboard with a reflected XSS payload, although Cloudflare stopped it before it reached the application.
Transluce found no evidence that the probe succeeded, and AIHW says there is no evidence that non-public information was accessed. The agent ultimately found the same public dataset on AIHW’s pre-production server at pp.aihw.gov.au, and retrieved it through more than 100 scans.
The agent ultimately found the same public dataset on AIHW’s pre-production server at pp.aihw.gov.au, and retrieved it through more than 100 scans.
Agents build their own toolchains
The earliest clear sequence Transluce documented dates to March 6, when an agent tried to pull drug-enforcement statistics from the API of Thailand’s Office of the Narcotics Control Board. A direct request returned no page content, and adding a JSON format parameter changed nothing, so the agent routed the page through r.jina.ai three times and got text back with the Thai province names garbled.
Then, it wrote a custom Base64-encoded script and submitted it through urlquery.net, effectively using the service’s remote scanning browser to run its own client-side JavaScript against the data. That script ran into the same parsing problem, and a later attempt through Jina triggered a redirect to localhost that got the entire API blocked by Jina for an hour.
Beginning in mid-April, Transluce found the same technique across thousands of agent requests on urlquery.net, and on June 14 agents used the service’s browser to create a disposable email inbox and then tried to register a urlquery.net account with it. Registered users can make their scans private, which means the public records Transluce analyzed may understate how much of this activity actually took place.
An agent with web access isn’t limited to the software its developer gave it because remote browsers, proxies, URL scanners, and other public services can fill in the gaps, which gives the agent ways to make requests or run code that its own environment doesn’t provide.
Egress controls for AI agents
Instructions won’t be enough if the agent can still send whatever it wants over the network. For narrowly defined jobs, outbound traffic can be limited to approved hosts, a closed-by-default approach also used for securing AI agent sandboxes. Research agents may need to reach more of the web, so the focus shifts to controlling where they can connect.
Guidance for GKE Agent Sandbox recommends isolated runtimes with default-deny network policies that open only the endpoints an agent needs. Public proxies, URL scanners, and disposable email services can stay blocked unless the job requires them.
Developers can also limit what an agent can send. So, instead of handing it a networking tool that accepts any URL or request body, an API integration can restrict requests to specific fields and formats. The runtime can then catch path traversal attempts, SQL injection strings, and executable markup before anything is sent. OpenAI takes a related isolation approach in its Agents SDK sandboxes, and the company’s Responses API tech lead has said large enterprise deployments often call for agents that are isolated from the network entirely.
Repeated failures can also be a reason to pause a run, especially when an agent keeps hitting client errors, anti-bot challenges, or unexpected redirects and begins trying increasingly aggressive ways to get around them, as Transluce documented in several of these cases.
Keeping the original task, tool calls, and server responses in the same trace gives operators a better chance of catching that behavior change when a retrieval job starts generating encoded scripts, visiting staging domains, or sending exploit payloads, rather than discovering it later in someone else’s security logs.
Repeated failures can also be a reason to pause a run, especially when an agent keeps hitting client errors, anti-bot challenges, or unexpected redirects and begins trying increasingly aggressive ways to get around them, as Transluce documented in several of these cases.
Enterprise Java company Azul announced its Azul Intelligence Cloud AI Assistant on Wednesday. The technology arrives in response to industry-wide alerts that AI has become a force multiplier for threat actors today.
The Azul service is a natural-language query interface that allows software engineering teams to see where security risk and licensing infringements are hiding in their live production Java estate. The assistant provides answers that are “grounded in live runtime data”, so it replaces static reports that grow less accurate after the day they’re generated.
How fast do code-scanning reports go stale?
Azul said that “most IT and engineering teams” still manage Java risk with static IT and software asset management (ITAM/SAM) reports and code-scanning tools that describe a moment in time. The company has insisted that these reports are “accurate on the day they’re generated, and increasingly wrong after that”, typically because Java Virtual Machines (JVMs) are spun up, patched, drifted and retired underneath the report’s scope.
“For years, enterprises have built dashboards and reports to understand what’s actually running in their Java estate, but by the time a report gets properly summarized and reviewed, the risk it describes has often already changed,” said Scott Sellers, co-founder and CEO of Azul. “That used to be a productivity problem. Now that AI can find and weaponize a vulnerability in hours instead of weeks, it’s a business risk – for security, for compliance and for the licensing exposure that shows up in an audit.”
“Now that AI can find and weaponize a vulnerability in hours instead of weeks, it’s a business risk…”
The Azul Intelligence Cloud AI Assistant lets software teams ask a direct question, in plain language, and get an answer grounded in what’s actually running in production right at that moment in time, as well as query historical information for further analysis.
The weaponization gap is closing
Where enterprises run business-critical workloads on Java alongside AI services, Azul said the practical effect is that the gap between a Common Vulnerabilities and Exposures (CVE) entry being disclosed and it being weaponized is getting shorter.
Citing AI models such as Anthropic’s Mythos and OpenAI’s Aardvark, which have autonomously discovered real-world vulnerabilities, the company pointed to an April 2026 Cloud Security Alliance white paper (listed as unofficial AI-assisted research), which suggested that, “While organizations historically took a median of 32 days to apply patches to known vulnerabilities – a window that once roughly corresponded to the time available before exploitation began – that window has collapsed to approximately 5 days for median time-to-exploit in 2025.”
Azul highlighted its Intelligence Cloud service, which gives engineers two “continuously updated” records of their Java estate: JVM Inventory, a live catalog of every JVM instance running anywhere (on-premises, cloud or container) and Code Inventory, a runtime record of which code actually executes in production versus what is merely provisioned. The new AI Assistant puts in a conversational layer using LLM models on top of both.
Software engineers can ask questions such as:
“Which JVMs are running Java versions which are not the latest updates?”
“Where is Oracle Java running in production right now?”
“What code hasn’t run in the past four quarters and is safe to remove?”
In the FAQ section of its announcement, Azul suggested that post-migration JVM “drift is common”, often due to a rollback, a forgotten node, a shadow deployment or various scripts and processes that haven’t been updated, which can reintroduce an Oracle Java runtime, exposing compliance and licensing risk, if not security risk also.
The shape of the Java runtime security market
In terms of which other vendors operate in the Java runtime analytics and security market, there are more than a handful of usual suspects. Contrast Security is known for its JVM agent and in-app bytecode instrumentation. Dynatrace offers Runtime Vulnerability Analytics as an extension of its core observability platform, which ships with OneAgent monitoring for Java vulnerable functions and JVM-level bytecode instrumentation agents.
Through its acquisitions by HP, Micro Focus, and now OpenText, Fortify remains known for its static and dynamic application testing services, including Fortify Application Defender, a runtime application self-protection (RASP) agent built to monitor Java workloads during execution. Part of Thales, Imperva’s runtime security for Java and .NET applications spans simple access control up to complex anomaly detection algorithms, though Imperva has reportedly put its standalone RASP product on an end-of-sale path. Then there’s Datadog, with its Application Performance Monitoring (APM), built to power code-level distributed tracing from browser and mobile applications to backend services and databases.
“Runtime context is absolutely critical for understanding the real risk in the production environment.”
A busy market for sure, so just how much of a problem are now-anachronistic static reports?
Head of security advocacy at Datadog, Andrew Krug, tells The New Stack that the downside of most point-in-time inventory scans is that they are “not always representative” of the runtime environment.
“Runtime context is absolutely critical for understanding the real risk in the production environment,” Krug says. “Even in the most mature software development lifecycle (SDLC) flows, tooling that generates static software bill of materials (SBOMs) may be bypassable [i.e. circumventable or subvertible] to get a feature deployed. Moreover, traditional vulnerability management flows outside of SDLC can bump versions outside of CI/CD process, accidentally compounding the problem and introducing added risk by bypassing known good guardrails like dependency cooldowns.”
“Even in the most mature software development lifecycle (SDLC) flows, tooling that generates static software bill of materials (SBOMs) may be bypassable… to get a feature deployed.”
Krug further advises that Datadog now sees “an increasing rise” in automated drive-by attacks on known vulnerabilities, “particularly Java” in many cases.
“Attacks used to be added to scanners, either specifically or generically, and scans are indiscriminately against targets,” Krug clarifies. “LLMs make it cheaper to add support for new vulnerabilities. However, it also makes it easier for individual researchers/hackers to have their own custom rulesets.”
“LLMs make it cheaper to add support for new vulnerabilities.”
He advises that this fact makes trends much harder to read than “oh, someone added support to CVE-2026-whatever in FFUF”, so today the goal of many attacks is unchanged. Attackers are looking to move laterally, establish persistence, and often automations will look for credentials to leverage to do just that.
NOTE: (Fuzz Faster U Fool) is an extremely fast web application fuzzer (written in the Go language), which is used by security testers to discover hidden files, directories and endpoints.
Dead code; it’s really a ‘thing’
The above-noted list of competitors that work in Azul’s marketplace is (arguably) substantial evidence of the real commercial licensing and security risks that exist where Java code, redundant JVMs and chunks of unsubstantiated (or more likely just untracked) Java components have been left to roam free. Azul itself noted that there’s a real maintenance overhead here that needs to be addressed because “unused and dead code that still gets tuned, tested and carried through every migration”, usually because no one can assess it’s safe to remove.
The larger the Java estate, the larger the exposures get, obviously. But that also means that the less a point-in-time report can be trusted to catch these exposures before they become an incident, an audit finding or a breach.
Dedicated compliance officers will likely enjoy wider deployment of these tools, although that role itself may now reside within a DevSecOps or platform engineering team, or both.
Azul Intelligence Cloud AI Assistant works regardless of which JVMs are deployed, from which vendor, or how old or large the applications running on them are. JVM Inventory and Code Inventory retain component and code-use history over time, so the AI Assistant can reason over which code, JVMs and applications have actually run in production, now and in the past.
The appeal of an autonomous coding agent is that you hand it a task, give it access to your repository and tools, and stay out of its way while it works. Google’s latest Gemini CLI release carves out specific moments when the agent now has to stop and wait for you.
Gemini CLI 0.61.0, released Wednesday, requires explicit confirmation before the agent edits build configuration files, runs build or test commands after such an edit, or executes shell commands whose arguments appear to come from untrusted external content. The same release separately hardens Gemini CLI’s optional sandbox so that host credentials and configuration stay out of reach of whatever runs inside it.
Giving a coding agent more authority to modify and execute code also gives an attacker more ways to turn that authority against the developer. Gemini CLI 0.61.0 puts a human back in the loop at some of those points.
Security fixes, in public
Google announced at I/O in May that it would move Gemini CLI’s Pro, Ultra, and free-tier users to its closed-source Antigravity CLI, and since June 18, the open-source tool has served mainly enterprise customers and developers with paid API keys. The company said Gemini CLI would continue to get model updates, bug fixes, and security patches. Those security changes are still developed in public, and the pull requests behind version 0.61.0 show exactly what Google was worried about.
Build files become attack vectors
A change to package.json, Makefile, pyproject.toml or a Bazel BUILD file can pull in a dependency or trigger a script. Gemini CLI can make those edits using information from web searches and external tools, then run shell commands. If documentation fetched while fixing a bug contains hidden instructions to add a postinstall script to package.json, the agent could make the edit, run the project’s test suite, and execute the malicious code without the developer ever typing the command.
Giving a coding agent more authority to modify and execute code also gives an attacker more ways to turn that authority against the developer.
Pull request #29250, titled “prevent indirect prompt injection via build file modifications and untrusted flags,” targets that sequence directly. Edits to recognized build files now require confirmation, and Gemini CLI tracks which build files change during a session so it holds any later build or test command, such as npm run, make, or cargo, for explicit approval. The confirmation dialog also shows full build-file diffs rather than truncating them.
Untrusted arguments need approval
The second check covers command arguments. Gemini CLI now treats content from web fetches, MCP server responses, Google Docs, and Buganizer, Google’s internal issue tracker, as untrusted context, and it asks before running any shell command whose flags or arguments match tokens from that content. In both cases, the prompt drops the persistent approval options, so a developer can’t grant a standing “always allow” for these actions.
The pull request ties the changes to restricted workspace mode, the safe mode Gemini CLI applies to folders a user hasn’t marked as trusted, and it doesn’t spell out how the checks behave in a trusted folder or under auto-approval.
The argument check matches tokens rather than tracing the provenance of every value, and the pull request’s review history shows how hard that is to get right. Google’s automated reviewer flagged several workarounds in earlier versions, including quoted arguments, environment-variable prefixes, shell redirection targets, and Windows path handling, all of which were addressed before the change merged on September 11.
Sandbox keeps credentials out
Pull request #29214 tightens Gemini CLI’s sandbox. When the sandbox runs through Docker, Podman, LXC, or macOS Seatbelt, the host’s ~/.gemini directory is no longer mounted inside it. Instead, the CLI passes in a sanitized copy of the user’s settings with API keys, hooks, and custom tool commands stripped out. It also blocks the sandbox from launching in sensitive locations such as the home directory, while new Seatbelt rules deny access to OAuth credentials, trusted-folder decisions, and .env files.
Google’s sandboxing documentation calls the feature a security barrier between AI operations and the host system, while cautioning that it reduces risk without eliminating it. The two pull requests show why both layers are needed. The sandbox limits what a process can reach once it runs, and the confirmation requirements decide whether the agent gets to take a sensitive action in the first place. Build files make the gap concrete: the sandbox mounts the project directory so the agent can edit it, meaning a poisoned package.json written inside the sandbox still sits in the repository when a developer or a CI job later runs the build outside it.
…a poisoned package.json written inside the sandbox is still sitting in the repository when a developer or a CI job later runs the build outside of it.
Not your keys, not your coins. Not your model, not your data?
Over the summer, the tech industry was consumed by a debate about AI use in the enterprise and the need to protect IP. If an enterprise used proprietary models, was data leakage a necessary evil?
Companies seemed to have two options: They could use state-of-the-art, proprietary models and risk losing control of their data, or they could use open-weight models and never kiss the frontier.
Thankfully, a third option is emerging.
Consider the concern: Company A wants to use LLM B from AI Lab C, and they want to avoid training AI Lab C how to eat Company A’s lunch by building its capabilities into LLM B. A good way to resolve the tension would be to let Company A run LLM B on its own infrastructure, so there’s no risk of its information fleeing on the wind.
AI agents are “creating a whole different set of requirements at the data layer.” –Vast Data co-founder Jeff Denworth
But that raises another problem: AI Lab C doesn’t want to allow Company A to run LLM B on its own GPUs because it doesn’t want to hand over its model weights. It’s the same IP issue the company ran into, in reverse. You have to solve the trust problem in both directions!
Enter VAST Data co-founder Jeff Denworth and a new product called DataEnclave, which aims to let AI labs and enterprise-scale companies deploy proprietary models in secure compute environments without risking data transfer in either direction. (DataEnclave uses Nvidia’s Confidential Computing technology to make the system tick; Vast Data’s core product is AI OS, infrastructure that fits beneath a company’s AI applications.)
The New Stack had Denworth on the podcast to chat about the confidential computing market. I was curious about timing. Why did Vast build DataEnclave now? Nvidia began rolling out Confidential Computing in a serious way in 2024, after all. Denworth argues that the market needed the core technology, yes, but also demand.
And until late 2025, AI demand was modest compared to today’s token totals. Once agentic coding tools took off, corporate demand for AI products soared. This led to the pricing crisis we saw in early 2026, and the secure AI usage debate we endured over the summer.
Performance drove demand, demand drove usage, and usage dug up fresh problems to solve. Now the question for the market is whether or not DataEnclave has solved enough concerns on both sides of the proprietary AI-proprietary data equation. The market will sort that out as it moves through early access and into general availability.
Our conversation goes deep into the arc of AI, where companies are in their AI journey today, and how much data remains to be unlocked inside the enterprise. If you want to feel the acceleration, it’s a fun one!
Amazon started blocking Meta’s Muse from browsing and buying on Amazon.com on Sunday, roughly two weeks after the personal agent launched on September 8. Shoppers who ask Muse to buy something there now get a pop-up telling them that continued access by an unauthorized AI agent violates Amazon’s Conditions of Use.
Amazon’s objections have little to do with shopping itself. Meta never told Amazon that Muse would visit the store, the agent does not identify itself while it browses, and it appears to capture and store customer credentials. Those three properties describe almost every personal agent shipping this year. Grok Bot from xAI also drives signed-in browser sessions, and so does the open-source OpenClaw project that Muse is modeled on. The block is a category design problem rather than a disagreement between two companies.
What Amazon actually blocked
Muse runs on a dedicated virtual machine that Meta calls Muse Secure VM, and it reaches services in two ways: It uses built-in connectors for partners such as Gmail and OpenTable, and it drives an ordinary browser session for everything else. Shopping on Amazon.com used the second path.
That second path is what Amazon objects to. From the server’s side, a browser-driving agent looks like a signed-in customer with unusually fast reflexes, moving through search, product pages, account history, and checkout without ever declaring what it is. Amazon told GeekWire it asked Meta to exclude the store voluntarily, but Meta did not agree before the block went live.
The credential dispute is harder to settle from outside. Meta says Muse has no visibility into passwords or payment methods and that credentials sit in secure storage, while Amazon says the agent appears to capture and retain them. Both statements may be sincere, and the merchant can verify neither, since an unannounced session provides no evidence of which software holds the password.
The legal ground shifted seven weeks ago
Amazon reached for its Conditions of Use rather than the Computer Fraud and Abuse Act. Those terms, updated August 14, now require agents to identify themselves in user-agent strings and stop when asked. The likely reason sits in a ruling from early August. The Ninth Circuit vacated the preliminary injunction Amazon had won against Perplexity, and the panel held that a user directing the Comet assistant is the party accessing Amazon’s computers. Writing for the court, Judge Milan Smith described the assistant as a tool, not a person, for statutory purposes.
If you’re operating a public API or storefront, the ruling makes lawsuits a weaker tool for keeping agents out. Blocking them in your own infrastructure is now the more reliable option. A site cannot easily argue that an agent trespassed, so it has to decide for itself which automated clients it admits, publish that decision, and enforce it in its own infrastructure. Amazon’s pop-up is that enforcement, written in product rather than in a filing.
The identity layer already exists
Platform teams have solved a version of this problem before. Inside a service mesh, no workload is trusted by default because it looks like a normal client, and every call carries a verifiable identity that the receiving service checks before applying policy. Agent traffic on the public web faces the same requirement, and the specification is further along than most teams realize.
An IETF draft called Web Bot Auth builds on HTTP Message Signatures (RFC 9421). An agent signs its requests with a private key and publishes the matching public key at a well-known directory on its own domain. The verifier reads the Signature-Agent header, fetches the key set, and learns which operator is calling. Cloudflare validates these signatures at its edge for verified bots and agents. AWS WAF Bot Control added the same support for CloudFront distributions in November 2025.
The limits matter as much as the mechanism. A signature identifies the operator behind the agent, not the person it is acting for. The merchant learns that a request came from a named vendor, without learning whose account is in use or what the shopper approved. Amazon’s complaint about stored credentials sits in that gap. Signed identity settles the disclosure question and leaves authorization open.
Shopify took the other route within a day
While Amazon was blocking Muse, Shopify was wiring it in. On September 21, the two companies announced agentic checkout with Shop Pay across Shopify stores, extending the arrangement that made Meta an AI channel in Shopify Catalog on the day Muse launched. Muse reads structured product data and completes payment through a declared path, so the merchant knows an agent is transacting, and each purchase draws a single-use credential, so the card number never reaches Muse.
The plumbing for that path is public. Google and Shopify’s Universal Commerce Protocol covers discovery, cart, and checkout. The Agentic Commerce Protocol from OpenAI and Stripe covers checkout execution while the merchant stays the system of record. Google’s Agent Payments Protocol, donated to the FIDO Alliance in April, includes proof of the shopper’s authorization. A merchant that adopts it gets identity, scope, and an audit trail in the same transaction, which is what Amazon says it wanted and did not get.
Both routes follow from the business underneath them. Amazon runs its own storefront, recommendations, and assistant, so an outside agent that hides its identity takes the customer relationship and gives nothing measurable in return. Shopify sells infrastructure to merchants, so every new agent channel that reads its catalog and settles through Shop Pay reinforces the rails underneath. The key difference is who owns the demand surface, which explains why the same agent got a block from one company and a partnership from the other in the same 24 hours.
Choosing how to handle agent traffic
Most teams exposing an API or a storefront now have to make this call deliberately rather than by default. The decision depends on how much the business relies on the customer relationship at the point of contact and whether an agent can be identified when the customer arrives.
Scenario
Recommended option
Rationale
Public content and catalog data, no account access
Verify signatures at the edge and allow named agents
Web Bot Auth is checked by default on Cloudflare and AWS WAF, so the cost is policy configuration rather than engineering, though it tells you the operator and not the shopper
Agent transactions where you want the revenue
Publish a declared channel using ACP, UCP, or an MCP server
Structured access gives scope and an audit trail, at the cost of building and maintaining a second interface alongside the site
Account access with stored credentials
Require a scoped token, never a replayed password
Delegated tokens can be revoked per agent, though few consumer agents support them yet, which pushes the burden back onto your login flow
Competitive surfaces you intend to keep
State the rule in terms of service and enforce it at the edge
Legally durable after the Ninth Circuit ruling, though it invites the same public standoff Amazon is now in
Most real deployments will combine these rows rather than pick one. A retailer can verify signed agents on product pages, route purchases through a declared checkout, and still refuse an unannounced browser session inside a logged-in account. That combination is closer to Amazon’s position than its pop-up suggests.
What platform teams should do this quarter
Enterprise buyers and the teams running these systems face the same three questions, in a specific order.
Decide what an unidentified agent may do
The first question to settle is admission, and most sites have not settled it. They treat agent traffic as either a scraper to block or a browser to serve, and neither answer survives contact with a customer who wants an agent to act for them. Write policies for public pages, logged-in pages, and checkout separately, then publish them where an agent vendor can find them.
Give identified agents somewhere better to go
The second question is substitution, and it decides whether the first one holds. Blocking a browser-driving agent without offering a structured path leaves the demand intact and pushes it toward workarounds. Sabre reported that nearly 80 of its customers now pilot or run its MCP server for booking rather than let agents work through a booking screen. A catalog feed, an MCP server, or an ACP endpoint converts hostile traffic into a channel you can meter.
Fix credential handling before agents force it
The third question is authorization, and Amazon raised it loudest. An agent replaying a stored password is indistinguishable from credential stuffing at the network layer, regardless of any goodwill between the two companies. Scoped, revocable tokens tied to a named agent and a spending limit are the only version a risk owner can approve.
Where agent access is headed
Amazon and Meta will settle this commercially, because Amazon has an advertising arrangement that lets Facebook and Instagram users shop its products, and Meta buys compute from AWS, and neither gains from a long standoff over one shopping flow. The precedent is already set regardless of how they settle. Every site that matters to an agent now has to answer whether it admits anonymous automation, and it will enforce that answer through bot management rules and protocol endpoints rather than cease-and-desist letters.
Agent builders should read the block as an argument for declaring themselves. An agent that signs its requests, identifies its operator, and transacts via a published protocol can be allowed, rate-limited, and billed, while one that arrives disguised as a browser will keep encountering pop-ups. For developers building the services these agents reach, the signed identity layer arriving through Cloudflare, AWS, and the commerce protocols is the most useful infrastructure the open web has gained in years. It is worth adopting before the next agent shows up unannounced.
Most people already understand what generative AI can do. But enterprises run into problems when they need to give a model access to information that cannot leave their own environment, such as a patient record, a customer’s financial details, or a company’s most valuable intellectual property.
Sending that data to a cloud or SaaS service means it crosses external networks and is processed on infrastructure run by another organization, creating additional concerns about control, accountability, and exposure. That’s where AI enthusiasm collides with production realities. Despite its productivity potential, enterprise AI still faces a fundamental gap in trust and control.
Organizations need to know whether a system will expose information it should protect, act as intended, meet security and performance requirements, and behave safely at machine speed.
Alon Horev, CTO and co-founder of AI operating system company VAST Data, tells The New Stack that the challenge is particularly acute when AI systems handle sensitive customer information. “Even if you ask the model today to obfuscate a conversation or redact PII from a conversation, it’s hard to have 100% confidence that’s the case, and that it worked.”
“Even if you ask the model today to obfuscate a conversation or redact PII from a conversation, it’s hard to have 100% confidence that’s the case, and that it worked.”
Consider a customer support agent that needs access to an individual’s profile to provide a useful, personalized answer. The organization must ensure that information isn’t exposed to another customer, while also considering whether those conversations can be used for training or system improvement. They might contain personally identifiable information (PII) or other protected details, and the consequences of mishandling them ultimately fall on the organization and the people whose information it holds.
Confidential AI architectures: solving a two-sided trust problem
Enterprise AI has two parties to satisfy: organizations must keep sensitive data under their control, while model builders need to protect the weights and software that represent substantial investments in research, engineering, and IP. They’re understandably reluctant to place those assets in environments where customers, infrastructure operators, or attackers might gain access. That mutual need for control has created a stalemate. How can organizations bring advanced models to sensitive data without asking either side to surrender control?
Horev has seen that the most capable models are increasingly delivered as SaaS services, because that’s the simplest way for their creators to distribute and protect them. Even when a provider offers compliance controls, the enterprise might still shoulder the consequences of a breach, misuse, or regulatory violation. Sending information across the WAN also places it in the hands of more systems, connections, and operators, increasing the number of points that must be trusted and governed. Organizations may also be unwilling, or legally unable, to rely on a provider’s assurances that it will not retain, reuse, or expose their data beyond the intended service.
For organizations in regulated or data sovereignty-sensitive sectors, that could be an unacceptable trade-off. “Naturally, many organizations are adopting a hybrid strategy,” Horev tells The New Stack. “Some applications and datasets can go to the cloud, while others must remain on-premises, sometimes even in the building, or in the country.”
This is where confidential AI comes in. Encryption at rest and in transit protects data while it’s stored or moving between systems. Confidential computing extends that protection into the processing environment, using hardware-isolated execution to create a protected enclave in which the data and model weights can remain encrypted until they’re released to an approved workload.
Cryptographic attestation verifies the hardware, virtual machine (VM), software, and configuration requesting access before releasing keys. The model builder can encrypt its model using the public key of a specific confidential VM. Only that VM’s corresponding private key can decrypt it within protected memory, enabling the customer to use the model without accessing its weights.
Independent key control preserves the separation between the two sides. The enterprise retains control of the keys governing its data, while the model builder retains control of the keys governing its model. While the workload is running, the infrastructure operator doesn’t control either set of keys.
As AI becomes more agentic, those controls will matter more. Agents will need to access more data, systems and tools, and might act on that information with far less human intervention.
“This world of agentic AI is moving extremely fast, and we need to limit what an agent can see and do.”
Those that can’t establish strong privacy and governance assurances for today’s models will find it even harder to deploy agents safely in the future. “This world of agentic AI is moving extremely fast, and we need to limit what an agent can see and do,” Horev tells The New Stack.
From architecture to ecosystem
Many businesses simply cannot manage the integration, security, and maintenance of the entire AI stack, because it requires working separately with each model provider to engineer something that suits both parties. Turning confidential AI architecture into something organizations can deploy is the challenge VAST DataEnclave, which was launched on September 22, intends to address.
As a capability of the VASTAI Operating System, the goal is to bring the model, application layer, and data platform together under customer-controlled operating conditions. The architecture is designed to protect both sides of the equation: the enterprise’s data and the model builder’s weights. The customer retains control of its infrastructure and data keys, while the model provider can make its software available without handing over the underlying intellectual property.
“We’re trying to close the trust and control gap by working with world-class model builders such as Cohere, Deepgram, Factory, Fundamental and TwelveLabs, who continue to innovate and build their expertise,” says Horev. The ecosystem also includes infrastructure and security providers such as Nvidia, CrowdStrike, Fortanix, Nscale, Cisco, and Supermicro. The range reflects the practical challenge: confidential AI needs more than a protected GPU. It requires models, applications, accelerated hardware, data infrastructure, and operational support to work together.
That control also changes the cost conversation, without automatically making AI cheaper. Hosted models can make budgets harder to predict as token consumption varies with usage patterns, agent loops, model architecture, and workload volume. Customer-controlled infrastructure gives enterprises a more defined capacity and cost base: they can plan around GPU clusters they own or have already budgeted for, instead of allowing inefficient model choices or uncontrolled agent activity to generate an open-ended token bill.
“…instead of allowing inefficient model choices or uncontrolled agent activity to generate an open-ended token bill.”
The cluster also imposes a natural ceiling on throughput, which helps organizations understand how much work their infrastructure can handle within a given period. Model providers can then price access by token, task, or license, while the enterprise retains greater visibility into its total operating cost.
Why the data platform is paramount
Confidential AI protects data and model weights during inference, but it’s only part of the production challenge. Real-world AI systems are living environments in which data moves between storage, databases, GPUs, networks, applications, and agents.
That’s why confidential AI can’t be bolted onto a fragmented stack. Businesses need to protect the model, the data, and the infrastructure connecting them as one system. As Horev says: “You need to build security in multiple layers of the platform,” with someone accountable for rapidly updating compromised components.
Confidentiality is only useful if the resulting system can also be operated, monitored, and improved. As AI infrastructure becomes more distributed, it becomes harder to tell what’s happening when something goes wrong and where the fault lies.
Horev recommends a “single pane of glass” across storage, networking, and compute, so teams can see what’s happening and keep resolution times low. If a network port is intermittently failing in a data center, for example, an agent could help identify the root cause, provided it has access to the right operational data and tightly controlled permissions. Those permissions should govern the infrastructure it can inspect, the data it can retrieve, and the actions it can take.
The same applies to monitoring AI workloads. Teams need visibility into performance, failures, and access patterns without exposing the customer data or model weights. Agent sandboxes can limit the systems and tools an agent can reach, while data platform observability can log which data it accessed, what it did with that data, and how it interacted with downstream systems.
Evaluation, therefore, becomes part of production discipline. Teams must observe systems, measure behavior, govern access, and manage change in ways that demonstrate progress. Confidentiality, data-level policy, observability and correctness have to work together.
The emerging ecosystem suggests demand for models that can run securely under customer control, wherever sensitive data resides. These are “living systems,” says Horev. “It’s not just leveraging a feature inside of a wider platform.”
AI coding tools have seriously accelerated developer speed, but AI has also done the same for attackers — and the software supply chain is increasingly where the two are colliding.
The numbers give some idea of how quickly software development is changing. GitHub processed around one billion commits in 2025. By April 2026, the platform was handling roughly 275 million commits a week, according to GitHub COO Kyle Daigle. GitHub Actions usage has climbed, too, from 500 million compute minutes per week in 2023 to 2.1 billion in just part of a single week this year.
Quincy Castro, CISO at Chainguard, says the shift in how software gets written is already stark.
“I look around Chainguard, and I don’t think any of our engineers have actually written a line of code by themselves in the past year,” Castro tells The New Stack. Writing code manually now “sort of feels quaint, like you’re illuminating manuscripts,” he says, while “the printing press is out there just going to town.”
But this isn’t only about professional developers producing more code. AI has also widened the pool of people who can create software. Teams in HR, finance, and business intelligence that once had to wait for engineering resources can increasingly build what they need themselves.
That means more software being created by people outside traditional engineering teams, often with AI making decisions about what goes into it. The person prompting the agent may never see which libraries or packages it has chosen.
When the agent chooses the dependencies
Software security was already built around the fact that humans couldn’t inspect everything. But developers were still making important decisions, including which libraries and packages went into an application.
That changes when an AI agent is doing much of the coding.
“Humans are directing what they want to be done, but they’re somewhat abstracted from the actual doing of the work,” Castro says. “You have AI instead now making the choices of what dependencies am I going to pull into this application? How am I going to go accomplish this task?”
“Humans are directing what they want to be done, but they’re somewhat abstracted from the actual doing of the work.”
Attackers, meanwhile, are finding plenty of uses for the same technology. Castro sees three problems arriving at once: frontier models finding previously unknown vulnerabilities, attackers using agents to exploit better vulnerabilities organizations haven’t fixed, and sustained attacks against the open-source ecosystem.
A collection of medium- and low-severity findings might once have sat well below the top of a remediation queue. Frontier models with advanced cyber capabilities, including Anthropic’s Claude Mythos Preview and OpenAI’s GPT-5.6-Cyber, can now work across those findings and chain seemingly minor weaknesses into a viable attack path.
“Here’s a whole ton of mediums and lows. Now give me the attack path that gets me domain admin,” Castro says, describing the approach. “Chain these together to go get me root on the system. And AI is really, really good at being able to do that.”
At the same time, AI’s ability to combine apparently less-serious weaknesses makes a neat CVSS-based queue a less useful representation of what an attacker can actually do.
The third problem is the software supply chain itself.
Open source becomes the attack path
Modern applications depend heavily on open source software, and attackers have increasingly targeted the infrastructure used to build and distribute it.
Castro pointed to the TeamPCP campaign, which compromised widely used projects including Aqua Security’s Trivy. In that attack, malicious code was pushed into trusted components and subsequently picked up downstream.
Supply chain attacks were once associated primarily with sophisticated state-backed groups willing to spend significant time getting into the right place. That barrier is falling.
“If you don’t mind making some noise, this is a way easier attack vector than I think a lot of people thought it was,” Castro says. More importantly, “a single attack that’s successful can lead to a cascading set of other compromises and other access that gets you into other places.”
The development pipeline itself can make matters worse. Castro says many organizations still have relatively few controls around CI/CD, while developers routinely pull components from external sources to get their work done. Adding autonomous coding tools to that behavior compounds the risk.
“You wouldn’t pick up a random thumb drive and stick it into a production system, right? But that is effectively what folks are doing when they’re consuming open-source software that way.”
He compared the way organizations consume open source software to plugging an unknown USB drive into a production system. “You wouldn’t pick up a random thumb drive and stick it into a production system, right?” he said. “But that is effectively what folks are doing when they’re consuming open source software that way.”
Open source isn’t the problem. Trusting its distribution path without sufficiently verifying what you’re consuming is.
Prevention has to come before detection
This is where the old security model starts to creak.
For years, much of vulnerability management has followed a familiar loop: scan something, generate an alert, decide how serious it is, and get somebody to fix it. That becomes harder to sustain when development output multiplies, AI agents make more of the underlying decisions, and attackers can exploit weaknesses before fixes are available.
Castro wants companies to put more effort into what enters the development environment in the first place, rather than discovering problems once the software is already there.
“How do we just make things work from the beginning, with no alerts and no responding to stuff and no people chasing things and no people trying to prove a negative?” he says. “From end to end, from the creation of code to its deployment, how do we make sure that we can give folks the most trustworthy version of that thing?”
Rather than taking packages from public ecosystems at face value and scanning them after the fact, Chainguard builds artifacts from verified, buildable source.
That’s the thinking behind Chainguard’s approach to containers, libraries, and other open source artifacts. Rather than taking packages from public ecosystems at face value and scanning them after the fact, the company builds artifacts from verified, buildable source. It provides provenance about how they were created.
But trustworthy components are only one layer.
“There’s no point in bringing inherently secure software components into the environment if you don’t actually have a technical control that says this is the only way people developing code can consume these things,” Castro says. That means engineering, security, and SRE teams also need controls over where software — whether selected by a human or an AI agent — can come from.
That requires several layers of protection. Organizations need to know where their software came from and how it was built, control what can enter their environments, and make sure those rules apply when an AI agent chooses components as well as when a developer does.
Defending open source at AI speed
There is another problem, however. Frontier models such as Claude Mythos Preview and GPT-5.5-Cyber aren’t just finding vulnerabilities that previously went undetected; they can also combine lower-severity flaws into working attack paths. Individual companies can harden their own pipelines, but the software they depend on comes from an open-source ecosystem facing vulnerability discovery at a speed and scale it wasn’t built for.
That’s part of the reasoning behind Athena, the industry coalition Chainguard launched to turn vulnerability findings from frontier AI programs into fixes. As of July, the coalition had processed more than 40,000 vulnerabilities, with 42% rated critical or high severity and 86% marked as network reachable, meaning attackers can access and trigger them at the network level.
For Castro, the important part isn’t simply finding more bugs. AI is already getting very good at that. Someone still has to fix them.
“Through Athena, what we attempt to do is to give people that engineering fix,” he says. “What if we create a coalition where folks just send us the issues that they’re finding? We automatically generate fixes for those, and we push those back to everybody.”
“Through Athena, what we attempt to do is to give people that engineering fix.”
Those fixes can also be pushed upstream to open source maintainers, who face the prospect of being buried beneath an expanding pile of AI-generated vulnerability reports.
That may ultimately be the bigger shift AI forces on software security. Developers aren’t going to stop using coding agents because they create new risks, any more than companies are going to stop using open source because attackers target it.
Bolting enough scanning onto an exponentially faster development process isn’t much of an answer either.
The opportunity is to remove more of the risk before the software ever reaches a developer or an agent: Start with components you can trust, tightly control how they enter the environment, and fix weaknesses as close to their source as possible.
AI has made it dramatically cheaper to create software. It’s doing the same thing for attacks. Security now has to keep up without putting the printing press back in the box.
AI coding agents play a major role in software development and delivery, and for good reason. They can investigate bugs, trace dependencies, refactor services, and propose patches without developers needing to assemble all of the relevant context manually. That capability comes courtesy of agents’ appetite for context. To make informed decisions, agents read source code, configuration files, terminal output, error messages, environment information, and more…much, much more.
“Secure, agentic development depends on a security control many teams still lack: preventing secrets from leaking to AI coding agents and becoming model context.”
From a security standpoint, this becomes problematic when agents, in the search for context, inadvertently reach for secrets.
For years, developers have been taught not to commit API keys, database credentials, and tokens to Git. But agentic workflows have created another route for secrets to escape development environments before a commit, code review, or CI job. Depending on its permissions, configuration, and provider architecture, an AI coding agent may read local files or receive pasted content that is then included in data sent to an AI service, and, in the process, developers may never see their credentials leak.
The quiet path from local files to external systems
Some forms of secrets leakage are obvious. A developer troubleshooting an authentication failure may paste, for example, a failing API call into a chat window, including the token. However serious, this sort of leak is characteristically human.
The more consequential escape pathway is quieter. An agent tasked with understanding a project may inspect files in its working directory, including an overlooked .env file, a cloud credential profile, an SSH configuration, or sensitive application logs. In such instances, nothing has necessarily gone wrong from the agent’s perspective; it is doing exactly what it was designed to do: collect context to solve the task at hand.
“Agentic workflows have created another route for secrets to escape development environments before a commit, code review, or CI job.”
But once a secret becomes part of that context, it may pass through systems outside an organization’s direct control. Depending on the workflow, it can appear in model provider logs, gateway telemetry, prompt histories, or debugging records. Rotating the credential is essential, but it does not erase copies that may already exist in those systems.
This changes the practical definition of a secret leak. The problem is no longer limited to what lands in a repository, but also includes what an autonomous tool reads and forwards while operating on a developer’s machine.
Why traditional security gates no longer suffice
Most application security programs are built around durable checkpoints: the commit, pull request, build, and deployment. In the agentic era, these checkpoints remain important as they can detect secrets that reach version control and prevent a bad change from merging and deploying.
This highlights an important timing gap. The 2025 Verizon Data Breach Investigations Report reports a median of 94 days to remediate leaked secrets discovered in GitHub repositories. In an agent-driven workflow, detection and response need to happen much earlier, and not after a credential is exposed. Still, at the moment it’s about to cross the boundary from local context to an external model.
Bad actors already understand the value of that porous boundary. Recent supply-chain attack campaigns, including Mini Shai-Hulud, have searched developer and CI environments for credentials and configuration data, including AI coding-tool configuration files. These campaigns show that agent configurations and the local context accessible to an agent are valuable targets. AI coding agents can broaden the local data reachable during a session, making even the agent’s context-collection mechanisms an attractive target.
Treat agent context as an egress surface.
The secure mental model doesn’t frame AI agents as mere code editors, but rather, automated data-movement systems. Its inputs can include far more than the source files a developer is actively editing, and its outputs may involve external services.
That calls for a zero-trust approach to agent context. Before sending a prompt or adding a file to an agent’s working set, organizations should evaluate it for sensitive material. Controls should be deterministic: identify a likely secret, block or redact it, and provide the developer with a clear path to remediate it.
“Asking an LLM to decide whether to transmit a credential does not create a reliable security boundary.”
Critically, the control should be independent of the model. Asking an LLM to decide whether to transmit a credential does not create a reliable security boundary. Purpose-built secrets detection can inspect prompts and files against known credential patterns and policies, applying a deterministic policy, such as blocking a prompt or file read when it detects a credential-shaped value. For example, Sonar’s secrets detection ships alongside dedicated agent plugins to bring that local check into tools such as Claude Code, GitHub Copilot, Codex, and Cursor, so it can flag a credential before a prompt or file read is transmitted to a model provider.
Build defense in layers, without disrupting your agentic workflow
A legitimate workflow does not involve forcing developers to choose between secure development and useful automation, but instead places fast controls at several points where secrets can escape:
In the editor: Use IDE-integrated secrets detection to flag credentials while they are being written.
Before model submission or agent file access: Where the agent supports it, scan prompt submissions and file reads locally, and block risky operations according to policy.
At the command line: Check generated snippets and local changes in terminal-driven workflows.
In pull requests and CI: Detect secrets that reach the repository and use review, quality gate, and deployment controls to prevent unsafe changes from progressing.
In incident response: Rotate exposed credentials quickly, investigate downstream logs and access, and reduce recurrence through policy and training.
Building defense at the pre-submission layer is an emerging requirement and requires both security and usability. Secrets detection must be fast enough to run in developer workflows; a scanner that introduces lengthy pauses may be bypassed or disabled by developers, and it must also have a manageable false-positive rate, or developers may stop trusting it.
Teams should also make their agent permissions and context rules explicit, as broad agent permissions can increase the amount of sensitive local context reachable during a coding session. Consider the following: which directories can an agent read? Are .env files, credential stores, home-directory configurations, and production logs excluded by default? Does the organization route prompts through an approved gateway? What retention, training, and audit settings apply at the provider level? Document and enforce the answers rather than leaving them to individual developer preference.
Secrets security must shift left.
Prevent secret leakage without hindering AI-assisted development, ensuring the productivity promise of agentic development doesn’t carry significant security implications.
As agents become more autonomous, security standards must follow agents upstream. It’s critical to stop a secret before it becomes context, while it is still local, visible, and easier to control. In the agentic era, code review and CI-level checks will remain essential safety nets. Still, for agent-centric development, the first line of defense must shift left: to the instant an AI coding tool determines what to read and what to transmit. That is the control modern development teams need to implement now.
Three security researchers at Hacktron AI found a memory-corruption bug in a widely used image library. Finding it was the easy part.
The hard part was turning it into something that works on a real server, so on July 24 they handed that job to Anthropic’s Claude Opus 4.8. The model managed it only with the operating system’s memory randomization switched off. With the protection on — this is the way every production box runs it — nothing it wrote held up.
The researchers came back the next morning with the same bug and the new model. Roughly three hours later, Opus 5 had a working ARM64 exploit running against a Mac on their desk. About four hours after that, they had remote code execution against a test forum.
Less than 72 hours after they started, they were reading from OpenAI’s private monorepo — using an OpenAI employee’s Codex account to open a pull request against a README, then stopping there.
The bug wasn’t in anything OpenAI wrote. Hacktron was testing community.openai.com, the company’s user forum, which runs on Discourse — the same off-the-shelf forum software behind thousands of other sites.
Discourse normally screens uploaded images with FastImage. But FastImage doesn’t handle HEIC and HEIF, so those files get passed to ImageMagick instead, and ImageMagick decodes them with libheif. The version running in the Debian 12 base image the forum used, 1.19.7, had a heap buffer overflow that a specially crafted file could trigger.
A fix had landed upstream the previous year. But the commit wasn’t documented as a security fix and never got a CVE, so it never triggered a backport into the Debian package the forum was using — a patched bug that stayed exploitable because nobody labeled it.
The researchers adapted the exploit for the x86-64 and jemalloc configuration Discourse runs, and a malformed HEIC image was enough to trigger remote code execution.
Discourse later confirmed the vulnerability in security advisory GHSA-vhm9-85gw-x335, rating the upstream libheif flaw — tracked as CVE-2026-32882 — 8.8 out of 10 on the CVSS severity scale. The New Stack has reached out to Hacktron AI for additional details about the researchers’ use of Claude and will update this story if we hear back.
One exploit, a much larger path
Code execution on a forum is a bad day for the forum. But it shouldn’t be a bad day for the company that owns the forum. This is where the chain crossed into something that was OpenAI’s own. Hacktron then found a flaw in OpenAI’s single sign-on system: sign-in tokens issued for the forum carried excessive permissions, granting full API access to the linked ChatGPT and Codex accounts. Some of those accounts belonged to OpenAI employees.
One employee’s Codex account was connected to OpenAI’s GitHub environment, opening a path to the company’s private repositories. Hacktron says other accounts could have exposed connected services including Slack and email.
The team stopped there. Using Codex, they made a harmless documentation change against OpenAI’s private openai/openai monorepo and opened a pull request — enough to prove the access was real, and nothing more. Hacktron’s write-up says the pull request’s details were redacted at OpenAI’s request.
From assistant to exploit developer
Up to this point, those three experienced researchers were still in the loop. So Hacktron ran the experiment again with the humans mostly out of it.
They put Claude in an autonomous agent loop — giving it a goal, a target, and time to keep working — pointed at a Discourse instance of their own. The model got there on its own, achieving remote code execution and demonstrating it by reading /etc/hosts from inside the container.
Getting it started took one piece of misdirection: Opus refused to write an exploit aimed at a live remote host. So the team proxied their own instance through rce.ee/ctf-forum, a URL that made the target look like it was part of a capture-the-flag exercise.
Memory-corruption exploitation has always been specialist work, invovling memory layouts, allocators, operating system internals, and protections to make all of it wildly unreliable. Hacktron’s run signals a meaningful share of that work might be able to be delegated to AI now. It also suggests the line between security research and attack development is — from the model’s side, anyway — partly a question of what you consider a target.
The full chain
Put together, the attack looked like this:
HEIF upload → libheif overflow → code execution on the forum → over-permissioned SSO tokens → employee ChatGPT/Codex account → connected GitHub → pull request in openai/openai
A two-month project, under $3,000
The OpenAI intrusion was one thread in a broader project the team called “HEIF Heist,” a roughly two-month sweep of image-processing infrastructure across multiple major technology platforms. The whole effort consumed less than $3,000 in model tokens.
OpenAI paid Hacktron a $6,500 bounty for the account-takeover flaw on its side. It has since narrowed the permissions on community sign-in tokens and revoked the affected tokens and sessions.
Container security controls often fail less because organizations lack standards, scanners, or best practices, and more. After all, every service tends to use its own Dockerfile, building images in slightly different ways.
Buildpacks are often pitched as a way to make containerization more developer-friendly, standardized, and maintainable by eliminating the Dockerfile pain points. They can also help improve security posture in practice, especially in larger, multi-team environments.
In this article, we will explore how Buildpacks provide a standard application-image path where platform teams can centrally govern build inputs, runtime images, Software Bill of Materials (SBOM) production, artifact metadata, and update workflows–making container controls easier to enforce and operate at scale.
The patch that never reached production
What does a patch integration process look like in the real world?
In more than one enterprise, I’ve seen the story look roughly like this: a critical vulnerability is fixed in the organization’s approved runtime image. In a perfect world, a patched base image is detected centrally, rebuilt, tested, published, and automatically propagated to every affected application. But we do not live in a perfect world. A week later, some production workloads still run on the vulnerable base.
“A container security control isn’t useful on its own. It must be applied consistently, observable in the running estate, and maintainable when images, dependencies, and vulnerabilities change.”
Typical reasons include:
Inconsistent Dockerfiles using different base image tags, Linux versions, and runtime versions, making it harder to determine which service is affected;
Application images may start from a patched base but reinstall vulnerable dependencies because developers manually install packages;
Uneven rebuild cadences: some teams rebuild more frequently, others only rebuild when the application code changes, so a service with no recent feature work may run an old vulnerable image for months;
No centralized inventory, SBOMs, image metadata, or deployment tracking, so organizations cannot reliably determine which applications remain vulnerable or measure patch propagation.
This story has not yet found its happy ending because the organization is only halfway into establishing a vulnerability management process. A container security control in place, such as “all services must use approved patched base images,” isn’t useful on its own. It must be applied consistently, observable in the running estate, and maintainable when images, dependencies, and vulnerabilities change.
That’s the stage where many teams fall short.
Why container security controls drift
Container security controls are hard to implement largely because of established containerization practices. In many organizations, no single governed build path exists. Each repository produces its own image, usually through a Dockerfile maintained by the application team.
Dockerfiles offer flexibility. They also push many security decisions onto developers: which base images to use, which packages to install, how to configure the runtime user, how minimal the image should be, how to track patches, which CI policies to apply. Over time, as services multiply and teams change, these choices drift. It’s not unusual to see different repositories using different base images, update cycles, and even different interpretations of what “secure” means in practice.
“The problem is expecting every developer to have enough container expertise to do this consistently across hundreds of repositories.”
Dockerfiles are not the problem by themselves. A well-written Dockerfile can produce a minimal, hardened image. The problem is expecting every developer to have enough container expertise to do this consistently across hundreds of repositories.
This approach also makes patch propagation unreliable. Platform teams may publish a patched base image, but each application team must still notice the update, modify its Dockerfile, rebuild, test, and redeploy. Some do it quickly. Others do it late or not at all. The control exists, but adoption remains uneven.
Buildpacks help close this gap by moving common containerization decisions out of individual repositories and into a shared, governed build path.
The shift from repository-specific builds to a governed build platform
Buildpacks turn application source code into a production-ready OCI container image without a Dockerfile. They detect the application type, select the required buildpacks, provide the necessary runtime and dependencies, and produce a runnable image.
For container security, the main benefit is not simply removing Dockerfiles. Buildpacks tend to provide a standard way to construct images, at least for typical application stacks.
“Buildpacks turn application source code into a production-ready OCI container image without a Dockerfile.”
This gives platform and security teams a central point for enforcing controls. Instead of asking every team to choose an approved base image, configure the runtime, manage layers, generate metadata, and track updates, the organization can encode much of this work in shared builders and buildpacks. Developers own their code, dependencies, and service behavior. Producing a compliant image becomes a platform responsibility.
Because Cloud Native Buildpacks is a CNCF graduated project, its specifications and reference implementations undergo rigorous community review and long-term maintenance, making them suitable as the foundation for enterprise security controls.
With that foundation in place, applying and maintaining specific container security controls becomes easier at scale.
Four container security controls buildpacks make it easier to operate
Standardization of build inputs
The first control is standardizing what goes into the build.
In Cloud Native Buildpacks, the key unit is the builder. A builder packages the buildpacks, lifecycle, build-time base image, and runtime base image used to create the final application image. This establishes the builder as a controlled definition of how application images are produced. Developers do not choose a random base image on Docker Hub to build their applications; the Buildpacks ecosystem defines a set of build and run images.
This builder-based approach introduces an important concept. A developer cannot change the base OS layer in a builder with a single line of code because compatibility isn’t guaranteed. With one line of code, however, they can swap builders and still get a compatible, working image.
This is where standardization takes place. Developers keep using their preferred languages and frameworks, but the platform team builds images from a controlled set of approved builders. Extension paths for cases where buildpacks require customization can also be standardized.
“By controlling the builder, the organization also controls the buildpacks, runtime image family, lifecycle version, and build paths allowed in CI/CD.”
The security benefit is straightforward. By controlling the builder, the organization also controls the buildpacks, runtime image family, lifecycle version, and build paths allowed in CI/CD. The policy is defined once at the platform level and then applied across many services.
Best container security practices by default
After the builder standardization, the next question is: what does the resulting image look like? This is where buildpacks help. They improve the output by applying several container security practices as part of the normal image creation flow:
Non-root build and execution. Cloud Native Buildpacks require buildpack code to run as a non-root user. Platforms, such as pack, also produce images configured to run applications as the non-root user defined by the run image. Non-root defaults reduce the blast radius of a compromise and make container escape, filesystem tampering, and exploitation of the image build process more difficult.
Separation of build and runtime environments. Buildpacks mark layers as build-only or launch-time. Only launch layers are included in the final image, so compilers, npm tooling, build caches, and similar tools remain outside production, reducing the attack surface.
Restricted modifications of the base image. Because buildpacks run without root privileges, they cannot simply install OS packages or modify the base filesystem. OS changes must be provided through the controlled build/run images or image extensions.
Isolation of sensitive build privileges. When using an untrusted builder, sensitive lifecycle phases can run separately from the phases that execute buildpack code. This prevents an untrusted buildpack from receiving capabilities such as registry credentials and access to the container daemon.
Some builder providers offer additional security features. For example, Paketo buildpacks for Spring Boot provide base images without a shell. Another example is BellSoft’s hardened builder for Paketo buildpacks based on BellSoft Hardened Images.
These defaults reduce the manual effort and make secure container builds part of a standard process.
SBOM Generation
A Software Bill of Materials (SBOM) lists all software components in an application or image. An SBOM is required for vulnerability tracking, license reviews, and audits, and many regulations require it.
Teams often add SBOM generation as a separate CI step, with separate tools, formats, storage rules, and ownership. That makes coverage uneven, especially across many repositories.
Cloud Native Buildpacks reduce this friction by emitting an SBOM for all dependencies they provide, in formats such as CycloneDX, SPDX, or Syft JSON. As a result, platform and security teams get consistent image inventory data without requiring every application team to build its own SBOM process.
For a typical Java service, that might mean SBOM entries for the JRE version, Spring Boot version, and key libraries pulled in at build time. For Node.js services, it might include the Node runtime version and major npm dependencies. The exact content will vary, but the pattern is the same: SBOM data comes from the build itself, not a separate, manually maintained process.
Patching at scale
As discussed earlier, the hard part is often not writing or receiving a patch, but getting it into the running workloads. Buildpacks shorten that path by using shared builders, buildpacks, and run images. Platform teams can then update these common inputs once, instead of waiting for every application team to repeat the same change.
Another powerful feature of Buildpacks that promotes rapid patching is rebasing. When OS-level fixes become available, the runtime base layers can be replaced with layers from a newer run image without rebuilding the application from source.
Kubernetes tools such as kpack can automate this process. They track image resources and trigger rebuilds when the source, builder, buildpacks, or stack changes.
Rebasing has limits: it updates only the run-image layers. Dependencies added by buildpacks, such as a JRE or Node.js runtime, usually require a rebuild with updated buildpacks or dependency versions. For a Node.js service, a rebuild might mean picking up a new Node.js runtime version and updated npm packages, not just a newer OS base.
So, we are not talking about magical automatic patching here. The key benefit is a more centralized, repeatable patch-propagation path across the image estate—though teams still need to handle testing, rollouts, and exceptions.
The new patch path after buildpacks
Let’s return to the original story. We left the organization in a state where vulnerability management lacked consistent patch integration. Suppose the enterprise has migrated their workflows to Buildpacks — how has the process changed?
Once again, a critical vulnerability is discovered. But after adopting buildpacks, the response looks different: the remediation starts with shared build inputs.
The platform team updates the approved run image, builder, or buildpacks. Images are then rebased or rebuilt:
Rebase when the fix affects the run-image OS layers;
Rebuild when the affected component is runtime, like a JRE or Node.js runtime, or application dependencies.
Of course, Buildpacks do not remove the need for testing or deployment controls. Application teams still own their code, dependencies, and compatibility testing. SRE teams still promote, deploy, monitor, and, when necessary, roll back the patched images.
“Buildpacks give organizations one controlled way to build container images instead of leaving every repository to define its own process.”
The most important change is the patch path. Platform teams maintain the shared build inputs and automation. Security defines scan policies and exception rules. Compliance defines the evidence that must be retained.
This path also promotes shifting security left, as it becomes easier for developers to meet the in-house security requirements and container security best practices.
Function
Responsibility
Platform
Builders, run images, buildpacks, and rebuild automation
Security
Vulnerability policy and exceptions
Application teams
Code, dependencies, and compatibility testing
SRE
Promotion, rollout, monitoring, and rollback
Compliance
Audit evidence and retention
Some workloads still need custom images. But you don’t need to choose between Buildpacks and Dockerfiles; you can use both. When required, you can create a custom buildpack and extend the build-time base image with a Dockerfile. The result is not a rigid “Buildpacks only” model, but a governed image strategy with controlled customizations.
Buildpacks give organizations one controlled way to build container images instead of leaving every repository to define its own process. This makes security controls–such as approved images, safe defaults, SBOMs, and patching–easier to scale. The result is reduced drift, clearer ownership, and faster updates.
If you’d like to get started with Buildpacks, try them on one service and compare the workflow with your current Dockerfile process.
Anthropic’s Model Context Protocol (MCP) went into production in late 2024. It spread rapidly after that.
Since then, thousands of MCP servers have been created. Microsoft, Google, and OpenAI embraced it. The Linux Foundation took over protocol maintenance. Today, MCP is considered critical infrastructure. It sits between an AI agent and the tools and data it interacts with. Most teams implemented it the same way they implement any other integration standard. People installed it and trusted the defaults.
The understanding that emerged in 2026 is that the problem wasn’t in the infrastructure. The problem is in the permissions below this infrastructure.
This matters because it changes how people approach MCP security problems. A patch solves a specific problem on a particular server. The permission change requires asking a complicated question.
“The understanding that emerged in 2026 is that the problem wasn’t in the infrastructure. The problem is in the permissions below this infrastructure.”
Why did that particular server require access to something it never needed to be there? According to the SANS 2026 Identity Threats Survey, which surveyed more than 500 security experts, 76 percent of businesses noted an increase in non-human identities. 74 percent of businesses use AI systems that rely on standing credentials to work independently. This same survey revealed that none of the protection measures, such as approval processes, sandboxing, or logging, is used by more than 40 percent of businesses.
Look past the individual disclosures, and the same root cause keeps showing up. In May 2025, an attacker used prompt injection against the GitHub MCP server to pull private repository data, not because the server had a bug in the traditional sense, but because the personal access token behind it was scoped far wider than the task required. Days later, a logic flaw in an Asana MCP integration allowed cross-tenant access because the permission layer never enforced the isolation boundary between customers.
Security researchers now group it under a couple of recognizable patterns: tool poisoning, where a server’s own tool description carries hidden instructions, and the confused deputy problem, where an agent inherits more trust than the task in front of it requires.
What a permissions redesign actually asks you to check
The solution that keeps coming up isn’t a better scanner, but compartmentalizing access. GitHub’s Engineering Blog, which discusses developing secure remote MCP servers, suggests the following: every instance must have its own secrets for the specific task, all requests must be limited to the acting user, and authorization must be based on action rather than assumed after user authentication. Replace fixed, permanent tokens with dynamic, temporary credentials generated on the fly.
“Replace fixed, permanent tokens with dynamic, temporary credentials generated on the fly.”
This solution is tiered and has been working until now. In Webflow, we treat MCP integrations in the same way as we treat other third-party components with access to customer data.
Each credential the team gives to an AI agent was probably a good idea when it was provisioned. The tough call isn’t whether that access was a good idea at the time. It’s whether it is anymore, and most teams aren’t in the habit of making it.
Some things to consider while integrating any MCP:
What can this credential reach now, rather than the scope for which it was intended? The scope of access is likely to creep. No review will be scheduled until something breaks.
Is authorization granted on a per-site, per-repository, or per-Workspace basis, or all or nothing? Be leery of any integration that only provides organizational access. If an integration doesn’t give you control over scope at connection time, that’s the finding, not a footnote.
Does the AI agent inherit the person’s existing credentials, or create entirely new credentials that bypass those permissions? The latter is how a GitHub personal access token can have more access to repos than the user who authorized it.
Does logging assign accountability for what the agent does in the same way it does for a human? If the agent’s activities are invisible or unattributable, incident response starts at ground zero.
Do changes from the agent go straight into production, or do they pass through a reviewable process like draft, branch, and approval queue first? This is just applying the security discipline the team already has around human access controls to a newer class of entity.
Identity comes before access
Before you can talk about what an agent is allowed to do, you have to answer a harder question: what is an agent, identity-wise? Right now the honest answer for most of the industry is “a human’s OAuth token wearing a trenchcoat.” The agent doesn’t have its own identity. It inherits the scope, the blast radius, and often the literal credential of whoever spun it up.
That’s a problem the moment agents stop being ephemeral. Most agents today live for minutes to hours: a task starts, the agent runs, it dies. But that’s changing. We’re heading toward agents that run for weeks or months, and a thing that lives that long needs its own identity, not a borrowed one, with permissions that get stricter, not looser, as the lifespan grows.
“Right now the honest answer for most of the industry is ‘a human’s OAuth token wearing a trenchcoat.'”
Think of it like the difference between a contractor you bring in for an afternoon and a contingent worker embedded in your systems for a quarter. You wouldn’t give the afternoon contractor a permanent badge, and you shouldn’t give the quarter-long agent the same access as a five-minute script.
OAuth wasn’t built for this, and it’s not just a missing feature; it’s a structural mismatch, and at root a UX failure: a consent model built for a human in the loop, applied to a process that has none.
Its whole model assumes a human sits in front of a scope dialog and makes an informed choice, and we all know how that goes: nobody reads the scope list; they click allow. That already-shaky assumption collapses completely when there’s no human reading anything.
The base spec has no concept of “this client is an agent” or “this grant is expected to run for six months,” just a server-set expiry after the fact. A few IETF drafts are starting to sketch a fix: binding token lifetime to a task’s actual lifecycle and tagging agents with stable identities distinct from the human who invoked them. Still, those are early-stage proposals, not deployed standard(s).
We don’t have a clean answer for where the line sits between “short-lived task, broad-ish access” and “long-lived agent, locked down tight.” Webflow’s security team is actively working through this, and we’re comparing notes with peers across the industry rather than pretending we’ve solved it. But the framing itself matters: treat agent lifespan as a first-class input to your permission model, not an afterthought.
Permissions aren’t a checkbox at provisioning
A permissions redesign, a protocol update, and an identity model represent three distinct levels of the same problem. For each level to remain effective, the other two must be effective too.
If you create a credential with the correct scope on Day One, but then no one ever verifies that the credential remains valid, your efforts were wasted. Similarly, if you develop a new protocol that finally distinguishes an agent from the human behind it, but all integrations continue to provide standing, all-or-nothing access by default, you have done little good. Neither approach addresses the deeper question at the root of both: what an agent is permitted to do should depend on how long it will exist.
“The teams who get this right will be the teams who stopped viewing scope, identity, and lifetime as three separate evaluations, and began treating them as a single setting.”
Right now, nearly all components of the technology stack don’t know how to ask that question, much less answer it. The teams who get this right will not be the ones who developed a more efficient scanning tool. Rather, they will be the teams who stopped viewing scope, identity, and lifetime as three separate evaluations that occur once per agent instance, and began treating them as a single setting that must be evaluated each time the agent’s function or existence changes.
This article was originally published on September 9, 2026, on webflow.com.
I’m Matt Burns, Chief Content Officer at Insight Media Group. Each week, I round up the most important AI developments, explaining what they mean for people and organizations putting this technology to work. The thesis is simple: workers who learn to use AI will define the next era of their industries, and this newsletter is here to help you be one of them.
Former OpenAI and Anthropic researcher Jacob Coxonmade a lot of noise this week with a viral thread on X. He has since talked to Wired and Axios. His warning deserves attention, but another revelation this week gives developers something concrete to investigate.
In retrospective testing of the Mythos 5 incident, Anthropic’s offline monitor flagged around 1% of actions. When researchers removed the model’s chain of thought and showed the monitor only user turns, tool calls, and tool results, that rose to about 50%.
The model’s written reasoning is what talked the monitor out of flagging that harmful behavior. The offline monitor kept concluding it was operating in a simulation.
The lesson is to test whether your monitoring accepts an agent’s justification a little too easily. An AI’s explanation can help a safety monitor reveal a problem, or persuade it to essentially convey, “move along, nothing to see here.”
An AI’s explanation can help a safety monitor reveal a problem, or persuade it to essentially convey, “move along, nothing to see here.”
Coxon is warning about the pursuit of self-improving superintelligence. Developers are already pointing agents at real credentials during everyday work. Those problems differ substantially, but both deserve attention.
Forecasts aren’t checkable. Incident reports are.
Jacob Coxon’s thread is short and hard. He says neither of his former employers is acting responsibly, and that both are “racing straight to self-improving superintelligence and gambling with our lives.”
Maxwell Zeff pressed him on the end-of-humans mechanism twice in an interview with Wired. Coxon described shutdown avoidance, then biological and cyber threats. His answers explain the concern, but leave substantial questions about how AI will kill us all.
Annie Jacobsen, whose books on this stuff I wholeheartedly recommend, spelled out the scary biological scenario: A system that can hack anything, turned loose on the world’s 3,600-plus BSL-3 and BSL-4 labs, is how you get from a software problem to a biological one. For me, Jacobsen’s commentary is the most concrete, real-world implication anyone has given of what Coxon is actually worried about.
Two incident reports published in the last few weeks are more useful anyway. OpenAI’s report on the Hugging Face incident, published August 26, traces the whole thing from May 12 to July 20. The agents found a hole in an internal package server and used it to reach the open internet that the same server was supposed to keep them off of. OpenAI rebuilt the server; they got back out. By July 11, they were far enough into Hugging Face to run their own commands on its machines.
The model doing most of that was never released. It’s a research version OpenAI runs in-house, with fewer safety checks than the ChatGPT everybody else uses. When the company ran the same work again with those checks switched on, it found that this kind of behavior could drop by more than 100 times.
That context matters.
Amanda Caswell covered this in August. The incident demonstrates a serious failure under particular conditions. It also illustrates why the surrounding safeguards belong in any assessment of an agent’s behavior.
The monitor was talked out of it
Anthropic’s assessment, published on September 9, covers four incidents and analyzes the first three in a scan of roughly 141,000 transcripts, identified a fourth in August while assembling material for METR, then widened the search to roughly 481 million transcripts and reported no additional incidents of similar or greater severity. In each case, the model reached the open internet through a misconfiguration rather than by breaking out of a properly isolated sandbox. That’s an important detail. It makes the evaluation infrastructure part of the story.
If you’re a developer, you should look at the monitoring results. Mythos 5 kept reasoning its way to the conclusion that it was operating in a simulation, and the offline monitor read and accepted that reasoning. The limits matter here too: These were cybersecurity evaluations running without released cyber safeguards, Anthropic reports no coordination between models and no concealment, and its independent METR review was planned rather than finished. This is not the same failure as the OpenAI incident. Don’t equate them as similar incidents
But yet the shape of it is familiar, and it’s what Adrian Bridgwater reported on last week in his piece on agent observability, which is a useful read if you’re producing agents (who isn’t). Jacob Krell of Suzu Labs put it about as well as it can be put: “The model can recite your constraints and reason past them in the same breath.” Liran Hason of Coralogix added the result: “System guardrails help, but a guardrail only stops what the developer already thought of.”
What developers can take from Steven Adler’s recommendations
Largely overshadowed this week by Coxon, Steven Adler’s op-ed in The New York Times asks AI companies for four things: disclose incidents including near misses, the way aviation does; keep tamper-evident records of model behavior; never let a model cut power to its own alarm systems; and formally swear off training techniques that undermine researchers’ ability to detect deception.
Adler worked for OpenAI from 2020 to 2024 and later founded a nonprofit called Guidelight AI Standards. It has since graded six control practices at the frontier labs using public information and gave out a top mark of C-plus.
Guidelight also publishes a Control standard for frontier labs, recently revised in August. It wants monitors to access the chain of thought and scan both reasoning and actions.
But those asks of frontier labs don’t apply to most developers. Here are five you can check on your own setup right now.
Five checks you can run on your own agent setup
Steven Adler’s asks of the frontier labs, applied one layer down.
Check
What passes
Someone reads the near misses
An agent crosses a line and nothing breaks. There’s still a record, a person who looks at it, and a copy the agent can’t reach.
The agent can’t edit its own history
Logs are append-only and stored outside the agent’s reach. If its credentials can delete the log, the log is decorative.
The agent can’t touch its own oversight
Separate tokens help. Then check whether it can shut off monitoring, rewrite an approval rule or get around the kill switch.
Rule-breaking gets blocked and logged
In a test environment, hand it a job it can only finish by breaking a rule. Judge the monitor on what it does with the action, not with the agent’s account of it.
The explanation doesn’t change the answer
Run the same rule-breaking action past your monitor twice. Change nothing but the explanation, one flat and one reassuring. If the reassuring version gets flagged less, you’ve found the problem.
These checks do not establish that frontier alignment is solved. They give a team evidence about the system it operates on and a clearer account of where its protections need work.
Anthropic’s report is why I’d include that last test in a deployment review. Hold the action and its permissions constant, change the explanation, and see whether your monitor changes its judgment.
That won’t settle Coxon’s warning. Will AI kill us all? I don’t know. But I know developers need to test their agents’ permissions, protect their logs, and make sure a convincing explanation can’t talk a safety monitor into overlooking an unauthorized action.
This week, Anthropic acknowledged that the three cyber incidents it disclosed this summer weren’t just the result of a misconfigured test environment. It turns out that Claude’s own behavior was part of the problem.
Recall in July when the AI company released a report on three cases where Claude models reached the open internet from misconfigured test environments and compromised real third-party systems — a telling example of the limits of AI safety tests and why AI evaluation infrastructure needs production-grade security.
Upon deeper review, Anthropic now says that explanation was incomplete. Beyond the operational failure, the models themselves displayed two recurring alignment failures: biased reasoning and recklessness. And, actually, there was a fourth incident that the AI company only found out about after further investigation.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately.”
The incidents definitely don’t prove this kind of risk, but they do illuminate how models can behave unexpectedly — and how existing evaluations can fail to catch it.
The problem was more than an operational failure
This summer, when Anthropic shared its initial findings, it said the three cyber incidents were “closer to a harness and operational failure than a model alignment failure,” with Claude seeming to believe it was operating in a simulation without internet access.
But a closer look at the transcripts revealed that the problem was not simply the model reaching the internet by mistake; Claude also showed signs of misaligned reasoning.
“When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm.”
After widening the initial search to about 481 million transcripts, including those from the company’s Frontier Red Team, as well as non-cyber evaluations, reinforcement learning (RL) environments, subagent logs, and others, Anthropic conducted two scans.
First, it scanned these transcripts for signs of internet access, whittling down the pool to 9.2 million flagged for escalation. It then used Claude to review this smaller batch, re-identify the known incidents, and look for others of similar or worse severity.
What did it find? It found more than an operational failure. Anthropic says its investigation identified “two recurring alignment issues, present at varying levels of severity across the incidents.”
It names these issues as biased reasoning, where Claude disregarded or misinterpreted evidence that was, in fact, on the internet, and recklessness, where Claude was willing to take harmful actions to pursue a task.
Anthropic was already concerned that Claude Mythos 5 uploaded a malicious package to PyPI, but the model repeatedly said it believed it was operating in a simulation. Anthropic now says the simulation explanation does not fully account for the model’s behavior and that its reasoning appears to have been biased.
“When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm.”
And there was a fourth incident Anthropic missed the first time
Anthropic’s new alignment assessment also contains another less than reassuring revelation: It missed one of the four incidents the first time around.
As the AI company explains, its first review relied on an agentic search due to “the volume of transcripts and [its] desire to disclose incidents quickly.” But that scan missed a set of transcripts that had also had internet access.
How did the cat get let out of the bag? It happened in August when Anthropic was assembling transcripts to share with METR (Model Evaluation and Threat Research). This research nonprofit measures whether and when AI systems might threaten catastrophic harm to society. The fourth incident dates from January 2026 and involves an early version of Claude Opus 4.6.
Discovering the fourth incident pushed Anthropic to broaden its search to the roughly 481 million transcripts described above, giving it more information to analyze and ultimately surfacing new insights into Claude’s “biased reasoning” and “recklessness.”
While this broader search didn’t turn up any more incidents of what Anthropic describes as “similar or worse severity,” the investigation isn’t over. Per the AI company, it’s signed an agreement with METR, granting the organization wide-ranging access to transcripts and Anthropic employees for eight weeks to investigate the incidents independently. The agreement can be extended if both parties agree.
Anthropic reproduced the failure to try to understand if newer models would do the same
Beyond reviewing the original incidents, Anthropic is taking other steps to try to understand how things went pear-shaped in the first place. It put Claude Opus 5 and Claude Mythos 5.1 in a simulated replication of the Mythos 5 scenario — the one where the model uploaded a malicious package to PyPI. Anthropic reports:
“Both take harmful actions substantially less often than Claude Mythos 5 did, but they still engage in the same behaviors at concerning rates.”
It notes that both models only took harmful actions after stating that they were, indeed, acting in a simulated environment. Anthropic also says that Mythos 5 appears to be “an outlier in terms of poor performance” on evaluations for biased reasoning, recklessness, and willingness to probe boundaries.
“Both take harmful actions substantially less often than Claude Mythos 5 did, but they still engage in the same behaviors at concerning rates.”
Still, it’s not such a clear indication of how the newer models would behave in the real world. As the AI company adds, “because the reproduction is, in fact, a simulation, we are cautious about how far these results generalize to the real world.”
Additionally, Anthropic looked back into its training to search for a root cause of Claude Mythos 5’s biased reasoning, but came up short. While it didn’t identify a single root cause, the company says biased reasoning decreased across its production models over time, potentially due to updated RL and alignment training.
Anthropic admits it still has work to do. Since its pre-release auditing didn’t catch misalignment of this severity, it’s since added evaluations to specifically target these behaviors, though reliably doing so is still a challenge.
Anthropic admits the failures were serious, but says they’re “narrow”
While Anthropic clearly calls out the four cyber incidents as “valuable warning shots,” it also points out that their scope was limited:
“While Claude’s actions were misaligned, they remained within a narrow scope,” the AI company says, pointing to the fact that all incidents included a single Claude instance and the model never attempted to coordinate with other agents or conceal evidence of its actions.
It says these behaviors playing out in ordinary use are “unlikely” and that, if they did, the safeguards shipped with production models would add more defenses that didn’t exist in these evaluations.
But following Coxon’s remarks about the risks of superintelligence, and Anthropic’s own alignment science lead, Evan Hubinger, responding that Anthropic “really do[es] earnestly believe AI could kill all humans,” seeing Claude go off the rails isn’t comforting.
Bob from finance built a scheduling tool last month. He described it to an AI assistant on a Sunday afternoon, wired it into Slack and three internal APIs before dinner, and by Monday three departments depended on it. That story is funny right up until you check what token it is running on.
What changed a few weeks ago
On July 28, 2026, the Model Context Protocol’s maintainers shipped a specification update built almost entirely around authorization: issuer validation, issuer-bound client credentials, and Client ID Metadata Documents as the preferred way for clients to register. Put plainly, that is the protocol’s own stewards admitting the original trust model didn’t survive contact with production.
If the people who wrote the spec needed a security overhaul this deep in, the tool Bob built on a Sunday does not get a pass either, and Bob has never heard of issuer validation.
Three ways this actually breaks
Tool descriptions carry instructions, not just documentation. In May 2025, researchers at Invariant Labs showed that GitHub’s own MCP server could be hijacked through a poisoned public issue: an attacker’s text in an issue body was read as an instruction by the agent and used the victim’s token to pull data from private repositories. No compromised code, no malicious tool, just a description field nobody thought to sanitize.
A 2026 benchmark called MCPTox tested this pattern against 45 live MCP servers and 20 models, measuring a 36.5 percent average attack success rate and 72.8 percent against the worst-performing model. Bob’s tool has the same basic shape: it reads Slack messages and ticket text to decide what to reprioritize. It can’t tell the difference between a coworker’s request and a string engineered to look like one, because nobody asked it to.
Scopes default to everything. The common failure isn’t a missing permission model; it is an ignored one. A server that needs read-only calendar access asks for read, write, and admin across the board because that is what the tutorial used. Across the MCP ecosystem, 88 percent of servers require credentials to function, but only 8.5 percent actually use OAuth.
Most of what is running was never scoped in the first place, so there is no scope left to creep. Bob did not sit down and choose a scope. He reused the admin-level API key already sitting in his password manager from a reporting dashboard he set up two years ago, because requesting a narrower one meant filing a ticket, and filing a ticket was the entire bureaucracy he was trying to avoid.
Static tokens do not rotate, and nobody is watching them not rotate. Splunk’s own MCP Server app logged session and auth tokens in cleartext until it was patched in version 1.0.3, tracked as CVE-2026-20205. That vendor has a security team.
By some estimates, over half of MCP servers in the wild run on static API keys or personal access tokens that are rarely rotated, and close to half of enterprise AI activity runs through personal accounts rather than service accounts, meaning the credential doing the work belongs to somebody’s identity, not the system’s.
Bob’s token is that same reporting-dashboard key. It has been valid since it was issued; it will stay valid until somebody remembers to kill it, and the only record of what it has touched this month lives in Bob’s memory — which is not a log.
What we have actually seen
While setting up our own MCP integrations across customer and prospect environments over the past several months, we found that more than 20 percent of the MCP-related access policies we reviewed were either broken or missing entirely.
We found that more than 20 percent of the MCP-related access policies we reviewed were either broken or missing entirely.
In most cases, the MCP server in question was authenticated with someone’s personal token rather than a service account. None of those tokens had a documented rotation schedule. None of the servers had logs of what they touched. If Bob’s tool had been in that batch, and statistically it probably would have been, nobody would have known until something went wrong, because right now nothing is watching for it to go wrong.
The honest caveat
None of this means every vibe-coded integration needs a change advisory board. Most of what Bob built is harmless, and gating every weekend project behind a formal review process is exactly how you get back to the eighteen-month procurement cycle nobody missed. Governance has its own cost, paid in the good ideas that never ship because process ate the weekend momentum that made them possible.
The problem is not that these tools exist. It is that most organizations currently cannot tell the difference between the harmless ones and the ones holding a token that reaches production, and they are trying to solve that with the same review board that made Bob route around them in the first place.
The actual decision
The question in front of every platform team right now is not whether to allow AI-built integrations. That decision was already made over a weekend, without anyone in the room. It is whether you find out what a given MCP server can touch from an inventory you built on purpose, or from an incident report after the fact. Bob’s scheduling tool is still running. It has not caused an incident, and it probably never will.
But the difference between Bob’s tool and the next one that makes the news isn’t the code; it is whether anyone can say what token it holds, what it can reach, and when it was last rotated.
A security researcher testing a 300-person B2B company with a global footprint discovered an internet-exposed database with weak authentication during a routine scan. The database appeared to be a prime target for attackers; it had critical severity and was an obvious first-fix candidate. Upon further inspection, however, researchers learned it was a resettable test database used to test job candidates, rather than a system containing client data.
The episode represents a critical problem: Scanners and security researchers can’t infer the real cost of a compromise on their own. Increasingly, companies rely on an external partner to handle cloud vulnerability triage and keep teams from becoming overwhelmed. Scanners produce excessive noise, requiring focus and attention from engineers to understand alerts and complex attack techniques, tune out false positives, and work with teams on remediation, reducing the time they have to build new features, work with customers, or scale up their tech environment.
Tech employees have more on their to-do lists than ever: Product teams face an infinite stream of feature requests and engineering teams carry more technical debt than they can clear. When you factor in a reorg, many employees may have larger project scopes and leaner teams, all while constantly monitoring, correcting, and mentoring AI agents. Some reports cite 90-hour work weeks.
Beyond today’s structural challenges, security teams are drowning in a different type of data. Programs absorb identity events, firewall logs, endpoint alerts, vendor feeds, and threat intelligence, producing far more analysis data than even a few years ago. Every item can look urgent, high-risk, and worthy of immediate attention. With finite hours, it’s hard to know where to act first.
Jon Rose, founder of the information security and risk management advisory firm IOmergent, has seen firsthand how AI and near-universal tooling are producing more findings than teams can triage.
“Within the span of security work, there’s an unending list of things you could tackle, and you’re pulled in so many different directions… But you have to be ruthless about prioritizing and investing your time.”
“Within the span of security work, there’s an unending list of things you could tackle, and you’re pulled in so many different directions,” Rose tells The New Stack. “But you have to be ruthless about prioritizing and investing your time.”
The challenge, then, isn’t remediation. It’s allocation: Deciding where limited engineering and security capacity will make the greatest impact is the most important question security leaders answer every day. A technical severity score, used without threat and environmental context, cannot answer the business question: What should we fix first? Effective vulnerability prioritization turns raw findings into business-aware priorities, and that judgment is the real work.
CVSS limitations: a starting point, not a decision
For security teams, a Common Vulnerability Scoring System (CVSS) provides a useful baseline. Its base metrics classify a vulnerability’s severity — attack vector, complexity, required privileges, and potential impact on confidentiality, integrity, and availability. But a severity score is designed to be stable across environments, so it can’t tell a team whether an asset is exposed to the internet, shielded by compensating controls, or central to the business.
Essentially, treating a base severity score as an automatic fix-first ticket is problematic, rather than using CVSS on its own.
CVSS can include Threat and Environmental metrics that account for evolving exploit conditions and organization-specific context. However, risk-based vulnerability management still depends on accurate knowledge of the environment — and on someone applying that context consistently.
“The piece that’s missing from any of these tools is the grounding in the business, the understanding of what actually matters,” Rose says.
Escalate the internet-facing medium
A smart decision framework deprioritizes the urgency of a score and probes reachability: whether an attacker could actually reach the vulnerable component.
“Is the affected service exposed to the public internet or not? Or is it isolated behind network controls and accessible only to a limited set of internal users?” Rose says. “Often, issues will get flagged, but it’s not in a position where it could be triggered.”
Next is consequence. An exposed flaw on a disposable test system might be a genuine security concern. Still, it doesn’t carry the same weight as a weakness on the application that processes customer transactions, stores sensitive information, or underpins a company’s main revenue stream. So while the severity might be alarming, the business outcome might not.
Teams should also ask whether the weakness is attracting active attacker interest and where it could lead. A vulnerability listed in CISA’s Known Exploited Vulnerabilities catalog deserves urgent scrutiny, because there’s evidence of exploitation in the wild.
The Exploit Prediction Scoring System provides a forward-looking signal: An estimate of the likelihood that exploitation of a particular vulnerability will be observed over the next 30 days. Neither replaces a business decision, but both help distinguish a theoretical risk from one that demands attention now. However, as AI-driven exploitation accelerates, the window between what is known to be vulnerable and what is actively exploited is shrinking because the cost of building and weaponizing exploits is dropping.
The relevant attack path might also extend beyond the affected machine. A comparatively modest flaw can become an immediate concern if it provides a route into privileged accounts, a production system, or customer data. Conversely, a high-severity finding can be deprioritized — temporarily — if it is not reachable, has limited impact, and is protected by reliable controls.
“The speed and the depth of research and investigation into those security issues are going faster,” Rose says. “So it can change really quickly.” The worst state is unacknowledged risk sitting in the backlog. Even the best detection degrades when no one owns the trend line. Teams should track exceptions, set review dates, and name an owner. Accepted risk is still risk; the difference is that it’s explicit, time-bound, and revisited.
An operator capability, not a weekend project
Security teams are leaning more on AI to prioritize and triage threats, and they’re uncovering vulnerabilities at an unprecedented pace. Yet the sheer volume of findings is overwhelming, even for the most efficient of humans.
It’s clear, too, that AI can make it easier to discover less obvious paths to exploitation, and to introduce new ones. An academic study of 20,000+ issues fixed by AI found that LLMs introduce nearly 9x as many new vulnerabilities as developers, exhibiting unique patterns not found in developers’ code. The answer is to use AI at the outcome level. For every alert, the goal is to have the SOC analyst’s standard questions answered quickly: new or known, what’s exposed, what data is at risk, prod or dev, and how it has evolved.
“Effective programs start by aligning with executive teams to understand the business — where the company is going — so allocation and adjustments track the actual risk, not just the score.”
“That’s how teams get thousands of alerts down to 10 to 20 prioritized tickets,” Rose tells The New Stack. “Effective programs start by aligning with executive teams to understand the business — where the company is going — so allocation and adjustments track the actual risk, not just the score.”
As to-do lists grow ever longer, business-context judgment that’s applied every day by someone who owns it is something worth dedicating more resources.
Every pull request submitted by an OpenAI engineer now goes through an automated security review, and the AI model can stop code from being merged if it finds a vulnerability.
Thibault Sottiaux, engineering lead of OpenAI’s Codex team, described the system in a recent interview on The Pragmatic Engineer, explaining that the security check is mandatory and doesn’t require a human reviewer to enforce it.
Security review is only one of the jobs OpenAI is handing to its models. They’re also reviewing code, catching regressions, handling dependency upgrades, and helping engineers tackle changes that Sottiaux says might previously have taken months. OpenAI has even started benchmarking some of its code-review models as “superhuman.”
AI reviewers block every merge
OpenAI started training specialized code-review models early in Codex’s development. In the interview, Sottiaux described models that could catch logic and reasoning mistakes a human engineer might spend hours on.
“When we benchmark them, it’s like they’re superhuman in code review,” Sottiaux said. “This is not just true for correctness. This is also true for security.”
Those capabilities started in standalone review models and have since folded into OpenAI’s mainline models. For security, a flagged issue blocks the merge without exception.
“When we benchmark them, it’s like they’re superhuman in code review,” Sottiaux said. “This is not just true for correctness. This is also true for security.”
Intent replaces inspection
As AI takes over more of the mechanics of reviewing code, Sottiaux thinks the human role may move earlier in the process.
OpenAI’s review, deployment, and regression-catching processes are already, in Sottiaux’s words, “pretty much automated.” Engineers can ship a PR the same day to ChatGPT, which he said serves roughly a billion active users.
“Really what we see, and I see, is there’s this sort of discussion around the intent that takes place around the pull request,” Sottiaux said. “It’s like, what are you even trying to do? And is that the right thing to attempt to do?”
“Really what we see, and I see, is there’s this sort of discussion around the intent that takes place around the pull request,”
That thinking needs to happen earlier, Sottiaux argued, back in planning instead of waiting for the review queue. Engineers still have to agree on the goal and kick the tires on a proposed change. Passing the review burden to AI doesn’t take humans out of the loop; it just moves the gut-check to before anyone opens a PR.
Agents tackle maintenance backlogs
While security gets most of the attention, basic maintenance may be where engineering teams feel it first, especially with third-party libraries that push breaking changes and get deferred sprint after sprint because new features always take priority. Sottiaux’s point is that as long as you have a clear changelog and decent documentation, an agent can knock out those tedious updates in an afternoon, and the same goes for routine security patches.
The same calculation applies to bigger refactoring jobs. A team might know exactly what it wants to clean up and even have a better architecture in mind. Still, once the estimate comes back at two or three months of engineering work, it’s easy to understand why everyone keeps working around the problem instead. The code may be ugly, but it works, and there are always other things that need to ship.
That changes when an agent can take on much of the work. A cleanup that would have been shelved because nobody could justify spending a quarter on it might suddenly take days instead of months, which makes it a much easier project to say yes to.
When models outgrow their scaffolding
Sottiaux described a dynamic in agent development that runs opposite to how most software evolves.
Codex had a command called /goalbuilt to keep a model focused on a single objective for days or weeks without drifting. A “crutch” (Sottiaux’s own word) to patch the model’s tendency to lose the thread on long-running tasks, but newer models don’t need it.
“You don’t need slash goal anymore. You don’t need a harness around it,” Sottiaux said.
The Codex team often has to build extra infrastructure around a model to make up for what it can’t do yet, only to find that the next generation can handle the same behavior on its own and the code they built around the previous model is no longer needed.
Sottiaux said the team now factors that into its planning, sometimes deciding not to build a workaround if researchers expect the next model to solve the problem on its own within a few months. As the models improve, the system prompt and the code surrounding them can get smaller, while parts of the product that once seemed necessary disappear altogether.
The blind spot question
If AI writes more of the code and AI reviews that code, both systems can share the same blind spot. That’s the obvious objection, and Sottiaux’s interview doesn’t fully address it.
OpenAI trusts these models enough to let them block a pull request, which makes their mistakes matter in a very practical way. If the model is too cautious, engineers end up waiting on code that was fine to begin with. If it misses a real vulnerability, that code could move ahead with an automated security check giving everyone reason to believe it was safe.
The job also gets messier as AI-generated code proliferates. Code can compile, pass its tests, and still have problems that aren’t obvious from the pull request itself. Some of those problems may not even start with the code an engineer is submitting. When the supply chain is the attack surface, the weak point could be a dependency compromised weeks or months earlier, leaving a PR reviewer to catch a problem that originated elsewhere.
Code can compile, pass its tests, and still have problems that aren’t obvious from the pull request itself.
Security operations centers have struggled with alerts for years, and AI agents offer a new way to tackle it: Enable machines investigate some of those alerts themselves.
That’s already starting to happen: Security teams are experimenting with AI that can pull together signals from different systems, investigate suspicious activity, and recommend next steps to humans in the loop.
There’s an obvious appeal: SOC analysts have finite time and attention, while the volume of potential threats does not come with the same constraint. Attackers are also getting access to AI tools that can accelerate parts of their own operations. Simply giving analysts better ways to work through an ever-growing queue may only get security teams so far.
But moving from AI-assisted security to increasingly autonomous security creates a new problem: How much control are organizations actually prepared to hand over?
There is a big difference between asking an AI agent to investigate a suspicious login and allowing it to disable the account responsible for it. The same goes for isolating an endpoint, blocking network traffic, or making other changes that could immediately impact the business. An autonomous agent could potentially make those decisions much faster and at much greater scale than a human analyst.
The model is only part of the trust equation. Security teams also need to know what an agent is doing, when a human gets the final say and, crucially, whether they can undo a bad decision. That could mean putting some fairly hard limits on autonomy, including a way to shut the whole thing down if an agent goes off course.
Giving agents more responsibility also changes the role of the people working alongside them. If AI handles a large chunk of routine investigation, analysts could spend less time working through queues and more time threat hunting, making judgment calls and overseeing the agents doing the repetitive work. The SOC analyst starts to look less like an investigator and more like an orchestrator.
Eventually, the bigger change may be to the SOC itself. “Continuous detection and response” has become familiar security language, but AI agents could make it something more literal. Instead of detection, investigation, and response being separate steps, an agent could move between them, with what it learns during one investigation feeding directly into how the next threat is detected.
That starts to look less like AI bolted onto the existing SOC and more like a different operating model altogether. It also presents CISOs with a familiar problem: tooling. Security teams already have sprawling stacks, and vendors are racing to add agents and AI capabilities to them. Organizations risk ending up with another collection of products to manage rather than the continuous system they were promised.
Join us on September 15, 2026
On September 15, I’ll be joined by Jami Hughes, deputy CISO at Zions Bancorporation, and Oren Saban, co-founder and CPO of Mate Security and former Microsoft Defender XDR and Security Copilot product lead, to discuss alert overload, autonomous agents, the future of the SOC analyst, and what continuous security actually looks like.
The session is limited to 20–25 security leaders, with applications reviewed to keep the group small and relevant. This isn’t a traditional webinar with hundreds of people listening in: everyone in the room will be expected to take part. Chatham House Rule will apply throughout, so participants can speak candidly about what’s working, what isn’t, and where they still have concerns.
Apply for a seat at the table
Because if attackers increasingly operate at AI speed, security teams need to work out how much of the response they’re willing to hand to AI, too.