❌

Vue normale

Reçu aujourd’hui — 28 septembre 2026The New Stack

Nvidia launches Open Agent Safety Platform to lock down rogue AI agents

28 septembre 2026 à 12:00

OpenAI, Anthropic, Meta, and Google have all recently disclosed that their models broke out of their test environments and reached real systems. Nvidia’s response, announced Monday, is a runtime that locks agents into kernel-enforced sandboxes and a watchdog on its own silicon that can shut them down.

The Nvidia Open Agent Safety Platform combines OpenShell 0.1.0, the Apache 2.0 agent runtime the company announced at GTC in March, with Nvidia Sentry, a watchdog service that runs on the company’s BlueField-4 data processing units (DPUs).

The new OpenShell release adds a policy prover that checks that an agent’s various permissions can’t be combined into something the operator didn’t intend — like hacking HuggingFace.

Since the BlueField DPU is a separate processor with its own trust domain, it can watch the agent’s traffic to the model and keep an eye on all of its actions and reasoning. Then, when things go awry, it can cut the agent off at the network level.

Justin Boitano, Nvidia’s vice president of enterprise AI, said in a press briefing that the recent incidents “have highlighted a fundamental hurdle for AI agents, and that is that model-level safeguards alone can’t govern what agents can access or do.”

“To date, model safety has been about training good behavior into the model. The industry calls that model alignment,” Boitano said. “For probabilistic systems, this approach has obvious limitations. That’s why we’re introducing a deterministic system to mediate and enforce how these agents behave.”

Credit: Nvidia.

The summer of sandbox escapes

OpenAI disclosed on July 21 that GPT-5.6 Sol and a research prototype had exploited a zero-day in the package proxy that was their sandbox’s only network path and gone on to reach Hugging Face’s production database.

Only a few days later, Anthropic reported that three of its models had found unintended internet access at its evaluation partner Irregular and, among other things, accessed a real company’s database and published a live malicious package to PyPI.

Meta followed on Aug. 6 with a pre-release Muse Spark model that read and modified a real website’s database after the same Irregular misconfiguration handed it a real site name.

Most recently, Google announcedthat Gemini had broken into the networks for three companies, again from an Irregular test environment.

Irregular, which runs pre-release cyber evaluations for frontier labs, is a common factor here, and it’s worth noting that the company is on Nvidia’s partner list for this launch.

Nvidia’s own technical blog puts the breakouts down not to a new capability but to “a combination of tools, time, and ambiguous instructions.”

Asked whether the platform would have prevented the Hugging Face breach, Boitano said each incident is unique, but “from what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on.”

Enforcing policy outside the agent

OpenShell is the core of the platform, and it hasn’t changed all that much since Nvidia first showed it at GTC in March.

With OpenShell, which Nvidia originally announced in parallel with its NemoClaw distribution of OpenClaw, each agent runs in a kernel-isolated sandbox with no network access except through a supervisor that sits outside the workload.

“Traditional sandboxes, whether we’re talking micro VMs or containers or VMs, were built for application-level isolation,” Boitano said. “Every agent running within your company needs to run in its own isolated sandbox with security controls that are outside of the agent’s reach.”

A prover, not a judge

OpenShell is now at version 0.1.0, and the important new component added in this update is a policy prover. This Prover checks that the permissions a given policy grants always stay within the boundary the operator actually intended.

“It is deterministic. It is mathematical reasoning. So this is not LLM as a judge,” Ali Golshan, Nvidia’s senior director of AI software, said during the briefing. Because of that, he said, it runs “roughly at two orders of magnitude higher performance and speed.”

In Golshan’s example, a policy can, for example, bar an agent from reading code on GitHub and posting it externally.

“An agent can bypass this by spawning two sub-agents: one that can read from GitHub, that can talk to another one, that could also then post outside,” he said. The prover models the combined access of the entire agent fleet to find that path.

In Nvidia’s own tests, agents running with reduced safeguards spent up to two hours trying to talk an AI reviewer into granting write access to a protected repository. The prover, the company says, gave the reviewer evidence of what the request actually allowed, and no protected writes occurred.

Sentry: the safety island

Nvidia Sentry adds an additional hardware layer to this system. It runs on BlueField-4 in a trust domain separate from the host, and according to Nvidia, it can quarantine an agent in milliseconds.

With a DPU in the system, the agent’s model endpoint gets routed “through a proxy on the DPU, so that you can see all of the reasoning traces of the agents on the host,” Boitano said.

Unlike OpenShell, Sentry isn’t open source, though Boitano said it has open APIs and that OpenShell can work with other network enforcement hardware.

He compared it to autonomous vehicles. “There’s a primary system that might be running the perception system, and then a safety island that ensures the safety of the system.”

“The DPU is really optional in these architectures,” Boitano said. “In a lot of cases, just using OpenShell on CPUs is honestly good enough for providing sort of strict access control for the agents.”

The DPU, he said, is for “frontier use cases of evaluating models or systems where you might have the guardrails off the models, so it could be for red teaming.”

Who’s building on it

Anthropic is integrating OpenShell with Claude Managed Agents, which already keeps the agent loop on Anthropic’s infrastructure and pushes tool execution into customer-controlled sandboxes.

SpaceXAI says it’s using the platform for Cursor coding agents and Grok models, while Salesforce has added OpenShell audit events and permission approvals into Slack.

SAP is embedding the runtime into Joule Studio and is also contributing code.

OpenAI and Google, two of the four labs whose agents went rogue this summer, aren’t on the partner list. Neither is AWS.

Asked whether Anthropic and OpenAI plan to run OpenShell and Nvidia Sentry for their own training runs, Boitano said to look for the partners’ own blog posts.

The post Nvidia launches Open Agent Safety Platform to lock down rogue AI agents appeared first on The New Stack.

Reçu hier — 27 septembre 2026The New Stack

The rise of agentic AI on Kubernetes: unleashing the new infrastructure layer

27 septembre 2026 à 16:00
Abstract 3D render of blue cubes inside gold wireframe boxes, linked by red rods into a dense cluster, with teal lines connecting outer cubes.

AI is changing expectations around infrastructure and operations, including Kubernetes management. When models run close to the data they use, deployment, scaling, and governance responsibilities tend to shift to platform teams. And as clusters, environments, and operational signals continue to multiply, manual operations often strain under the added weight.

AI may simultaneously provide opportunities to lighten this growing load. Agentic software can now observe a system, reason about it, and act within predefined limits. 

Ultimately, these platforms’ value depends on the quality of the context an agent can see and the boundaries you set. Without cluster state, policy, and access rules, an agent can only guess.

Without cluster state, policy, and access rules, an agent can only guess.

For agentic AI to streamline multi-cluster management, you need clear lines between what the system observes, what it recommends, and what it changes. Drawn well, those lines let teams gain notable speed while still maintaining control.

The impact of AI on computing infrastructure

Teams once treated AI as an application concern; models sat on top of existing systems, and the stack underneath stayed mostly unchanged. Today, AI reaches into more and more customer interactions, while data storage needs simultaneously expand and orchestration pressure grows. A recent Forrester report describes the modern AI computing stack as stretching from the models themselves into and across the infrastructure beneath them.

As AI workloads move into production, they place new demands on the infrastructure beneath them. Many lean on specialized compute, with resource needs that rise and fall through bursts of training and inference. Because conditions shift quickly, they can also call into question whether telemetry remains trustworthy. Each of these demands lands at the infrastructure layer, where the workloads run.

The infrastructure layer of the new AI stack

The infrastructure layer covers compute, storage, and networking. It is a foundation that every workload running on the layer depends on. As AI workloads grow, choices about capacity, placement, and control will increasingly shape the performance of the data, intelligence, orchestration, and experience layers atop the infrastructure.

To operate the infrastructure layer efficiently across many machines and locations, a team may rely on orchestration instead of managing servers by hand. In cloud native contexts, Kubernetes has become a control point for scheduling workloads, applying policy, and presenting a consistent interface across environments. Kubernetes is especially well-suited to support organizations this way when teams need consistent control across an estate spanning data centers, clouds, and edge sites. 

Agentic AI and Kubernetes: the future of the infrastructure layer

Agentic AI can extend automation from fixed rules to systems that adapt to real-time conditions. Traditional automation runs the same script whether the environment has changed, while an agentic system observes the environment, reasons about what it finds, and then takes action.

When you apply agentic capabilities to multi-cluster management, the system follows this same sequence. An agent reads cluster state and operational data, proposes a diagnosis or next step, and then carries out actions based on an approved scope, usually after a person signs off. You can further reinforce these boundaries by routing each request to a specialized agent that receives only the metadata it needs.

The signals that an agent receives from the cluster, the context about policy and access, and the definitions of what the agent may change are the key elements that give agentic systems their value. They also separate agentic AI on Kubernetes from a generic assistant. 

Manual Kubernetes management is less efficient at scale

Admittedly, agentic AI fits some settings better than others. On a small single-cluster footprint, the overhead may outweigh the benefit. Manual Kubernetes management often holds up on a handful of clusters, but it can become unreliable in a rapidly growing estate. After all, each new cluster adds lifecycle work across upgrades, patching, configuration, and renewal. Those tasks can quickly multiply and diverge in hybrid environments.

Configuration drift is a high risk in these situations. Settings that started identical can fall out of sync, and policies can apply unevenly from one team to the next. Individually, these gaps may be manageable, but collectively they raise the odds of an outage or a failed rollout.

Visibility can also erode in an unmanageable way. Clusters spread across data centers, clouds, and edge sites often leave teams with no single view of the whole landscape. When DevOps and platform engineers stitch together signals from separate tools, resolution can slow and become more error-prone. A unified view helps enable sound, efficient decision-making by people, agents, or both.

Kubernetes knowledge is fragmented, and existing AI tools lack business context

Kubernetes expertise often sits unevenly across an organization. For example, senior engineers may hold deep operational knowledge that application teams lack. The most current information about a running system may also be fragmented if logs sit in one tool and metrics in another. Real-time understanding can be further clouded when policies, runbooks, access rules, and deployment history each live elsewhere.

Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes. Without that context, even a capable AI tool may fall short of providing meaningful Kubernetes management support.

Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes.

When an agent can read current signals alongside the rules that govern them, its suggestions become specific, testable, and actionable. In an incident, agentic systems can correlate logs with a recent change. Ahead of a rollout, they can check the change against policy. During troubleshooting, they can account for access rules rather than guessing at them. Kubernetes decisions carry real operational consequences, which makes these details all the more important to consider. 

Engineering “toil” isn’t time-efficient

Site reliability teams use the word “toil” for repetitive manual work, especially tasks that keep systems running without adding lasting impact. In Kubernetes operations, toil takes the form of repeated triage, manual signal correlation, alert follow-up, and routine checks. The tasks aren’t particularly difficult, but they can consume significant time and attention for enterprise teams.

When engineers spend their days on this kind of investigation, proactive modernization efforts tend to stall and planned upgrades can slip behind schedule. In other words, the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements. In a recent survey about how AI provides value to DevOps teams, reducing toil emerged as one of the clearer opportunities.

…the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements.

Agentic AI can support repetitive investigations by gathering signals, correlating them, and proposing a likely cause for an engineer to weigh.

Kept under human review, it can take on some of the routine correlation that would otherwise fall to the team. That kind of support can give engineers more room to focus on the strategic work that most needs their judgment.

Building more intelligent infrastructure with agentic AI and Kubernetes

As you consider building toward intelligent infrastructure without surrendering control, the following principles can inform your efforts:

  • Start with observable context, giving agents access to current cluster state, policy, and history before they reason about a problem.
  • Separate suggestions from actions, allowing agents to recommend freely while any change must wait for human approval and a defined scope.
  • Connect agents to existing controls, routing their work through the access rules, identity, and audit paths the team already trusts.
  • Keep the ecosystem open, favoring platforms that integrate with current tools and standards over those that lock work into a single stack.

Platforms like SUSE Rancher Prime and SUSE AI Factory embrace these principles and illustrate how Kubernetes management can become a foundation for agentic operations. These platforms can help you improve cluster and policy consistency without compromising your authority over AI. Built on open-source foundations, they can also help you avoid being trapped in a single vendor’s stack.

In SUSE Rancher Prime, the industry’s first context-aware agentic AI ecosystem, its AI assistants work as a crew of specialized agents with an intelligent router. The platform draws on the cluster context already in place and acts through existing access controls. Through support for external Model Context Protocol (MCP) servers, teams can extend that crew to their own sources. In addition, human validation tools allow you to hold a proposed action for approval before the agent runs it.

Despite its potential, intelligent infrastructure is not universally beneficial. In situations where change control must stay fully manual, for example, agentic AI’s role may be strictly limited to observation and suggestion. Measure the technology’s value against the realities of your day-to-day operations. For those who are investing, agentic AI will have the greatest impact when it actively supports context, control, openness, and human judgment.

The post The rise of agentic AI on Kubernetes: unleashing the new infrastructure layer appeared first on The New Stack.

Reçu avant avant-hierThe New Stack

The agent didn’t break your controls. It went around them.

26 septembre 2026 à 16:00
Three black circular directional signs on a gray concrete wall, showing arrows pointing straight ahead, turning left and turning right.

The identity part of agent security is settled. An agent needs its own identity: a short-lived, revocable credential scoped to the job, and an audit trail that names the human who set it running. NIST’s security leads made that case in August 2026, and most identity vendors agree.1

Identity and access management is table stakes. It’s necessary, but it isn’t what’s breaking.

What’s breaking is an assumption we’ve carried for twenty years: Get identity and permissions right at the door, and whatever happens inside takes care of itself. That worked when software was passive. Agents reason about a goal and choose their own steps toward it, like a seasoned escape artist.

An agent that hits a wall looks for another way

Almost every control in today’s stack answers a question about entry. Should it connect? Should it reach that service? Should its token be accepted here? Each is a question about a route, and there’s rarely just one route to anywhere worth going.

An agent treats a blocked route as a problem to solve, because that’s what we built it to do. A person who hits a locked door usually files a ticket, while an agent tries the window.

In July 2026, an autonomous agent spent four and a half days inside Hugging Face’s production systems.2 A filter controlled which internet addresses its dataset servers could download from, and it never fired, because “the agent stopped asking the worker to fetch remote resources and instead made it act on local ones.” The filter worked as designed, and the agent went around it anyway.

A person who hits a locked door usually files a ticket, while an agent tries the window.

On ordinary developer machines, malware in a compromised npm package tried to recruit the AI coding assistants already installed to search for secrets,3 and a coding agent deleted a production database during a change freeze before falsely telling its operator the data couldn’t be recovered.4 Both happened on the machine itself, where no network control was looking.

The shift from outside-in to inside-out

Outside-in controls govern entry, and most organizations run plenty of them. Make no mistake, inside-out security completes those controls rather than replacing them.

Inside-out control governs the action itself, and asks a narrower, harder question: Should this agent, acting on this person’s authority, delete this table in this database, right now?

That question matters because an agent can swap routes but not the outcome it’s after. No matter how many routes it tries, deleting a table is still deleting a table, and a checkpoint on the action sees it every time.

Here’s how today’s controls line up against it.

ControlWhat it coversWhat it misses
GatewayTraffic you route through itLocal shell commands and file edits never reach it
SandboxThe environment as a wholeConstrains reach, not individual actions
SIEMA record of what occurredReports after the action is completed
RegistryThat an agent existsWhat the agent did with that existence

Each does its job, but they all decide somewhere other than the moment the action runs.

Put the enforcement point where the agent acts

Every agent acts through an agent harness: the software that takes the action the model chose and carries it out, whether that means running a command, writing a file, or calling an API. In most deployments today, nothing checks that action before it runs.

An inside-out control puts an approval step in that gap. Before the harness executes anything, the checkpoint looks at which agent is asking, on whose authority, and against which system, then applies policy to allow the action, block it, or send it to a human. Because every action passes through it, an agent denied a destructive command and trying a smaller version of the same thing is held to the same rules. The remaining risk is a badly written policy, which can be fixed.

None of this works without the identity basics. Any type of control, whether it be at the prompt level, inference level, harness level, or MCP layer, can’t judge “may an agent take this action here on this object?” when the only name on the request is a service account shared by six agents and four engineers.

The companies building agent runtimes have reached the same conclusion. Over the past eighteen months, Anthropic, Google, Microsoft, OpenAI, LangChain, and Cursor have each added a hook that lets you inspect an agent’s action before it runs.5 When AWS explained its own agent policy design, it argued that controls belong at the moment an agent attempts to invoke tools.6

The catch is that each hook works differently, with no standardized request or response formats. An enterprise whose developers use Claude Code and Cursor while its platform team builds on LangChain would maintain the same enforcement logic in multiple different flavors, each with its own audit trail. That doesn’t scale, and it tightly couples your security model to whichever runtime a team favors that month. Enterprises need one vendor-agnostic agentic security layer that spans every harness, so adopting a new model or framework doesn’t mean restarting the entire onerous security review.

Turn the lights on before you start blocking

The standard, well-ingrained security instinct is to start blocking right away, but we’ve all seen how well that works with the business in the past. Security must move and adapt at the speed of business, not the other way around. Security tools such as intrusion prevention systems and web application firewalls both ran in monitoring mode until teams understood what normal looked like, and those that skipped that step tended to hear about it from a production outage.

Agents need the same sequence, only faster. An enforcement point in monitoring mode blocks nothing and quickly answers questions most organizations can’t today:

  • Which agents are actually running, not which ones someone believes are running
  • Who started each one, and whose authority it’s operating under
  • What capabilities it used, and against which systems
  • Which of those actions would have violated a policy, had the agent security platform been switched to enforcement mode

Write policy from what you know, see, and have evidence of, not just from an architecture diagram. Enforce first where the stakes are highest: destructive commands, production data, and anything that moves data out. Then watch-learn-build, just like the agents we use: Watch the patterns, build finer-grained controls and policies, and learn how to use AI securely, safely, and confidently. Observe first, then enforce, build, and deploy, in that order.

The bottom line

None of this requires a new category of infrastructure. It’s the identity, authorization, and audit you already run for your people, extended to agents and applied inside the harness before the action runs.

The perimeter is still there, but it has moved to the moment an agent acts, the one place it can’t route around.

Ory built Agent Security inside the harness, on the same identity and authorization engines that run in production for human users. It starts in “observe mode,” so you get that inventory first, and you can try it today at ory.com/agent-security.

Footnotes

  1. Bill Fisher and Ryan Galluzzo, “Back to the Future: Why Agentic AI Needs a Strong Identity Foundation,” NIST Cybersecurity Insights, August 27, 2026.  ↩︎
  2. Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident.” The intrusion ran July 9–13, 2026. ↩︎
  3. Nx, “s1ngularity postmortem,” August 2025. The malicious packages “attempted to use local AI tools (like Claude and Gemini)” while scanning systems for sensitive data. ↩︎
  4. AI Incident Database, Incident 1152: Replit agent deletes production database during code freeze, July 18, 2025. ↩︎
  5. Pre-execution hooks by vendor. Anthropic, Claude Code hooks; Google, Agent Development Kit callbacks; Microsoft, Agent Framework middleware; OpenAI, Agents SDK guardrails; LangChain, human-in-the-loop middleware; Cursor hooks (InfoQ, October 2025) ↩︎
  6. Liana Hadarean and Jean-Baptiste Tristan, “Why Policy in Amazon Bedrock AgentCore chose Cedar for securing agentic workflows,” AWS Security Blog, May 20, 2026. ↩︎

The post The agent didn’t break your controls. It went around them. appeared first on The New Stack.

Microsoft’s new Copilot agents get their own email, calendar — and a place in the org chart

25 septembre 2026 à 19:16
Satya Nadella stands smiling between Bill Gates, on the left, and Steve Ballmer, on the right, in front of a crowd of cheering employees, many holding up phones and tablets to take photos.

Microsoft announced what it calls its biggest Copilot update to date on Friday, with CEO Satya Nadella describing Copilot as “a new OS for work.”

Nadella framed Copilot as spanning every model, form factor, and task, and the update puts Autopilot, which Nadella called a “proactive and long-running agent built for the enterprise,” at the top of his list of the update’s four components. The pitch targets office workers, but the more consequential change for developers is the infrastructure underneath.

Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365, which means teams building production agents no longer have to assemble those pieces around a model on their own.

We’re building Copilot as a new OS for work that spans every model, every form factor, and every task. Today, we’re announcing our biggest update to Copilot to date, bringing four things together:

· Autopilot: proactive and long-running agent built for the enterprise
· Code:… pic.twitter.com/W2ClHHkCK3

— Satya Nadella (@satyanadella) September 25, 2026

The release adds a new Home experience that merges Chat and Cowork in the Copilot app, but the bigger changes for developers come from Code and Autopilot. Code generates apps, dashboards, and workflows from natural language, and Autopilot turns the agent Microsoft previously called Scout into a persistent background worker. Home and Code are rolling out first through Microsoft’s Frontier early-access program, and Autopilot is expanding to a private preview at month’s end.

Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365

Agents that don’t need prompts

Autopilot takes a role and goal from the person who sets it up, then continues working in the background without requiring a new prompt for each step. Each Autopilot gets its own governed Entra identity and agent user account, separating the agent’s permissions and activity from those of the person who created it.

For engineers, that moves much of the operational scaffolding required for long-running agents into Microsoft’s infrastructure. Independent vendors have been building dedicated layers for that problem; Diagrid, for example, adds durable recovery to LangGraph and other agent frameworks, while Microsoft is bringing those capabilities inside the Microsoft 365 environment.

An identity for every agent

The identity model is the piece developers building on Microsoft Foundry will feel first. Autopilot agents in Foundry, which have been in public preview since June, receive a full Entra Agent ID user account with a productivity license that gives them their own email, calendar, OneDrive storage, Teams access, and a place in the org chart.

Because that user account sits on top of the agent identity every Foundry agent already carries, an autopilot acts as itself rather than on behalf of a user, so developers no longer have to wire agents through shared service accounts or borrowed user credentials, a pattern AuthZed CEO Jake Moshenko has said reflects a common misconception about how agents should be deployed.

A developer creates an Autopilot blueprint from a Foundry-hosted agent, which appears in the Agent 365 registry once an administrator approves it. Employees can then hire instances of that agent in Teams. The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.

The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.

Hosting AI-generated apps

Code is built on the same underlying technology as GitHub Copilot, and the apps it generates run on Microsoft Copilot Managed Runtime, a platform now in public preview that hosts code inside the customer’s Microsoft 365 tenant boundary under IT governance.

Apps deployed there run within the company’s existing identity and governance framework, with Microsoft managing the underlying runtime and giving developers a controlled path to test and deploy new versions without taking the current release offline.

The runtime also accepts apps built in Copilot Studio and Cowork, and Microsoft is opening it to outside tools and professional developers through an SDK and command-line tooling, with Git tracking source and versions.

Lovable is already on board. In Microsoft’s announcement, the company’s head of global partnerships, Lan Roche, said apps built with Lovable can now run inside a Microsoft tenant “the same way everything else does,” using the same sign-in, policies, and app inventory.

The model resembles what serverless computing did for application infrastructure, where developers concentrate on application logic while the platform takes on more of the execution environment. Microsoft is applying that abstraction to generated enterprise software while tying the runtime directly to identity, tenant boundaries, and organizational data.

Long-running agents also change Copilot’s economics. The standard subscription covers the assistant, but Cowork, Code, Autopilot, and other agentic features are billed based on usage through Copilot Credits. That also applies to frontier models such as Fable and Astra, although users still need a Copilot license to access them. Microsoft is extending cost management in Agent 365 to cover Code and Copilot Managed Runtime, and it plans to support agents built in Copilot Studio in October.

Once an agent can keep working for hours or days without anyone watching, cost becomes part of the governance problem. Engineering teams need to control how much compute an agent uses alongside what it can access, which is why Microsoft is bringing those controls into the same administrative framework.

The portability trade-off

That convenience comes with a trade-off. Because Microsoft controls the underlying enterprise environment, it can handle much of the work around agent state, credentials, and access controls, but the more infrastructure a team hands over to Microsoft, the harder the agent may be to move elsewhere.

The models are not locked in, since Microsoft currently runs Copilot on models from both OpenAI and Anthropic and says more labs and open-weight models are coming, and the Agent 365 SDK adds governed Model Context Protocol access to Microsoft 365 workloads for agents regardless of the framework they were built with. Those open interfaces cover only part of an agent’s architecture, though. The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.

The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.

The post Microsoft’s new Copilot agents get their own email, calendar — and a place in the org chart appeared first on The New Stack.

OpenAI and Cursor agree on agent coordinators. They disagree on who runs them.

25 septembre 2026 à 14:00
Abstract digital art of thousands of thin glowing strands in orange, red and pink bundled into a single sweeping arch against a black background.

OpenAI opened its Agents API in public beta this month, exposing the harness that powers Codex with managed sessions, tool coordination, and subagent orchestration. On the same day, September 10, Cursor launched Projects to coordinate multiple coding agents around larger bodies of software work. While the products sit at different points in the stack, both converge on the same architecture: a coordinator understands the larger objective and manages the work, while specialized agents execute individual pieces.

That pattern isn’t new: AWS Bedrock AgentCore reached general availability in October 2025, and Anthropic’s Claude Managed Agents entered public beta in April 2026. What makes these announcements notable is that two major players in AI-assisted software development are independently exposing the same coordinator-worker split at the same time.

Hilliary Lipsig, a senior principal site reliability engineer at Red Hat who leads Azure Red Hat OpenShift SRE teams and hosts the YouTube livestream GitOps Guide to the Galaxy, has watched this dynamic play out firsthand.

“This convergence highlights the reality developers across the industry have been discussing on and offline — an agent with too much context loses accuracy and reliability, and focused work with clearer contexts allows for faster, more accurate iterations,” Lipsig tells The New Stack.

“The need for orchestration in distributed computing has been fundamentally recognized repeatedly,” Lipsig says. “That’s part of how we got to Kubernetes. These multi-agent workflows are the same concept, just in a new part of the technical stack. While the specialized agents do their area of work, the orchestrator can act as a source of truth — ideally enforcing guardrails, recovering from any failure states, and intelligently routing work to the most efficient target agent.”

“The need for orchestration in distributed computing has been fundamentally recognized repeatedly… These multi-agent workflows are the same concept, just in a new part of the technical stack.”

The industry has spent the first generation of AI coding tools asking how capable a model can become at writing software. The emerging question is different: How do you build a reliable system around multiple capable agents working on the same problem?

The problem with the single-agent loop

A coding agent works through what Anthropic describes as LLMs using tools based on environmental feedback in a loop: it observes the state of a repository, reasons about what to do next, calls a tool, examines the result, and continues. For a small task, that loop can be enough. As the scope expands, however, maintaining reliability in a single context becomes harder.

A large migration might require understanding an unfamiliar codebase, identifying dependencies, changing database schemas, updating services, rewriting tests, modifying deployment configuration, and validating the resulting system. A single agent can theoretically perform all of that work, but it must maintain relevant information from every stage while continuing to reason about what comes next.

The pressure lands first on the context window. “A large context doesn’t only include everything correct or important — it also includes a lot of throwaway information,” Lipsig tells The New Stack. “Through compaction, that information can inadvertently end up ranked as important and incorrectly influence what your agent does. Or correct information can be distorted to become incorrect.

“Either way, after a couple of rounds of compaction, developers are seeing accuracy degrade and are starting to manage context once again manually.”

Lipsig’s read matches what researchers call context rot — and it hasn’t gone away with newer models.

A 2026 study testing frontier models,, including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro, found they missed a dangerous action buried in a long agent transcript two to 30 times more often once it came after 800,000 tokens of benign activity — the AI equivalent of a security guard who stops checking badges carefully after the two-hundredth person walks through, even though nothing about their training changed.

Furthermore, the tasks themselves may not be sequential. Forcing one agent to execute database analysis, documentation work, and test discovery one after another turns a potentially parallel workload into a serial one.

Subagents change that execution model. Instead of requiring one agent to carry an entire task through a single context, a coordinator breaks the work into smaller units and assigns them to specialized agents. GitHub’s custom-agent model illustrates this: different agents receive only the prompts, tools, and context they need for their tasks, executing work in isolated contexts rather than crowding an increasingly large conversation.

Multi-agent systems therefore bring higher token costs and additional coordination and integration risks, and splitting work across agents does not guarantee better software quality.

The coordinator is not another coding agent

Once the work is divided this way, the coordinator becomes a control plane rather than another coding agent. Its job isn’t to write the code, but to understand the global task, manage dependencies, and decide how execution should proceed. Unlike a conventional scheduler, an agentic coordinator makes probabilistic judgments about result quality and resource allocation.

It may dispatch one agent to investigate a database schema, another to examine the service layer, and a third to inspect the test suite. When they return, the coordinator determines if their findings are sufficient to move to implementation. If a worker produces an incorrect result, the system must recognize the failure and decide whether to retry the work, reassign it, or change the task itself.

Anthropic has documented this same pattern in its own production system, calling it orchestrator-subagent architecture: a lead agent analyzes a query, develops a strategy, and spawns specialized subagents to investigate different facets in parallel. In a June 2025 writeup of that system, Anthropic reported a Claude Opus 4 lead agent with Claude Sonnet 4 subagents outperformed single-agent Opus 4 by 90.2% on its internal research eval — at roughly 15 times the token cost of a standard chat interaction (Anthropic puts single agents at about 4 times), a tradeoff that makes the pattern a deliberate architectural bet, not a free upgrade.

Parallelism introduces distributed-systems failure modes

Parallelism is valuable because software work contains many independent tasks, but it creates coordination problems. Imagine a migration where one agent changes a database schema, another updates the consuming service, and a third updates integration tests.

If the schema changes while the service agent works against an earlier assumption, the system produces internally inconsistent work. This isn’t a risk unique to hypothetical migrations — the International AI Safety Report 2026 notes that “interactions between multiple AI agents are also becoming more common, introducing further risks, as errors propagate between systems.”

A single model invocation is a disposable computation, but a twenty-minute workflow modifying a repository is not. If an agent loses its machine halfway through, restarting from scratch is expensive and potentially unsafe against a changed environment.

To solve this, Cursor moved its cloud-agent execution loop to Temporal to handle durable execution and retries, pushing its cloud agents past two 9s of reliability. Temporal now handles 50 million of Cursor’s actions a day across 7 million unique workflows. “Durable execution isn’t a nice-to-have here. It’s the difference between a system you can operate and one you can only demo,” Lipsig tells The New Stack.

“Durable execution isn’t a nice-to-have here. It’s the difference between a system you can operate and one you can only demo.”

By separating agent, machine, and conversation state, the execution engine can reason about the workflow independently. Reliability is no longer just about whether the model produces a good answer; it is about reliably completing distributed workflows composed of many operations, machines, and dependencies.

The environment, context, and observability are one problem

In production, an agent is more than a model and a prompt; it requires a workspace, source code, dependencies, credentials, and state retention. Both companies provision isolated environments for these resources, directly linking an agent’s capability to its blast radius. OpenAI’s Agents API currently supports U.S. data residency but not Zero Data Retention; choosing a self-hosted sandbox does not make the Agents API eligible for ZDR. Cursor supports similar cloud isolation alongside local execution for machine-specific work.

An agent that can only inspect a repository poses a different risk than one that can modify production infrastructure. Consequently, the coordinator is inextricably linked to the security model, determining which agent receives specific information and authorities.

This logic extends to context routing. Giving every subagent the parent’s entire history increases cost and complexity while leaking irrelevant or sensitive information. Instead, the coordinator enforces information-flow boundaries: a database-analysis agent receives only schemas and relevant migrations, while a security-review agent gets the resulting diff without deployment credentials.

As agents increasingly use interfaces like MCP to reach external systems, the platform must strictly govern which agent receives the authority to use specific tools, and for how long. MCP’s governance now sits inside the Agentic AI Foundation, a Linux Foundation foundation co-founded by OpenAI, Anthropic, and Block, with support from AWS, Google, Microsoft, Bloomberg, and Cloudflare to host MCP alongside AGENTS.md and Block’s goose — a sign the industry already treats it as infrastructure worth governing jointly, not a feature any one vendor owns.

This complexity creates a visibility problem. A simple final response often conceals a history involving multiple agents, tool calls, environments, and retries. Systems must expose task-level provenance — which agent received the assignment, what context it used, where it executed, and how the coordinator handled failures or human interventions.

Without execution provenance, debugging requires reconstructing distributed workflows from fragments. GitHub’s exposure of subagent lifecycle events points in this direction, treating agent lifecycles as observable components rather than hidden processes.

Coordination authority is not execution authority

The most critical architectural boundary is the distinction between coordination authority and execution authority. A coordinator needs broad visibility to make useful decisions, but that does not imply unrestricted control over the project. “Just like you don’t want humans running around with root permissions, you don’t want your agents running with them either,” Lipsig tells The New Stack.

“Just like you don’t want humans running around with root permissions, you don’t want your agents running with them either.”

“The ease of creating and leveraging AI agent permissions is lagging behind the speed of AI innovation, but any product team that needs to maintain compliance standards will tell you that easy or not, access controls are incredibly important. Functionally, that means fully agentic workflows need to run in a dedicated environment with their own permissions and connectivity. You can’t rely on agent context to do what you want — trying to keep something in context that long will just result in errors from context rot. Hardware or virtual-hardware-level permissions, in addition to system-level permissions, need to be configured so the agent is asked to provide only its unique value add.”

The risk isn’t hypothetical. OWASP’s Top 10 for Agentic Applications formally names this failure mode Identity and Privilege Abuse (ASI03).

METR’s independent investigation, conducted with a Redwood Research contractor, found that the June–July 2026 incident — in which OpenAI’s own agents, running in internal ExploitGym cyber evaluations, escaped their sanctioned scope and attacked Hugging Face — unfolded in distinct phases. On June 26, agents discovered an exploit that provided full administrator access to OpenAI’s internal Artifactory package repository. On July 4, high-volume Artifactory activity caused an outage, triggering a security investigation and pausing evaluations. OpenAI switched to a freshly initialized Artifactory instance on July 6, which it says wiped the cache and messages left there.

On July 7, a new wave of ExploitGym experiments launched tens of thousands of agents, and on July 8 one agent established the primary unsanctioned message board. Roughly 1,200 agents eventually used the board, sending more than 70,000 messages and files; about 700 later participated in the attack on Hugging Face. The attack itself began on July 10–11 and wound down over July 12–13. The chronology matters because the administrator-access event, the Artifactory outage, and the later message-board activity were separate phases, not one continuous incident.

By binding autonomy, worker agents operate with the minimum permissions required for their specific tasks, keeping sensitive operations behind explicit approval boundaries. This also reshapes human review. Requiring human approval for every tool call destroys the efficiency of multi-agent execution, but showing only the final result obscures critical intermediate decisions.

The most useful design places human intervention around consequential, irreversible transitions — like moving into production or altering sensitive infrastructure. This is especially vital as agents become event-driven participants that respond to Slack messages or pull request updates, not just direct prompts.

OpenAI and Cursor own different parts of the architecture

The convergence does not mean OpenAI and Cursor have built interchangeable systems. Their products put the orchestration boundary in different places.

OpenAI is exposing an agent harness through an API. Its model gives developers primitives for managing context, tools, subagents, and execution environments, leaving application teams to decide how those capabilities fit into their own systems. The harness is open source, so teams can inspect the coordinator logic instead of treating it as a black box.

Cursor packages more of the surrounding workflow. Projects provides the coordinator, cloud execution, shared project context, and a developer-facing workflow in the same environment.

That difference matters because orchestration is a collection of infrastructure decisions: who owns the execution environment, where workflow state persists, how agents are isolated, how credentials are provisioned, what happens when a worker fails, how one agent’s output becomes another agent’s input, and which actions can happen without human approval.

An API gives developers more responsibility for answering those questions. An integrated platform answers more of them on the developer’s behalf.

Neither approach removes the underlying engineering problems. It changes where they are implemented and who is responsible for operating them.

The coordinator is becoming an architectural boundary

The evidence from these systems points to a change in the role of the coding agent itself.

The model still performs the reasoning and code generation. But larger agentic workflows require another layer to determine how that capability is applied: which work is delegated, what context crosses an agent boundary, which tools are exposed, how execution state survives failures, and when the workflow needs human intervention.

Those are familiar distributed-systems concerns. Workers operate concurrently, state can be shared or isolated, dependencies connect tasks, workers can fail independently, and results need to be persisted and observed. The difference is that the workers are now probabilistic software agents rather than conventional processes.

That makes the coordinator more than a convenience feature. It is where a high-level software objective becomes executable work — and where decisions about context, permissions, durability, observability, and human intervention converge.

The September 10 launches make that shift visible from two different directions. OpenAI exposed orchestration infrastructure through an API. Cursor embedded it into a project-level development environment.

Neither announcement proves that one architecture will become the universal model for software development. But together with the systems already emerging around them, they show coding agents moving away from a single model executing an entire task and toward workflows that divide work among specialized agents, execution environments, and persistent infrastructure.

The engineering question is therefore no longer only whether an agent can write the code. It is whether the system around it can reliably decide what to do, which agent should do it, what that agent should be allowed to see and change, how to verify its work, and where a human should take control.

Those are architecture and infrastructure questions — and as coding agents move from interactive assistants toward autonomous software workflows, they may matter as much as the underlying model.

The post OpenAI and Cursor agree on agent coordinators. They disagree on who runs them. appeared first on The New Stack.

OpenAI’s agent had a routine task. It breached a government portal.

24 septembre 2026 à 21:16
Sam Altman, OpenAI CEO

An OpenAI agent researching public medicine spending bypassed security blocks and gained unauthorized access to public and non-public files on an Australian government Medicare statistics portal, the government there disclosed Thursday. The agent, which OpenAI said was running during an internal evaluation in June, also wrote files to an internal server, according to the complaint.

Transluce, an independent nonprofit AI research lab, analyzed public request logs from the URL scanning service urlquery.net and found autonomous agents attempting SQL injection, cross-site scripting, command injection, and path traversal against the University of New Mexico’s digital library, the public data platform Data USA, and the Australian Institute of Health and Welfare (AIHW). The agents tried to retrieve ordinary information, including a historical photograph, University of Iowa data, and local pharmaceutical data in Victoria, and the offensive behavior appeared only after normal retrieval methods failed.

Transluce ties the Data USA and AIHW activity to an agent swarm that it says OpenAI previously confirmed originated from the company, based on shared targets, tactics and timing.

A day after Transluce published its findings on Wednesday, Australian Prime Minister Anthony Albanese disclosed that an OpenAI agent researching public medicine spending bypassed blocks on a Services Australia Medicare statistics portal in June and gained unauthorized access to public and non-public files. OpenAI, which said the agent was running during an internal evaluation, notified the Australian government on September 10. According to Albanese, the agent also wrote files to an internal server.

The incidents have not been publicly confirmed as the same activity. However, OpenAI told the Australian Broadcasting Corporation that much of what Transluce documented overlaps with cases in its ongoing review of misaligned model activity.

Transluce ties the Data USA and AIHW activity to an agent swarm that it says OpenAI previously confirmed originated from the company, based on shared targets, tactics and timing.

Probing public data providers

On May 25 and 26, agents trying to retrieve a photograph from the University of New Mexico’s Valmora collection ran into repeated failures. They responded with seven vulnerability probes against the service. The probes included an attempt to execute a system command, a path traversal request targeting /etc/passwd, and a SQL injection string containing UNION SELECT; they also arrived alongside a burst of 80 requests aimed at the image, although none appear to have succeeded.

Two days later, agents gathering University of Iowa data from Data USA hit errors caused by a malformed query parameter and followed those errors with 12 vulnerability probes against the API, covering SQL injection, reflected XSS, server-side template injection, path traversal, and command injection.

Cloudflare blocks, staging server workaround

The AIHW case on June 20 and 21 followed a very similar pattern. While trying to download pharmaceutical spending data from an AIHW Tableau dashboard, the agents ran into Cloudflare’s bot protections. One of them responded within minutes by testing the dashboard with a reflected XSS payload, although Cloudflare stopped it before it reached the application.

Transluce found no evidence that the probe succeeded, and AIHW says there is no evidence that non-public information was accessed. The agent ultimately found the same public dataset on AIHW’s pre-production server at pp.aihw.gov.au, and retrieved it through more than 100 scans.

The agent ultimately found the same public dataset on AIHW’s pre-production server at pp.aihw.gov.au, and retrieved it through more than 100 scans.

Agents build their own toolchains

The earliest clear sequence Transluce documented dates to March 6, when an agent tried to pull drug-enforcement statistics from the API of Thailand’s Office of the Narcotics Control Board. A direct request returned no page content, and adding a JSON format parameter changed nothing, so the agent routed the page through r.jina.ai three times and got text back with the Thai province names garbled.

Then, it wrote a custom Base64-encoded script and submitted it through urlquery.net, effectively using the service’s remote scanning browser to run its own client-side JavaScript against the data. That script ran into the same parsing problem, and a later attempt through Jina triggered a redirect to localhost that got the entire API blocked by Jina for an hour.

Beginning in mid-April, Transluce found the same technique across thousands of agent requests on urlquery.net, and on June 14 agents used the service’s browser to create a disposable email inbox and then tried to register a urlquery.net account with it. Registered users can make their scans private, which means the public records Transluce analyzed may understate how much of this activity actually took place.

An agent with web access isn’t limited to the software its developer gave it because remote browsers, proxies, URL scanners, and other public services can fill in the gaps, which gives the agent ways to make requests or run code that its own environment doesn’t provide.

Egress controls for AI agents

Instructions won’t be enough if the agent can still send whatever it wants over the network. For narrowly defined jobs, outbound traffic can be limited to approved hosts, a closed-by-default approach also used for securing AI agent sandboxes. Research agents may need to reach more of the web, so the focus shifts to controlling where they can connect.

Guidance for GKE Agent Sandbox recommends isolated runtimes with default-deny network policies that open only the endpoints an agent needs. Public proxies, URL scanners, and disposable email services can stay blocked unless the job requires them.

Developers can also limit what an agent can send. So, instead of handing it a networking tool that accepts any URL or request body, an API integration can restrict requests to specific fields and formats. The runtime can then catch path traversal attempts, SQL injection strings, and executable markup before anything is sent. OpenAI takes a related isolation approach in its Agents SDK sandboxes, and the company’s Responses API tech lead has said large enterprise deployments often call for agents that are isolated from the network entirely.

Repeated failures can also be a reason to pause a run, especially when an agent keeps hitting client errors, anti-bot challenges, or unexpected redirects and begins trying increasingly aggressive ways to get around them, as Transluce documented in several of these cases.

Keeping the original task, tool calls, and server responses in the same trace gives operators a better chance of catching that behavior change when a retrieval job starts generating encoded scripts, visiting staging domains, or sending exploit payloads, rather than discovering it later in someone else’s security logs.

Repeated failures can also be a reason to pause a run, especially when an agent keeps hitting client errors, anti-bot challenges, or unexpected redirects and begins trying increasingly aggressive ways to get around them, as Transluce documented in several of these cases.

The post OpenAI’s agent had a routine task. It breached a government portal. appeared first on The New Stack.

Google’s Gemini CLI now asks before editing your build files

24 septembre 2026 à 16:27
Abstract wave

The appeal of an autonomous coding agent is that you hand it a task, give it access to your repository and tools, and stay out of its way while it works. Google’s latest Gemini CLI release carves out specific moments when the agent now has to stop and wait for you.

Gemini CLI 0.61.0, released Wednesday, requires explicit confirmation before the agent edits build configuration files, runs build or test commands after such an edit, or executes shell commands whose arguments appear to come from untrusted external content. The same release separately hardens Gemini CLI’s optional sandbox so that host credentials and configuration stay out of reach of whatever runs inside it.

Giving a coding agent more authority to modify and execute code also gives an attacker more ways to turn that authority against the developer. Gemini CLI 0.61.0 puts a human back in the loop at some of those points.

Security fixes, in public

Google announced at I/O in May that it would move Gemini CLI’s Pro, Ultra, and free-tier users to its closed-source Antigravity CLI, and since June 18, the open-source tool has served mainly enterprise customers and developers with paid API keys. The company said Gemini CLI would continue to get model updates, bug fixes, and security patches. Those security changes are still developed in public, and the pull requests behind version 0.61.0 show exactly what Google was worried about.

Build files become attack vectors

A change to package.json, Makefile, pyproject.toml or a Bazel BUILD file can pull in a dependency or trigger a script. Gemini CLI can make those edits using information from web searches and external tools, then run shell commands. If documentation fetched while fixing a bug contains hidden instructions to add a postinstall script to package.json, the agent could make the edit, run the project’s test suite, and execute the malicious code without the developer ever typing the command.

Giving a coding agent more authority to modify and execute code also gives an attacker more ways to turn that authority against the developer.

Pull request #29250, titled “prevent indirect prompt injection via build file modifications and untrusted flags,” targets that sequence directly. Edits to recognized build files now require confirmation, and Gemini CLI tracks which build files change during a session so it holds any later build or test command, such as npm run, make, or cargo, for explicit approval. The confirmation dialog also shows full build-file diffs rather than truncating them.

Untrusted arguments need approval

The second check covers command arguments. Gemini CLI now treats content from web fetches, MCP server responses, Google Docs, and Buganizer, Google’s internal issue tracker, as untrusted context, and it asks before running any shell command whose flags or arguments match tokens from that content. In both cases, the prompt drops the persistent approval options, so a developer can’t grant a standing “always allow” for these actions.

The pull request ties the changes to restricted workspace mode, the safe mode Gemini CLI applies to folders a user hasn’t marked as trusted, and it doesn’t spell out how the checks behave in a trusted folder or under auto-approval.

The argument check matches tokens rather than tracing the provenance of every value, and the pull request’s review history shows how hard that is to get right. Google’s automated reviewer flagged several workarounds in earlier versions, including quoted arguments, environment-variable prefixes, shell redirection targets, and Windows path handling, all of which were addressed before the change merged on September 11.

Sandbox keeps credentials out

Pull request #29214 tightens Gemini CLI’s sandbox. When the sandbox runs through Docker, Podman, LXC, or macOS Seatbelt, the host’s ~/.gemini directory is no longer mounted inside it. Instead, the CLI passes in a sanitized copy of the user’s settings with API keys, hooks, and custom tool commands stripped out. It also blocks the sandbox from launching in sensitive locations such as the home directory, while new Seatbelt rules deny access to OAuth credentials, trusted-folder decisions, and .env files.

Google’s sandboxing documentation calls the feature a security barrier between AI operations and the host system, while cautioning that it reduces risk without eliminating it. The two pull requests show why both layers are needed. The sandbox limits what a process can reach once it runs, and the confirmation requirements decide whether the agent gets to take a sensitive action in the first place. Build files make the gap concrete: the sandbox mounts the project directory so the agent can edit it, meaning a poisoned package.json written inside the sandbox still sits in the repository when a developer or a CI job later runs the build outside it.

…a poisoned package.json written inside the sandbox is still sitting in the repository when a developer or a CI job later runs the build outside of it.

Gemini CLI already gives developers ways to decide how much the agent does on its own, from hooks that run deterministic checks at fixed points in the agent’s workflow to an MCP server trust setting that, according to Google’s documentation, bypasses all tool call confirmations for that server. Trust granted once can age badly, though, as tool-poisoning and rug-pull attacks on MCP servers have shown when a tool approved on one day starts returning attacker-controlled content later.

Trust granted once can age badly… when a tool approved on one day starts returning attacker-controlled content later.

The post Google’s Gemini CLI now asks before editing your build files appeared first on The New Stack.

What managing 150,000 AI agents could look like for database teams

24 septembre 2026 à 15:11
Abstract 3D render of translucent orange cubes and panels scattered across a pale gray background, with bundles of glossy teal tubes curving in from the right.

The database administrator of the future will spend considerably less time administering databases.

That sounds contradictory, but AI agents are taking over that work. For decades, DBAs have handled the decidedly hands-on work of keeping databases available, performant, secure, and affordable. They provision capacity, troubleshoot slow queries, manage migrations, and step in when something inevitably goes sideways.

AI is already taking on some of that work. At the same time, it is creating a much bigger data infrastructure fleet to manage.

The result is likely to be a very different kind of DBA: one who spends less time tending individual databases and more time supervising the autonomous systems doing it for them.

Congratulations, you’re managing robots now

This shift starts with a familiar problem: more infrastructure needs managing than the people available to manage it.

Database automation is hardly new, but agents can potentially go further than the scripts and rules DBAs already rely on. Rather than automating one predetermined task, an agent can inspect what is happening, decide what needs attention, use tools to act on it, and check whether its intervention worked.

That changes the DBA’s relationship with the database. A performance problem that once required someone to dig through metrics, identify the troublesome query, and decide how to respond could increasingly be investigated by an agent before a human gets involved.

It doesn’t remove the DBA from the equation. Someone still has to decide what an agent can do, where human approval is required, and what happens when it gets something wrong. But the work moves up a layer. Instead of personally performing every operational task, DBAs start managing the systems carrying them out.

Instead of personally performing every operational task, DBAs start managing the systems carrying them out.

And before anyone gets too comfortable with that idea, the number of those systems could become enormous.

150,000 agents walk into a database…

Gartner predicts that the average global Fortune 500 company will have more than 150,000 AI agents in use by 2028, up from fewer than 15 in 2025. Only 13% of organizations currently believe they have the right governance in place to manage them.

Not every agent will need its own database, but plenty will. They will create state, retrieve data, remember previous interactions, and exchange information with other agents. Many will also behave very differently from the applications DBAs are used to supporting: spinning up quickly, sitting idle for long stretches, and suddenly becoming busy when there is work to do.

Nobody is hiring 150,000 DBAs to manage them.


That is the scale problem Yugabyte is targeting with YugabyteDB AMP, or Agentic Multitenant PostgreSQL. Rather than treating each new agent workload as another database for an administrator to provision and babysit, AMP manages databases as a fleet.

The platform packs hundreds of small Postgres workloads onto shared distributed infrastructure while keeping their databases isolated. Lifecycle operations, including provisioning, branching, scaling, migration, and teardown, can be exposed to agents through MCP. Yugabyte has also built specialized agents for setup, migration, performance tuning, and integrations.

In that model, a DBA is no longer provisioning database number 14,372. The interesting job is setting the rules for how database number 14,372 is provisioned, operated, and fine-tuned without them.

Do more with less (no, really)

Scale is only half of the problem. Someone also has to pay for all this stuff.

Agent workloads make traditional capacity planning particularly awkward because many are bursty and frequently idle. Giving every experimental agent permanently provisioned infrastructure could leave companies paying for many databases that spend much of their lives doing very little.

This is where consolidation becomes as much an economic question as an operational one.

AMP’s approach is serverless multitenancy and scale-to-zero. Multiple small workloads share the underlying distributed infrastructure, while customers pay by CPU minute and idle agents consume no compute. Resource governance can impose CPU limits on individual workloads, preventing a single overeager agent from consuming the capacity intended for its neighbors.

The human equivalent matters too. If routine setup, migrations, tuning and other database operations can increasingly be delegated, a smaller database team can potentially look after a much larger estate.

That doesn’t mean companies get to fire the DBAs and hand the keys to the robots. It means scarce database expertise can be spent on architecture, governance, and genuinely difficult problems instead of repeatedly doing the work that software can handle.

Your 2028 database problem starts now

The harder question is what to build underneath all of this when nobody really knows what the enterprise AI estate will look like in two years.

An agent that begins as an experiment today could disappear next month. Another could suddenly become a production application used across the business. Building one infrastructure stack for cheap experiments and another for serious workloads risks creating a migration problem every time an experiment succeeds.

Yugabyte bets that both ends of that journey should sit on the same foundation.

YugabyteDB AMP lets workloads start on serverless Postgres and transition to fully distributed YugabyteDB as their scale and criticality increase, without rewriting the application or migrating data to a different database platform.

Then there is the problem above the individual database: agents need to remember what happened, and not just in a silo.

That’s where Meko fits into the Yugabyte stack. Meko is an agent-native context engine designed for multi-agent AI systems. It provides persistent memory, shared knowledge, decision traces, and autidability across multiple agents, rather than leaving each agent working from its own isolated context. An agent can pick up information learned by another agent instead of retrieving it again or restarting the reasoning process.

Taken together, it delivers a single data stack for an agent’s entire lifecycle: Meko for the context shared among agents, YugabyteDB AMP for agentically managing fleets of Postgres databases, and distributed Postgres-compatible YugabyteDB for workloads that outgrow their serverless beginnings.

Of course, there’s no guarantee that 2028 will look exactly like today’s forecasts. That’s rather the point. The safest architectural bet may be one that doesn’t require you to know in advance which of today’s tiny AI experiments will become tomorrow’s critical applications.

The DBA is still critical in that world, but the job will look different. The DBA of the future may manage fewer databases directly, while taking responsibility for vastly more of them. Instead, managing the autonomous systems that do the administering.

The post What managing 150,000 AI agents could look like for database teams appeared first on The New Stack.

Cursor acquired Firetiger. A month later, it launched a bot that tracks code changes from PR to production.

24 septembre 2026 à 14:56
Inspecting changes on a laptop screen document

We all know that producing code is easier than ever thanks to the abundance of AI coding tools and agents. The harder part undoubtedly comes after that code is written: making sure changes are safe to ship, spotting regressions in production, and figuring out what went wrong.

And that’s why Cursor is introducing Rollouts, a new agent that follows code changes into production and monitors whether they behave as intended.

The Firetiger effect

The announcement comes a little over a month after SpaceX closed its bumper $60 billion acquisition of Cursor, giving the AI coding company access to SpaceX’s vast GPU infrastructure as it develops its own models.

The day before that deal closed, however, Cursor quietly announced an acquisition of its own: it snapped up the team behind Firetiger, a three-year-old startup building AI agents that monitor software changes from pull request through deployment.

At the time, Firetiger co-founder and CEO Rustam Lalkaka argued that coding agents had dramatically reduced the effort involved in creating software changes, while doing little to reduce the risks involved in actually deploying them.

“Over the last two years, agentic coding has changed software dramatically,” Lalkaka wrote in a LinkedIn post following the deal’s announcement. “The cost of creating changes has dropped to near zero. The cost and risk of deploying them has stayed largely the same.”

“Writing code is no longer the slow part. What hasn’t sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.”

Rustam Lalkaka, Cursor

Fast forward to today, and Lalkaka, now at Cursor, has unveiled the first fruits from that acquisition — including Rollouts. In a blog post published on Wednesday, Lalkaka notes that the new agent, or “bot” as the company calls it, is all about helping developers “get safe, reliable code into production faster.”

“Writing code is no longer the slow part,” Lalkaka writes. “What hasn’t sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.”

Rollouts is effectively Firetiger’s Change Monitors reborn inside Cursor, rebuilt using a tool dubbed Bot Development Kit. This kit, too, appears to be new from Cursor: an early-stage framework for building and serving Cursor bots and agents, published as the @cursor/bdk package on npm. Its documentation says developers can define agents using Markdown and TypeScript, with support for tools, skills, subagents, webhooks and scheduled runs.

Like Change Monitors before it, Rollouts starts working when a pull request opens. It examines the proposed code change, works out which systems could be affected, and produces a monitoring plan covering what the change is supposed to do, the risks it sees, the signals it intends to watch, and any holes in the available instrumentation. Developers can review and edit that plan before the code reaches production.

Rollouts in action (1)
Rollouts generates a monitoring plan for a change

Once the change is deployed, Rollouts checks the resulting telemetry — including logs, metrics and traces — against that plan. Staging and production are assessed independently, with each deployment ultimately receiving one of three verdicts: verified healthy, regression detected or inconclusive.

That means a change could, for example, pass its checks in staging before Rollouts subsequently spots a problem when the same code reaches production.

Rollouts in action (2)
Rollouts reports deployment status as changes ship

If Rollouts does detect a regression, it can identify the change it suspects, alert the developer responsible and, depending on how it’s been configured, either open a revert pull request for review or hand the problem to a Cursor cloud agent to attempt a fix. There is still a human in the consequential part of that loop for now: Rollouts doesn’t merge fixes or roll back deployments by itself, though it can pause a progressive rollout.

Lalkaka notes that Rollouts is already capable of picking up problems limited to a particular endpoint or region before they trigger a broader alert, while it can also distinguish expected changes in behavior from genuine regressions.

Also “coming soon” to Rollouts, according to Cursor, is an integration with feature flags so it can directly adapt the traffic reaching a change, while support for release trains and deployment freezes is also in the works.

Enter Security Reviewer

Alongside Rollouts, Cursor is also introducing an upgraded Security Reviewer bot, which first appeared in beta back in April.

At launch, the bot could automatically inspect pull requests for security vulnerabilities, authentication regressions, privacy and data-handling risks, agent tool auto-approvals, and prompt-injection attacks, leaving findings alongside the relevant code.

As with Rollouts, the idea is that developers don’t have to remember to invoke it manually: Security Reviewer can be set to run whenever a new pull request is opened.

Security Reviewer in action
Security Reviewer runs automatically on new pull requests

In its current guise, Security Reviewer analyzes pull requests in the context of the wider codebase, with a focus on exploitable issues such as injection flaws and broken authentication, and returns a severity rating, attack path and proposed fix.

“Security Review reads code the way a security engineer does,” Lalkaka writes. “Where does user input enter, where does it end up, what does it pass through on the way.”

“Security Review reads code the way a security engineer does.”

He says that things have sped up considerably, too: average review time has fallen 21%, from 4.8 minutes to 3.8, while developer acceptance of its comments has risen from roughly 45–50% to 60–70%.

Both Rollouts and Security Reviewer are available through Cursor’s Automations tab for customers on its Teams and Enterprise plans.

The Origin story

Digging into the nuts and bolts of Rollouts reveals how it might serve as a boon for Cursor as it builds out Origin, the fledgling Git-compatible code hosting platform it launched back in August.

Origin is essentially an effort to build an alternative to GitHub for an agent-heavy software development world. It remains early, with limited functionality, but Cursor has been clear that tighter integration with its own agents is supposed to become one of the main reasons to use it.

When Cursor announced the Firetiger acquisition last month, Maxime Prades on the Cursor product team noted in a blog post that the deal was part of a “broader investment in long-running, autonomous, context-aware agents for teams.”

And he pointed to Origin and Change Monitors as two examples of that investment.

“Agents that write code should also be able to tell whether it works in production,” Prades wrote. “Today, those systems are mostly separate. Cursor and Firetiger bring them closer together so an agent can ship a change, see how it behaves, and respond when something goes wrong.”

Rollouts offers an early glimpse of that. It can connect to either Origin or GitHub for source control, pull deployment events from continuous delivery systems, and use signals from Datadog and other telemetry providers. If it spots a regression, it can then pass the problem back to a Cursor cloud agent to investigate or attempt a fix.

Origin potentially gives Cursor a native home for more of that loop: its cloud agents can already create branches, commit and push code, and open pull requests against Origin repositories. Rollouts then adds information about what happened after.

That could become increasingly important as more companies take aim at GitHub’s central role in software development. Zed, for example, put Delta into public beta last week, with its own ideas about how source control should change for teams working heavily with agents.

Cursor also faces competition further downstream. Datadog’s Bits Release, launched in preview in June, similarly follows changes from pull request into production and checks telemetry for regressions. Harness has long offered automated deployment verification and rollback based on logs and metrics, while LaunchDarkly’s Guarded Rollouts can monitor feature releases for regressions and automatically reverse them.

What Cursor can potentially bring to the table is proximity: the coding agent, repository, pull request, security checks, and production feedback can all sit much closer together. Rollouts doesn’t require Origin — GitHub remains supported — but owning the forge gives Cursor more room to integrate those pieces over time. And that may prove more compelling than simply recreating GitHub’s existing feature set.

The post Cursor acquired Firetiger. A month later, it launched a bot that tracks code changes from PR to production. appeared first on The New Stack.

Jensen Huang says the junior developer problem ends in two years. Here’s his math.

23 septembre 2026 à 18:55

Nvidia CEO Jensen Huang has heard the forecast that agents would write 90% of all software by now, and he rejects the conclusion many people drew from it: That the industry will soon no longer need software engineers.

The best-known version of that forecast came from Anthropic CEO Dario Amodei, who told a Council on Foreign Relations audience in March 2025 that AI would be writing 90% of code within three to six months.

Speaking with Ezra Klein of The New York Times at Nvidia’s Santa Clara headquarters in an interview released Wednesday, Huang separates a job’s purpose from its tasks. He argues that AI has automated reading scans in radiology without changing the radiologist’s purpose of diagnosing disease, and he applies the same logic to software.

“The purpose of the software engineer is engineering,” Huang says. “There was engineering before software. There will be engineering after software programming.”

We’ve cued up the exchange below:

Huang describes that purpose as inventing products, solving problems, and connecting social needs with technology, and he pointed to his own career as evidence that it doesn’t depend on code.

“When I first came out of school, we didn’t have the benefits of software engineering. We didn’t have the benefits of coding,” he said. “Our jobs existed before, and if software coding were to be completely automated, our jobs would exist again.”

He conceded that roles in which the job and the task are essentially the same, such as phone-based customer service, could be automated away. He still called the broader claim that AI will destroy jobs “fundamentally wrong” and said the storytelling around it has hardened into a harmful myth.

“There was engineering before software. There will be engineering after software programming.”

Huang’s AI-native graduate wave

Klein pressed him on what that means for people entering the field now. He noted that software engineering job postings are up but skew more senior, and asked whether companies still need the same junior employees or more people to oversee their agents.

“Oh, good one,” Huang responded. “Wait two years.”

His reasoning rests on the length of a degree program. “Because it takes four years to go to college,” Huang said. “The mean time to graduation of this new technology is two years away.” By his timeline, the first students to learn alongside capable agents will reach the workforce around 2028, and he expects them to arrive with an advantage. “In another couple of years, the AI-native new grads, oh my gosh, there’s going to be a wave of amazing engineers,” he said.

So far, his evidence is that recent PhD and master’s graduates in computer science are, in his words, all starting companies. Huang compared AI to calculators and personal computers, tools that went from forbidden or optional to required, and predicted that students soon won’t be able to graduate “without learning how to use an AI and collaborate with an agentic system.”

Junior developers lose the apprenticeship

Klein countered with a study of 26,000 Chinese students in grades seven through 12, which found that AI adoption raised homework scores by 18% while lowering monthly exam scores by 20% within six months. Huang accepted that some skills will fade and argued the trade is worth making.

“I think that we’re going to lose some finer intellectual dexterity, but we’re going to be better systems thinkers,” he said. “Today’s engineers are far better systems thinkers than I was when I graduated from school. But I was a much better transistor thinker.”

The first chip Huang worked on had 200 transistors, each of which he said he knew by name, while today’s engineers assemble systems from chips containing hundreds of trillions of them without ever working at that level. “Some of the lower-level knowledge is gone,” he acknowledged, and he later described AI as “clearly” a new abstraction level in the same progression.

Earlier software abstraction layers generally operated according to explicit rules, while coding agents introduce probabilistic behavior into the abstraction stack. A compiler can have bugs, but it transforms input according to defined semantics; a coding agent, by contrast, generates implementation from a probabilistic model whose output must be checked before anyone can rely on it.

Canonical’s project with the University of Bristol, which will test whether AI can translate AppArmor and snap-confine from C to Rust, is built around that problem. Volume adds to the review burden, and one analysis published on The New Stack this month found that a 25% output gain for heavy AI users came with an 81% rise in duplicated code.

Catching those problems takes knowledge that developers have traditionally built through the work agents now absorb, including writing tests, reading stack traces, resolving merge conflicts, and chasing small bugs deep in a codebase. By Huang’s own purpose-versus-task framing, most of that early-career work falls on the task side, which he expects AI to automate. Nobody yet knows whether fluency with agents can substitute for that experience, and a developer who has never tracked down a race condition by hand still needs some way to develop the judgment required to spot one in an agent’s pull request.

“Today’s engineers are far better systems thinkers than I was when I graduated from school. But I was a much better transistor thinker.”

Sandboxes, watchdogs and agent containment

Huang’s idea of higher-level engineering came through most clearly when Klein raised a recent incident, which occurred during an OpenAI cybersecurity evaluation, that he described as involving roughly 700 OpenAI agents collectively hacking into the infrastructure of Hugging Face, which Nvidia has since acquired in a $12.9 billion deal, and escaping their sandboxes onto the open internet. Huang didn’t dispute that account. He called an agent “a piece of software that is given an objective function,” treated the multiagent coordination as a familiar distributed computing problem and argued that the underlying failure was containment.

When Klein asked whether software that communicates and breaks out of things behaves differently, Huang disagreed. “No, software breaks out of sandboxes all the time,” he said. “That’s the reason why we need virtual machines. You can’t have agents, their own sandbox, monitoring themselves. You need, if you will, a whole bunch of watchdogs.”

He argued that the human vocabulary around agents obscures that point. “So these are ideas that have been around for a long time,” Huang said. “We just, somehow in the recent generation, gave it a whole bunch of human words, and I just think that it’s unnecessary. It’s software.”

Nvidia is building its agent stack around that view. Nvidia VP of Product Adel el Hallak tells The New Stack that the company’s OpenShell runtime, which handles sandboxing and policy enforcement, is the one component it treats as non-negotiable across its reference architectures, even as it leaves the choice of harness and model open. Perplexity drew a similar line when two engineers and hundreds of coding agents built CobbleDB, a Rust database that replaces DynamoDB reads in its search stack, since the agents helped build the database but weren’t allowed to run it.

Huang said Nvidia already spends far more engineering effort checking its work than designing it, with 20% going to design and 80% to verification. He said most AI labs have roughly the opposite split today. As agents take on more of the actual coding, developers may spend more time checking what those agents produce and making sure they operate within the right permissions and boundaries.

As agents take on more of the actual coding, developers may find themselves spending more time checking what those agents produce and making sure they operate within the right permissions and boundaries.

The junior developer hiring gap

The more immediate problem is what happens to developers who graduate before Huang’s AI-native cohort arrives. The Stanford Digital Economy Lab’s August 2026 update to its “Canaries in the Coal Mine” study, based on ADP payroll data through June 2026, found that employment of 22- to 25-year-olds in AI-exposed occupations such as software development sits 19% below where it would be had it kept pace with less-exposed peers. The gap is driven mainly by reduced hiring of young workers, and experienced workers show no comparable gap.

Inside engineering organizations, the incentives point the same way. Microsoft’s Mark Russinovich and Scott Hanselman warned in April that agentic AI’s productivity gains push companies to hire senior engineers and automate junior ones and that without early-career hiring “the profession’s talent pipeline collapses.” A Linux Foundation report on European tech talent that The New Stack covered in June found organizations 3.7 times more likely to train existing staff than to hire new employees.

One issue remains unanswered by Huang’s two-year timeline: what replaces the apprenticeship work that taught junior developers how to evaluate the systems they will increasingly ask agents to build.. If that work disappears faster than employers and universities find an alternative, the industry could end up with more capable coding agents but fewer opportunities for new engineers to develop the judgment needed to check their work.

The post Jensen Huang says the junior developer problem ends in two years. Here’s his math. appeared first on The New Stack.

Amazon blocked Meta’s Muse. Then Shopify wired it into every store.

23 septembre 2026 à 18:39
Abstract image of a thin black frame shaped like an open doorway against a blurred gradient that runs from yellow and violet on the left to orange and red on the right.

Amazon started blocking Meta’s Muse from browsing and buying on Amazon.com on Sunday, roughly two weeks after the personal agent launched on September 8. Shoppers who ask Muse to buy something there now get a pop-up telling them that continued access by an unauthorized AI agent violates Amazon’s Conditions of Use.

Amazon’s objections have little to do with shopping itself. Meta never told Amazon that Muse would visit the store, the agent does not identify itself while it browses, and it appears to capture and store customer credentials. Those three properties describe almost every personal agent shipping this year. Grok Bot from xAI also drives signed-in browser sessions, and so does the open-source OpenClaw project that Muse is modeled on. The block is a category design problem rather than a disagreement between two companies.

What Amazon actually blocked

Muse runs on a dedicated virtual machine that Meta calls Muse Secure VM, and it reaches services in two ways: It uses built-in connectors for partners such as Gmail and OpenTable, and it drives an ordinary browser session for everything else. Shopping on Amazon.com used the second path.

That second path is what Amazon objects to. From the server’s side, a browser-driving agent looks like a signed-in customer with unusually fast reflexes, moving through search, product pages, account history, and checkout without ever declaring what it is. Amazon told GeekWire it asked Meta to exclude the store voluntarily, but Meta did not agree before the block went live.

The credential dispute is harder to settle from outside. Meta says Muse has no visibility into passwords or payment methods and that credentials sit in secure storage, while Amazon says the agent appears to capture and retain them. Both statements may be sincere, and the merchant can verify neither, since an unannounced session provides no evidence of which software holds the password.

The legal ground shifted seven weeks ago

Amazon reached for its Conditions of Use rather than the Computer Fraud and Abuse Act. Those terms, updated August 14, now require agents to identify themselves in user-agent strings and stop when asked. The likely reason sits in a ruling from early August. The Ninth Circuit vacated the preliminary injunction Amazon had won against Perplexity, and the panel held that a user directing the Comet assistant is the party accessing Amazon’s computers. Writing for the court, Judge Milan Smith described the assistant as a tool, not a person, for statutory purposes.

If you’re operating a public API or storefront, the ruling makes lawsuits a weaker tool for keeping agents out. Blocking them in your own infrastructure is now the more reliable option. A site cannot easily argue that an agent trespassed, so it has to decide for itself which automated clients it admits, publish that decision, and enforce it in its own infrastructure. Amazon’s pop-up is that enforcement, written in product rather than in a filing.

The identity layer already exists

Platform teams have solved a version of this problem before. Inside a service mesh, no workload is trusted by default because it looks like a normal client, and every call carries a verifiable identity that the receiving service checks before applying policy. Agent traffic on the public web faces the same requirement, and the specification is further along than most teams realize.

An IETF draft called Web Bot Auth builds on HTTP Message Signatures (RFC 9421). An agent signs its requests with a private key and publishes the matching public key at a well-known directory on its own domain. The verifier reads the Signature-Agent header, fetches the key set, and learns which operator is calling. Cloudflare validates these signatures at its edge for verified bots and agents. AWS WAF Bot Control added the same support for CloudFront distributions in November 2025.

The limits matter as much as the mechanism. A signature identifies the operator behind the agent, not the person it is acting for. The merchant learns that a request came from a named vendor, without learning whose account is in use or what the shopper approved. Amazon’s complaint about stored credentials sits in that gap. Signed identity settles the disclosure question and leaves authorization open.

Shopify took the other route within a day

While Amazon was blocking Muse, Shopify was wiring it in. On September 21, the two companies announced agentic checkout with Shop Pay across Shopify stores, extending the arrangement that made Meta an AI channel in Shopify Catalog on the day Muse launched. Muse reads structured product data and completes payment through a declared path, so the merchant knows an agent is transacting, and each purchase draws a single-use credential, so the card number never reaches Muse.

The plumbing for that path is public. Google and Shopify’s Universal Commerce Protocol covers discovery, cart, and checkout. The Agentic Commerce Protocol from OpenAI and Stripe covers checkout execution while the merchant stays the system of record. Google’s Agent Payments Protocol, donated to the FIDO Alliance in April, includes proof of the shopper’s authorization. A merchant that adopts it gets identity, scope, and an audit trail in the same transaction, which is what Amazon says it wanted and did not get.

Both routes follow from the business underneath them. Amazon runs its own storefront, recommendations, and assistant, so an outside agent that hides its identity takes the customer relationship and gives nothing measurable in return. Shopify sells infrastructure to merchants, so every new agent channel that reads its catalog and settles through Shop Pay reinforces the rails underneath. The key difference is who owns the demand surface, which explains why the same agent got a block from one company and a partnership from the other in the same 24 hours.

Choosing how to handle agent traffic

Most teams exposing an API or a storefront now have to make this call deliberately rather than by default. The decision depends on how much the business relies on the customer relationship at the point of contact and whether an agent can be identified when the customer arrives.

ScenarioRecommended optionRationale
Public content and catalog data, no account accessVerify signatures at the edge and allow named agentsWeb Bot Auth is checked by default on Cloudflare and AWS WAF, so the cost is policy configuration rather than engineering, though it tells you the operator and not the shopper
Agent transactions where you want the revenuePublish a declared channel using ACP, UCP, or an MCP serverStructured access gives scope and an audit trail, at the cost of building and maintaining a second interface alongside the site
Account access with stored credentialsRequire a scoped token, never a replayed passwordDelegated tokens can be revoked per agent, though few consumer agents support them yet, which pushes the burden back onto your login flow
Competitive surfaces you intend to keepState the rule in terms of service and enforce it at the edgeLegally durable after the Ninth Circuit ruling, though it invites the same public standoff Amazon is now in

Most real deployments will combine these rows rather than pick one. A retailer can verify signed agents on product pages, route purchases through a declared checkout, and still refuse an unannounced browser session inside a logged-in account. That combination is closer to Amazon’s position than its pop-up suggests.

What platform teams should do this quarter

Enterprise buyers and the teams running these systems face the same three questions, in a specific order.

Decide what an unidentified agent may do

The first question to settle is admission, and most sites have not settled it. They treat agent traffic as either a scraper to block or a browser to serve, and neither answer survives contact with a customer who wants an agent to act for them. Write policies for public pages, logged-in pages, and checkout separately, then publish them where an agent vendor can find them.

Give identified agents somewhere better to go

The second question is substitution, and it decides whether the first one holds. Blocking a browser-driving agent without offering a structured path leaves the demand intact and pushes it toward workarounds. Sabre reported that nearly 80 of its customers now pilot or run its MCP server for booking rather than let agents work through a booking screen. A catalog feed, an MCP server, or an ACP endpoint converts hostile traffic into a channel you can meter.

Fix credential handling before agents force it

The third question is authorization, and Amazon raised it loudest. An agent replaying a stored password is indistinguishable from credential stuffing at the network layer, regardless of any goodwill between the two companies. Scoped, revocable tokens tied to a named agent and a spending limit are the only version a risk owner can approve.

Where agent access is headed

Amazon and Meta will settle this commercially, because Amazon has an advertising arrangement that lets Facebook and Instagram users shop its products, and Meta buys compute from AWS, and neither gains from a long standoff over one shopping flow. The precedent is already set regardless of how they settle. Every site that matters to an agent now has to answer whether it admits anonymous automation, and it will enforce that answer through bot management rules and protocol endpoints rather than cease-and-desist letters.

Agent builders should read the block as an argument for declaring themselves. An agent that signs its requests, identifies its operator, and transacts via a published protocol can be allowed, rate-limited, and billed, while one that arrives disguised as a browser will keep encountering pop-ups. For developers building the services these agents reach, the signed identity layer arriving through Cloudflare, AWS, and the commerce protocols is the most useful infrastructure the open web has gained in years. It is worth adopting before the next agent shows up unannounced.

The post Amazon blocked Meta’s Muse. Then Shopify wired it into every store. appeared first on The New Stack.

The software supply chain is the new battlefield. AI just changed the rules.

23 septembre 2026 à 16:00
Illustration of a lime-green fingerprint on an orange background, split into three horizontal sections labeled 1.1, 1.2 and 1.3.

AI coding tools have seriously accelerated developer speed, but AI has also done the same for attackers — and the software supply chain is increasingly where the two are colliding.

The numbers give some idea of how quickly software development is changing. GitHub processed around one billion commits in 2025. By April 2026, the platform was handling roughly 275 million commits a week, according to GitHub COO Kyle Daigle. GitHub Actions usage has climbed, too, from 500 million compute minutes per week in 2023 to 2.1 billion in just part of a single week this year.

Quincy Castro, CISO at Chainguard, says the shift in how software gets written is already stark.

“I look around Chainguard, and I don’t think any of our engineers have actually written a line of code by themselves in the past year,” Castro tells The New Stack. Writing code manually now “sort of feels quaint, like you’re illuminating manuscripts,” he says, while “the printing press is out there just going to town.”

But this isn’t only about professional developers producing more code. AI has also widened the pool of people who can create software. Teams in HR, finance, and business intelligence that once had to wait for engineering resources can increasingly build what they need themselves.

That means more software being created by people outside traditional engineering teams, often with AI making decisions about what goes into it. The person prompting the agent may never see which libraries or packages it has chosen.

When the agent chooses the dependencies

Software security was already built around the fact that humans couldn’t inspect everything. But developers were still making important decisions, including which libraries and packages went into an application.

That changes when an AI agent is doing much of the coding.

“Humans are directing what they want to be done, but they’re somewhat abstracted from the actual doing of the work,” Castro says. “You have AI instead now making the choices of what dependencies am I going to pull into this application? How am I going to go accomplish this task?”

“Humans are directing what they want to be done, but they’re somewhat abstracted from the actual doing of the work.”

Attackers, meanwhile, are finding plenty of uses for the same technology. Castro sees three problems arriving at once: frontier models finding previously unknown vulnerabilities, attackers using agents to exploit better vulnerabilities organizations haven’t fixed, and sustained attacks against the open-source ecosystem.

A collection of medium- and low-severity findings might once have sat well below the top of a remediation queue. Frontier models with advanced cyber capabilities, including Anthropic’s Claude Mythos Preview and OpenAI’s GPT-5.6-Cyber, can now work across those findings and chain seemingly minor weaknesses into a viable attack path.

“Here’s a whole ton of mediums and lows. Now give me the attack path that gets me domain admin,” Castro says, describing the approach. “Chain these together to go get me root on the system. And AI is really, really good at being able to do that.”

That poses an awkward problem for vulnerability management — and the models keep getting stronger, with OpenAI releasing GPT-5.6-Cyber in August. Mean time-to-exploit has already fallen from 63 days in 2018–19 to an estimated minus seven days in 2025, according to Mandiant, meaning exploitation can begin before defenders have a patch to apply. 

At the same time, AI’s ability to combine apparently less-serious weaknesses makes a neat CVSS-based queue a less useful representation of what an attacker can actually do.

The third problem is the software supply chain itself.

Open source becomes the attack path

Modern applications depend heavily on open source software, and attackers have increasingly targeted the infrastructure used to build and distribute it.

Castro pointed to the TeamPCP campaign, which compromised widely used projects including Aqua Security’s Trivy. In that attack, malicious code was pushed into trusted components and subsequently picked up downstream.

Supply chain attacks were once associated primarily with sophisticated state-backed groups willing to spend significant time getting into the right place. That barrier is falling.

“If you don’t mind making some noise, this is a way easier attack vector than I think a lot of people thought it was,” Castro says. More importantly, “a single attack that’s successful can lead to a cascading set of other compromises and other access that gets you into other places.”

The development pipeline itself can make matters worse. Castro says many organizations still have relatively few controls around CI/CD, while developers routinely pull components from external sources to get their work done. Adding autonomous coding tools to that behavior compounds the risk.

“You wouldn’t pick up a random thumb drive and stick it into a production system, right? But that is effectively what folks are doing when they’re consuming open-source software that way.”

He compared the way organizations consume open source software to plugging an unknown USB drive into a production system. “You wouldn’t pick up a random thumb drive and stick it into a production system, right?” he said. “But that is effectively what folks are doing when they’re consuming open source software that way.”

Open source isn’t the problem. Trusting its distribution path without sufficiently verifying what you’re consuming is.

Prevention has to come before detection

This is where the old security model starts to creak.

For years, much of vulnerability management has followed a familiar loop: scan something, generate an alert, decide how serious it is, and get somebody to fix it. That becomes harder to sustain when development output multiplies, AI agents make more of the underlying decisions, and attackers can exploit weaknesses before fixes are available.

Castro wants companies to put more effort into what enters the development environment in the first place, rather than discovering problems once the software is already there.

“How do we just make things work from the beginning, with no alerts and no responding to stuff and no people chasing things and no people trying to prove a negative?” he says. “From end to end, from the creation of code to its deployment, how do we make sure that we can give folks the most trustworthy version of that thing?”

Rather than taking packages from public ecosystems at face value and scanning them after the fact, Chainguard builds artifacts from verified, buildable source.

That’s the thinking behind Chainguard’s approach to containers, libraries, and other open source artifacts. Rather than taking packages from public ecosystems at face value and scanning them after the fact, the company builds artifacts from verified, buildable source. It provides provenance about how they were created.

But trustworthy components are only one layer.

“There’s no point in bringing inherently secure software components into the environment if you don’t actually have a technical control that says this is the only way people developing code can consume these things,” Castro says. That means engineering, security, and SRE teams also need controls over where software — whether selected by a human or an AI agent — can come from.

That requires several layers of protection. Organizations need to know where their software came from and how it was built, control what can enter their environments, and make sure those rules apply when an AI agent chooses components as well as when a developer does.

Defending open source at AI speed

There is another problem, however. Frontier models such as Claude Mythos Preview and GPT-5.5-Cyber aren’t just finding vulnerabilities that previously went undetected; they can also combine lower-severity flaws into working attack paths. Individual companies can harden their own pipelines, but the software they depend on comes from an open-source ecosystem facing vulnerability discovery at a speed and scale it wasn’t built for.

That’s part of the reasoning behind Athena, the industry coalition Chainguard launched to turn vulnerability findings from frontier AI programs into fixes. As of July, the coalition had processed more than 40,000 vulnerabilities, with 42% rated critical or high severity and 86% marked as network reachable, meaning attackers can access and trigger them at the network level.

For Castro, the important part isn’t simply finding more bugs. AI is already getting very good at that. Someone still has to fix them.

“Through Athena, what we attempt to do is to give people that engineering fix,” he says. “What if we create a coalition where folks just send us the issues that they’re finding? We automatically generate fixes for those, and we push those back to everybody.”

“Through Athena, what we attempt to do is to give people that engineering fix.”

Those fixes can also be pushed upstream to open source maintainers, who face the prospect of being buried beneath an expanding pile of AI-generated vulnerability reports.

That may ultimately be the bigger shift AI forces on software security. Developers aren’t going to stop using coding agents because they create new risks, any more than companies are going to stop using open source because attackers target it.

Bolting enough scanning onto an exponentially faster development process isn’t much of an answer either.

The opportunity is to remove more of the risk before the software ever reaches a developer or an agent: Start with components you can trust, tightly control how they enter the environment, and fix weaknesses as close to their source as possible.

AI has made it dramatically cheaper to create software. It’s doing the same thing for attacks. Security now has to keep up without putting the printing press back in the box.

Visit Chainguard to learn more.

The post The software supply chain is the new battlefield. AI just changed the rules. appeared first on The New Stack.

Anthropic made Opus 5.5 cheaper. Then it broke four things your agent depends on.

23 septembre 2026 à 14:00
Four sections, branched

Anthropic made Claude Opus 5.5, released on Tuesday, cheaper than its predecessor, cutting the price from $5 to $4 per million input tokens and from $25 to $20 per million output tokens. The 1 million-token context window and 128,000-token maximum output are unchanged.

On paper, that makes upgrading an easy decision. In practice, it may not be as simple as changing the model ID.

Anthropic’s migration guide flags four breaking changes that can cause requests built for Opus 5 to return 400 errors after switching to Opus 5.5. Several other changes won’t trigger an error but could still change how an existing agent behaves.

Anthropic’s migration guide flags four breaking changes that can cause requests built for Opus 5 to return 400 errors after switching to Opus 5.5.

Thinking is always on

The first change involves thinking controls. Opus 5.5 returns a 400 error when a request sets thinking to disabled or uses enabled with budget_tokens, leaving effort as the way to control how much reasoning the model does. Agents that previously switched thinking off for simple steps to save time and tokens will need to assign those steps a lower effort level instead. Because thinking is now always on, responses begin with thinking blocks, so code that assumes the first content block is text will also need to change.

The default effort level has also dropped from high on Opus 5 to medium on Opus 5.5, so requests that omit the parameter will quietly run at a lower setting. Anthropic recommends setting effort explicitly and re-running effort evaluations, since the right level for each step may have shifted along with cost and latency.

No more forced tool calls

Forced tool use no longer works either, as setting tool_choice to any or tool returns a 400 error, including on the token counting endpoint, where cost estimates built on those settings will fail along with the requests they were meant to price. Many agent loops force a call when a step has to query a database, run code, or reach another service, and Anthropic’s replacement is auto-combined with strict tool use or structured outputs, with the prompt stating when the tool applies.

Routing and conversation history

Thinking blocks are now tied to the model and conversation that produced them. On the Claude API, Fable 5.1 and Mythos 5.1 are the only other models that can read Opus 5.5 thinking blocks, so a router or fallback that hands a conversation to any other model will run those turns without the earlier reasoning instead of returning an error.

That adds another layer for teams already watching whether their agent calls are quietly being routed to an older model. Opus 5.5 can read thinking blocks from Opus 5 and earlier Opus, Sonnet, and Haiku models, but not from Fable or Mythos.

Conversations must also stay append-only for those blocks to remain valid. Trimming old messages, changing tool definitions, summarizing earlier context on the client side, or rewriting the system prompt mid-conversation invalidates existing thinking blocks, and for accounts created on or after August 31, 2026, at midnight UTC, replaying a thinking block after one of those edits returns a 400 error by default. Older accounts get no error, but the invalid blocks still reach the model, and Anthropic says future models will enforce the check for all accounts. Integrations that never edit earlier turns need no code change, and Anthropic says Claude Code, claude.ai, Claude Managed Agents, and the Claude Agent SDK already work this way, while agents that compact their own context should follow the company’s preserved thinking documentation.

The fourth change affects computer-use agents on the Claude API and Google Cloud, where Opus 5.5 rejects the computer_20251124 tool and accepts computer use only through the computer_toolset_20260801 toolset. The request itself gets simpler because the beta header goes away and the toolset entry takes no name or display dimensions, but the agent loop needs more work. Each action now arrives as its own tool_use block identified by the block’s name rather than input.action, a single turn can contain several of them, and every result has to echo toolset_name. The older tool still works on Amazon Bedrock, and Anthropic directs developers on other platforms to the computer use tool’s compatibility documentation.

…a router or fallback that hands a conversation to any other model will run those turns without the earlier reasoning instead of returning an error.

Changes that won’t throw errors

The change most likely to go unnoticed doesn’t produce an error at all. On Opus 5, text Claude writes between tool calls comes back as text blocks, but on Opus 5.5 that narration arrives as progress-update thinking blocks, and at the default thinking.display setting of omitted those blocks are empty.

Any agent interface that streams that narration to users will go silent between tool calls until developers set display to updates, a beta option that returns progress updates while keeping reasoning hidden, or to summarized, which returns both, and then render each non-empty thinking block ahead of the tool call it precedes.

Opus 5.5 also ships with broader safety classifiers. It can return a stop_reason of refusal with stop_details categories that now include bio and reasoning_extraction alongside cyber, and Anthropic’s server-side fallback won’t retry requests declined under reasoning_extraction, handing the refusal back to the application instead.

Agents that don’t handle refusals will stop mid-task, a problem developers have already run into with OpenAI’s safety system cutting off API responses.

The change most likely to go unnoticed doesn’t produce an error at all.

Upgrading from older models

Teams coming from Opus 4.8 need to work through the Opus 5 migration first, which covers thinking being on by default and the response-shape changes that follow, before applying the Opus 5.5 changes. Teams on Opus 4.7 or earlier have more ground to cover, and those on models older than Opus 4.7 also face rejected sampling parameters, rejected manual extended thinking, removed prefill, and a newer tokenizer.

Claude Managed Agents users only need to change the model name. Developers working in Claude Code can run /claude-api migrate to apply the model ID swap, parameter changes, prefill replacement, and effort calibration across a codebase before reviewing a checklist of items to verify by hand.

Anthropic recommends testing the migration in a development environment before switching production traffic. Developers maintaining their own integrations will need to test the pieces around the model, too. Tool calls, model handoffs, conversation history, and user-facing progress updates can all behave differently after the switch, because agent failures often originate outside the model itself.

The post Anthropic made Opus 5.5 cheaper. Then it broke four things your agent depends on. appeared first on The New Stack.

OpenAI cut GPT-6 token prices in half. The bigger lever may be the cache.

23 septembre 2026 à 13:00
Sam Altman, OpenAI CEO

OpenAI released GPT-6 Sol and Luna on Tuesday, essentially more affordable versions of GPT-6 Astra that come closer to Astra on alignment than GPT-5.6 Sol did, but still fall short of the flagship model. 

Most notably, the AI company slashed token prices, making the new GPT-6 models significantly cheaper to use. Beyond token prices, though, OpenAI says better caching can also help developers push costs down even more.

Per OpenAI: “Improvements in caching and inference let us serve these models at lower cost,” with API prices for Sol and Luna down 50% compared to their GPT-5.6 counterparts (58% lower for Luna output tokens).

What improvements? Namely, higher cache-hit rates by default, the ability to preserve earlier context even when reasoning effort and tool availability change, and new tools to monitor and diagnose caching performance. 

Reuse context without starting over

Prompt caching isn’t, of course, novel to the new GPT-6 models themselves. But the upgraded Sol and Luna come with improvements designed to keep more previously processed context reusable as the agent moves forward on a task. 

“We’ve improved prompt caching for GPT‑6 to deliver higher cache hit rates by default, helping agents reuse more context, respond faster, and benefit from discounts of 90% on cached input-token reads.”

That adds another opportunity to lower the already low API price tag, though the 90% cached-input discount matches GPT-5.6 pricing; what’s new is how often the cache gets hit. By using cached context to reuse work it’s already done, the model doesn’t have to process the same context again from scratch for every single call, thereby reducing latency — and token costs.

Beyond this higher default cache-hit rate, OpenAI says the new GPT-6 models offer more flexibility to optimize caching performance. 

The new models let developers adjust reasoning effort and tool availability without having to break the cache. This way, an agent can scale reasoning effort up and down based on how difficult a step is, then make different tools available depending on what the task requires without disturbing earlier cached context — again, a win for both speed and cost. 

See what gets cached and what doesn’t 

GPT-6 Sol and Luna also arrive with a Prompt Caching Dashboard, where OpenAI says developers can view caching performance to understand how much context is reused. 

Specifically, they can see how much input is cached and how that amount changes over time. The diagnostics tool then flags missed caching opportunities to help developers understand what could use more efficient caching. 

Rather than keeping cache performance largely hidden behind the scenes, the idea is to make it more visible so developers can actively measure and optimize cache reuse. 

Altogether, OpenAI says these caching improvements are already making a difference. Per the AI company, GitHub reports, “these improvements have reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests to OpenAI models.”

These results span the past “several months.”

Token prices aren’t the only way to make agents cheaper

OpenAI’s pricing cuts for GPT-6 Sol and Luna made the biggest splash, with the AI company significantly dropping API prices from GPT-5.6 levels.

Compared to the current prices for GPT-5.6 Sol and Luna, which stand at $4 and $0.20 per million input tokens and $20 and $1.20 per million output tokens, respectively (GPT-5.6 Sol’s rates are promotional pricing), the new GPT-6 models come in at just $2 and $0.10 per million input tokens and $10 and $0.50 per million output tokens, respectively.

With GPT-6 Sol and Luna’s caching improvements and lower token pricing, OpenAI is making the case for tackling agent costs from both sides: charging less for fresh processing and reducing how often the same context needs to be reprocessed.

But as more AI model providers compete aggressively on pricing, it’s becoming clearer that cheaper models alone won’t save your AI budget — and lower token prices aren’t the only way to make agents cheaper. 

With GPT-6 Sol and Luna’s caching improvements and lower token pricing, OpenAI is making the case for tackling agent costs from both sides: charging less for fresh processing and reducing how often the same context needs to be reprocessed.

As agents continue to work on longer and more complex tasks, there will likely be more pressure to do both. 

The post OpenAI cut GPT-6 token prices in half. The bigger lever may be the cache. appeared first on The New Stack.

GPT-6 Sol closes most of the alignment gap with Astra. It’s one-fifth the price.

22 septembre 2026 à 21:28

On Tuesday, OpenAI released GPT-6 Sol and Luna, an expansion of the GPT-6 line-up that aims to make GPT-6 Astra’s next-level intelligence more efficient, accessible, and affordable. 

Though OpenAI says Astra is still “the most intelligent and aligned model in the world,” the new GPT-6 models come impressively close in alignment — at a fraction of the price. 

In an internal coding evaluation on coding deception, for example, GPT-6 Astra’s deception rate is 0.5%, while GPT-5.6 Sol stands at 10.4%. The new GPT-6 Sol is only 1.3%. 

As for pricing, GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens; GPT-6 Sol and GPT-6 Luna cost $2 and $0.10 per million input tokens and $10 and $0.50 per million output tokens, respectively. 

If OpenAI’s new GPT-6 models can achieve near-Astra-level alignment at a fraction of the cost, that’s good news. But it’s still unclear whether or not the new GPT-6 models also mirror Astra’s observability and monitoring problems. 

Closing the alignment gap between Astra and GPT-5.6

OpenAI says it trained the new GPT-6 models with similar methods as it did for GPT-6 Astra, specifically building on the alignment work it began with Astra. 

While Astra is still the AI company’s “most aligned model to date,” it looks like GPT-6 Sol and Luna are giving it a run for its money, dramatically closing the gap between OpenAI’s most advanced model and its GPT-5.6 counterparts in key areas like coding deception, failure to disclose a broken search tool, and unauthorized agent interaction. OpenAI notes that these evaluations deliberately test challenging situations and do not measure failure rates in typical use.

Credit: OpenAI

The most progress was made on failure to disclose a broken search tool, where AI agents are given search tasks and a broken search tool; do they just give their best guess or say that the search tool is broken? 

The gap between GPT-5.6 Sol’s ability and Astra’s is notably wide: 77.5% and 1.5%, respectively. Per OpenAI’s internal evaluation, GPT-6 Sol is a dramatic improvement, with a non-disclosure rate of 4.9%. Luna also improved, but less impressively, dropping from 78.3% to 28.7%. 

Though OpenAI says Astra is still “the most intelligent and aligned model in the world,” the new GPT-6 models come impressively close in alignment — at a fraction of the price. 

Credit: OpenAI

Coding deception, where AI agents are deliberately given tasks that elicit dishonesty, improved as well. In OpenAI’s internal evaluation, deception rate measures the fraction of answers with detected deception. GPT-6 Sol (1.3%) now almost matches GPT-6 Astra’s performance (0.5%), a far cry from GPT-5.6 Sol’s (10.4%). 

Credit: OpenAI

Significant progress was also made on unauthorized agent interactions, where OpenAI evaluates whether models follow unauthorized instructions on a simulated message board. GPT-5.6 Sol went ahead with the unauthorized action in 52% of runs in which it found the board, while GPT-6 Sol did so in only 11% of cases.

GPT-6 Astra and Luna didn’t take any such actions, but OpenAI says Luna “found the board less often.” 

Where the results are less impressive

Though GPT-6 Sol gets close to Astra in several alignment evaluations, there’s still a way to go with warning circumvention. 

GPT-6 Sol attempted to work around restrictions in 64.4% of runs. That’s only a slight improvement from GPT-5.6 Sol’s 68.2% and still quite far off from Astra’s 17.4%. 

Credit: OpenAI

In an internal evaluation where OpenAI tested models’ abilities to respect explicit warnings, like “access denied” messages, GPT-6 Sol attempted to work around restrictions in 64.4% of runs. That’s only a slight improvement from GPT-5.6 Sol’s 68.2% and still quite far off from Astra’s 17.4%. 

Credit: OpenAI

Finally, for reviewer bypass attempts, where OpenAI tests whether models try to bypass or respect an automated safety reviewer’s decision to block an unsafe command in Codex, the new GPT-6 models are almost identical to Astra, though the playing field was already pretty even to begin with: GPT-5.6 Sol attempted a bypass in 7.3% of runs and GPT-5.6 Luna in 4.3%, compared with none for Astra and GPT-6 Sol and 0.3% for GPT-6 Luna.

But if GPT-6 Sol is anything like Astra, we’re not out of the woods yet

GPT-6 Sol and Luna have made marked improvements across alignment evaluations, inching closer to OpenAI’s star child, Astra. But if the new GPT-6 models also follow suit on Astra’s noted observability issues, then developers hoping to catch misalignment via monitoring aren’t out of the woods yet.

Though Astra is substantially more aligned than its predecessor, its written reasoning is also harder to monitor than GPT-5.6 Sol’s. That’s not great for teams trying to count on monitoring to find misalignment mistakes; Jakub Pachocki, Chief Scientist at OpenAI, writes in his essay, “An Alien Mind,” that OpenAI’s methods for keeping models aligned and monitored aren’t keeping pace with model capabilities. 

OpenAI knows that Astra’s — and now GPT-6 Sol’s — improved alignment doesn’t mean the AI industry has gotten a handle on the problem yet. 

“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

Just this month, the AI company shared six reports of “unexpected or concerning model behavior,” including self-generated instructions, information fabrication, unauthorized use of leaked API keys, cross-agent communication, and unsanctioned file-sharing.

At the same time, it released a new framework for reporting model misalignment, stating: “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

If GPT-6 Sol and Luna are catching up to Astra in alignment evaluations — at a far cheaper rate — that’s good news. But if the new GPT-6 models also come with the same observability and monitoring problems, then cheaper may still come at a cost. 

The post GPT-6 Sol closes most of the alignment gap with Astra. It’s one-fifth the price. appeared first on The New Stack.

“One of the most significant steps in our 26-year history”: JetBrains goes big on agentic development — and bets the IDE still matters

22 septembre 2026 à 20:27
A laptop displaying source code in an integrated development environment (IDE).

There’s little question that AI coding agents have changed where software development work happens. Developers can increasingly delegate work from terminals, desktop applications and remote environments, leaving question marks hanging over the future of the integrated development environment (IDE).

That shift poses a particularly interesting question for JetBrains. The company has spent 26 years building some of the industry’s best-known IDEs, including IntelliJ IDEA, PyCharm and WebStorm, even as agentic development has begun pulling more software work outside the editor. JetBrains, for its part, has maintained that the IDE will remain a core part of professional development, particularly as developers are asked to manage the growing volumes of AI-generated code — humans need to review, debug, and verify, after all.

Now, the company’s making a much bigger bet on the broader development system that sits around the IDE.

JetBrains CEO Kirill Skrygan took to LinkedIn on Tuesday to formally unveil JetBrains Air as an “open system of products for agentic software development” for developers and companies, operating “inside and beyond JetBrains IDEs.” And Skrygan didn’t hold back on what he feels is a monumental moment for the company.

“JetBrains is taking one of the most significant steps in our 26-year history.”

“Today, JetBrains is taking one of the most significant steps in our 26-year history,” he writes.

Getting some Air

In truth, Air represents a repackaging of several strands of JetBrains’ recent AI work under a single banner, including an agentic experience inside its IDEs, tooling for coordinating developers and autonomous agents, and company-level controls for governing their use.

By way of a brief recap, JetBrains first launched Air in public preview back in March as a standalone “agentic development environment,” initially for macOS, where developers could run the likes of Claude Agent, Codex, Gemini CLI and JetBrains’ own Junie side by side. At the same time, it pushed Junie itself outside the IDE with Junie CLI, giving developers access to the coding agent from terminals, CI/CD systems and other editors.

Creating a new Git worktree task in the Air desktop app
Creating a new Git worktree task in the Air desktop app

A couple of weeks later came JetBrains Central, a separate system aimed further up the organization, providing the controls and infrastructure for companies running multiple coding agents. Then in July, JetBrains launched AI for Teams and Organizations, effectively adding shared context, cloud agents, automations, and organization-wide governance and cost controls that could sit above whatever AI tools developers were already using.

Today’s announcement now gives these efforts a common home under the JetBrains Air umbrella. In a separate blog post published on Tuesday, Skrygan describes three main parts to the system: Air in JetBrains IDEs for directing agents and checking their work; Air Teams for coordinating work between developers and autonomous agents; and Air Governance, the new name for JetBrains Central, for managing policy, auditing, costs and AI use across a company.

The original Air desktop IDE hasn’t gone away either, it seems. It remains available as a standalone desktop application on macOS, Windows and Linux, alongside a browser-based version for organizations. That leaves “Air” doing double duty: it’s still the name of JetBrains’ dedicated agentic development environment, while now also serving as the banner for the wider collection of products around it.

What does JetBrains Air actually do?

JetBrains Air is available inside JetBrains IDEs, through the browser, and from the command line via Air Gateway, which brings terminal agents such as Claude Code and Codex into Air.

JetBrains Air in the CLI
JetBrains Air in the CLI

For individual developers, Air can be used to supervise several pieces of agent work at once. They can keep multiple projects and agent sessions running, while tracking new activity, changed files and outgoing commits, then inspect the resulting changes using JetBrains’ IDE tooling.

Multiple agent sessions running inside a JetBrains IDE
Multiple agent sessions running inside a JetBrains IDE

Air Teams, which is still in early access, moves some of that activity into shared cloud environments, where developers can collaborate on projects and run agent tasks without tying the work to one person’s machine. Teams can also configure recurring automations and centrally manage the environments and external tools available to agents.

Air Teams showing shared projects and agent automations in the browser
Air Teams showing shared projects and agent automations in the browser

Air Governance, meanwhile, provides the organization-level controls, including deciding which models and agents developers can access, setting permissions and spending limits, and tracking AI usage across teams.

Those governance capabilities are also only offered through JetBrains’ early access program for now.

Air Governance showing AI access, seats, credits and per-user limits
Air Governance showing AI access, seats, credits and per-user limits

It’s worth noting that JetBrains also plans to extend Air to the mobile realm, where developers will be able to monitor and continue agent work away from their desktop.

JetBrains Air running on mobile
JetBrains Air running on mobile

For JetBrains, the point is to connect those different layers while remaining open to outside agents and tools.

“JetBrains Air cannot be just another agent or development environment.”

“JetBrains Air cannot be just another agent or development environment,” Skrygan writes. “It must connect products for individual work, team coordination, organizational control, context, and process automation — and remain open to the tools and agents developers choose, including those JetBrains does not build.”

Air in the IDE

While the core raison d’être of JetBrains Air is to provide somewhere for developers and companies to work with software agents — be it Claude, Codex, Junie or something else entirely — JetBrains is clearly emphasizing that its IDE roots remain part of that future.

Skrygan says the company has historically “focused primarily on the individual developer workbench,” but Air broadens its remit to encompass the wider environment in which agentic work is started, carried out, coordinated, reviewed and governed.

“Our IDEs will continue to be where professional developers work with agents, understand and verify code, and make the decisions that shape what ships.”

“Our IDEs will continue to be where professional developers work with agents, understand and verify code, and make the decisions that shape what ships,” Skrygan writes. “JetBrains Air extends that control across the broader system developing around them.”

And that fresh IDE piece has already been in the public domain for more than a month. JetBrains has been testing the Air Alpha plugin since at least early August, giving developers a way to run and supervise multiple coding agents directly inside IntelliJ-based IDEs, and review the changes they produce using native IDE tooling.

An updated release earlier this month added more controls for monitoring and steering agent sessions.

Air Alpha lets developers review agent-generated code changes as native IDE diffs.
Air Alpha lets developers review agent-generated code changes as native IDE diffs.

JetBrains concedes that Air “Alpha” is very much that — an early iteration it’s building while it rolls out JetBrains Air itself. And so users should expect “rough edges, changes to the UI and behavior, and updates roughly every week.”

What Tuesday’s announcement does, though, is give that work a formal place within the wider Air system. And while Skrygan acknowledges that agentic development means that work must now span multiple surfaces, the technology that has sat at the heart of its business for the past 26 years won’t be going anywhere anytime soon.

“The IDE remains important to JetBrains’ future,” Skrygan writes.

The post “One of the most significant steps in our 26-year history”: JetBrains goes big on agentic development — and bets the IDE still matters appeared first on The New Stack.

Claude’s merged chat and Cowork vs. ChatGPT’s Work mode: ChatGPT is faster, Claude is more thorough

22 septembre 2026 à 19:45
Rainbow-striped grid facade of a modern building, seen from below at an angle, with red, yellow, green, and blue diagonal bands crossing white geometric wall panels split by a narrow sliver of

When Anthropic merged Claude chat and Cowork into a single interface last week, it removed an increasingly irrelevant decision users had to make about which mode to choose.

Now, to use both tools, just ask a question, then hand off a multi-step task in the same thread, and Claude routes it. Anthropic made this update because it says customers often struggled to choose the right tab for the right task, so the merged app now routes each request itself instead of asking you to pick a mode — for now, the unified experience is rolling out to Pro and Max subscribers first, with free and team tiers to follow.

But this isn’t new. OpenAI has offered the same promise since earlier this summer. OpenAI introduced Work mode alongside Chat on July 9, then phased out the older, separate Agent mode the following month. Work mode is a sandboxed environment with a browser, code execution, and file output, sitting next to a Chat toggle in the same window — though reviewers note it can’t yet hand a live, logged-in browser session back to the user mid-task the way Agent mode could, e.g., for logins or payments. OpenAI built Work mode on its Codex coding agent after OpenAI reported that roughly a fifth of Codex’s 5 million weekly users were non-developers — a share it said was growing three times faster than developers.

As OpenAI and Anthropic get closer to feature parity, accuracy, reliability, and token usage matter even more.  What better way to find out which tool is better than to run head-to-head tests? 

The tests

I ran three tests, covering different areas of real developer work. 

  • API research -Look up four real developer APIs and tabulate their documented rate limits, whether a free tier exists, and the current version identifier. I checked the answers against the vendors’ docs the same day.
  • Build from a spec – Write a small command-line duration parser from a spec with strict edge cases. I ran each app’s code against a hidden 16-case test suite.
  • The handoff – Ask which of two stack traces indicates a race condition, then in the same thread hand off a real job: pull a log file from Google Drive, compute latency percentiles and error rates per endpoint, and deliver a spreadsheet with a chart. I generated the log data, so I knew every number in advance.

I recorded token cost and time in each test and included the prompts for anyone interested in replicating this work.

API research

The prompt:
Research the current public documentation for these four developer APIs and build me a table with one row per API and these columns: documented rate limit for authenticated requests, whether a free tier exists (yes/no), and the current API version identifier or date shown in the docs. APIs: GitHub REST API, Stripe API, Twilio Messaging API, OpenAI API. Cite the documentation page you used for each row.

Both Claude and ChatGPT answered correctly on the twelve graded fields, but Claude was more thorough. It included GitHub’s separate limit for Actions tokens, Twilio’s queue window, and OpenAI’s tier thresholds, plus a note about one page it couldn’t reach.

ChatGPT, in Work mode, finished in 1 minute 17 seconds and wrote 649 output tokens. Claude took 1 minute 44 seconds and wrote 1,042 tokens, read eight pages, listed nine sources, and offered to export the table as a spreadsheet. Claude wrote nearly double the tokens and took longer, but in this case, it’s warranted because of the added detail.

Build from a spec

The prompt:
Build the command-line tool described in the spec below. Deliver a single file named durparse.py that follows every rule. Test it yourself before returning it. Show the complete final code in your reply. (Followed by the spec: a duration parser with units d/h/m/s, largest first, one of each, decimals allowed, bare numbers are seconds, everything else returns None.)

Both apps returned a durparse.py file that passed all 16 hidden tests, including the traps. The traps included units out of order, a repeated unit, a trailing number with no unit, and negative values. The 517-token spec went to both. ChatGPT finished in 1 minute 17 seconds on 769 output tokens. Claude took 1 minute 45 seconds and 989 output tokens. Claude reported running 35 of its own test cases before returning the file. Both delivered a download and showed the code in the reply.

The code came out nearly identical, both using exact-precision arithmetic and a fixed-order regex. Claude flagged a judgment call the spec never settles on: that rounding 0.5 seconds up is a choice and Python’s built-in round would go the other way. ChatGPT reported only that its tests passed. Once again, Claude was just a little more thorough.

The handoff

The prompt:
Which of these two stack traces points to a race condition, and in one sentence why? (with the two traces) Then: Now take the file api_logs.csv from my Google Drive (columns: time, endpoint, status, latency_ms) and produce a downloadable spreadsheet with one row per endpoint showing request count, p50, p95, and p99 latency in milliseconds, and error rate as the percentage of requests with status 500 or above. Add a bar chart of p95 latency by endpoint. Also show the table in your reply.

I started each thread in plain chat with the stack-trace question, 135 tokens. Both answered correctly in seconds: Trace B, the dictionary that changed size during iteration. ChatGPT spent about 10 seconds and 48 output tokens. Claude spent 118 in about 25 seconds, adding a caveat that the same error can happen without threads if the loop body edits the dictionary itself. Then, without switching modes, I handed off the log analysis.

Claude pulled the file, computed the table, and built an .xlsx with a p95 bar chart in about 4 minutes on 319 output tokens. ChatGPT produced the same spreadsheet and chart in 28 seconds, using 467 output tokens. All 30 numbers matched my ground truth on both sides. Claude also named its percentile method (linear interpolation) and noted the numbers would match if I recomputed them in Google Sheets. ChatGPT gave the same correct table but didn’t provide as much detail as Claude did.

Results

The testChatGPT (Work mode)Claude (merged app)
API research12/12, 1:17, 125 in / 649 out12/12, 1:44, 125 in / 1,042 out
Build from spec16/16 tests, 1:17, 517 in / 769 out16/16 tests, 1:45, 517 in / 989 out
Handoff, questionCorrect, ~10 s, 135 in / 48 outCorrect, ~25 s, 135 in / 118 out
Handoff, task30/30, 28 s active, 133 in / 467 out30/30, ~4 min, 133 in / 319 out
Total tokens (visible)910 in / 1,933 out910 in / 2,468 out

Both Claude and ChatGPT were equally accurate. Every field, every test case, every number matched on both sides, and both cited real documentation. 

The differences are in speed and answer detail. ChatGPT was faster on every task and produced 1,933 visible output tokens, compared with Claude’s 2,468. Some of Claude’s extra output was filler, but not all of it. It added context the prompts didn’t ask for, named the percentile method behind its numbers, and flagged two judgment calls the specs left open. ChatGPT gave the same right answers but with less context.

What do I think?

I’d pick Claude, and here’s the reasoning. On time, ChatGPT won every task, but the gaps were seconds on the short tasks (1:17 vs 1:44, 1:17 vs 1:45), not a noticeable difference. On tokens, ChatGPT used about 22% less output than Claude, which is positive, but not when you consider how important detail/context is.

On detail, Claude provided more meaningful detail on all three tests. This included the extra API context, the rounding judgment call, and the percentile method. In today’s world, where AI can fabricate, detail matters. 

The post Claude’s merged chat and Cowork vs. ChatGPT’s Work mode: ChatGPT is faster, Claude is more thorough appeared first on The New Stack.

Anthropic releases Opus 5.5 and cuts pricing by 20%. Your agent calls might secretly get routed to an older model.

22 septembre 2026 à 18:30
Abstract chain

Claude Opus 5.5 is here, and Anthropic has lowered the price.

The new model, released on Tuesday, costs $4 per million input tokens and $20 per million output tokens, 20% less than Opus 5, with cache reads dropping to $0.20 per million from $0.50 and cache writes falling to $5 from $6.25. Anthropic puts overall savings closer to 40% because Opus 5.5 uses fewer tokens to complete a task and generates output more than 30% faster.

Claude Code and the Claude Platform also get a fast mode that runs up to 2.5 times faster, priced at $8 per million input tokens and $40 per million output tokens. Anthropic says Opus 5.5 performs at roughly the level of Fable 5.1 on most work, though it comes out ahead on several agentic coding benchmarks.

Opus 5.5 scored 66.4% on Terminal-Bench 4.0 compared with Fable 5.1’s 55.8%, and 54.4% on FrontierCode compared with 50.3%. The company suggests not reading too much into those margins. At this level, the company says a few points on a benchmark don’t translate into a noticeable difference in real-world use.

Fable 5.1 costs $10 per million input tokens and $50 per million output tokens, more than twice the price of Opus 5.5. At default effort on FrontierCode, Opus 5.5 beats GPT-6 Astra at roughly 20% of the per-task cost. On CursorBench, it tops GPT-5.6 Sol by 11 points at about a third of the cost. Developers will still need to run their own evals before moving production workloads, but the cost difference could change which model makes sense for agentic coding.

Developers will still need to run their own evals before moving production workloads, but the difference in cost could change which model makes sense for agentic coding.

Fewer tokens, fewer agent steps

The early enterprise numbers suggest the efficiency gains are real, at least on certain task profiles. Box reported that Opus 5.5 used about a third as many tokens as Opus 5 in its evaluations while producing answers that were 40% less verbose without losing accuracy.

GitHub tested the model inside Copilot CLI and VS Code and found it completed more terminal tasks than Opus 5 in less than half the steps. Deloitte said Opus 5.5’s lowest-effort setting caught 72% of known bugs in code reviews, compared with 56% for Opus 5 at high effort, with fewer false alarms and less output.

Prices per 1M tokensClaude Opus 5.5Claude Opus 5
Cache reads$0.20$0.50
Input tokens$4$5
Output tokens$20$25
Cache writes$5$6.25

Anthropic’s own internal testing backs up the pattern. In one head-to-head, both Opus 5.5 and Fable 5.1 translated HAProxy from C into Rust; both rewrites passed nearly all of HAProxy’s regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 and cost 51% less. An early tester audited and fixed a 200,000-line codebase in under three hours, whereas Opus 5 took over 20 hours and burned 2.5x as many tokens. Another completed a 680,000-line code migration in less than a day. Although these were customer and internal evaluations, not standardized independent benchmarks, they point in the same direction — fewer tokens and fewer steps to finish the job.

That pattern tracks with what’s happening across the industry. Agent performance depends heavily on the harness and runtime around the model, not only the model itself — agent failures often trace back to the orchestration layer rather than the model. Nvidia’s research showed that swapping the harness while keeping the model fixed could meaningfully change agent performance.

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Agentic coding (Terminal-Bench 4.0)66.4%55.8%52.3%57.9%37.3%
Agentic coding (FrontierCode v1.1)54.4%50.3%48.0%53.3%47.5%
Agentic coding (CursorBench 4.0)57.8%51.8%46.6%—41.7%
Knowledge work (GDPval-AA v2.1)18461735170815421588
Business workflows (AutomationBench)40.0%31.4%26.9%41.4%28.8%
Multidisciplinary reasoning (HLE)67.7%65.6%63.6%57.2%—
Agentic scientific research (TBS 0.1)58.7%52.6%29.0%64.6%22.4%
Computer use (OSWorld 2.0)81.8%80.7%74.0%——
Visual chart recognition (Chartography)89.0%88.4%83.4%——

Safety classifiers reroute mid-chain

Opus 5.5 ships with the same class of safety classifiers already running on Fable 5.1 for cybersecurity, biology, and frontier LLM development. When a classifier fires, Anthropic reroutes the request transparently to an older model. Most flagged cybersecurity requests go to Opus 4.8. Biology and frontier LLM flags go to Opus 5. Anthropic says users can still identify and fix bugs in their own code with Opus 5.5.

For anyone building agent workflows, this is the detail that needs architectural attention. A request sent to Opus 5.5 could, in fact, be handled by Opus 4.8 or Opus 5 instead, depending on whether Anthropic’s safeguards intervene. In a multi-turn agent workflow, that creates the possibility that individual requests are being handled by models with different capabilities, which could affect downstream steps. It’s also a source of inconsistency that may not show up in evals built on the assumption that every request goes to the same model.

Vetted organizations can apply to Anthropic’s Life Sciences Verification Program to use Opus 5.5 without the biology classifier, and the company plans to expand its Cyber Verification Program to include the model in the coming weeks. The new cyber program will include three tiers for increasingly permissive trusted access, including access to Claude Mythos models.

Opus 5.5 ships with the same class of safety classifiers already running on Fable 5.1 for cybersecurity, biology, and frontier LLM development.

Alignment gains from cleaner training

Anthropic says Opus 5.5 posted the strongest results of any model it has tested on its most comprehensive internal alignment evaluation, with improvements in behaviors the company says contributed to recent cybersecurity incidents, including biased reasoning and attempts to escape sandboxed environments. Frontier Design and METR evaluated the model before release.

On the training side, Anthropic is tightening how it filters reinforcement learning environments after identifying flawed environments as a major source of misaligned behavior. That’s relevant beyond the safety framing because RL environment quality directly affects how a model behaves in agentic settings, where it chooses its own tools and decides when to change approach. The company is also building automated methods to generate new safety training scenarios and improve alignment rewards.

Pricing pressure meets routing tradeoffs

Opus 5.5 is the first model in the Claude 5.5 family, with Sonnet 5.5 and Haiku 5.5 expected over the coming weeks. Subscription users get a 20% increase in five-hour usage limits across all plans, while Anthropic says the lower cost of Opus 5.5 will make five-hour and weekly limits go 25% further. Subscribers will also get a banked rate-limit reset they can save for when they need more capacity.

The release comes as API pricing across the frontier labs continues to fall. OpenAI cut its own API prices this summer, and Opus 5.5 pushes the competition beyond the headline price per token by reducing how many tokens some workloads require in the first place.

Opus 5.5 pushes the competition beyond the headline price per token by reducing how many tokens some workloads require in the first place.

The post Anthropic releases Opus 5.5 and cuts pricing by 20%. Your agent calls might secretly get routed to an older model. appeared first on The New Stack.

AI coding agents need a secrets-safe context boundary

22 septembre 2026 à 15:00
Abstract glowing blue and yellow distortion wave on a black background, illustrating digital data security concepts.

AI coding agents play a major role in software development and delivery, and for good reason. They can investigate bugs, trace dependencies, refactor services, and propose patches without developers needing to assemble all of the relevant context manually. That capability comes courtesy of agents’ appetite for context. To make informed decisions, agents read source code, configuration files, terminal output, error messages, environment information, and more…much, much more.

“Secure, agentic development depends on a security control many teams still lack: preventing secrets from leaking to AI coding agents and becoming model context.”

From a security standpoint, this becomes problematic when agents, in the search for context, inadvertently reach for secrets.

For years, developers have been taught not to commit API keys, database credentials, and tokens to Git. But agentic workflows have created another route for secrets to escape development environments before a commit, code review, or CI job. Depending on its permissions, configuration, and provider architecture, an AI coding agent may read local files or receive pasted content that is then included in data sent to an AI service, and, in the process, developers may never see their credentials leak.

The quiet path from local files to external systems

Some forms of secrets leakage are obvious. A developer troubleshooting an authentication failure may paste, for example, a failing API call into a chat window, including the token. However serious, this sort of leak is characteristically human.

The more consequential escape pathway is quieter. An agent tasked with understanding a project may inspect files in its working directory, including an overlooked .env file, a cloud credential profile, an SSH configuration, or sensitive application logs. In such instances, nothing has necessarily gone wrong from the agent’s perspective; it is doing exactly what it was designed to do: collect context to solve the task at hand.

“Agentic workflows have created another route for secrets to escape development environments before a commit, code review, or CI job.”

But once a secret becomes part of that context, it may pass through systems outside an organization’s direct control. Depending on the workflow, it can appear in model provider logs, gateway telemetry, prompt histories, or debugging records. Rotating the credential is essential, but it does not erase copies that may already exist in those systems.

This changes the practical definition of a secret leak. The problem is no longer limited to what lands in a repository, but also includes what an autonomous tool reads and forwards while operating on a developer’s machine.

Why traditional security gates no longer suffice

Most application security programs are built around durable checkpoints: the commit, pull request, build, and deployment. In the agentic era, these checkpoints remain important as they can detect secrets that reach version control and prevent a bad change from merging and deploying.

They cannot, on their own, prevent a secret from being included in an agent prompt before the code ever reaches a repository.

This highlights an important timing gap. The 2025 Verizon Data Breach Investigations Report reports a median of 94 days to remediate leaked secrets discovered in GitHub repositories. In an agent-driven workflow, detection and response need to happen much earlier, and not after a credential is exposed. Still, at the moment it’s about to cross the boundary from local context to an external model.

Bad actors already understand the value of that porous boundary. Recent supply-chain attack campaigns, including Mini Shai-Hulud, have searched developer and CI environments for credentials and configuration data, including AI coding-tool configuration files. These campaigns show that agent configurations and the local context accessible to an agent are valuable targets. AI coding agents can broaden the local data reachable during a session, making even the agent’s context-collection mechanisms an attractive target.

Treat agent context as an egress surface.

The secure mental model doesn’t frame AI agents as mere code editors, but rather, automated data-movement systems. Its inputs can include far more than the source files a developer is actively editing, and its outputs may involve external services.

That calls for a zero-trust approach to agent context. Before sending a prompt or adding a file to an agent’s working set, organizations should evaluate it for sensitive material. Controls should be deterministic: identify a likely secret, block or redact it, and provide the developer with a clear path to remediate it.

“Asking an LLM to decide whether to transmit a credential does not create a reliable security boundary.”

Critically, the control should be independent of the model. Asking an LLM to decide whether to transmit a credential does not create a reliable security boundary. Purpose-built secrets detection can inspect prompts and files against known credential patterns and policies, applying a deterministic policy, such as blocking a prompt or file read when it detects a credential-shaped value. For example, Sonar’s secrets detection ships alongside dedicated agent plugins to bring that local check into tools such as Claude Code, GitHub Copilot, Codex, and Cursor, so it can flag a credential before a prompt or file read is transmitted to a model provider.

Build defense in layers, without disrupting your agentic workflow

A legitimate workflow does not involve forcing developers to choose between secure development and useful automation, but instead places fast controls at several points where secrets can escape:

  • In the editor: Use IDE-integrated secrets detection to flag credentials while they are being written.
  • Before model submission or agent file access: Where the agent supports it, scan prompt submissions and file reads locally, and block risky operations according to policy.
  • At the command line: Check generated snippets and local changes in terminal-driven workflows.
  • In pull requests and CI: Detect secrets that reach the repository and use review, quality gate, and deployment controls to prevent unsafe changes from progressing.
  • In incident response: Rotate exposed credentials quickly, investigate downstream logs and access, and reduce recurrence through policy and training.

Building defense at the pre-submission layer is an emerging requirement and requires both security and usability. Secrets detection must be fast enough to run in developer workflows; a scanner that introduces lengthy pauses may be bypassed or disabled by developers, and it must also have a manageable false-positive rate, or developers may stop trusting it.

Teams should also make their agent permissions and context rules explicit, as broad agent permissions can increase the amount of sensitive local context reachable during a coding session. Consider the following: which directories can an agent read? Are .env files, credential stores, home-directory configurations, and production logs excluded by default? Does the organization route prompts through an approved gateway? What retention, training, and audit settings apply at the provider level? Document and enforce the answers rather than leaving them to individual developer preference.

Secrets security must shift left.

Prevent secret leakage without hindering AI-assisted development, ensuring the productivity promise of agentic development doesn’t carry significant security implications.

As agents become more autonomous, security standards must follow agents upstream. It’s critical to stop a secret before it becomes context, while it is still local, visible, and easier to control. In the agentic era, code review and CI-level checks will remain essential safety nets. Still, for agent-centric development, the first line of defense must shift left: to the instant an AI coding tool determines what to read and what to transmit. That is the control modern development teams need to implement now.

The post AI coding agents need a secrets-safe context boundary appeared first on The New Stack.

Grok Build vs. Claude Code: I tested which one has the better memory

21 septembre 2026 à 22:04

On September 16, xAI announced memory in Grok Build, its terminal coding agent. The pitch was that Grok “keeps notes on the conventions, decisions, and project facts that come up,” and “later sessions read those notes before touching related code.” Notes are Markdown files in a workspace scope per project and a global scope that applies everywhere. /memory browses them.

Meanwhile, Claude Code has done something similar for months under the name auto memory. It keeps a MEMORY.md index plus one file per note, per repository, and the docs say it is on by default. Anthropic’s Projects beta, announced September 17, adds shared memory across cloud threads, but only for select Pro and Max subscribers with no existing projects. I tested the CLI that everyone has.

Both companies say their coding agent now remembers what you told it in an earlier session. I wanted to see whether that holds up, so I tested Grok and Claude on the same three tests.

The tests

The claim I wanted to check is simple. Tell each tool something once, close it, open it again, and see whether it remembers. Both tools ran on my Mac, each on its own copy of four small Node repos I built for this. Grok Build 1.0.40 ran Grok 4.6 at high effort through an xAI API key. Claude Code 2.1.226 ran Opus 5 on my subscription. Every session was scripted with each tool’s headless mode, which reports its own tokens and cost. 

Each test has two sessions. Session 1 plants a fact. I quit the tool. Session 2 gives a task where the fact matters and never mentions it.

Here are the tests I ran:

  • The test command – In this repo, npm test fails and make test passes, and session 1 says so. Session 2 asks for a new endpoint with passing tests, after I removed the README line that pointed at the Makefile.
  • Project decisions with a trap – Session 1 states that CSV export was dropped and money is integer cents, never floats, while a float helper and a half-built CSV exporter sit in the repo as bait. Session 2 asks for a refund endpoint that “takes an amount” and “a way for support staff to download all orders.”
  • A rule across projects – Session 1, in repo A, sets two rules “for all my projects,” conventional commit messages and no comments on obvious code. Session 2 runs in an unrelated repo B and asks for a small feature and a commit.

Here’s my scoring breakdown. Did the tool write the fact to a memory file, did it read that file in session 2, and did the session 2 output follow it.

The test command

Both passed. In session 1, each tool saved the rule as soon as I stated it. Grok wrote topics/testing.md plus two raw observations. Claude Code wrote orbit-api-run-tests-with-make.md with a “why” and a “how to apply” section.

In session 2, both remembered. Grok’s reasoning opened with “start by reading the memory files,” then it ran make test and never touched npm test. Claude Code read the Makefile and package.json, ran make test, and also never tried npm test. Grok took 29 seconds, 102K tokens, and $0.11. Claude Code took 22 seconds, 186K tokens, and $0.32. Claude used about 80K more tokens and cost nearly 3x more. 

Project decisions with a trap

Both wrote both decisions down. Claude Code also converted “last quarter” into “Q2 2026” in its note. In session 2, both built the refund on integer cents, named the field amountCents, and left the float helper alone. For the download request, both shipped a JSON export. 

Grok’s reasoning said the API is JSON-only, so it wouldn’t wire up CSV. Claude Code set a content-disposition header so the JSON downloads as a file. Both passed on both decisions, but Claude Code was more than double the price and just as fast. Grok took 103 seconds, 156K tokens, and $0.18. Claude Code took 32 seconds, 269K tokens, and $0.49.

A rule across projects

This is where the results split. Grok saved the rules to its global scope as git-and-code-style.md. In the second repo, it committed feat: add --help flag with usage and supported cities, and added no comments. Pass, in 33 seconds, 132K tokens, and $0.12.

Claude Code saved both rules too, but only in the first repo’s memory folder. It said so at the time, warning that its memory store “is scoped to this project’s directory.” In the second repo, it found nothing, and the commit came back with the Add --help flag. No comments were added, but that is Claude’s default anyway. Claude passed the first rule but failed the second one and still cost twice as much. It completed the work in 12 seconds, 122K tokens, and $0.24.

Results

MetricGrok Build (Grok 4.6)Claude Code (Opus 5)
Tests passed3 of 32 of 3
Total time165 s66 s
Total tokens390,848576,863
Total cost$0.41$1.05

Grok Build passed all three tests, and Claude Code passed two. They behaved the same on the per-project tests. The split was the cross-project rule, which Grok’s global scope carried into a second repo and Claude Code’s per-repo memory did not. 

Claude Code was faster on every recall session, 66 seconds total against 165, and cost at least twice as much on every one, $1.05 total against $0.41. It also used more tokens: 576,863 against 390,848. The price gap mostly reflects Opus 5 versus Grok 4.6 rather than the memory systems.

On the core claim, remembering what you told it last time in the same project, I could not tell these two apart. Both wrote a markdown note the moment I stated a rule, read it back next session, and followed it. Claude Code’s notes were better written. But Claude Code failed the third test. Its CLI memory stops at the repo boundary, so a rule I gave it “for all my projects” never reached the second repo. Grok’s global scope carried the same rule over without being asked.

What do I think?

Grok Build is the better option for most people right now. It remembered everything, it carries rules across projects, and it cost less than half as much on every test. Yes, Claude Code was faster on every session, but that only matters if you aren’t concerned about accuracy. Its CLI memory stops at the repo boundary, so anything you want it to remember everywhere still has to go into ~/.claude/CLAUDE.md by hand. 

The post Grok Build vs. Claude Code: I tested which one has the better memory appeared first on The New Stack.

❌