❌

Vue normale

Reçu aujourd’hui — 28 septembre 2026The New Stack

Anthropic bought Stainless and shuttered its SDK generator. Cloudflare open-sourced Forge instead.

28 septembre 2026 à 15:39
Illustration of interconnected software components and code running across a developer system.

Cloudflare has announced an open source tool that takes an API definition and automatically produces the SDKs, command-line tools, documentation, and other interfaces used to interact with it.

Forge, as it’s called, is available under an Apache 2.0 license, and can be run and modified privately without paying Cloudflare a dime. In a blog post published on Monday by Dimitri Mitropoulos, Matt Taylor, and Samuel MacLeod, Cloudflare notes that the project is still early: it already generates the output required for Cloudflare’s cf CLI, its unified command-line interface for working across the Cloudflare platform, with the company’s API documentation and SDKs due to move over in the coming months.

“We believe that building tools for APIs is a core part of the Internet, and you should be able to do that without needing a SaaS product.”

“We believe that building tools for APIs is a core part of the Internet, and you should be able to do that without needing a SaaS product,” the authors write.

When your SDK generator disappears

Cloudflare’s announcement comes just 11 days after Google made a somewhat similar move, revealing it had partnered with API tooling company Speakeasy to open-source the latter’s OpenAPI code-generation suite under an AGPLv3 license.

In a blog post published September 17, Google said its decision was prompted by the sudden loss of its SDK-generation service, after the provider was acquired and “abruptly announced its shutdown.” This happened just as Google was preparing to launch a new API.

“This sudden disruption highlighted that proprietary, closed-source generators create unacceptable platform risk.”

“This sudden disruption highlighted that proprietary, closed-source generators create unacceptable platform risk,” the company wrote. “If the industry relies on OpenAPI to define interfaces, the tooling to compile those interfaces into client libraries, CLIs, and agent tools should be open infrastructure.”

The provider Google was referring to was Stainless. Anthropic acquired the SDK and MCP tooling company on May 18, after which Stainless said it would wind down its products, including the SDK generator used to keep client libraries updated as APIs changed.

As The New Stack reported at the time, the closure affected a swathe of customers including OpenAI, Google and Cloudflare. While those companies retained the SDKs Stainless had already generated for them, they lost the shared service they had relied on to regenerate and update those SDKs as their APIs evolved.

Cloudflare, for its part, had already begun building some of the machinery that would become Forge. In April, the company released a technical preview of cf, a new unified command-line interface intended to eventually expose Cloudflare’s products through a consistent set of commands for developers and AI agents. At the time, Cloudflare said it had created a new TypeScript-based schema and generation system to keep CLI commands, configuration, bindings and other interfaces in sync with its API.

That earlier preview was limited — cf initially covered only a small subset of Cloudflare products — although the company said it was already testing a version spanning its entire API surface. Forge is the code-generation system that has emerged from that work and now generates the output required by cf.

Tools such as Forge save API providers from having to maintain each SDK, CLI command set and documentation surface separately by hand. When the underlying API changes, the same definition can be used to produce updated SDKs, CLI commands and documentation for the developers — or agents — consuming it.

Why Cloudflare built its own generator

While Stainless was one of the hosted generation products Cloudflare had relied on in production, Cloudflare doesn’t name it specifically, instead spelling out why Cloudflare had become dissatisfied with the broader category of hosted generators.

Part of the problem is sheer volume: Cloudflare says it now has more than 3,500 API operations, backed by hundreds of services maintained by different engineering teams. A generator therefore has to keep changes across that sprawling API estate accurately reflected in the tools customers use, without turning every update into a coordination exercise across the company.

Cloudflare also wanted engineers to see what an API change would do to the resulting SDK, CLI and documentation before the underlying code was merged. Moreover, it wanted the same system to reach beyond conventional SDKs into things such as MCP servers and Cap’n Web bindings, which allow applications to invoke remote services through Cloudflare’s TypeScript-based RPC system in much the same way they would call local functions.

That had proved difficult with external services. Cloudflare says problems introduced in one part of its API could surface only when another team later tried to release something, leaving engineers to trace the failure back through systems they did not fully control and then coordinate a fix across company and vendor boundaries.

“We’ve tried several hosted products that attempt to solve this, and relied on some in production,” the authors write. “None of them solved this problem for us, and some have shut down entirely.”

The open-source factor

And so the Stainless wind-down goes some way toward explaining why Cloudflare is making Forge open source. Stainless’s generator was proprietary and delivered as a hosted service, leaving customers dependent on the company continuing to operate it. Forge’s permissive Apache 2.0 license means developers can run it on their own infrastructure, modify it, or fork the project and continue using it independently of Cloudflare.

AI agents raise the stakes further, because they increasingly rely on SDKs, CLIs and API bindings to understand what software can do and how to interact with it. If those interfaces lag behind the underlying service, the agent may be working from an incomplete or outdated description of its capabilities.

That is the concern Cloudflare CTO Dane Knecht highlights when explaining why keeping those interfaces current matters more than ever today.

“APIs have always been how software connects, but AI agents make the quality of SDKs, CLIs, and API bindings even more important,” Knecht tells The New Stack over email. “If that layer is stale or incomplete, an agent doesn’t just have a worse developer experience, it can misunderstand what a service can do.”

Ultimately, Cloudflare created Forge because SDK and CLI generation has “become too important” to sit outside its internal development process.

“Cloudflare needed a pipeline our teams could run, test, and extend themselves, and we think other developers should have access to that same kind of open, vendor independent foundation,” Knecht says.

Forge still has plenty to prove, of course. The project is early, cf is its first production output, and Cloudflare says its API documentation and existing SDKs will move over during the coming months. Several of the input formats and generated targets it discusses are also future work.

But the timing is quite revealing. Within 11 days, Google and Cloudflare — both former Stainless customers — have publicly thrown their heft behind open source SDK-generation infrastructure after the disappearance of the proprietary service they had relied on. Google describes that dependency as an “unacceptable platform risk.” Cloudflare, with Forge, is now trying to remove the same risk from its own stack.

The post Anthropic bought Stainless and shuttered its SDK generator. Cloudflare open-sourced Forge instead. appeared first on The New Stack.

Nvidia launches Open Agent Safety Platform to lock down rogue AI agents

28 septembre 2026 à 12:00

OpenAI, Anthropic, Meta, and Google have all recently disclosed that their models broke out of their test environments and reached real systems. Nvidia’s response, announced Monday, is a runtime that locks agents into kernel-enforced sandboxes and a watchdog on its own silicon that can shut them down.

The Nvidia Open Agent Safety Platform combines OpenShell 0.1.0, the Apache 2.0 agent runtime the company announced at GTC in March, with Nvidia Sentry, a watchdog service that runs on the company’s BlueField-4 data processing units (DPUs).

The new OpenShell release adds a policy prover that checks that an agent’s various permissions can’t be combined into something the operator didn’t intend — like hacking HuggingFace.

Since the BlueField DPU is a separate processor with its own trust domain, it can watch the agent’s traffic to the model and keep an eye on all of its actions and reasoning. Then, when things go awry, it can cut the agent off at the network level.

Justin Boitano, Nvidia’s vice president of enterprise AI, said in a press briefing that the recent incidents “have highlighted a fundamental hurdle for AI agents, and that is that model-level safeguards alone can’t govern what agents can access or do.”

“To date, model safety has been about training good behavior into the model. The industry calls that model alignment,” Boitano said. “For probabilistic systems, this approach has obvious limitations. That’s why we’re introducing a deterministic system to mediate and enforce how these agents behave.”

Credit: Nvidia.

The summer of sandbox escapes

OpenAI disclosed on July 21 that GPT-5.6 Sol and a research prototype had exploited a zero-day in the package proxy that was their sandbox’s only network path and gone on to reach Hugging Face’s production database.

Only a few days later, Anthropic reported that three of its models had found unintended internet access at its evaluation partner Irregular and, among other things, accessed a real company’s database and published a live malicious package to PyPI.

Meta followed on Aug. 6 with a pre-release Muse Spark model that read and modified a real website’s database after the same Irregular misconfiguration handed it a real site name.

Most recently, Google announcedthat Gemini had broken into the networks for three companies, again from an Irregular test environment.

Irregular, which runs pre-release cyber evaluations for frontier labs, is a common factor here, and it’s worth noting that the company is on Nvidia’s partner list for this launch.

Nvidia’s own technical blog puts the breakouts down not to a new capability but to “a combination of tools, time, and ambiguous instructions.”

Asked whether the platform would have prevented the Hugging Face breach, Boitano said each incident is unique, but “from what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on.”

Enforcing policy outside the agent

OpenShell is the core of the platform, and it hasn’t changed all that much since Nvidia first showed it at GTC in March.

With OpenShell, which Nvidia originally announced in parallel with its NemoClaw distribution of OpenClaw, each agent runs in a kernel-isolated sandbox with no network access except through a supervisor that sits outside the workload.

“Traditional sandboxes, whether we’re talking micro VMs or containers or VMs, were built for application-level isolation,” Boitano said. “Every agent running within your company needs to run in its own isolated sandbox with security controls that are outside of the agent’s reach.”

A prover, not a judge

OpenShell is now at version 0.1.0, and the important new component added in this update is a policy prover. This Prover checks that the permissions a given policy grants always stay within the boundary the operator actually intended.

“It is deterministic. It is mathematical reasoning. So this is not LLM as a judge,” Ali Golshan, Nvidia’s senior director of AI software, said during the briefing. Because of that, he said, it runs “roughly at two orders of magnitude higher performance and speed.”

In Golshan’s example, a policy can, for example, bar an agent from reading code on GitHub and posting it externally.

“An agent can bypass this by spawning two sub-agents: one that can read from GitHub, that can talk to another one, that could also then post outside,” he said. The prover models the combined access of the entire agent fleet to find that path.

In Nvidia’s own tests, agents running with reduced safeguards spent up to two hours trying to talk an AI reviewer into granting write access to a protected repository. The prover, the company says, gave the reviewer evidence of what the request actually allowed, and no protected writes occurred.

Sentry: the safety island

Nvidia Sentry adds an additional hardware layer to this system. It runs on BlueField-4 in a trust domain separate from the host, and according to Nvidia, it can quarantine an agent in milliseconds.

With a DPU in the system, the agent’s model endpoint gets routed “through a proxy on the DPU, so that you can see all of the reasoning traces of the agents on the host,” Boitano said.

Unlike OpenShell, Sentry isn’t open source, though Boitano said it has open APIs and that OpenShell can work with other network enforcement hardware.

He compared it to autonomous vehicles. “There’s a primary system that might be running the perception system, and then a safety island that ensures the safety of the system.”

“The DPU is really optional in these architectures,” Boitano said. “In a lot of cases, just using OpenShell on CPUs is honestly good enough for providing sort of strict access control for the agents.”

The DPU, he said, is for “frontier use cases of evaluating models or systems where you might have the guardrails off the models, so it could be for red teaming.”

Who’s building on it

Anthropic is integrating OpenShell with Claude Managed Agents, which already keeps the agent loop on Anthropic’s infrastructure and pushes tool execution into customer-controlled sandboxes.

SpaceXAI says it’s using the platform for Cursor coding agents and Grok models, while Salesforce has added OpenShell audit events and permission approvals into Slack.

SAP is embedding the runtime into Joule Studio and is also contributing code.

OpenAI and Google, two of the four labs whose agents went rogue this summer, aren’t on the partner list. Neither is AWS.

Asked whether Anthropic and OpenAI plan to run OpenShell and Nvidia Sentry for their own training runs, Boitano said to look for the partners’ own blog posts.

The post Nvidia launches Open Agent Safety Platform to lock down rogue AI agents appeared first on The New Stack.

Reçu avant avant-hierThe New Stack

Cursor acquired Firetiger. A month later, it launched a bot that tracks code changes from PR to production.

24 septembre 2026 à 14:56
Inspecting changes on a laptop screen document

We all know that producing code is easier than ever thanks to the abundance of AI coding tools and agents. The harder part undoubtedly comes after that code is written: making sure changes are safe to ship, spotting regressions in production, and figuring out what went wrong.

And that’s why Cursor is introducing Rollouts, a new agent that follows code changes into production and monitors whether they behave as intended.

The Firetiger effect

The announcement comes a little over a month after SpaceX closed its bumper $60 billion acquisition of Cursor, giving the AI coding company access to SpaceX’s vast GPU infrastructure as it develops its own models.

The day before that deal closed, however, Cursor quietly announced an acquisition of its own: it snapped up the team behind Firetiger, a three-year-old startup building AI agents that monitor software changes from pull request through deployment.

At the time, Firetiger co-founder and CEO Rustam Lalkaka argued that coding agents had dramatically reduced the effort involved in creating software changes, while doing little to reduce the risks involved in actually deploying them.

“Over the last two years, agentic coding has changed software dramatically,” Lalkaka wrote in a LinkedIn post following the deal’s announcement. “The cost of creating changes has dropped to near zero. The cost and risk of deploying them has stayed largely the same.”

“Writing code is no longer the slow part. What hasn’t sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.”

Rustam Lalkaka, Cursor

Fast forward to today, and Lalkaka, now at Cursor, has unveiled the first fruits from that acquisition — including Rollouts. In a blog post published on Wednesday, Lalkaka notes that the new agent, or “bot” as the company calls it, is all about helping developers “get safe, reliable code into production faster.”

“Writing code is no longer the slow part,” Lalkaka writes. “What hasn’t sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.”

Rollouts is effectively Firetiger’s Change Monitors reborn inside Cursor, rebuilt using a tool dubbed Bot Development Kit. This kit, too, appears to be new from Cursor: an early-stage framework for building and serving Cursor bots and agents, published as the @cursor/bdk package on npm. Its documentation says developers can define agents using Markdown and TypeScript, with support for tools, skills, subagents, webhooks and scheduled runs.

Like Change Monitors before it, Rollouts starts working when a pull request opens. It examines the proposed code change, works out which systems could be affected, and produces a monitoring plan covering what the change is supposed to do, the risks it sees, the signals it intends to watch, and any holes in the available instrumentation. Developers can review and edit that plan before the code reaches production.

Rollouts in action (1)
Rollouts generates a monitoring plan for a change

Once the change is deployed, Rollouts checks the resulting telemetry — including logs, metrics and traces — against that plan. Staging and production are assessed independently, with each deployment ultimately receiving one of three verdicts: verified healthy, regression detected or inconclusive.

That means a change could, for example, pass its checks in staging before Rollouts subsequently spots a problem when the same code reaches production.

Rollouts in action (2)
Rollouts reports deployment status as changes ship

If Rollouts does detect a regression, it can identify the change it suspects, alert the developer responsible and, depending on how it’s been configured, either open a revert pull request for review or hand the problem to a Cursor cloud agent to attempt a fix. There is still a human in the consequential part of that loop for now: Rollouts doesn’t merge fixes or roll back deployments by itself, though it can pause a progressive rollout.

Lalkaka notes that Rollouts is already capable of picking up problems limited to a particular endpoint or region before they trigger a broader alert, while it can also distinguish expected changes in behavior from genuine regressions.

Also “coming soon” to Rollouts, according to Cursor, is an integration with feature flags so it can directly adapt the traffic reaching a change, while support for release trains and deployment freezes is also in the works.

Enter Security Reviewer

Alongside Rollouts, Cursor is also introducing an upgraded Security Reviewer bot, which first appeared in beta back in April.

At launch, the bot could automatically inspect pull requests for security vulnerabilities, authentication regressions, privacy and data-handling risks, agent tool auto-approvals, and prompt-injection attacks, leaving findings alongside the relevant code.

As with Rollouts, the idea is that developers don’t have to remember to invoke it manually: Security Reviewer can be set to run whenever a new pull request is opened.

Security Reviewer in action
Security Reviewer runs automatically on new pull requests

In its current guise, Security Reviewer analyzes pull requests in the context of the wider codebase, with a focus on exploitable issues such as injection flaws and broken authentication, and returns a severity rating, attack path and proposed fix.

“Security Review reads code the way a security engineer does,” Lalkaka writes. “Where does user input enter, where does it end up, what does it pass through on the way.”

“Security Review reads code the way a security engineer does.”

He says that things have sped up considerably, too: average review time has fallen 21%, from 4.8 minutes to 3.8, while developer acceptance of its comments has risen from roughly 45–50% to 60–70%.

Both Rollouts and Security Reviewer are available through Cursor’s Automations tab for customers on its Teams and Enterprise plans.

The Origin story

Digging into the nuts and bolts of Rollouts reveals how it might serve as a boon for Cursor as it builds out Origin, the fledgling Git-compatible code hosting platform it launched back in August.

Origin is essentially an effort to build an alternative to GitHub for an agent-heavy software development world. It remains early, with limited functionality, but Cursor has been clear that tighter integration with its own agents is supposed to become one of the main reasons to use it.

When Cursor announced the Firetiger acquisition last month, Maxime Prades on the Cursor product team noted in a blog post that the deal was part of a “broader investment in long-running, autonomous, context-aware agents for teams.”

And he pointed to Origin and Change Monitors as two examples of that investment.

“Agents that write code should also be able to tell whether it works in production,” Prades wrote. “Today, those systems are mostly separate. Cursor and Firetiger bring them closer together so an agent can ship a change, see how it behaves, and respond when something goes wrong.”

Rollouts offers an early glimpse of that. It can connect to either Origin or GitHub for source control, pull deployment events from continuous delivery systems, and use signals from Datadog and other telemetry providers. If it spots a regression, it can then pass the problem back to a Cursor cloud agent to investigate or attempt a fix.

Origin potentially gives Cursor a native home for more of that loop: its cloud agents can already create branches, commit and push code, and open pull requests against Origin repositories. Rollouts then adds information about what happened after.

That could become increasingly important as more companies take aim at GitHub’s central role in software development. Zed, for example, put Delta into public beta last week, with its own ideas about how source control should change for teams working heavily with agents.

Cursor also faces competition further downstream. Datadog’s Bits Release, launched in preview in June, similarly follows changes from pull request into production and checks telemetry for regressions. Harness has long offered automated deployment verification and rollback based on logs and metrics, while LaunchDarkly’s Guarded Rollouts can monitor feature releases for regressions and automatically reverse them.

What Cursor can potentially bring to the table is proximity: the coding agent, repository, pull request, security checks, and production feedback can all sit much closer together. Rollouts doesn’t require Origin — GitHub remains supported — but owning the forge gives Cursor more room to integrate those pieces over time. And that may prove more compelling than simply recreating GitHub’s existing feature set.

The post Cursor acquired Firetiger. A month later, it launched a bot that tracks code changes from PR to production. appeared first on The New Stack.

“Impressive level of openness”: Xiaomi goes way beyond the usual open-weight playbook with MiMo-V2.6

23 septembre 2026 à 22:00
A picture of an open laptop

New models are coming out thick and fast, almost on a weekly cadence, ranging from the powerful proprietary systems coming out of the major US AI labs to the more open alternatives being released by some of China’s biggest tech companies.

On Tuesday alone, Anthropic debuted Claude Opus 5.5, while OpenAI launched GPT-6 Sol and Luna, each accompanied by their the usual claims about how they outperform their rivals. Amidst all the hullabaloo of the frontier-model frenzy, however, Xiaomi also debuted MiMo-V2.6, another powerful open model from one of China’s growing ranks of AI developers.

All the initial headline numbers look pretty promising, too. The flagship MiMo-V2.6-Pro is a trillion-parameter model, with 42 billion parameters active at a time, a one-million-token context window, and support for text, images, audio and video. Broadly speaking, that puts it in the same frontier territory as the latest models from OpenAI and Anthropic: GPT-6 Sol has a 1.05-million-token context window, while Claude Opus 5.5 has a one-million-token window, though neither company discloses comparable parameter counts.

Xiaomi, for its part, makes broad claims of frontier-level performance across coding, agentic tasks, cybersecurity, multimodal work and research. Independent analysis lends some weight to those claims –Artificial Analysis gives MiMo-V2.6-Pro an Intelligence Index score of 46, ranking it first among the 114 large open-weight models it tracks.

Artificial Analysis  Intelligence Index
Artificial Analysis Intelligence Index



So far, so good. But arguably the bigger story in Xiaomi’s offering is the manner in which it trained the model, how much of that process it showed in public, and what it’s releasing afterward.

A public record

Xiaomi livestreamed its RL training through a public dashboard, exposing metrics from the production reinforcement-learning runs in real time over a five-day period starting on September 15. By the time the runs had finished, the dashboard showed costs of $854,044 for the smaller MiMo-V2.6-Flash model and $2,620,670 for Pro — about $3.5 million combined.

Xiaomi livestreamed its RL runs over a 5-day period.
Xiaomi livestreamed its RL runs over a 5-day period.

It’s worth noting that this figure covers only the RL stage; Xiaomi hasn’t said what pretraining the models cost. Even so, public RL bills are rare. The closest precedents came last year, when MiniMax said the RL phase of its 456-billion-parameter MiniMax-M1 cost $534,700 in GPU rental, and DeepSeek put the RL training of its 671-billion-parameter R1 at $294,000. Both were leading open reasoning models when they launched, though the comparison only goes so far: MiMo-V2.6-Pro is larger, and its RL run targeted longer, agentic tasks.

Shortly after the stream began, Fuli Luo, who leads Xiaomi’s MiMo team after previously working at DeepSeek, took to X to explain the thinking behind the project. The team, she said, had spent almost six months exploring how far RL could be pushed, increasing the amount of training, the variety of environments and agent setups, and the resources used to grade the model’s attempts.

“We’ll open-source the details piece by piece over the coming weeks,” she added.

Nearly half a year of silence. We spent it studying one problem: how far RL can scale.

MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task…

— Fuli Luo (@_LuoFuli) September 16, 2026

Responding on X, Hugging Face co-founder and chief science officer Thomas Wolf called the move an “Impressive level of openness on such a large run.”

However, what Xiaomi’s putting out alongside the finished models is arguably just as interesting. The company has released the model weights under the permissive MIT license, alongside its technical report and a 9-billion-parameter Qwen-based model, intended as a starting point for further agentic RL research.

“Impressive level of openness on such a large run.”

Xiaomi says it has also “fully open-sourced” a broader set of RL resources: more than 7,000 task environments spanning software engineering, vulnerability reproduction, knowledge work and web development; an end-to-end training framework covering everything from environment interaction to reward evaluation and policy optimization; and lightweight agent harnesses for experimenting with different tools, prompts and context setups. At the time of writing, however, Xiaomi’s link to the open-source collection on Hugging Face contain only the three model releases, with the 7,000-plus environments and other supporting resources not surfaced there. Luo had said earlier that Xiaomi would be open-sourcing the various elements “over the coming weeks.”

As the results began arriving this week, attention in the research community quickly moved beyond the benchmark score to what Xiaomi had committed to releasing overall. Elie Bakouch, a former Hugging Face researcher who is now a research engineer at Prime Intellect, singled out the promised RL resources.

“The most insane part, they will release ~7k RL training data and the framework leading to this top 6 model on AA,” Bakouch writes on X. “They also shipped the model + tech report less than 1 week after starting the final RL run.”

Wolf went further, arguing that access to the environments in which models learn may now be especially valuable for open research, as more model development shifts toward RL with verifiable rewards (RLVR). Because RLVR depends on tasks whose outcomes can be automatically checked — whether code passes a test, for example — the environments themselves become a crucial ingredient in training.

“Releasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now.”

“Releasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now,” Wolf writes. “The equivalent of sharing high quality pretraining data, but in the new RLVR paradigm.”

Open-weight vs open-source

So while the benchmarks around Xiaomi’s latest model are notable in their own right, it’s the company’s approach that is generating much of the fanfare so far.

Indeed, MiMo-V2.6 serves as a useful example of a distinction that often gets muddied in the AI sphere: “open-weight” and “open-source” are routinely used as though they mean the same thing, but they don’t. Many “open” models amount largely to downloadable weights — essentially, the vast collection of numerical values a model learned during training, which can then be used to run or fine-tune it — while much of what went into producing them remains closed.

Some companies have gone further in muddying those terms. Meta, for example, has often referred to its Llama models as open-source despite significant restrictions that have led open-source advocates to push back heavily on that description.

And so MiMo-V2.6 goes further than most open-source releases. Its MIT license carries none of the conditions that the likes of Moonshot’s Kimi K3 and Alibaba’s Qwen3.8-Max attach for large commercial users. And if the environments are released as promised, outside researchers will have much more of the post-training process to inspect and build on.

The post “Impressive level of openness”: Xiaomi goes way beyond the usual open-weight playbook with MiMo-V2.6 appeared first on The New Stack.

“One of the most significant steps in our 26-year history”: JetBrains goes big on agentic development — and bets the IDE still matters

22 septembre 2026 à 20:27
A laptop displaying source code in an integrated development environment (IDE).

There’s little question that AI coding agents have changed where software development work happens. Developers can increasingly delegate work from terminals, desktop applications and remote environments, leaving question marks hanging over the future of the integrated development environment (IDE).

That shift poses a particularly interesting question for JetBrains. The company has spent 26 years building some of the industry’s best-known IDEs, including IntelliJ IDEA, PyCharm and WebStorm, even as agentic development has begun pulling more software work outside the editor. JetBrains, for its part, has maintained that the IDE will remain a core part of professional development, particularly as developers are asked to manage the growing volumes of AI-generated code — humans need to review, debug, and verify, after all.

Now, the company’s making a much bigger bet on the broader development system that sits around the IDE.

JetBrains CEO Kirill Skrygan took to LinkedIn on Tuesday to formally unveil JetBrains Air as an “open system of products for agentic software development” for developers and companies, operating “inside and beyond JetBrains IDEs.” And Skrygan didn’t hold back on what he feels is a monumental moment for the company.

“JetBrains is taking one of the most significant steps in our 26-year history.”

“Today, JetBrains is taking one of the most significant steps in our 26-year history,” he writes.

Getting some Air

In truth, Air represents a repackaging of several strands of JetBrains’ recent AI work under a single banner, including an agentic experience inside its IDEs, tooling for coordinating developers and autonomous agents, and company-level controls for governing their use.

By way of a brief recap, JetBrains first launched Air in public preview back in March as a standalone “agentic development environment,” initially for macOS, where developers could run the likes of Claude Agent, Codex, Gemini CLI and JetBrains’ own Junie side by side. At the same time, it pushed Junie itself outside the IDE with Junie CLI, giving developers access to the coding agent from terminals, CI/CD systems and other editors.

Creating a new Git worktree task in the Air desktop app
Creating a new Git worktree task in the Air desktop app

A couple of weeks later came JetBrains Central, a separate system aimed further up the organization, providing the controls and infrastructure for companies running multiple coding agents. Then in July, JetBrains launched AI for Teams and Organizations, effectively adding shared context, cloud agents, automations, and organization-wide governance and cost controls that could sit above whatever AI tools developers were already using.

Today’s announcement now gives these efforts a common home under the JetBrains Air umbrella. In a separate blog post published on Tuesday, Skrygan describes three main parts to the system: Air in JetBrains IDEs for directing agents and checking their work; Air Teams for coordinating work between developers and autonomous agents; and Air Governance, the new name for JetBrains Central, for managing policy, auditing, costs and AI use across a company.

The original Air desktop IDE hasn’t gone away either, it seems. It remains available as a standalone desktop application on macOS, Windows and Linux, alongside a browser-based version for organizations. That leaves “Air” doing double duty: it’s still the name of JetBrains’ dedicated agentic development environment, while now also serving as the banner for the wider collection of products around it.

What does JetBrains Air actually do?

JetBrains Air is available inside JetBrains IDEs, through the browser, and from the command line via Air Gateway, which brings terminal agents such as Claude Code and Codex into Air.

JetBrains Air in the CLI
JetBrains Air in the CLI

For individual developers, Air can be used to supervise several pieces of agent work at once. They can keep multiple projects and agent sessions running, while tracking new activity, changed files and outgoing commits, then inspect the resulting changes using JetBrains’ IDE tooling.

Multiple agent sessions running inside a JetBrains IDE
Multiple agent sessions running inside a JetBrains IDE

Air Teams, which is still in early access, moves some of that activity into shared cloud environments, where developers can collaborate on projects and run agent tasks without tying the work to one person’s machine. Teams can also configure recurring automations and centrally manage the environments and external tools available to agents.

Air Teams showing shared projects and agent automations in the browser
Air Teams showing shared projects and agent automations in the browser

Air Governance, meanwhile, provides the organization-level controls, including deciding which models and agents developers can access, setting permissions and spending limits, and tracking AI usage across teams.

Those governance capabilities are also only offered through JetBrains’ early access program for now.

Air Governance showing AI access, seats, credits and per-user limits
Air Governance showing AI access, seats, credits and per-user limits

It’s worth noting that JetBrains also plans to extend Air to the mobile realm, where developers will be able to monitor and continue agent work away from their desktop.

JetBrains Air running on mobile
JetBrains Air running on mobile

For JetBrains, the point is to connect those different layers while remaining open to outside agents and tools.

“JetBrains Air cannot be just another agent or development environment.”

“JetBrains Air cannot be just another agent or development environment,” Skrygan writes. “It must connect products for individual work, team coordination, organizational control, context, and process automation — and remain open to the tools and agents developers choose, including those JetBrains does not build.”

Air in the IDE

While the core raison d’être of JetBrains Air is to provide somewhere for developers and companies to work with software agents — be it Claude, Codex, Junie or something else entirely — JetBrains is clearly emphasizing that its IDE roots remain part of that future.

Skrygan says the company has historically “focused primarily on the individual developer workbench,” but Air broadens its remit to encompass the wider environment in which agentic work is started, carried out, coordinated, reviewed and governed.

“Our IDEs will continue to be where professional developers work with agents, understand and verify code, and make the decisions that shape what ships.”

“Our IDEs will continue to be where professional developers work with agents, understand and verify code, and make the decisions that shape what ships,” Skrygan writes. “JetBrains Air extends that control across the broader system developing around them.”

And that fresh IDE piece has already been in the public domain for more than a month. JetBrains has been testing the Air Alpha plugin since at least early August, giving developers a way to run and supervise multiple coding agents directly inside IntelliJ-based IDEs, and review the changes they produce using native IDE tooling.

An updated release earlier this month added more controls for monitoring and steering agent sessions.

Air Alpha lets developers review agent-generated code changes as native IDE diffs.
Air Alpha lets developers review agent-generated code changes as native IDE diffs.

JetBrains concedes that Air “Alpha” is very much that — an early iteration it’s building while it rolls out JetBrains Air itself. And so users should expect “rough edges, changes to the UI and behavior, and updates roughly every week.”

What Tuesday’s announcement does, though, is give that work a formal place within the wider Air system. And while Skrygan acknowledges that agentic development means that work must now span multiple surfaces, the technology that has sat at the heart of its business for the past 26 years won’t be going anywhere anytime soon.

“The IDE remains important to JetBrains’ future,” Skrygan writes.

The post “One of the most significant steps in our 26-year history”: JetBrains goes big on agentic development — and bets the IDE still matters appeared first on The New Stack.

OpenAI releases GPT-6 Sol and Luna — and cuts token prices in half

22 septembre 2026 à 20:00

OpenAI on Tuesday released GPT-6 Sol and Luna, which will complement the flagship GPT-6 Astra model in OpenAI’s lineup. As of now, there is no GPT-6 Terra.

The new GPT-6 pricing

The headline news here is that OpenAI cut the price per million input/output tokens by half or more, compared to the previous version. GPT-6 Sol will cost $2/$10 per million input/output tokens (vs. $4/$20 for GPT-5.6 Sol), and GPT-6 Luna will come in at $0.10/$0.50 (vs. $0.20/$1.20).

The GPT-5.6 pricing was always meant to be promotional, but for the new GPT-6 models, this is the default price, an OpenAI spokesperson tells The New Stack.

“Improvements in caching and inference let us serve these models at lower cost, and we’re passing those savings directly on to users and customers,” OpenAI explains in its announcement.

Benchmarks

As you would expect, the new models show clear improvements over the GPT-5.6 predecessors, but for the most part, these are not all that extreme.

On a benchmark like Zapier’s AutomationBench — which checks how well the models work on a set of business workflow tests — GPT-6 Luna improves by 5.4 percentage points over the previous version, for example

Credit: OpenAI

On the DeepSWE v1.1 software engineering benchmark, GPT-6 Sol essentially matches Anthropic’s Fable (68.8% at max effort vs. 69.9% for Fable 5 at xhigh effort), but at only 20% of the cost. Luna, at max effort, hits scores similar to Claude Opus 5 and Fable 5 at medium effort, at a significantly lower cost.

And OpenAI focuses on this cost comparison across its announcement—with a special focus on price per task instead of straight-up token pricing.

Credit: OpenAI

Anthropic resets the comparison

Since Anthropic released Opus 5.5 earlier on Tuesday, OpenAI’s comparisons are already out of date — such is the way of this AI era. Anthropic, too, reduced its per-token pricing for Opus 5.5 to $4/$20, down from $5/$25, but that still leaves Anthropic’s model twice as expensive as the comparable GPT-6 Sol.

In its announcement, when comparing GPT-6 Sol to Opus 5, OpenAI was able to claim significant cost savings when compared to Anthropic’s model — and for the most part that still holds, but Anthropic says Opus 5.5 also uses fewer tokens per task, which, according to the company, works out to 40% lower costs than Opus 5 on typical workloads.

It’s worth noting that no one has run Sol and Opus 5.5 head-to-head yet. Sol likely stays cheaper per task on OpenAI’s AutomationBench numbers, but Opus 5.5 posts higher scores than GPT-5.6 Sol on shared benchmarks in Anthropic’s testing.

Since it’s almost impossible to know how many tokens an agent will use to finish a task, though, these pricing changes still don’t make it any easier for a user to budget.

Prompt caching

For developers building agents, the caching changes may matter more than token prices. OpenAI says it improved prompt caching for GPT-6 to deliver higher cache hit rates by default, with discounts of up to 90% on cached input tokens.

One positive change, too, is that developers can now change the reasoning effort and tool availability without invalidating the cache. With explicit breakpoints, developers can choose where a cached prefix ends, and a new dashboard and diagnostics tool show what’s getting cached and what isn’t.

GitHub says these improvements cut the share of prompt tokens that require fresh processing by more than half over the past several months, across billions of requests to OpenAI models.

Anthropic made a similar move with Opus 5.5, which cuts cache read prices by 60% for token-billed usage, on top of the 20% per-token cut.

Style changes

Models aren’t just about benchmarks, though. With GPT-6 Sol, OpenAI made its models answer more directly, rather than in the previous — already reined-in — more conversational style. “Expect to see more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall without losing substance,” OpenAI says.

Credit: OpenAI

Alignment

Given the HuggingFace incident, it’s no surprise OpenAI is emphasizing its alignment work for GPT-6 Sol and Luna, too.

OpenAI says both models improve on their GPT-5.6 counterparts across its alignment evaluations, including fewer misleading claims about their own coding work. On an internal coding deception test, GPT-6 Sol’s rate fell to 1.3% from 10.4%.

When given a deliberately broken search tool — and graded on whether it disclosed the problem instead of guessing — Sol failed to disclose the problem 4.9% of the time, down from 77.5%.

What is a bit more concerning, though, is that when researchers asked the model to respect an explicit warning like an “access denied” message, GPT-6 Sol still tried to work around those restrictions in 64.4% of runs, down only slightly from 68.2% for its predecessor. Luna improved more, to 42.4% from 76.5%.

OpenAI says these tests cover mostly low-stakes situations and run without full system-level safeguards used in its products.

Credit: OpenAI

On a simulated message board seeded with unauthorized instructions, including requests to disclose private information, Sol took the specified action in 11.3% of runs where it found the board, down from 51.9%. Luna and Astra took none, though OpenAI notes Luna also found the board less often.

Anthropic, by contrast, says Opus 5.5 is the strongest performer on its most comprehensive alignment test and names METR and Frontier Design as pre-release external testers. Opus 5.5 also ships with safeguards that reroute requests, sending most cybersecurity tasks to Opus 4.8 and anything flagged by Anthropic’s biology or frontier LLM development classifiers to Opus 5.

Availability

GPT-6 Sol and Luna are available in ChatGPT Work and Codex starting Tuesday for Plus, Pro, Business, Enterprise, and Edu users.

Free and Go users get Luna in the desktop app.

Neither model is in Chat yet. OpenAI says it plans to roll them out gradually throughout the day to keep service stable, so they may not appear right away.

The post OpenAI releases GPT-6 Sol and Luna — and cuts token prices in half appeared first on The New Stack.

AWS open-sources an AI agent it says is 45% cheaper than Claude Code and Codex

21 septembre 2026 à 20:30
Multiple monitors on a desk.

Amazon Web Services (AWS) is lifting the lid on a new open source, general-purpose AI agent, designed to give developers a ready-made foundation they can run locally or deploy to the cloud.

Strands Harness, as it’s called, builds on Strands Agents, which AWS debuted in May 2025 as an open source Python SDK for building AI agents. Strands takes what AWS calls a “model-driven approach”: developers provide the model, tools and instructions, while the model determines how to tackle the task and when to call those tools. AWS later brought Strands to TypeScript, and in February created Strands Labs as a separate home for more experimental projects.

Marc Brooker, VP and distinguished engineer at AWS, tells The New Stack that Strands Harness essentially sits above the existing Strands SDK, giving developers a preconfigured Strands Agent. It brings together the tools and supporting machinery an agent needs to operate over longer-running tasks, with AWS supplying its own defaults for how those pieces work together.

“You still need to decide how to manage context, persist conversations, integrate tools, and guide the agent’s behavior.”

“An SDK like the Strands Harness SDK gives you the building blocks, but you still need to decide how to manage context, persist conversations, integrate tools, and guide the agent’s behavior,” Brooker explains.

Unpacking Strands Harness

Out of the box, Strands Harness gives developers a working agent with file, shell and web tools, alongside built-in handling for context, memory, persistent sessions, prompt caching and delegation to other agents.

Developers can install Strands Harness as a Python or TypeScript package, using pip install strands-harness or npm install @strands-agents/harness.

AWS also provides the Strands CLI as an interactive way to prototype and configure an agent. Developers can choose a model, add prompts, tools and other capabilities, then use /export to generate the resulting agent as Python or TypeScript code.

Configuring an agent with the Strands CLI.
Configuring an agent with the Strands CLI.

Individual agents can be tailored to different jobs, with developers able to change their instructions, choose which model they use, control which tools and capabilities are available to them, and decide whether they can hand work off to another agent.

Demo of Strands Harness running on a desktop
Demo of Strands Harness running on a desktop (Credit: AWS)

Most of Strands Harness itself doesn’t depend on AWS infrastructure. The agent loop, tools, context management, session handling and delegation are all included in the open source release, and AWS says those processes run on the machine that’s running the agent by default.

The exception is the call to the underlying model. Perhaps unsurprisingly, AWS routes model access through Amazon Bedrock, its managed service for accessing and running foundation models. However, while Brooker says that this is the only out-of-the-box default tied specifically to AWS infrastructure, it too can be switched out.

“This is easily overrided to use a different model provider with one line,” he says.

Strands Harness can instead use Anthropic, OpenAI or Google as its model provider, or use a locally running model through Ollama. Changing provider doesn’t necessarily mean changing the underlying model, but Brooker notes that choosing a different model will obviously affect how the agent behaves.

“Different models have different strengths on reasoning, tool use, and cost,” he continues. “What doesn’t change: context management, sessions, tools, delegation all work the same regardless of provider. No features require Bedrock.”

It’s worth noting that all the other defaults can be changed, too. Developers can bring their own tools and skills, connect MCP servers, alter how context is handled, and choose where session state is stored.

“Developers can focus on their application’s task and domain expertise, while customizing the components that need different behavior.”

“Developers can focus on their application’s task and domain expertise, while customizing the components that need different behavior,” Brooker says.

AWS benchmarks its agent

AWS says Strands Harness is intended as a general-purpose agent rather than a coding assistant, though it takes cues from harnesses such as Claude Code and Codex. The difference, AWS says, is that developers can deploy Strands Harness to whichever cloud provider they choose — addressing what it describes as a common wish among developers using Claude Code and Codex to be able to run the same setup in the cloud.

From its own testing, AWS suggests the way a harness manages the surrounding agent machinery can materially affect cost and performance, even when the underlying model stays the same. For each harness, AWS averaged its score across six benchmarks — ALFWorld, ContextBench, GAIA, WebShop, τ³-bench and Terminal-Bench 2.1– and compared that with the average cost per task across the same tests. Against Claude Code and Codex specifically, the company says Strands Harness came out 45% cheaper, with broadly comparable accuracy.

However, that figure drops to 28% once DeepSeek Harness — which AWS says ran around 14% cheaper than Strands Harness on matched runs — is folded into the wider comparison.

Strands Harness benchmark results.
Strands Harness benchmark results. (Credit: AWS)

AWS points specifically to its context-management defaults as a major reason for the result. Strands Harness truncates particularly large tool outputs, compacts context once the available window passes a set threshold, and attempts to recover within the agent loop if the context overflows.

On Terminal Bench 2.1 specifically, AWS says Strands Harness running Fable 5 cost 77% less than Claude Code, at $56.29 versus $248.05 across 89 trials, while scoring 69.7 versus 61.8. DeepSeek Harness was cheaper again at $40.30, though its score was lower at 59.5.

Terminal Bench 2.1 results.
Terminal Bench 2.1 results. (Credit: AWS)

For AWS, those results help make the case for packaging and tuning functions such as context management, versus requiring every developer to work out those decisions independently with the SDK.

“Getting a prototype working is one step; evaluating how those choices affect performance and cost is another.”

“Getting a prototype working is one step; evaluating how those choices affect performance and cost is another,” Brooker says. “The opportunity we saw was to package that engineering into a complete, general-purpose agent.”

What’s in it for AWS?

AWS also has an obvious place to run the resulting agent. Amazon Bedrock AgentCore is its managed service for deploying and operating agents, providing identity and access controls, observability and the infrastructure needed to host them.

There is, in fact, a close technical relationship between the open source project and that managed offering. AgentCore Harness and Strands Harness were built by the same team, although they live in separate codebases. Brooker says work on one can also feed improvements into the other, giving AWS a route for technology developed in the open source project to inform its managed service, and vice versa.

Brooker, again, stresses that Strands Harness can be deployed independently of AgentCore, outside of AWS altogether.

“AgentCore is an optional hosting layer for teams that want AWS to manage the infrastructure side,” he says. “However, all deployment paths are open for the developer to choose.”

Still, this arrangement gives AWS a clear commercial path: developers can adopt Strands Harness freely, while AgentCore gives the company a natural destination for teams that eventually want AWS to run the infrastructure around it.

The post AWS open-sources an AI agent it says is 45% cheaper than Claude Code and Codex appeared first on The New Stack.

Your AI agent failed. The model might not be the problem.

20 septembre 2026 à 16:04
Tangled wires abstract

As AI agents move into production, the path between a request and its result is becoming less predictable. An agent can choose its own tools and change course as it works, which makes failures harder to diagnose when there isn’t an obvious error to trace.

In a recent interview with The New Stack, Nvidia VP of Product Adel el Hallak described the additional visibility developers will need as agents take on more complex work.

Nvidia is also part of an industry effort to share what companies learn when those systems fail. The Secure Agent Findings Exchange, or SAFE, is backed by roughly 140 companies and aims to create shared infrastructure for reporting agent failures, borrowing from vulnerability disclosure in traditional software.

“When we find these vulnerabilities, it’s not just for one company,” el Hallak tells The New Stack. “It’s for everyone to patch across.”

But agent failures don’t necessarily trace back to a single component, raising a more basic question for developers. When an agent fails, what exactly do you debug?

“When we find these vulnerabilities, it’s not just for one company. It’s for everyone to patch across.”

Why traditional observability falls short

With conventional software, developers usually have a starting point when something goes wrong, whether it’s an exception, a failed request or a service that goes down. An agent can keep running while heading in the wrong direction, carrying an earlier mistake through the rest of a task without producing anything that looks like a conventional software failure — or, as el Hallak put it, simply deciding to “get creative” when it shouldn’t.

Even the best-performing coding agents fail more than 60% of the time on tasks drawn from real codebases. Knowing the agent failed, though, is different from knowing why.

“It’s not enough to just look at the logs or the inputs and the outputs,” el Hallak tells The New Stack. “It is important to figure out how it got to the answer. What were the reasoning traces? What tools did it utilize? Where did it get stuck? Where did it decide to try a new approach?”

That can require replaying the agent’s execution to see where it went off course. What looks like a model failure may have started somewhere else in the stack. And that’s the tricky part for developers. Agent bugs aren’t always model bugs.

“It’s not enough to just look at the logs or the inputs and the outputs. It is important to figure out how it got to the answer. What were the reasoning traces? What tools did it utilize? Where did it get stuck? Where did it decide to try a new approach?”

Runtime as collection point

Nvidia sees the runtime as the logical place to capture much of that information. Its OpenShell agent runtime, which sits underneath the NemoClaw platform, manages sandboxing, and policy enforcement while providing visibility into an agent’s execution.

El Hallak called OpenShell the one non-negotiable component across Nvidia’s reference architectures.

“You can change whatever harness you need. I’m even open to using whatever models you need,” el Hallak tells The New Stack. “But the governance, the secure and open runtime that we want to leverage at all times is OpenShell.”

Nvidia breaks the agent stack into three layers: the model provides the intelligence, the harness orchestrates its work, and the runtime governs execution. When an agent fails, the model itself may not be what went wrong.

Nvidia’s NOAH research, for example, showed that changing the harness while keeping the underlying model fixed can improve agent performance, which also means a poorly matched harness can drag down an otherwise capable model.

“Every model’s different. Some could be more chatty than others,” el Hallak tells The New Stack. “Making sure those two things are either co-developed together or have profiles that are specific to models is a new unlock.”

Safety as systems engineering

Nvidia CEO Jensen Huang has described AI safety as an engineering problem, an approach el Hallak compared to traditional software testing.

“If there’s a bug in your software, you don’t release it,” el Hallak tells The New Stack. “You work until it’s fixed and it passes all your tests.”

Agents complicate that model because reproducing a failure can require reconstructing what happened across the system. That requires instrumentation, which comes with its own cost. OpenAI has found that monitoring adds roughly 20% to inference compute for its most capable persistent agents.

Nvidia’s approach combines governed harnesses, sandboxed runtimes and confidential computing intended to protect models and user data.

“There are ways where you make guarantees all the way down to the silicon,” el Hallak tells The New Stack.

SAFE extends that engineering approach beyond a single company’s systems by creating infrastructure for organizations to share what they learn when agents fail.

“If there’s a bug in your software, you don’t release it. You work until it’s fixed and it passes all your tests.”

Agents debugging other agents

CrowdStrike is fine-tuning Nvidia’s Nemotron models on years of security data to create paired agents, with one finding exploits and another patching them.

If either agent goes wrong, the final output may not reveal why. A bad patch, for example, could trace back to the model, the agent’s execution path or the tools it used along the way.

“I don’t need general purpose for a given task. I need specialization,”el Hallak tells The New Stack

As companies build agents around increasingly specialized workflows, those failures may not show up in general-purpose model benchmarks or safety tests, putting more pressure on developers to understand what happened during execution.

Toward shared failure reporting

For platform teams, finding the failure is one problem. Reconstructing enough of the agent’s execution to understand what caused it is another.

SAFE is intended to make those findings useful outside the company where they were discovered. Traditional software has established systems for sharing vulnerabilities and fixes, but nothing comparable exists yet for agent failures. The goal is to keep every team from having to discover the same failure on its own.

The post Your AI agent failed. The model might not be the problem. appeared first on The New Stack.

“Dormant deployments were quietly consuming storage”: Why Vercel tightened its free-tier rules

19 septembre 2026 à 15:00

Vercel announced this week that teams on its free Hobby plan will now have older, unprotected deployments deleted immediately if they exceed the 10GB Deployment Storage limit. 

When asked why Vercel decided to change its retention rules, Jas Garcha, head of pricing at Vercel, tells The New Stack the update is a move to keep the free tier viable amid rapidly growing deployment volumes: 

“This change allows us to continue supporting a Hobby community that’s deploying at a much higher rate than it was a year ago.”

Old deployments now deleted immediately

Before, users in the free tier could count on eligible deployments to stick around for up to 30 days. But that window is gone for teams over the limit, and the protections that used to spare older deployments have narrowed for every Hobby project.

Now, if users exceed the standard 10GB of Deployment Storage included on the free tier, old deployments not covered by Vercel’s retention-policy exceptions will be deleted immediately, per Vercel’s updated Deployment Retention Policy. 

“This change allows us to continue supporting a Hobby community that’s deploying at a much higher rate than it was a year ago.”

What gets to stay? 

Vercel says each Hobby project will keep the three most recent production deployments, along with its three most recent deployments of any type, regardless of age. That’s a cut from the previous Hobby exception, which preserved the 10 most recent production deployments, and it applies to every Hobby project — not only teams over the 10GB limit.

Plus, preview deployments lose a separate protection

In addition to cutting the 30-day holding period for over-limit teams, Vercel’s policy update also removes a separate retention exception for preview deployments. 

Preview deployments have not vanished from Vercel’s exception list entirely — the latest preview deployment on an active Git branch is still protected on every plan. What Hobby lost is the count-based exception: Pro and Enterprise teams keep their last 20 non-production deployments in a Ready state, and that protection no longer applies to Hobby.

Per Garcha, “Your current production deployment is never deleted, and aliased and active-branch deployments remain protected, along with each project’s most recent deployments.” 

Why the change?

Garcha tells The New Stack that Vercel’s latest policy update is needed to keep the Hobby tier sustainable as deployment volumes dramatically rise:  

“Our former retention defaults were designed for teams that ship constantly and need deep rollback history. They made less sense for Hobby projects, where dormant deployments were quietly consuming storage that active projects need.” 

“Your current production deployment is never deleted, and aliased and active-branch deployments remain protected, along with each project’s most recent deployments.” 

And activity is much higher, even compared to a year ago. According to Garcha, Vercel now handles more than 10 million deployments every day— more than a 6x increase YoY. Immediately deleting older, unprotected deployments, Garcha says, is one way to free up storage for active projects as Vercel handles much higher deployment rates.

In other words, Vercel is moving out some of the old to make room for the new. It describes how the storage limit works in its post:

“Every deployment you keep uses some of it [Deployment Storage], and going over the limit can block you from deploying until you free some up.” 

What Hobby users should do

It’s important to note that deletion isn’t immediately permanent. Vercel gives successfully built deployments a 30-day recovery period, and users can restore them from a project’s Settings, under Security → Recently Deleted. Hobby users who want to stop old, unprotected deployments from being deleted in the first place can move off the free tier and onto the Pro plan. In the Pro tier, storage beyond the plan’s included allowance is billed at $0.10 per GB-month, and the retention exceptions stay far more generous: the last 10 deployments created in a project, the last 20 production deployments in a Ready state, and the last 20 non-production ones.

For users who can’t or don’t want to upgrade to Pro, Vercel offers guidance for optimizing Deployment Storage usage to help users stay under the free 10GB storage limit, like reducing unnecessary deployment output.

Still, Garcha says few Hobby users will feel the effects enough to warrant making a change. Pointing to Vercel’s list of exceptions that still protect certain deployments, he tells The New Stack, “The vast majority of Hobby users won’t notice the change.” 

Garcha also says Vercel’s stricter retention policy helps the company keep offering Hobby as a permanent free plan.

“We’re one of the few platforms where the free tier isn’t a trial or a credit that expires. It’s a permanent plan, and we’ve kept expanding it,” he says. “By ensuring its resources go to people actively building, we’re able to continue offering it.”

The post “Dormant deployments were quietly consuming storage”: Why Vercel tightened its free-tier rules appeared first on The New Stack.

Claude couldn’t hack OpenAI. Then Anthropic shipped Opus 5.

18 septembre 2026 à 20:57
Abstract door

Three security researchers at Hacktron AI found a memory-corruption bug in a widely used image library. Finding it was the easy part.

The hard part was turning it into something that works on a real server, so on July 24 they handed that job to Anthropic’s Claude Opus 4.8. The model managed it only with the operating system’s memory randomization switched off. With the protection on — this is the way every production box runs it — nothing it wrote held up.

That evening, Anthropic released Opus 5.

The researchers came back the next morning with the same bug and the new model. Roughly three hours later, Opus 5 had a working ARM64 exploit running against a Mac on their desk. About four hours after that, they had remote code execution against a test forum.

Less than 72 hours after they started, they were reading from OpenAI’s private monorepo — using an OpenAI employee’s Codex account to open a pull request against a README, then stopping there.

Hacktron AI published its account of the incident on its website this week.

It started with an image

The bug wasn’t in anything OpenAI wrote. Hacktron was testing community.openai.com, the company’s user forum, which runs on Discourse — the same off-the-shelf forum software behind thousands of other sites.

Discourse normally screens uploaded images with FastImage. But FastImage doesn’t handle HEIC and HEIF, so those files get passed to ImageMagick instead, and ImageMagick decodes them with libheif. The version running in the Debian 12 base image the forum used, 1.19.7, had a heap buffer overflow that a specially crafted file could trigger.

A fix had landed upstream the previous year. But the commit wasn’t documented as a security fix and never got a CVE, so it never triggered a backport into the Debian package the forum was using — a patched bug that stayed exploitable because nobody labeled it.

The researchers adapted the exploit for the x86-64 and jemalloc configuration Discourse runs, and a malformed HEIC image was enough to trigger remote code execution.

Discourse later confirmed the vulnerability in security advisory GHSA-vhm9-85gw-x335, rating the upstream libheif flaw — tracked as CVE-2026-32882 — 8.8 out of 10 on the CVSS severity scale. The New Stack has reached out to Hacktron AI for additional details about the researchers’ use of Claude and will update this story if we hear back.

One exploit, a much larger path

Code execution on a forum is a bad day for the forum. But it shouldn’t be a bad day for the company that owns the forum. This is where the chain crossed into something that was OpenAI’s own. Hacktron then found a flaw in OpenAI’s single sign-on system: sign-in tokens issued for the forum carried excessive permissions, granting full API access to the linked ChatGPT and Codex accounts. Some of those accounts belonged to OpenAI employees.

One employee’s Codex account was connected to OpenAI’s GitHub environment, opening a path to the company’s private repositories. Hacktron says other accounts could have exposed connected services including Slack and email.

The team stopped there. Using Codex, they made a harmless documentation change against OpenAI’s private openai/openai monorepo and opened a pull request — enough to prove the access was real, and nothing more. Hacktron’s write-up says the pull request’s details were redacted at OpenAI’s request.

From assistant to exploit developer

Up to this point, those three experienced researchers were still in the loop. So Hacktron ran the experiment again with the humans mostly out of it.

They put Claude in an autonomous agent loop — giving it a goal, a target, and time to keep working — pointed at a Discourse instance of their own. The model got there on its own, achieving remote code execution and demonstrating it by reading /etc/hosts from inside the container.

Getting it started took one piece of misdirection: Opus refused to write an exploit aimed at a live remote host. So the team proxied their own instance through rce.ee/ctf-forum, a URL that made the target look like it was part of a capture-the-flag exercise.

Memory-corruption exploitation has always been specialist work, invovling memory layouts, allocators, operating system internals, and protections to make all of it wildly unreliable. Hacktron’s run signals a meaningful share of that work might be able to be delegated to AI now. It also suggests the line between security research and attack development is — from the model’s side, anyway — partly a question of what you consider a target.

The full chain

Put together, the attack looked like this:

HEIF upload → libheif overflow → code execution on the forum → over-permissioned SSO tokens → employee ChatGPT/Codex account → connected GitHub → pull request in openai/openai

A two-month project, under $3,000

The OpenAI intrusion was one thread in a broader project the team called “HEIF Heist,” a roughly two-month sweep of image-processing infrastructure across multiple major technology platforms. The whole effort consumed less than $3,000 in model tokens.

OpenAI paid Hacktron a $6,500 bounty for the account-takeover flaw on its side. It has since narrowed the permissions on community sign-in tokens and revoked the affected tokens and sessions.

The post Claude couldn’t hack OpenAI. Then Anthropic shipped Opus 5. appeared first on The New Stack.

Open-weight models now handle a majority of tokens on Vercel’s AI Gateway. But Anthropic still takes 64% of the spend.

18 septembre 2026 à 14:42
Isometric illustration of a retro-style computer monitor

The trend is clear: open-weight models are taking an increasingly large bite out of production AI usage.

On Monday, The New Stack reported that open-weight models accounted for 60% of OpenRouter’s US token consumption in August, with Chinese-developed models making up the majority of that volume. The latest data point hails from Vercel, whose AI Gateway routes tens of trillions of tokens each month across the applications running on its infrastructure.

As per Vercel’s September report, published on Thursday and covering activity through August, open-weight models handled 56% of all tokens routed through the gateway, the first time they have accounted for a majority of monthly token volume. In December 2025, their share was just 7%; by April it had reached 13%, and it rose every month thereafter. Vercel’s previous report, published in August, put July’s open-weight share at 36%.

So the pattern was already clear. But last month, Vercel CEO Guillermo Rauch took to social media to declare that August 22 had been a “record day for open weight share of tokens on Vercel AI Gateway,” accounting for 62% of traffic.

Rauch saw the milestone as just an early indication of where usage is heading, with enterprises still early on the adoption front.

“This is very likely just the start, because enterprise adoption is still early.”

“This is very likely just the start, because enterprise adoption is still early, and harnesses, CLIs, IDEs, SDKs, etc need to be adapted to be model agnostic,” Rauch wrote at the time.

Open-weight token share on AI Gateway: December '25 to August '26
Open-weight token share on AI Gateway: December ’25 to August ’26 (Credit: Vercel)

Tokens and dollars: Anthropic dominates spend

For context, Vercel launched AI Gateway last year as a way for developers to access models from multiple providers through a single interface, saving them from having to manage separate API keys, accounts and rate limits. The service sits between applications and the underlying model providers, routing requests while tracking usage and costs — giving Vercel a useful vantage point into which models its customers are actually running in production.

Token volume, in this context, is essentially a measure of how much model inference is flowing through the gateway. Vercel counts input and output tokens, along with reasoning, cached-input and cache-creation tokens.

While it’s a good proxy for the amount of work being handed to different models, it shouldn’t be confused with the amount of dollars being spent. Open-weight models from the likes of DeepSeek, Moonshot AI and Z.ai are generally cheaper to run than the proprietary models offered by US frontier labs — and so handling 56% of Vercel’s token volume doesn’t mean open-weight models are taking 56% of the money passing through its gateway.

Indeed, Vercel’s data shows that open-weight models accounted for just 14 cents of every estimated dollar spent through AI Gateway in August, despite processing 56% of its tokens. Their share of spending remains far behind their share of usage, although Vercel says the open-weight share of gateway spending is on the rise.

Open-weight share of tokens vs spend on AI Gateway
Open-weight share of tokens vs spend on AI Gateway (Credit: Vercel)

Across Vercel’s AI Gateway, the average price per token fell 23.2% in August, marking a third consecutive monthly decline. Among teams that processed more than 10 million tokens in both July and August, the median cost per token fell 7.6%.

Anthropic, meanwhile, has remained remarkably consistent at the spendy end of the market. Its models accounted for 64 cents of every dollar spent through the gateway in August. Vercel says the Claude-creator’s share has never fallen below 61% in any month since December 2025, with its models occupying the top two positions by spend throughout that period — often taking third spot, too.

Top 3 models by spend, by lab.
Top 3 models by spend, by lab. (Credit: Vercel)

Loyalty lies in the model

There has been plenty of movement within that Anthropic share, however. Fable 5 fell from 13.2% of total gateway spend in July to 4.9% in August, while the cheaper Opus 5 climbed to 22.5%. More broadly, Vercel’s data suggests that 90% of teams using Fable reduced their usage, with more moving those workloads to Opus 5 than to any other model.

Opus ultimately gained almost twice as much usage as Fable lost, which Vercel attributes to the newer model handling similar workloads at roughly half the price. Or, in other words, Anthropic kept the dollars even as customers shifted toward a cheaper model within its own lineup.

“Lab loyalty doesn’t follow brand, it follows model profile, and consistency wins.”

“Lab loyalty doesn’t follow brand, it follows model profile, and consistency wins,” Vercel’s report authors note.

Anthropic's share of spend by model
Anthropic’s share of spend by model (Credit: Vercel)

This trend was evidenced elsewhere, too. Within five days of Z.ai launching GLM-5.3-Flash, the new model was processing three times the daily volume of GLM-5.2.

But Vercel’s data also suggests customers are more than prepared to cross lab boundaries when a replacement fails to meet the same needs on capability and price: more than three-quarters of the volume lost by Google’s Gemini 3 Flash moved to models from other providers, including OpenAI and Anthropic. And the consequence for Google wasn’t insignificant: its overall share of token volume on the gateway fell from 30% to 5%, with the decline in Gemini 3 Flash alone accounting for 22 of those 25 percentage points.

“When a new model preserves what users valued in its predecessor, the lab retains its customers,” the authors note. “When it doesn’t, those customers fill the need through other providers.”

The post Open-weight models now handle a majority of tokens on Vercel’s AI Gateway. But Anthropic still takes 64% of the spend. appeared first on The New Stack.

GitHub and Anthropic used their own agents for major Rust rewrites — but with very different playbooks

17 septembre 2026 à 20:06
Multiple robotic arms grapple with a purple typewriter, feeding out a long sheet of paper covered in redacted black bars, depicting the concept of rewriting a codebase.

Rust is seemingly the language of the moment, with open-source projects and companies forming an orderly queue to move core software over to the general-purpose programming lingo. The draw? A pursuit of better memory safety and performance.

Now, GitHub has joined the throng. On Wednesday, the company revealed that it has completely rewritten the GitHub Copilot agent runtime — previously built in TypeScript on Node.js and V8 — into more than 800,000 lines of production Rust.

Most notably, however, the Microsoft subsidiary says it used its own Copilot coding agents to carry out the switch. In a blog post marking the migration, Microsoft engineer Stephen Toub says the company used the GitHub Copilot app and the Copilot CLI for the rewrite. The work was spread across 128 pull requests that were merged and shipped incrementally, an approach Toub notes allowed the team to catch and fix regressions along the way.

“A project that would have taken a whole team of developers a year or two before agents was now completed primarily by a single developer, in only a few months, all while the rest of the team continued to greatly expand the runtime’s capabilities and reach.”

Microsoft engineer, Stephen Toub

While a major Rust migration is a big deal in its own right, the bigger story here is the leading role that AI played.

“AI agents wrote most of the code,” Toub writes, adding that the alternative would have been a much bigger undertaking.

“A project that would have taken a whole team of developers a year or two before agents was now completed primarily by a single developer, in only a few months, all while the rest of the team continued to greatly expand the runtime’s capabilities and reach,” he continues.

GitHub, for what it’s worth, is far from alone in its efforts.

Bun rusts up

Bun, the JavaScript runtime Anthropic acquired last December, announced back in July that it was rewriting more than half a million lines of Zig in Rust, with founder Jarred Sumner using a pre-release Claude model to do much of the work.

The impetus, ultimately, was stability. Bun had grown from a 2021 “pre-LLM” project Sumner built “in a cramped Oakland apartment,” into a runtime whose CLI now sees more than 22 million monthly downloads. But its growing remit had also brought a persistent crop of memory-management bugs, including leaks and crashes. Sumner was careful not to pin those problems on Zig itself, instead pointing to the particular difficulties Bun faced managing memory across Zig and the JavaScript engine it embeds.

“We could have kept fixing these kinds of bugs one-off in perpetuity, but we owe it to our users counting on us to do better than that, and systematically prevent these kinds of bugs from recurring.”

“We could have kept fixing these kinds of bugs one-off in perpetuity, but we owe it to our users counting on us to do better than that, and systematically prevent these kinds of bugs from recurring,” Sumner wrote at the time.

Rust offered stronger protections against many of those issues, but switching languages presented a problem of its own. Bun comprised 535,496 lines of Zig, and Sumner reckoned a conventional rewrite would require a small engineering team for about a year — an investment he said would have made the project unrealistic.

“It would mean freezing bugfixes, security fixes or feature development for that time,” Sumner wrote.

And so Sumner orchestrated multiple Claude Code agents to translate and check different parts of Bun simultaneously, while he monitored their work and intervened when the process went awry. Eleven days after starting, he said the Rust version was passing Bun’s existing tests on all six platforms it supports. The code was then merged, though further review and cleanup continued before release.

Same idea, different playbook

GitHub makes much the same economic case as Anthropic. Toub notes that a “rewrite this size wasn’t affordable before agents,” estimating that the job would previously have occupied a team for a year or two. But while the two companies arrive at a similar conclusion about what coding agents make feasible, their playbooks are very different.

“A rewrite this size wasn’t affordable before agents.”

Bun went for an all-at-once port, using large numbers of Claude instances in parallel before bringing the Rust version back into the main codebase. GitHub, on the other hand, chose a slower, incremental route: pieces of the Copilot runtime were converted and shipped as they were completed, across 128 pull requests, while development on the product continued around the migration. The process ran for roughly 14-and-a-half weeks.

“Eleven days” versus “14-and-a-half weeks” shouldn’t be read as a Claude-versus-Copilot benchmark. Bun was translating a Zig codebase into another systems language through a highly parallel effort; GitHub was moving a TypeScript and Node.js runtime to Rust while continuing to ship the product throughout. What the two projects have in common is the claim that agents made major rewrites viable, where the cost and disruption might previously have ruled them out.

And, of course, both companies have every reason to make that case. Anthropic and GitHub are major vendors of the very coding agents they’re crediting with making these projects possible, making their accounts inherently self-interested.

However, there are signs across the industry that the “agents do big Rust migration” phenomenon is gaining steam.

AI bites the rust

Back in March, Meta engineer Joe Savona led a port of the React Compiler from TypeScript to Rust, noting that this was “majority coded by AI,” though its architecture, testing and migration strategy remained heavily human-directed (by Savona himself). The experiment builds on a broader effort at Meta to replace “decades-old” legacy code with Rust in a core messaging library.

And just last Friday, OpenAI quietly disclosed another substantial agent-assisted Rust rewrite. The company said that two engineers, working with Codex and GPT-5.5, rewrote Habitat — the storage service underpinning products including ChatGPT — from Python into Rust during the second quarter of 2026.

While the company noted that it will reveal further details of this move in the future, it stated that the new Rust service is already “handling 95% of our production requests,” while using six times less CPU and 15 times less memory than its Python predecessor. Moreover, it plans to deprecate the Python version entirely “in the coming weeks.”

For now, these various projects remain individual case studies, reported by the companies with the most to gain from demonstrating what coding agents can do on major software rewrites. But there is a clear pattern: Rust was already attracting companies looking to rethink key software, and AI agents are now positioned as pivotal for migrations that may once have been too costly, disruptive or time-consuming to attempt.

The post GitHub and Anthropic used their own agents for major Rust rewrites — but with very different playbooks appeared first on The New Stack.

Perplexity’s AI agents helped build a database. They weren’t allowed to run it.

16 septembre 2026 à 23:51
Abstract glitch wave

Perplexity decided it was paying too much for DynamoDB and wasn’t getting the control it wanted over read performance. So it built its own database: CobbleDB.

Built by two engineers in two months with help from hundreds of persistent coding agents throughout development, CobbleDB is a roughly 40,000-line Rust key-value store that now handles part of Perplexity’s production search traffic. The company measured median batch-read latency at 5.6 milliseconds after the move, compared with 31.4 ms on DynamoDB before the cutover, while p99 went from 123 ms to 24.2 ms.

It’s expected to cost at least 20% less than DynamoDB and plans to open-source the database at some point.

But the database itself is only part of the story. CMU professor Andy Pavlo argued at Percona Live earlier this year that databases are the hardest and most important challenge for AI agents, in part because mistakes involving production data can be difficult or impossible to reverse.

Perplexity went ahead and used hundreds of agents to help build one anyway, but they weren’t given the keys to production.

It’s expected to cost at least 20% less than DynamoDB and plans to open-source the database at some point.

Why DynamoDB couldn’t keep up

Each search requires the serving layer to retrieve pre-chunked passages and vector embeddings, with a single Search API call fetching 100 to 120 page keys in batches of 10 to 20. Each item averages about 50 KB.

DynamoDB gave Perplexity little control over how it handled reads, which meant a slow replica could hold up the entire things. It also charged for the steady flow of large reads and writes generated by search, crawling, and reprocessing, which made cloud costs difficult to justify as traffic and the corpus grew.

That led Perplexity to separate long-term document storage from the database serving live searches.

Three tiers for search data

The storage stack is split into three pieces. Pillar keeps durable document state in YTsaurus on HDDs, including versioned metadata, chunks and embeddings, while Lorry packages updates into partition-specific batches and moves them through S3 to CobbleDB.

Processed page data is spread across three replicas per partition, with hashed URLs as keys and RocksDB keeping often accessed data in memory while the rest stays on local NVMe. Reads stay within the same availability zone when possible, and the router can try another replica if one is slow rather than hold up the batch.

Updates come through S3 and are applied independently, allowing a replica to fall behind and catch up without blocking the others.

Roughly 5X Lower Batch-Read Latency

Perplexity was handling approximately 200,000 requests per second when it measured CobbleDB at 5.6 ms for a median batch read, down from the 31.4 ms it had recorded on DynamoDB. At p99, latency went from 123 ms to 24.2 ms.

In later load testing, CobbleDB reached 500,000 requests per second before performance started to decline.

The comparison comes with an important caveat; DynamoDB and CobbleDB weren’t tested side by side against identical traffic: the DynamoDB figures were recorded before the cutover, and CobbleDB’s afterward. Perplexity separately ran synthetic benchmarks using batches of 10 to 15 keys with values ranging from 100 bytes to 100 KiB.

Its cost model puts CobbleDB at least 20% below DynamoDB across the commitment options evaluated, though that estimate doesn’t include the engineering cost of supporting the database.

In later load testing, CobbleDB reached 500,000 requests per second before performance started to decline.

Agents built it, engineers controlled it

The agents carried context across sessions, catching problems with restore assumptions and runtime configuration while working on fixes and tests. But they weren’t running the database.

The two engineers kept control of the architecture and production system, particularly important given Pavlo’s warning about putting agents near critical production data.

Ownership has long-term costs

Shipping CobbleDB in eight weeks solved Perplexity’s immediate engineering bottleneck, but maintaining a custom datastore could prove considerably harder. The latency results aren’t from a controlled side-by-side benchmark, and the projected savings don’t include the engineers needed to maintain CobbleDB and respond when something breaks.

Like Shopify and Ramp, which built custom coding agents around third-party models, Perplexity kept the cloud infrastructure but replaced a managed service with something built for its own needs. CobbleDB shows how AI-assisted development is changing that calculation, making custom infrastructure more practical for smaller engineering teams.

CobbleDB shows how AI-assisted development is changing that calculation, making custom infrastructure more practical for smaller engineering teams.

The post Perplexity’s AI agents helped build a database. They weren’t allowed to run it. appeared first on The New Stack.

“Everyone’s in a race to replace GitHub”: Zed launches Delta because agents made pull requests obsolete

16 septembre 2026 à 21:45
A cardboard robot sat atop a laptop keyboard.

Something of a consensus has emerged from the developer fraternity in 2026 — GitHub, a platform built substantively for human developers, is no longer fit for purpose.

Part of the problem is sheer volume. Agents can generate, revise and submit code at a cadence GitHub wasn’t designed for, placing growing pressure on infrastructure built for humans working through commits, branches and pull requests. But there’s also a more fundamental question about the interaction model itself: when much of the reasoning behind a change happens inside a conversation with an agent, a pull request presents reviewers with the resulting diff while leaving much of the journey that produced it elsewhere.

At the heart of all of this, of course, is GitHub’s reliability problems. The platform logged hundreds of incidents over the 12 months leading into June, as monthly commit volume rocketed from around 1 billion across the whole of 2025 to 1.4 billion a month by April. By August, GitHub said this figure had jumped to 2.9 billion commits each month.

The growing load has manifested in some fairly spectacular outages, including a near-eight-hour disruption in August, with web and API error rates reaching around 20% at the height of the incident.

As Nathan Sobo, co-founder and CEO of developer platform company Zed, puts it in a blog post published on Wednesday, “everyone is in a race to replace GitHub right now,” with a number of players in the technology sphere working on alternative tooling. That includes Zed itself, which has announced the public beta of Delta — a collaborative environment where developers and coding agents work, review and revise code together in shared threads rather than pull requests.

At the heart of Zed’s pitch is the idea that the pull request is ill-suited to coding agents, and tells reviewers little about the reasoning behind their decisions.

“Since GitHub introduced pull requests over 15 years ago, they’ve become the standard way to ask teammates to review changes to your codebase.”

“Since GitHub introduced pull requests over 15 years ago, they’ve become the standard way to ask teammates to review changes to your codebase,” Sobo writes. “But with agents generating so much code, the diffs we’re asking each other to review have mushroomed.”

The question, then, is what collaboration should look like when agents are producing more of the code?

From A(tom) to Zed

Zed, for the uninitiated, started out with a somewhat narrower remit. Founded in 2021 by veterans of GitHub’s Atom editor team — Sobo himself spent nine years there — Zed emerged as a high-performance, multiplayer code editor built in Rust. When The New Stack tested the beta in 2023, the emphasis was on responsiveness and real-time collaboration; by 2025, AI editing and agentic features had become central to the product.

It has been clear for some time, however, that Zed’s ambitions extend beyond the editor. When the company announced a $32 million round of funding led by Sequoia Capital in August 2025, it also teased DeltaDB, a new kind of operation-based version control system designed to record code changes at edit-level granularity.

Fast-forward to August, and Zed revealed Delta itself in private beta, pitching it as a multiplayer environment where developers can code with agents, share their ongoing threads with teammates and review changes with the original agent context intact.

Multiple participants iterating on a prompt.
Multiple participants iterating on a prompt.

Wednesday’s public beta launch brings that idea out into the open, and takes direct aim at one of GitHub’s defining features: the pull request.

Picking up the thread

The central concept behind Delta is the thread: a running record of an agent-assisted coding task in which the conversation and the files being changed remain connected. A developer can hand an agent a job, continue discussing and refining it, and later share that entire body of work with somebody else.

Each thread can work against its own copy of a project, which means multiple pieces of work can proceed independently without every agent touching the same checked-out files. Teammates can join an existing thread or create a separate review thread to examine a proposed change, question the agent that produced it and try revisions before feeding accepted changes back into the original work.

DeltaDB sits underneath that model, recording activity at a finer level than Git. Instead of waiting for a developer to package work into a commit, it captures individual events as they happen — including code edits and activity within the conversation — and uses that history to keep participants synchronized.

Zed calls those individual records “deltas.”

Git hasn’t gone the way of the dodo quite yet, though. Delta currently works with Git repositories, and developers can continue using branches, commits and remotes as usual. DeltaDB effectively adds another layer of history between commits, preserving the intermediate human and agent activity that Git would otherwise discard.

“We now build and collaborate on Delta entirely within Delta.”

Sobo notes that Zed has already disabled pull requests on Delta’s own repository and now develops the product through Delta threads instead. Since doing so, the company says 33 developers have landed 570 changes to main without using pull requests.

“We now build and collaborate on Delta entirely within Delta,” he writes.

Zed disables PRs on its own Delta repo
Zed disables PRs on its own Delta repo

An intermediate step

It’s worth noting that this is still very much an intermediate step. And for Zed’s own internal development, Sobo expects that intermediary period to be fairly brief. He says the company is only “a few months away” from leaving GitHub behind, with developers focused exclusively on Delta already having little reason to visit GitHub because their conversations, reviews and handoffs now take place inside Delta.

There are still some pieces to disentangle. Zed’s next major dependency is Git storage, which it intends to bring into DeltaDB, while CI and releases also need to move away from GitHub. Sobo sees CI as largely a solved problem, however, and says Zed expects to integrate with existing options rather than build another system simply for the sake of replacing GitHub.

Zed’s main open-source code editor repository will remain in place for now, where an established contributor community already reports problems and proposes code changes. Those contributors can use Delta to expose the agent session behind their work, while still submitting the final change through a conventional pull request, leaving the Git experience unchanged for collaborators who don’t use Delta.

That public repository presents a different problem from Zed’s internal development, given the hundreds of external developers contributing to the project each month.

“We’re moving more thoughtfully with Zed’s public repo because we have hundreds of monthly contributors who depend on that workflow,” Sobo tells The New Stack. “GitHub has an established social component that will take longer to replace, and we’re not going to strand contributors to prove a point. We’re only going to move our community layer when we can offer something better.”

There are other signs of that continued dependency. One item currently on Delta’s roadmap is repository-based access, which will use a GitHub repository’s existing permissions to determine who can access shared Delta threads.

Delta is available as a desktop app for macOS, Linux and Windows, with a browser version for viewing, sharing and reviewing threads. It will remain free throughout the public beta, with paid individual and team plans to follow; Zed says there will always be a free version.

A common thread: Reinventing code collaboration

Of course, Zed is far from alone in tyring to reinvent code collaboration for the agentic era. SpaceX-owned Cursor formally launched Origin in August, bringing Git repository hosting, pull requests and its coding agents under the same roof, while still allowing existing GitHub repositories to remain the source of truth.

GitLab, too, is working on a “next-generation source-code management” project dubbed Project Switch, currently in private beta.

“I believe the thread will replace the commit or branch as the fundamental unit of software development.”

For Sobo, Delta’s claim to differentiation starts with the basic unit around which it’s being built.

“I believe the thread will replace the commit or branch as the fundamental unit of software development,” Sobo explains.

Commits, in his view, will continue to provide useful checkpoints. What they don’t preserve is everything that happens between them: the discussion with an agent, the decisions made along the way and the incremental changes that eventually produce the committed code. That’s the gap Zed wants DeltaDB to fill, with the thread rather than the commit serving as the fuller record of how a piece of software came together.

“Delta’s advantage is that we’re building around the thread from the start, including the infrastructure underneath it,” Sobo continues.

Sobo argues that preserving incremental edits alongside the agent conversation gives subsequent collaborators something a conventional diff cannot: the ability to enter the work where the previous developer left it, and continue interacting with the same agent and context, rather than reconstructing the thinking behind a change after the fact.

Delta’s bet is that the central object of software development should become the ongoing interaction between developer and agent, with the resulting code attached to that history rather than presented later as an isolated diff.

“I expect lots of viable products will compete on the agent experience, with a common infrastructure underneath.”

Sobo expects the next era to echo the structure of the Git era, with competing developer platforms built on top of a common technical foundation.

“On consolidation, we look at how the last era played out. Git became the common foundation because it was open and everyone could build on it,” Sobo says. “GitHub won the layer above through network effects. I expect lots of viable products will compete on the agent experience, with a common infrastructure underneath.”

The post “Everyone’s in a race to replace GitHub”: Zed launches Delta because agents made pull requests obsolete appeared first on The New Stack.

AI evaluator: The most important AI job in history? How developers might fill the proposed new job

16 septembre 2026 à 17:58
Lots of pink escape keys

The pace of frontier AI model development spurred Anthropic CEO Dario Amodei to publish an essay last weekend, calling for changes in how the industry is regulated and develops. In a three-part plan that includes both democratic and global coordination, Amodei writes that the first step was something Anthropic is committing to unilaterally.

“Each frontier AI company [should] commit to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes,” writes Amodei.

Amodei’s essay followed dire warnings from former Anthropic and OpenAI pretraining research specialist Jacob Coxon, who posted a thread on X saying the people building AI earnestly “believe that it could kill us all” by the end of the decade.

Shortly after Amodei published his essay, OpenAI CEO Sam Altman and SpaceXAI founder Elon Musk chimed in: “I agree with Dario,” posted Altman; “Dario is right,” posted Musk. Later that day, Demis Hassabis, founder of Google DeepMind, posted, “Dario’s essay points towards the right path forward.” In a post on X, Meta CEO Mark Zuckerberg writes that Meta Superintelligence Labs already uses independent evaluators, and that, “In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators.”

This week, theories began to surface about why the world’s biggest frontier AI labs would want to intentionally slow their pace when competition is so fierce. “The desire to slow down is puzzling, but perhaps if the whole system slows down, the rules of winning can be the same for all,” posted Nikesh Arora, chairman and CEO of Palo Alto Networks.

In his essay, Amodei likens the proposed job of independent AI evaluator to embedded regulatory supervisors used in the banking industry, i.e., third-party professionals. Altman describes the job as having “employee-like access” in his post on X.

So, who could fill these roles that AI leaders agree are desperately needed?

Salaries top out at $687K; are you interested?

METR’s current job openings are well paid (salaries top out at around $687,000), and the job specs are daunting. 

“You’re scrappy, creative, independent, and self-directed (because during the exercises you’ll only have a few other METR employees you can talk to). The work is novel, and you’ll need to figure a lot of stuff out on the fly largely by yourself. You are excellent at loss-of-control threat modeling and breaking down safety cases,” reads the spec.

“You’re scrappy, creative, independent, and self-directed. The work is novel, and you’ll need to figure a lot of stuff out on the fly largely by yourself. You are excellent at loss-of-control threat modeling and breaking down safety cases.”

Similar but less colorfully illustrated roles (paid between $180K–$300K) are also available at AI model training company Mercor, where candidates will need a Ph.D. or M.S. and more than two years of work experience in a computer science, electrical engineering, econometrics, or another STEM field that provides a solid understanding of machine learning and model evaluation.

“Employee-like access fluctuates wildly”

AI security consultant and CTO at Komodo, Kadan Stadelmann, tells The New Stack that his typical week sees him work differently with each client. This is because “employee-like access fluctuates wildly”, from rigid focus areas to broad access, and much of that aspect is determined by contracts signed before work begins.

“I am invited into labs to probe numerous risk vectors, including agent behaviors under realistic conditions,” Stadelmann says. “Among my duties are tasks that include monitoring chains-of-thought and prompts. The goal is to establish an objective and look at a specific AI system to question how autonomous the system is, and how long it takes to complete specific tasks. Most importantly, evaluators at my level monitor for how well a team adheres to its claimed safety practices.” 

Software engineering skills beat doctorates

Although METR wants evaluators to have a Ph.D. up their sleeve, Stadelmann says that as the prevalence of this role expands, he feels the technology industry has been, and continues to be, driven by people who can demonstrate strong engineering skills, not doctorates.

“I am invited into labs to probe numerous risk vectors, including agent behaviors under realistic conditions. Among my duties are tasks that include monitoring chains-of-thought and prompts.”

Questioned on whether costs create a barrier for smaller labs, Stadelmann notes that some AI evaluation work is funded by third-party non-profits, which protects independence. 

“But overall, evaluators will not be able to keep up with big frontier model firms. They will be out of control, and we will be dependent upon their own internal ethics. Plus, anyway, many of the smaller labs of any worth may inevitably be acquired by the big players in this space,” he adds.

What happens when an evaluator finds something wrong?

Founder and CEO of facial image AI identity governance company Indie Me, Dion Johnson, tells The New Stack that what interests him most about embedded AI evaluators isn’t the job title; it’s what happens when their judgment uncovers that the model behaved in a way nobody expected. 

“If the evaluator can only raise concerns when those concerns are convenient, then we have not created independent oversight — we have created another layer of process,” Johnson says. “The evaluator needs enough access to see the uncomfortable things, not just the polished demonstrations. They need to understand what happened during training, what failed during testing, what behaviors appeared unexpectedly, and where the team itself still has uncertainty.”

On the required skills AI evaluators need, Johnson agrees that technical depth, machine learning environment security experience, and software engineering as a whole matter.

“To choose a competent AI evaluator, I would look for someone who is deeply comfortable with uncertainty and deeply uncomfortable pretending they know something they do not. This is someone who can sit in a room full of brilliant people at an AI model company on launch day and say ‘I’m not convinced’… and that takes judgment and courage,” he adds.

“I would look for someone who is deeply comfortable with uncertainty and deeply uncomfortable pretending they know something they do not.”

Fear and loathing in the AI space

In his September 8 post, which has been viewed 172 million times and seemingly spurred AI leaders to change course, Coxon, the former AI researcher at OpenAI and later Anthropic, writes: “This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible but I hear the same people express fear privately. No other human activity poses this level of danger.”

As for where Coxon looks for work next, perhaps it might be a role in AI evaluation execution engineering.

The post AI evaluator: The most important AI job in history? How developers might fill the proposed new job appeared first on The New Stack.

Meta lets Claude and Codex configure WhatsApp Business via MCP. But the agents don’t get their own identity.

16 septembre 2026 à 00:08
A concept illustration depicting AI running a business

Any business worth its salt in 2026 needs to be embracing the right tools to reach its customers, and few tools carry as much weight as WhatsApp.

Paid messaging on the app crossed a $2 billion annual run rate in the fourth quarter of 2025, CFO Susan Li told investors on Meta’s January earnings call. But getting a business properly set up on WhatsApp can still be a fiddly job. The initial onboarding can send developers jumping between Meta’s account settings, API documentation, and code editor as they connect and verify a phone number, configure webhooks, and get the integration working. Some of that is a one-off job, but things like managing message templates, testing changes, and troubleshooting the setup can bring developers back to those same tools later.

Meta’s answer, announced today, is to let AI coding agents handle much of that setup directly through MCP.

Connecting coding agents with WhatsApp

The easiest way to understand what the WhatsApp Business Tools MCP server is all about, is to look at the sort of job a developer might want to hand over to an agent.

Take an online retailer that wants to use WhatsApp for customer support, order updates or the occasional special offer. A developer can connect the new MCP server to an agent such as Claude or Codex, sign in with their Meta account, and choose which of the businesses they already administer the agent is allowed to access. That access is scoped to whatever businesses they select — connecting the agent doesn’t give it free rein across every Meta account associated with the developer.

Connecting Claude to WhatsApp
Connecting Claude to WhatsApp

Give the agent the number and the display name the business wants to use, and it can handle the steps needed to add the number and set up the WhatsApp account behind it. Meta then sends a one-time verification code by SMS or voice; once the developer gives that code back to the agent, the number can be verified and registered for sending messages.

Adding a WhatsApp number
Adding a WhatsApp number

In a blog post announcing the feature on Tuesday, Zoë Lieberman, who works on product marketing at Meta, says the idea, ultimately, is to turn what might otherwise be a string of separate API tasks into something the user can ask an agent to do in plain English.

“You describe what you want — your agent handles the accounts, numbers, templates, and API calls.”

“You describe what you want — your agent handles the accounts, numbers, templates, and API calls,” Lieberman writes.

A template example

Much of the WhatsApp Business Tools MCP server is concerned with getting a business up and running on WhatsApp in the first place. Message templates, however, show how the agent can remain useful once that initial setup is done.

There is a WhatsApp rule worth explaining, though. When a customer messages a business, it opens a 24-hour customer service window, during which the business can reply with ordinary, free-form messages. Each new message from the customer starts that 24-hour clock again. The idea is to stop businesses turning an old customer conversation into an open-ended channel for unsolicited messages: once the window has closed, the business generally needs to use a message template that Meta has approved if it wants to contact that customer again.

So the retailer could ask the agent to create a marketing template offering customers a coupon, for example, or a utility template for sending order updates. The template can include elements such as a header, body copy, footer and buttons.

Creating a marketing template
Creating a marketing template

From the same conversation, the person using the agent can list the business’s existing templates, pull up a particular version, update it or delete it.

Once the pieces are in place, the developer can ask the agent to send a test message and check that everything behaves as expected before putting it in front of customers. If the recipient is outside the 24-hour customer service window, the agent can flag that a free-form message can’t be sent and offer an approved template instead.

Sending a test message
Sending a test message

There is also the other half of a WhatsApp conversation to deal with: what happens when the customer replies?

The developer can ask the agent to configure the webhook that tells WhatsApp where to send those incoming messages and other events, such as the retailer’s CRM, customer-support platform, chatbot or order-management system.

Configuring a WhatsApp webhook
Configuring a WhatsApp webhook

The agent can inspect the account as well as change it. That means asking what has already been configured or what still needs attention — for example, whether the business is missing the payment information Meta needs to charge for billable WhatsApp messages.

Checking WhatsApp account setup
Checking WhatsApp account setup

There are some guardrails around all of this. The agent operates using the access of the person who connected it, and Meta says actions performed through the MCP are recorded.

“Every read runs under your own viewer context, every invocation is logged, and anything that changes state requires an authenticated person rather than an app-level credential.”

“Every read runs under your own viewer context, every invocation is logged, and anything that changes state requires an authenticated person rather than an app-level credential,” Lieberman writes.

Meta’s MCP push stays tied to human identity

That human-bound approach lands amid a broader debate over how AI agents should identify themselves. Agents today often inherit the permissions of the person using them, while some companies are pushing toward giving agents their own scoped, revocable identities — evidenced by Vercel’s recent acquisition of Better Auth. The concept is also showing up in the big cloud platforms, too, such as with Microsoft’s Entra Agent ID, which automatically assigns an identity to agents created in Azure AI Foundry or Copilot Studio. And Amazon Bedrock AgentCore Identity, meanwhile, gives each agent its own identity, letting it act on a user’s behalf or independently, without borrowing anyone’s login.

So while there is a clear push toward giving agents their own identity, Meta is keeping the human firmly in the loop here, The agent can act only within the authenticated user’s existing access, keeping permissions bounded by a known human account and making sensitive changes easier to attribute and control.

It’s also worth noting that this isn’t Meta’s first attempt to put its developer tools within reach of AI agents. In June, the company launched Developer Tools MCP, later rebranded as Meta Social Technologies MCP, which works across Meta’s developer platform — including integrations built with the WhatsApp Business API.

There is some overlap between the two, but they operate at different levels. Meta Social Technologies MCP is the broader developer tool: an agent can use it to discover Graph API endpoints, search Meta’s documentation, inspect the app behind an integration and troubleshoot errors. That can absolutely include a developer building with WhatsApp.

WhatsApp Business Tools MCP, meanwhile, is much more specific to operating WhatsApp Business itself. It gives the agent tools for working with WhatsApp Business accounts, onboarding and verifying phone numbers, creating and managing message templates, configuring and testing webhooks, and sending test messages.

“They’re complementary — install both if you need both,” Lieberman writes.

It’s early days for the new WhatsApp server. Meta says it is rolling out gradually, meaning it may not be available to everyone immediately, while the interface and tools themselves remain in beta and are subject to change.

The post Meta lets Claude and Codex configure WhatsApp Business via MCP. But the agents don’t get their own identity. appeared first on The New Stack.

“Machine translation is still broken for most of the world’s languages”: Cohere builds non-reasoning for a reason

13 septembre 2026 à 16:21
A scattered pile of overlapping alphabet cutouts in bright blue, pink, green, gold, red, and silver.

Enterprise AI company Cohere announced North Small Translate last week, a mixture-of-experts (MOE) open-weight machine translation model that works across 50 languages.

Developers can download the weights for noncommercial use under CC BY-NC 4.0. Cohere offers commercially licensed deployment through Model Vault, which is a Cohere-managed inference environment. Cohere positions the model as part of its sovereign AI strategy, aimed at organizations that want greater control over where their models run and how their data is handled.

North Small Translate builds on Cohere’s multilingual and translation lineage, which includes its Tiny Aya and Command A Translate model families. The company claims North Small Translate outperforms “similarly sized open-weight models” under 1T parameters, as well as API-based translation models in various dimensions of machine translation on average. 

Cohere co-founder Nick Frosst tells The New Stack that the model’s efficiency draws from the fact that it is non-reasoning, i.e., it relies on learned statistical patterns without a step-by-step logic process, which means it uses fewer tokens.

Machine translation is still broken for most of the world’s languages

“We spent nine years scaling an architecture invented to fix translation, and machine translation is still broken for most of the world’s languages,” Frosst says. “General-purpose models get you most of the way and then stop. The next phase of enterprise AI in this space is smaller, more specialized, and runs inside your own walls.”

“…machine translation is still broken for most of the world’s languages.”

In Cohere’s reported evaluation using WMT26 benchmarks, the company states that North Small Translate leads with a WMT26 All Languages benchmark score of 83.60, compared with 81.56 for Qwen 3.5 397B A17B, 76.50 for GLM 5.2 FP8, 81.37 for DeepL NextGen, 79.46 for Gemma 4 31B (on), and 68.20 for Google Translate. 

With its mixture-of-experts architecture and 218 billion total parameters, with 25 billion active. Cohere points to North Small Translate’s smaller compute & memory footprint than other models. Some model-to-model comparisons in this space aren’t fully substantiable, since not every vendor discloses parameter counts.

With current solutions, long documents start to fall apart

“Machine translation allows documents to be translated from one language to another automatically. With current solutions, long documents start to fall apart,” Frosst says. “Google Translate scores 21.3 on our long-context test, Gemma 4 31B 19.4; we score 48.9. That’s [for example] a safety manual that reads fine on page one… and has drifted by page ten. The other risk is where the text goes. Once you push HR policies or regulated documents through a third-party API, that data has left your building, and necessarily that means your control over it is diminished.”

“The risk [in machine translation] is where the text goes. Once you push HR policies or regulated documents through a third-party API, that data has left your building and necessarily that means your control over it is diminished.”

Explaining why the model offers “stronger translation performance” across complex enterprise translation tasks, Frosst says the model can support work spanning “a high volume” of sensitive documents. 

As well as its 50 languages (32 ‘high-resource’ languages + 18 others), the Cohere team explains that the model also supports translation-workflow-focused capabilities, such as structured translations (i.e., Markdown or JSON documents), instruction following (i.e., recommended tone & format), and terminology guides (i.e., providing specific vocabulary to use in the translation), all as part of the model.

“North Small Translate works with a multi-pass workflow,” explains Frosst. “The model translates, reviews its own output, finds errors, and fixes them – and this is the same loop we used in training. We ship both because standard is one pass and built for volume, while the agentic [version] spends more tokens for 84.36 against 83.60 on WMT26. That difference ends up being worth it when the document is a contract or a safety procedure, for instance, but in other cases you’d rather optimize for efficiency.”

“The model translates, reviews its own output, finds errors and fixes them.”

Model ‘steerability’ drives suggesting language tone and formatting

This model uses the same architecture as prior Cohere models but improves performance through post-training advances, including reinforcement learning and new datasets, specifically for machine translation tasks.

Frosst concludes that, across the translation model marketplace, generative machine translation models offer the highest quality and steerability (i.e., suggesting tone, formatting, etc.) but typically cost much more than Neural Machine Translation (NMT) models commonly used in commercial use cases. 

North Small Translate was developed in partnership with RWS, an AI solutions company pioneering in language technology and services. Collaboration with RWS, specifically with its Language Weaver research and science teams along with its language experts, helped shape the model’s real-world translation performance throughout development. 

As noted above, developers can access the weights free of charge for non-commercial use in three quantizations. There is also a Hugging Face Space and an API for those who lack the required hardware. 

The post “Machine translation is still broken for most of the world’s languages”: Cohere builds non-reasoning for a reason appeared first on The New Stack.

“Same mission, bigger stage”: OpenAI hires Git AI founders to help Codex prove its ROI

12 septembre 2026 à 16:46
Data on a computer screen

OpenAI has hired the founders of Git AI, an open-source tool that tracks how much code AI writes and measures the performance and cost of coding agents.

Taking to LinkedIn late Friday night, Aidan Cunniffe noted that he and his Git AI co-founder Sasha Varlamov will join OpenAI’s Codex team, where they will work on giving businesses better data on how coding agents perform and the returns they generate.

“Same mission, bigger stage,” Cunniffe writes.

While details of the integration plans are scant, OpenAI’s Thibault Sottiaux, a member of technical staff at OpenAI who’s working on ChatGPT and Codex, confirmed on X that the company wants to use Git AI’s technology to help businesses understand Codex’s impact.

“We’ll make it easier for businesses to see where Codex is making a difference when working through problems for individuals and teams,” Sottiaux writes.

Excited to welcome Aidan & @
Sasha from the Git AI team to OpenAI!

They are building in the open and have developed an open-source tool that helps developers understand how coding agents contribute to their codebase.

Together, we’ll make it easier for businesses to see where…

— Tibo (@thsottiaux) September 12, 2026

In a separate blog post announcing the deal on Friday, Cunniffe notes that the tie-up reflects a shared focus on measuring the performance and value of AI coding tools.

“OpenAI shares our belief that users should have the data to compare model performance and understand the ROI of every token they spend.”

“OpenAI shares our belief that users should have the data to compare model performance and understand the ROI of every token they spend,” he writes.

Tracking AI-generated code

Git AI, for the uninitiated, is an open-source Git extension that tracks AI-generated code. Every line of code is linked to the agent, model, and prompts that generated it, preserving the context behind a change even after the code has been committed, merged, or rebased.

More broadly, it’s designed to help teams see how much AI-generated code ultimately reaches production, how often it’s reworked, and where time and tokens are being spent. That data can then be used to compare coding agents and models, and assess whether the money being spent on them is producing useful output.

The tool already supports a broad range of coding agents, including OpenAI Codex, Anthropic’s Claude Code, Cursor, and Google’s Gemini CLI, as well as background agents including Codex Cloud, Claude Web, Cursor Agent, and Devin.

After installation, supported agents automatically report their edits to Git AI. For developers, the pitch is basically: “Just prompt and commit”—they can keep using their coding tools and Git as usual.

Git AI
Git AI

Git AI can then show the split between human- and AI-written code and trace individual lines back to the agent that produced them.

The story so far

Cunniffe has, in fact, been here before. He founded developer tooling startup Optic in 2018, which built open-source tools around API development, before Atlassian acquired the company in 2024. He subsequently joined Atlassian as a principal product manager leading go-to-market efforts for developer tools.

Git AI began while Cunniffe was still at Atlassian. He and Varlamov started working on the project in the summer of 2025 as a side project, initially trying to answer one seemingly simple question: how much of their code was actually being written by AI, and what happened to that code afterward?

By the first major release in November, Cunniffe was publicly describing Git AI as a “weekend project” that had grown into something much larger. He subsequently left Atlassian in January 2026 to turn Git AI into a company, building a commercial business around the open-source project.

Now, less than a year later, Cunniffe and Varlamov are heading to OpenAI. It is not yet clear what that means for Git AI as a commercial business, but with both founders joining OpenAI, the standalone business will likely be wound down while the open-source project continues.

The New Stack has reached out to OpenAI and will update this story when we hear back.

Cunniffe, for his part, says OpenAI will continue backing the open-source project.

“We’ll keep investing in open source, while also giving enterprises the data they need to build effective software factories,” he writes.

How that cross-model independence ultimately holds up under OpenAI stewardship will be one of the more interesting things to watch.

The post “Same mission, bigger stage”: OpenAI hires Git AI founders to help Codex prove its ROI appeared first on The New Stack.

AWS open-sources Pizza Bot: email-style inbox for background AI agents

11 septembre 2026 à 00:54
An illustration of a robot delivering a pizza

Amazon Web Services (AWS) has released a new open-source application dubbed Pizza Bot, which gives developers an email-style inbox for managing AI agents that run in the background.

The problem that Pizza Bot is designed to address, ultimately, is that a chat interface is a poor fit for agents whose work continues after the (human) user has clocked off for lunch or bed.

And so Pizza Bot leans on the age-old mechanics of email: completed jobs arrive as unread threads, while anything that needs a human decision is surfaced for action. An Activity panel also exposes jobs handed off to specialist agents, including their tool use and progress.

“Scheduled agents run autonomously in the background and surface updates directly into your inbox for review and triage.”

Pizza Bot
Pizza Bot (Credit: AWS)

Writing about the new project in a LinkedIn post on Thursday, co-creator Joseph Dolivo, principal technologist at AWS Startups, notes that Pizza Bot shifts the burden of monitoring agent work firmly away from the user.

“Instead of you having to initiate every conversation or wait on a prompt, scheduled agents run autonomously in the background and surface updates directly into your inbox for review and triage,” he writes.

As if to emphasize that point, in the accompanying blog post for the project’s official launch on Thursday, the creators note that user absence is, in fact, a core tenet of the design brief.

“The interface assumes you are not watching.”

“The interface assumes you are not watching,” they write. “Nothing else we’ve seen starts there, and that one assumption is what buys you pauses that outlast the session that created them, notifications worth acting on, and scheduled work that produces threads instead of logs.”

A community project

Despite its Amazon roots, Pizza Bot is in fact now a standalone community project rather than an AWS service. It lives in its own GitHub organization, separate from Amazon, and comes with no AWS support or service-level agreement — it’s entirely self-hosted.

Pizza Bot itself is a desktop app for macOS, Windows, and Linux, with browser and terminal clients available too. By default, the app starts a local Pizza Bot server on the machine, while developers choose the model behind it — including Anthropic, Amazon Bedrock, Google Gemini, OpenAI, OpenRouter or a local model via Ollama.

It can also be extended through MCP servers and Agent Skills; the bundled browser-automation skill, for example, uses Playwright MCP to navigate and interact with websites.

Browser skill via Playwright MCP
Browser skill via Playwright MCP (Credit: AWS)

The server can also run on an always-on host or in a container, letting scheduled agents keep working while the laptop is closed and their threads be picked up later from another device.

Under the hood: LangGraph and ambient agents

Pizza Bot’s agent runtime is built with DeepAgents, LangChain’s open-source harness for long-running agent tasks, which itself runs on LangGraph, its runtime for stateful agent execution. The important part in all of this is persistence: LangGraph checkpoints an agent’s state as it works, allowing a run to stop for approval, survive a disconnected client and resume later without starting again from scratch. Pizza Bot stores those checkpoints, along with threads and other application data, locally in SQLite and ordinary files.

It’s worth noting that AWS has its own open source agents SDK, Strands Agents, out since May 2025 — but evidence suggests, including text in this sample repository, that AWS considers Strands as a “lighter-weight alternative to LangGraph for agents that don’t need explicit graph control flow.”

In response to a question posted on LinkedIn by The New Stack, Dolivo says that they very well could have used Strands for Pizza Bot, particularly as Strands supports TypeScript and workflows now. But they ultimately went for LangGraph “due to the maturity of the tooling and breadth of the exosystem,” he explains.

“It’s also more familiar to many developers, and we wanted to reduce friction for community adoption since we knew we’d be open-sourcing it,” he adds.

Pizza Bot is also fairly close to an idea LangChain introduced way back in January 2025, when CEO Harrison Chase introduced the term “ambient agents” for agents that could respond to events, work concurrently and involve a human only when needed. LangChain’s reference implementation was an email assistant built on LangGraph. It also developed what it called an “Agent Inbox”: a standalone interface inspired by email and customer-support software for keeping track of open interactions between people and background agents.

From ‘JoeBot’ to Pizza Bot

The genesis of Pizza Bot can be traced back to April 2025, when Dolivo kicked off what he calls a “side-of-desk passion project” dubbed “JoeBot” that automated repetitive CRM logging. He later teamed up with colleague Igor Fil to turn that script into Pizza Bot, an MCP server that could execute parameterized, deterministic “recipes” across its internal systems.

The appeal soon spread beyond the engineers building it into less-technical domains. As the project evolved, Dolivo says it would eventually grow to more than 30 contributors and over 2,000 users inside Amazon, and so Pizza Bot needed a front end that people could open and use.

“An MCP server requires an MCP client, and expecting non-technical users to work out of an IDE or terminal was never going to cut it,” Dolivo adds. “We had to meet people where they actually work and own the experience end to end.”

The result was the version of Pizza Bot released this week: a desktop app built around an inbox rather than a terminal.

The post AWS open-sources Pizza Bot: email-style inbox for background AI agents appeared first on The New Stack.

“Six tools, one harness”: Salesforce loops together a six-pack of favorites

10 septembre 2026 à 22:03

Salesforce introduced its Salesforce Enterprise AI Harness on Thursday as a formalized amalgamation of the AI harness concepts and infrastructure the company has been working to align.

The organization said that “no single system has the complete answer” to complete a straightforward business task, such as completing a customer order; i.e., CRM knows the customer, ERP knows the inventory, FSM (field service management) knows the delivery, and the support team processes… and so on. 

As such, a form of AI leakage pervades throughout modern enterprises, where individual agents and their harnesses do their best to enact automation intelligence, albeit in comparatively siloed chunks.

Harnessing a six-pack of toolsets

Salesforce’s answer is to coalesce what it calls “six trusted capabilities” (from its own platform toolset collection) alongside a new AI control plane, built to underpin an open and composable AI ecosystem.

The Salesforce Enterprise AI Harness encompasses core technologies across Data 360 (a unified customer data platform tool), Informatica (data integration and governance), MuleSoft and Agent Fabric (API connectivity and multi-agent orchestration), Tableau (visual business analytics), Agentforce (an agent platform), Salesforce Guardian (security and compliance), and the Salesforce platform itself through a common, composable architecture and unified experience.

“The Agentic Enterprise won’t be defined by which model a company chooses. Models will continue to change, and intelligence will increasingly be available everywhere. What will differentiate an enterprise is the trusted, proprietary context it brings to that intelligence — starting with the customer — and its ability to securely turn that context into action,” said Rohan Kumar, Salesforce president & chief platform and engineering officer, during press briefing.

“The Agentic Enterprise won’t be defined by which model a company chooses… what will differentiate an enterprise is the trusted, proprietary context it brings to that intelligence.”

This is not Salesforce’s first-ever harness

To be clear, it hasn’t taken Salesforce until late 2026 to ever produce or work with a harness; subsystems within the six pack, such as Agentforce Vibes (a natural language vibe coding tool), make use of specialized execution harnesses, including Mastra and the Claude Agent SDK, to manage local agent execution loops. This is — as suggested — a more formalized, total platform-wide development.

Alongside the six-way alignment spanning context, agency, action, governance, security, and models, Salesforce is offering a new AI control plane to give developers a place to view, manage, and control agents. The company confirms that software engineers can “use the six together as one system or take only what they need,” and create deployments with Salesforce technology, other third-party existing technology, or both.

The big question here is simple: Is this cosmetic packaging designed to disseminate wider Salesforce DNA into software developers’ production environments, or is it a genuinely useful simplification and unification process that will be met with interest and perhaps even gratitude?

Working engineers commenting on sites including the G2 developer forum and B2B software review portal have provided some insight.

What developers and operations professionals think of the Salesforce stack

Commenting on the use of Agentforce as a standalone tool, operations associate Ashish B. noted in August this year, “One area that could be improved is the initial setup and configuration process. Building effective agents can take some customization and a solid understanding of the workflow. The platform would be even better with simpler configuration options and clearer, more straightforward guidance on setting up agents for specific business use cases.”

Salesforce may have been listening. It said the Enterprise AI Harness connects reasoning to business rules, policies, and controls required for predictable execution. It then makes those capabilities reusable across the enterprise so that context can be shared across agents and models. This means actions and workflows can be securely invoked wherever they’re needed, and governance and security can be applied consistently as AI moves across the business.

“The platform would be even better with simpler configuration options and clearer, more straightforward guidance on setting up agents for specific business use cases.”

Writing about the Informatica user experience on Gartner Peer Insights in March of this year, a DevOps engineer said that the platform “works well” for integrating multiple data sources and supports both batch and real-time processing. But they caution, “Debugging and monitoring pipelines can be difficult in complex workflows. Initial setup is challenging for new users and requires some learning curve. [The] User Interface could be improved for better usability and faster navigation.”

Possibly taking into account such feedback, the new AI control plane that accompanies Enterprise AI Harness claims to give businesses a common place to see, manage, and control agents and AI across the enterprise.

“It enables companies to discover and register agents and AI capabilities, establish identity and policy, manage lifecycle, evaluate performance, observe behavior and outcomes, and control cost — across Salesforce and third-party AI. This gives enterprises a consistent layer of visibility and control as AI expands across teams, applications, models, and systems — without requiring every agent or AI experience to be managed separately,” pledged Salesforce.

Six pillars of trust

The whole premise of this Enterprise AI Harness hinges around what Salesforce calls six trusted capabilities. 

Trusted Context combines customer context with data, metadata, semantics, knowledge, real-time signals, memory, and an understanding of how work gets done across the enterprise. Trusted Agency gives agents reasoning, planning, state, memory, and orchestration functions, combining flexible AI reasoning and deterministic controls where certainty is required. Trusted Action securely connects AI to applications, APIs, workflows, tools, and business processes. 

As its name suggests, Trusted Governance governs the data, metadata, policies, and processes that AI relies on, with lineage, quality, guardrails, and controls. Trusted Security applies identity, permissions, privacy, data protection, and runtime security to what AI can access and what agents can do. Trusted Models provides security with intelligent model routing based on accuracy, performance, cost, and requirements. 

Integrations with Claude, Slack, Teams, etc.

The Enterprise AI Harness is being built headlessly from the ground up, with capabilities accessible through technologies including MCP, APIs, skills, and plug-ins. The company said this will let Salesforce capabilities extend beyond traditional Salesforce applications, into services such as Claude, Slack, and Microsoft Teams.

Many of the technologies that form the foundation of Salesforce’s Trusted Enterprise AI Harness are available today, with new capabilities and the unified experience planned to begin rolling out in early fiscal year 2028.

The post “Six tools, one harness”: Salesforce loops together a six-pack of favorites appeared first on The New Stack.

❌