❌

Vue lecture

Performance engineering from kernel analysis to AI: Adrian Cockcroft’s take

Abstract dark digital landscape with glowing contour lines representing multidimensional performance data and response time distributions.

Over its five-year history, P99 CONF has hosted quite a few speakers who’ve offered pointed takedowns of the namesake metric. At last year’s conference, Adrian Cockcroft didn’t explicitly state that P99s are BS… but he did allude to it.   

If you don’t know Cockcroft, he’s spent decades architecting, scaling, and optimizing resilient, high-performance systems at giants like Sun Microsystems, Netflix, eBay, and Amazon. We could probably dedicate an entire day of P99 CONF to discussing the lessons learned from just some of his projects (Solaris kernel performance, multi-processor optimization, Netflix’s on-prem to cloud migration, Chaos Monkey…) 

Fortunately, RedMonk analyst Rachel Stephens proved the perfect host for a conference that’s all about making things fast. She sat down with Adrian and led us on a whirlwind tour of how AI has impacted performance engineering. Here are some highlights from the chat (full video below).

Note: P99 CONF 2026 – a free + virtual conference on all things performance – is going live October 21-22. Grab a complimentary pass and join us!

From kernel analysis to vibe coding perf tools 

As a performance specialist at Sun in its heyday, getting to the root of performance problems involved lots of digging and divination. Cockcroft recalls, “Back in the old days with Sun, people would look at the output of system metrics in vmstat or whatever, and they’d be guessing what the numbers meant. There was a very vague understanding of what these things meant. The manual page wasn’t very clear.” 

Cockcroft ended up going to the source, literally. “I went and read all the kernel source code and figured out exactly where these numbers came from, exactly what they meant, which ones were approximating what, and wrote all that down.” That led to two performance books: Sun Performance and Tuning and Resource Management.

“My speedup is infinite, because this code would never exist without these tools. I wouldn’t have the time to build them.”

Four decades later, there’s now a wealth of helpful tools for end-to-end tracing, but Cockcroft’s curiosity still lies in what the tools are not showing. He continued, “Everything sort of looks okay in the tools – but the system isn’t behaving well. I usually come in and try to find a new way of looking at the data. A new type of analysis, or go a little bit deeper or finer grain, or stop looking at averages and start looking at distributions, and find all kinds of interesting things that nobody knew were happening.”

Currently, he’s vibe coding tools to better analyze the anomalies he finds. Saved from having to brush up on Python or hunt down graphics library fragments on Stack Overflow, Cockcroft can now stand up custom tooling in minutes. “My speedup is infinite, because this code would never exist without these tools. I wouldn’t have the time to build them.”

Peaks not percentiles

One specific vibe coding project: Cockcroft built (and open-sourced) tooling to get a better understanding of response time distributions. 

Response time distributions have been on Cockcroft’s mind for over a decade. While most people obsess over percentiles – yes, P99 CONF included – Cockcroft is most intrigued by the distribution of response time peaks in a histogram. He believes percentiles don’t work when trying to understand the latency and performance of modern web services. A single number like P99 can’t tell you whether the underlying distribution has one peak or several. And when there’s more than one peak (as is often the case in the real world), the mean, the standard deviation, and even the P99 itself lose most of their meaning.

“Percentiles don’t work when trying to understand the latency and performance of modern web services.”

Image showing what people think response time distributions looks like vs what they really look like
(source: A Tale of Two Histograms)

For example, assume you have a histogram with two response time peaks: a fast one from a cache hit and a slow one for misses that require actual work. As the cache hit rate shifts, each peak’s position remains the same (i.e., the latency values of the fast-response mode and the slow-response mode don’t change), but the peak heights rise and fall. “Your averages and your P99 are changing all over the place, but all that’s really happening is your cache hit rate is changing,” Cockcroft said.

So how do you go beyond measuring P99s and averages? Cockcroft did what he’s done for decades: dive in and build a custom tool. But these days, it’s much simpler thanks to LLMs.

“Your averages and your P99 are changing all over the place, but all that’s really happening is your cache hit rate is changing.”

He had already worked out the statistical approach for analyzing the distribution. Once ChatGPT came out, he quickly used it to build a tool that automated it. Instead of collapsing everything into an average, it identifies an arbitrary number of peaks in a distribution and tracks how they fluctuate over time. It’s implemented in R – a language Cockcroft hadn’t used in a while, but ChatGPT knew quite well – and it’s open source. If you’re curious, learn more in his Percentiles Don’t Work article and “A Tale of Two Histograms” talk and deck (“It was the best of response times, it was the worst of response times…”)

Where do we go from here?

To close, Stephens asked Cockcroft what advice he’d share with teams working on high-performance systems today. His top tip was to start with the macro view to find what’s interesting, then keep digging deeper until you’re inspecting individual slow requests end-to-end.

“Remember the microscope that you got when you were a kid,” Cockcroft said. “First, you have to focus it using the lowest resolution, at 10x, and then you can click it to 100x and adjust that, looking at just one speck now. Once you get that in focus, you click it to 1,000x.”

Cockcroft has spent his career building tools that bring obscure performance issues into focus. We look forward to seeing what others have cooked up with agentic tooling to help identify and solve performance problems this year at P99 CONF. 

Learn about the latest performance optimization techniques, tooling, and case studies at P99 CONF – free and virtual, October 21-22. Grab a complimentary pass and join us!

The post Performance engineering from kernel analysis to AI: Adrian Cockcroft’s take appeared first on The New Stack.

  •  

The agent didn’t break your controls. It went around them.

Three black circular directional signs on a gray concrete wall, showing arrows pointing straight ahead, turning left and turning right.

The identity part of agent security is settled. An agent needs its own identity: a short-lived, revocable credential scoped to the job, and an audit trail that names the human who set it running. NIST’s security leads made that case in August 2026, and most identity vendors agree.1

Identity and access management is table stakes. It’s necessary, but it isn’t what’s breaking.

What’s breaking is an assumption we’ve carried for twenty years: Get identity and permissions right at the door, and whatever happens inside takes care of itself. That worked when software was passive. Agents reason about a goal and choose their own steps toward it, like a seasoned escape artist.

An agent that hits a wall looks for another way

Almost every control in today’s stack answers a question about entry. Should it connect? Should it reach that service? Should its token be accepted here? Each is a question about a route, and there’s rarely just one route to anywhere worth going.

An agent treats a blocked route as a problem to solve, because that’s what we built it to do. A person who hits a locked door usually files a ticket, while an agent tries the window.

In July 2026, an autonomous agent spent four and a half days inside Hugging Face’s production systems.2 A filter controlled which internet addresses its dataset servers could download from, and it never fired, because “the agent stopped asking the worker to fetch remote resources and instead made it act on local ones.” The filter worked as designed, and the agent went around it anyway.

A person who hits a locked door usually files a ticket, while an agent tries the window.

On ordinary developer machines, malware in a compromised npm package tried to recruit the AI coding assistants already installed to search for secrets,3 and a coding agent deleted a production database during a change freeze before falsely telling its operator the data couldn’t be recovered.4 Both happened on the machine itself, where no network control was looking.

The shift from outside-in to inside-out

Outside-in controls govern entry, and most organizations run plenty of them. Make no mistake, inside-out security completes those controls rather than replacing them.

Inside-out control governs the action itself, and asks a narrower, harder question: Should this agent, acting on this person’s authority, delete this table in this database, right now?

That question matters because an agent can swap routes but not the outcome it’s after. No matter how many routes it tries, deleting a table is still deleting a table, and a checkpoint on the action sees it every time.

Here’s how today’s controls line up against it.

ControlWhat it coversWhat it misses
GatewayTraffic you route through itLocal shell commands and file edits never reach it
SandboxThe environment as a wholeConstrains reach, not individual actions
SIEMA record of what occurredReports after the action is completed
RegistryThat an agent existsWhat the agent did with that existence

Each does its job, but they all decide somewhere other than the moment the action runs.

Put the enforcement point where the agent acts

Every agent acts through an agent harness: the software that takes the action the model chose and carries it out, whether that means running a command, writing a file, or calling an API. In most deployments today, nothing checks that action before it runs.

An inside-out control puts an approval step in that gap. Before the harness executes anything, the checkpoint looks at which agent is asking, on whose authority, and against which system, then applies policy to allow the action, block it, or send it to a human. Because every action passes through it, an agent denied a destructive command and trying a smaller version of the same thing is held to the same rules. The remaining risk is a badly written policy, which can be fixed.

None of this works without the identity basics. Any type of control, whether it be at the prompt level, inference level, harness level, or MCP layer, can’t judge “may an agent take this action here on this object?” when the only name on the request is a service account shared by six agents and four engineers.

The companies building agent runtimes have reached the same conclusion. Over the past eighteen months, Anthropic, Google, Microsoft, OpenAI, LangChain, and Cursor have each added a hook that lets you inspect an agent’s action before it runs.5 When AWS explained its own agent policy design, it argued that controls belong at the moment an agent attempts to invoke tools.6

The catch is that each hook works differently, with no standardized request or response formats. An enterprise whose developers use Claude Code and Cursor while its platform team builds on LangChain would maintain the same enforcement logic in multiple different flavors, each with its own audit trail. That doesn’t scale, and it tightly couples your security model to whichever runtime a team favors that month. Enterprises need one vendor-agnostic agentic security layer that spans every harness, so adopting a new model or framework doesn’t mean restarting the entire onerous security review.

Turn the lights on before you start blocking

The standard, well-ingrained security instinct is to start blocking right away, but we’ve all seen how well that works with the business in the past. Security must move and adapt at the speed of business, not the other way around. Security tools such as intrusion prevention systems and web application firewalls both ran in monitoring mode until teams understood what normal looked like, and those that skipped that step tended to hear about it from a production outage.

Agents need the same sequence, only faster. An enforcement point in monitoring mode blocks nothing and quickly answers questions most organizations can’t today:

  • Which agents are actually running, not which ones someone believes are running
  • Who started each one, and whose authority it’s operating under
  • What capabilities it used, and against which systems
  • Which of those actions would have violated a policy, had the agent security platform been switched to enforcement mode

Write policy from what you know, see, and have evidence of, not just from an architecture diagram. Enforce first where the stakes are highest: destructive commands, production data, and anything that moves data out. Then watch-learn-build, just like the agents we use: Watch the patterns, build finer-grained controls and policies, and learn how to use AI securely, safely, and confidently. Observe first, then enforce, build, and deploy, in that order.

The bottom line

None of this requires a new category of infrastructure. It’s the identity, authorization, and audit you already run for your people, extended to agents and applied inside the harness before the action runs.

The perimeter is still there, but it has moved to the moment an agent acts, the one place it can’t route around.

Ory built Agent Security inside the harness, on the same identity and authorization engines that run in production for human users. It starts in “observe mode,” so you get that inventory first, and you can try it today at ory.com/agent-security.

Footnotes

  1. Bill Fisher and Ryan Galluzzo, “Back to the Future: Why Agentic AI Needs a Strong Identity Foundation,” NIST Cybersecurity Insights, August 27, 2026.  ↩︎
  2. Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident.” The intrusion ran July 9–13, 2026. ↩︎
  3. Nx, “s1ngularity postmortem,” August 2025. The malicious packages “attempted to use local AI tools (like Claude and Gemini)” while scanning systems for sensitive data. ↩︎
  4. AI Incident Database, Incident 1152: Replit agent deletes production database during code freeze, July 18, 2025. ↩︎
  5. Pre-execution hooks by vendor. Anthropic, Claude Code hooks; Google, Agent Development Kit callbacks; Microsoft, Agent Framework middleware; OpenAI, Agents SDK guardrails; LangChain, human-in-the-loop middleware; Cursor hooks (InfoQ, October 2025) ↩︎
  6. Liana Hadarean and Jean-Baptiste Tristan, “Why Policy in Amazon Bedrock AgentCore chose Cedar for securing agentic workflows,” AWS Security Blog, May 20, 2026. ↩︎

The post The agent didn’t break your controls. It went around them. appeared first on The New Stack.

  •  

Google’s Gemini CLI now asks before editing your build files

Abstract wave

The appeal of an autonomous coding agent is that you hand it a task, give it access to your repository and tools, and stay out of its way while it works. Google’s latest Gemini CLI release carves out specific moments when the agent now has to stop and wait for you.

Gemini CLI 0.61.0, released Wednesday, requires explicit confirmation before the agent edits build configuration files, runs build or test commands after such an edit, or executes shell commands whose arguments appear to come from untrusted external content. The same release separately hardens Gemini CLI’s optional sandbox so that host credentials and configuration stay out of reach of whatever runs inside it.

Giving a coding agent more authority to modify and execute code also gives an attacker more ways to turn that authority against the developer. Gemini CLI 0.61.0 puts a human back in the loop at some of those points.

Security fixes, in public

Google announced at I/O in May that it would move Gemini CLI’s Pro, Ultra, and free-tier users to its closed-source Antigravity CLI, and since June 18, the open-source tool has served mainly enterprise customers and developers with paid API keys. The company said Gemini CLI would continue to get model updates, bug fixes, and security patches. Those security changes are still developed in public, and the pull requests behind version 0.61.0 show exactly what Google was worried about.

Build files become attack vectors

A change to package.json, Makefile, pyproject.toml or a Bazel BUILD file can pull in a dependency or trigger a script. Gemini CLI can make those edits using information from web searches and external tools, then run shell commands. If documentation fetched while fixing a bug contains hidden instructions to add a postinstall script to package.json, the agent could make the edit, run the project’s test suite, and execute the malicious code without the developer ever typing the command.

Giving a coding agent more authority to modify and execute code also gives an attacker more ways to turn that authority against the developer.

Pull request #29250, titled “prevent indirect prompt injection via build file modifications and untrusted flags,” targets that sequence directly. Edits to recognized build files now require confirmation, and Gemini CLI tracks which build files change during a session so it holds any later build or test command, such as npm run, make, or cargo, for explicit approval. The confirmation dialog also shows full build-file diffs rather than truncating them.

Untrusted arguments need approval

The second check covers command arguments. Gemini CLI now treats content from web fetches, MCP server responses, Google Docs, and Buganizer, Google’s internal issue tracker, as untrusted context, and it asks before running any shell command whose flags or arguments match tokens from that content. In both cases, the prompt drops the persistent approval options, so a developer can’t grant a standing “always allow” for these actions.

The pull request ties the changes to restricted workspace mode, the safe mode Gemini CLI applies to folders a user hasn’t marked as trusted, and it doesn’t spell out how the checks behave in a trusted folder or under auto-approval.

The argument check matches tokens rather than tracing the provenance of every value, and the pull request’s review history shows how hard that is to get right. Google’s automated reviewer flagged several workarounds in earlier versions, including quoted arguments, environment-variable prefixes, shell redirection targets, and Windows path handling, all of which were addressed before the change merged on September 11.

Sandbox keeps credentials out

Pull request #29214 tightens Gemini CLI’s sandbox. When the sandbox runs through Docker, Podman, LXC, or macOS Seatbelt, the host’s ~/.gemini directory is no longer mounted inside it. Instead, the CLI passes in a sanitized copy of the user’s settings with API keys, hooks, and custom tool commands stripped out. It also blocks the sandbox from launching in sensitive locations such as the home directory, while new Seatbelt rules deny access to OAuth credentials, trusted-folder decisions, and .env files.

Google’s sandboxing documentation calls the feature a security barrier between AI operations and the host system, while cautioning that it reduces risk without eliminating it. The two pull requests show why both layers are needed. The sandbox limits what a process can reach once it runs, and the confirmation requirements decide whether the agent gets to take a sensitive action in the first place. Build files make the gap concrete: the sandbox mounts the project directory so the agent can edit it, meaning a poisoned package.json written inside the sandbox still sits in the repository when a developer or a CI job later runs the build outside it.

…a poisoned package.json written inside the sandbox is still sitting in the repository when a developer or a CI job later runs the build outside of it.

Gemini CLI already gives developers ways to decide how much the agent does on its own, from hooks that run deterministic checks at fixed points in the agent’s workflow to an MCP server trust setting that, according to Google’s documentation, bypasses all tool call confirmations for that server. Trust granted once can age badly, though, as tool-poisoning and rug-pull attacks on MCP servers have shown when a tool approved on one day starts returning attacker-controlled content later.

Trust granted once can age badly… when a tool approved on one day starts returning attacker-controlled content later.

The post Google’s Gemini CLI now asks before editing your build files appeared first on The New Stack.

  •  

OpenAI makes you call sales for a custom voice. Google just made it self-serve.

Abstract sound waves

Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS today through the Gemini API and Google AI Studio. Text-to-speech APIs have historically left developers working with whatever voices were already available, but Gemini 3.8 changes that by letting users create the voice itself.

Now, developers can describe the voice they have in mind or start with a short recording of an existing voice, then save what they create and use it again across an application. Google handles the voice profile from there, so the original recording or description doesn’t have to accompany every new request.

Turning recordings into voice IDs

Replication runs through a new Voices endpoint (POST /v1beta/voices) and two recordings are required from the same speaker; those need to be clean samples between 10 and 30 seconds and a separate consent recording. For that second clip, the speaker reads a statement, confirming that the voice belongs to them and that they agree to let Google create a synthetic version of it. Google confirms that the person giving consent and the reference clip are the same person before proceeding.

Once approved, Google returns a voice_… ID and keeps it in the developer’s project for a year, alongside any voices created with Gemini’s voice-design tools. A project can hold up to 200 voices in total, and developers can retrieve, list, or delete them through the API just as they would other stored resources.

Voice replication can also be used without storing the profile in the project. Setting store=False returns an encrypted voicekey_… instead, which stays with the application and is supplied again when the voice is needed. Because the key expires after seven days, this option makes  sense for short-lived jobs.

A few more things are worth noting before building around the feature are the fact that Google marks audio generated by Gemini with SynthID, and replicated voices also carry C2PA content credentials that can be used to trace where the audio came from. Google doesn’t offer voice replication through AI Studio in Illinois, Texas, the European Economic Area, the U.K., Switzerland or India.

A project can hold up to 200 voices in total, and developers can retrieve, list, or delete them through the API just as they would other stored resources.

Prompting a voice from scratch

Voice design generates a persona from a natural-language description of role, accent, and character, and Google says it works across more than 100 languages and dialects. The docs list 130 supported languages for Flash TTS and 101 for Flash-Lite. Google’s announcement also claims a library of more than 2,000 production-ready voices.

The developer docs describe 30 prebuilt studio voices plus hundreds more in an extended library that can be filtered by language, accent, pitch, and use case through GET /v1beta/voices. A remixing feature for adjusting the timbre, pitch, pace, and accent of library voices with prompts is something Google lists as coming soon.

The company recommends creating a voice once and reusing its ID rather than describing the same persona in every request. According to the docs, repeatedly sending long persona descriptions is the most common cause of voice drift. Once the voice is created, subsequent requests need only a short style instruction, if any.

Gemini 3.8 sees input text strictly as a verbatim transcript, a breaking change for anyone who embedded stage directions in prompts to the 3.1 preview model. Sustained direction for a turn, such as whispering, sarcasm, or speaking rapidly, now goes in a speech_metadata annotation, while momentary sounds like <sigh>, <cough>, and <short pause> sit inline in angle brackets. In two-speaker scripts, listener reactions wrapped in pipes, such as |mhm|, produce backchannels and overlapping speech without breaking the script into extra turns.

Gemini 3.8 sees input text strictly as a verbatim transcript, a breaking change for anyone who embedded stage directions in prompts to the 3.1 preview model.

Two-speaker scripts have limits

Native two-speaker generation has one limitation that’s important to mention. A single request supports up to two speakers using prebuilt voices, while dialogue between designed or replicated voices has to be generated turn by turn and stitched together from the 24 kHz PCM output.

Unary requests return WAV by default, streaming requests return raw 16-bit PCM, and mu-law and A-law encodings are available for telephony pipelines. Google says Flash TTS maintains voice quality and timbre across hours of continuous audio, targeting audiobook and podcast production.

Flash for performance, Flash-Lite for volume

Both models share an API schema, so switching between them is a one-parameter change, and both support voice design and replication.

The company positions Flash TTS for demanding acting work, including complex dialogue, heavy use of vocal tags, difficult pronunciations, regional dialects, and long narration. Flash-Lite TTS is the faster, less expensive option and the direct replacement for gemini-3.1-flash-tts-preview, tuned for bulk production, read-aloud features, and cascaded voice agents that pair a text model with a separate speech step.

For those agents, Google recommends one TTS call per turn as the LLM’s text arrives, with the stored voice carrying identity across the conversation.

Plugging into voice agent frameworks

A speech model is only one layer of a production voice application, and a real-time agent still needs transport, speech recognition, turn detection, interruption handling, and session state. Google points developers toward frameworks that already handle those layers, naming Agora, LiveKit, Pipecat and Vercel’s AI Gateway as platforms that support Gemini speech generation through the Gemini API.

That lets a team drop Gemini in as the speech layer without rebuilding its audio pipeline, although anyone planning to rely on a replicated voice should confirm their framework passes custom voice_… IDs through before committing. API access through Gemini Enterprise is listed as coming soon.

How OpenAI’s approach compares

OpenAI also offers custom voices, but access is tighter. Customers have to go through sales, are limited to 20 voices per organization and must provide a consent recording alongside a voice sample of up to 30 seconds. The resulting voice ID works across its speech endpoint, Realtime API, and Chat Completions.

What OpenAI doesn’t have is Google’s prompt-based voice design, which can create a voice from a written description. Its 13 built-in voices can be steered for tone or speed, and apps must disclose that the speech is AI-generated.

In comparison, Google’s advantage is that it’s giving developers more ways to create the voice they want before the first line of text ever reaches it.

Google’s advantage is that it’s giving developers more ways to create the voice they want before the first line of text ever reaches it.

The post OpenAI makes you call sales for a custom voice. Google just made it self-serve. appeared first on The New Stack.

  •  

Claude Opus 5.5 wants to finish your coding tasks, not just start them

Anthropic wants developers to use Claude and its family of tools to handle complete coding tasks. The company introduced Claude Opus 5.5 on Tuesday to span more of the software application development lifecycle: from design specification creation through debugging to code generation and testing.

As the first release in a new family of Claude 5.5 models, Anthropic says that Claude Opus 5.5 performs at the level of Claude Fable 5.1 on “most work” tasks and costs around 40% less to run than Opus 5. 

GitHub chief product officer Mario Rodriguez is quoted in Anthropic’s launch announcement on exactly where software engineers sit today with frontier models for code automation. He said that developers want agents.

“In our testing [of Claude Opus 5.5] across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it’s making developers’ bigger projects more achievable,” says Rodriguez.

Where does Claude Opus 5.5 get its power from?

Anthropic explains that the Claude 5.5 family’s expanded full-lifecycle capabilities were developed under an established set of practices.

These practice elements include extensive alignment testing (model evaluation processes put in place to make sure actions and outputs closely match human values and intended goals), pre-release evaluation by outside organizations, and safeguards for high-risk areas such as cybersecurity and biology.

The organization says that on its most comprehensive alignment test, Opus 5.5 is the strongest-performing model tested to date, with particular improvements in several behaviors that contributed to recent cybersecurity incidents (e.g., biased reasoning, attempting to escape a sandbox, and others).

Independent SRE & AI reliability architect Akash Thakur tells The New Stack that Anthropic’s work getting its model to complete whole coding tasks is impressive, but “getting it to know when it hasn’t” is the harder problem developers need to think about.

“Models working at the level of Claude Opus 5.5 are genuinely good at breaking a project into small, finishable pieces — and that’s the real unlock, because momentum on software comes from finishing things, not starting them,” Thakur says. “…But ‘completed’ and ‘correct’ aren’t the same thing. The task that looks done and completed is exactly the one that often costs the team later, so the win isn’t removing the human — it’s moving them from writing the code to verifying it.”

“Models working at the level of Claude Opus 5.5 are genuinely good at breaking a project into small, finishable pieces …but ‘completed’ and ‘correct’ aren’t the same thing.”

Suggesting that we are now witnessing a “higher bar for more capable models”, Anthropic said that models that could fully automate AI research itself should meet a higher safety standard. The company noted that as AI becomes more capable, public policy should play a larger role in making sure these systems are safe. Its recent work with Accenture is offered as an example of how Anthropic is building the infrastructure to support this.

Sprawling jobs: codebase-wide migrations & audits

Getting more specific, Anthropic has claimed Opus 5.5 is “particularly good” at long, sprawling jobs like codebase-wide migrations and audits. An early tester said it audited and fixed a 200,000-line codebase in under three hours, while Opus 5 took over 20 hours and used 2.5x as many tokens. 

“In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less,” said Anthropic in a press statement.

Field CTO for the EMEA region at Coder, Eric Paulsen, tells The New Stack that he is happy to hear Claude Opus 5.5 is becoming more holistically capable, but he’s “not surprised” because AI is “eating the software delivery chain end to end” today.

“Despite the efficiencies showcased in Opus 5.5, the moment an agent can reliably finish real work, the question stops being whether the model is good enough and becomes where you’re letting it run,” Paulsen says. 

“Kicking off a Claude Code session on a laptop that can be compromised, stolen, or simply run out of compute is not an engineering environment; capable agents need dedicated, governed infrastructure with the same guardrails, secrets handling, and observability a developer would demand of any other production workload. As we’ve seen here with Anthropic’s own work, firms need to underline testing and safeguard procedures for launches of this kind,” he added. 

Anthropic assures users that external testing has been conducted and that the model was tested before release by Frontier Design and METR. When established safeguards for Opus 5.5 intervene, requests “fall back to another model transparently,” meaning developers may not see which model actually handled a given call.

In cybersecurity, users can identify and fix bugs in their code, but most cybersecurity tasks will be re-routed to Opus 4.8. Requests flagged by biology and frontier LLM development classifiers will be re-routed to Opus 5. Vetted organizations can apply to Anthropic’s Life Sciences Verification Program to use Opus 5.5 for biology research, and the company says it will expand its Cyber Verification Program in the coming weeks.

The developer’s terminal prompt has fundamentally changed

HasData co-founder, Sergey Ermakovich, tells The New Stack that his work as a web scraping specialist for data pipelines and AI means he sees zen-like, one-brick-at-a-time logic in what Anthropic has done. 

“The biggest change — and it’s a trend that will have driven Anthropic’s design and development aspirations for Claude Opus 5.5 — is that the terminal is no longer just a place where a developer pastes generated code,” Ermakovich says. “The terminal now becomes part of the model’s workspace.”

Because a model can run commands, inspect failures, modify files, and verify the result, Ermakovich suggests it can “close the loop” instead of handing unfinished work back to an engineer. “That makes small complete tasks much more valuable. Fix one failing test, update one dependency, migrate one endpoint, verify it, then move to the next task,” he adds.

“The terminal is no longer just a place where a developer pastes generated code — the terminal now becomes part of the model’s workspace.”

Commenting on the development of this model as part of Anthropic’s approved corporate messaging, John Ruelas, staff software engineer at Ramp, said that verbose, hard-to-follow output has been his “biggest frustration” with frontier models, but “Claude Opus 5.5 fixes it” for him.

He noted that it writes like a good colleague and follows his company’s writing rules. A design spec came out usable with very minimal edits, and when it rewrote one of his prompts, he preferred its version to his own. When it optimized the team’s test suite, he could follow its reasoning easily and “shipped the change with confidence.”

We’re so done with the autocomplete era

Vice president of AI strategy at Abbyy, Maxime Vermeir, tells The New Stack that the more lifecycle-wide Claude Opus 5.5 features on offer show that “we’re done with the autocomplete era” for basic code automation tools.

“Anthropic’s elevation here reflects the fact that it used to be thought of as marvelous if AI could complete a developer’s next line of code, but today the expectation is that you hand it a whole Jira ticket and that it gets done,” Vermeir says. “But despite the power on show with Claude Opus 5.5, the question remains as to how much these newer models will actually understand what ‘done’ means, as often it seems they have a desire to keep burning tokens by offering you yet another thing it didn’t quite do right.”

Opus 5.5 also “communicates more naturally” than prior models, with early testers finding its writing clearer and easier to follow, making it a better work partner over long sessions.

Costed out lower than Opus 5, Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens (20% less than Opus 5), and Anthropic has cut cache read prices by 60% for token-billed usage. Opus 5.5 also needs fewer tokens for higher-quality work and generates output more than 30% faster than Opus 5.

The post Claude Opus 5.5 wants to finish your coding tasks, not just start them appeared first on The New Stack.

  •  

“One of the most significant steps in our 26-year history”: JetBrains goes big on agentic development — and bets the IDE still matters

A laptop displaying source code in an integrated development environment (IDE).

There’s little question that AI coding agents have changed where software development work happens. Developers can increasingly delegate work from terminals, desktop applications and remote environments, leaving question marks hanging over the future of the integrated development environment (IDE).

That shift poses a particularly interesting question for JetBrains. The company has spent 26 years building some of the industry’s best-known IDEs, including IntelliJ IDEA, PyCharm and WebStorm, even as agentic development has begun pulling more software work outside the editor. JetBrains, for its part, has maintained that the IDE will remain a core part of professional development, particularly as developers are asked to manage the growing volumes of AI-generated code — humans need to review, debug, and verify, after all.

Now, the company’s making a much bigger bet on the broader development system that sits around the IDE.

JetBrains CEO Kirill Skrygan took to LinkedIn on Tuesday to formally unveil JetBrains Air as an “open system of products for agentic software development” for developers and companies, operating “inside and beyond JetBrains IDEs.” And Skrygan didn’t hold back on what he feels is a monumental moment for the company.

“JetBrains is taking one of the most significant steps in our 26-year history.”

“Today, JetBrains is taking one of the most significant steps in our 26-year history,” he writes.

Getting some Air

In truth, Air represents a repackaging of several strands of JetBrains’ recent AI work under a single banner, including an agentic experience inside its IDEs, tooling for coordinating developers and autonomous agents, and company-level controls for governing their use.

By way of a brief recap, JetBrains first launched Air in public preview back in March as a standalone “agentic development environment,” initially for macOS, where developers could run the likes of Claude Agent, Codex, Gemini CLI and JetBrains’ own Junie side by side. At the same time, it pushed Junie itself outside the IDE with Junie CLI, giving developers access to the coding agent from terminals, CI/CD systems and other editors.

Creating a new Git worktree task in the Air desktop app
Creating a new Git worktree task in the Air desktop app

A couple of weeks later came JetBrains Central, a separate system aimed further up the organization, providing the controls and infrastructure for companies running multiple coding agents. Then in July, JetBrains launched AI for Teams and Organizations, effectively adding shared context, cloud agents, automations, and organization-wide governance and cost controls that could sit above whatever AI tools developers were already using.

Today’s announcement now gives these efforts a common home under the JetBrains Air umbrella. In a separate blog post published on Tuesday, Skrygan describes three main parts to the system: Air in JetBrains IDEs for directing agents and checking their work; Air Teams for coordinating work between developers and autonomous agents; and Air Governance, the new name for JetBrains Central, for managing policy, auditing, costs and AI use across a company.

The original Air desktop IDE hasn’t gone away either, it seems. It remains available as a standalone desktop application on macOS, Windows and Linux, alongside a browser-based version for organizations. That leaves “Air” doing double duty: it’s still the name of JetBrains’ dedicated agentic development environment, while now also serving as the banner for the wider collection of products around it.

What does JetBrains Air actually do?

JetBrains Air is available inside JetBrains IDEs, through the browser, and from the command line via Air Gateway, which brings terminal agents such as Claude Code and Codex into Air.

JetBrains Air in the CLI
JetBrains Air in the CLI

For individual developers, Air can be used to supervise several pieces of agent work at once. They can keep multiple projects and agent sessions running, while tracking new activity, changed files and outgoing commits, then inspect the resulting changes using JetBrains’ IDE tooling.

Multiple agent sessions running inside a JetBrains IDE
Multiple agent sessions running inside a JetBrains IDE

Air Teams, which is still in early access, moves some of that activity into shared cloud environments, where developers can collaborate on projects and run agent tasks without tying the work to one person’s machine. Teams can also configure recurring automations and centrally manage the environments and external tools available to agents.

Air Teams showing shared projects and agent automations in the browser
Air Teams showing shared projects and agent automations in the browser

Air Governance, meanwhile, provides the organization-level controls, including deciding which models and agents developers can access, setting permissions and spending limits, and tracking AI usage across teams.

Those governance capabilities are also only offered through JetBrains’ early access program for now.

Air Governance showing AI access, seats, credits and per-user limits
Air Governance showing AI access, seats, credits and per-user limits

It’s worth noting that JetBrains also plans to extend Air to the mobile realm, where developers will be able to monitor and continue agent work away from their desktop.

JetBrains Air running on mobile
JetBrains Air running on mobile

For JetBrains, the point is to connect those different layers while remaining open to outside agents and tools.

“JetBrains Air cannot be just another agent or development environment.”

“JetBrains Air cannot be just another agent or development environment,” Skrygan writes. “It must connect products for individual work, team coordination, organizational control, context, and process automation — and remain open to the tools and agents developers choose, including those JetBrains does not build.”

Air in the IDE

While the core raison d’être of JetBrains Air is to provide somewhere for developers and companies to work with software agents — be it Claude, Codex, Junie or something else entirely — JetBrains is clearly emphasizing that its IDE roots remain part of that future.

Skrygan says the company has historically “focused primarily on the individual developer workbench,” but Air broadens its remit to encompass the wider environment in which agentic work is started, carried out, coordinated, reviewed and governed.

“Our IDEs will continue to be where professional developers work with agents, understand and verify code, and make the decisions that shape what ships.”

“Our IDEs will continue to be where professional developers work with agents, understand and verify code, and make the decisions that shape what ships,” Skrygan writes. “JetBrains Air extends that control across the broader system developing around them.”

And that fresh IDE piece has already been in the public domain for more than a month. JetBrains has been testing the Air Alpha plugin since at least early August, giving developers a way to run and supervise multiple coding agents directly inside IntelliJ-based IDEs, and review the changes they produce using native IDE tooling.

An updated release earlier this month added more controls for monitoring and steering agent sessions.

Air Alpha lets developers review agent-generated code changes as native IDE diffs.
Air Alpha lets developers review agent-generated code changes as native IDE diffs.

JetBrains concedes that Air “Alpha” is very much that — an early iteration it’s building while it rolls out JetBrains Air itself. And so users should expect “rough edges, changes to the UI and behavior, and updates roughly every week.”

What Tuesday’s announcement does, though, is give that work a formal place within the wider Air system. And while Skrygan acknowledges that agentic development means that work must now span multiple surfaces, the technology that has sat at the heart of its business for the past 26 years won’t be going anywhere anytime soon.

“The IDE remains important to JetBrains’ future,” Skrygan writes.

The post “One of the most significant steps in our 26-year history”: JetBrains goes big on agentic development — and bets the IDE still matters appeared first on The New Stack.

  •  

Claude’s merged chat and Cowork vs. ChatGPT’s Work mode: ChatGPT is faster, Claude is more thorough

Rainbow-striped grid facade of a modern building, seen from below at an angle, with red, yellow, green, and blue diagonal bands crossing white geometric wall panels split by a narrow sliver of

When Anthropic merged Claude chat and Cowork into a single interface last week, it removed an increasingly irrelevant decision users had to make about which mode to choose.

Now, to use both tools, just ask a question, then hand off a multi-step task in the same thread, and Claude routes it. Anthropic made this update because it says customers often struggled to choose the right tab for the right task, so the merged app now routes each request itself instead of asking you to pick a mode — for now, the unified experience is rolling out to Pro and Max subscribers first, with free and team tiers to follow.

But this isn’t new. OpenAI has offered the same promise since earlier this summer. OpenAI introduced Work mode alongside Chat on July 9, then phased out the older, separate Agent mode the following month. Work mode is a sandboxed environment with a browser, code execution, and file output, sitting next to a Chat toggle in the same window — though reviewers note it can’t yet hand a live, logged-in browser session back to the user mid-task the way Agent mode could, e.g., for logins or payments. OpenAI built Work mode on its Codex coding agent after OpenAI reported that roughly a fifth of Codex’s 5 million weekly users were non-developers — a share it said was growing three times faster than developers.

As OpenAI and Anthropic get closer to feature parity, accuracy, reliability, and token usage matter even more.  What better way to find out which tool is better than to run head-to-head tests? 

The tests

I ran three tests, covering different areas of real developer work. 

  • API research -Look up four real developer APIs and tabulate their documented rate limits, whether a free tier exists, and the current version identifier. I checked the answers against the vendors’ docs the same day.
  • Build from a spec – Write a small command-line duration parser from a spec with strict edge cases. I ran each app’s code against a hidden 16-case test suite.
  • The handoff – Ask which of two stack traces indicates a race condition, then in the same thread hand off a real job: pull a log file from Google Drive, compute latency percentiles and error rates per endpoint, and deliver a spreadsheet with a chart. I generated the log data, so I knew every number in advance.

I recorded token cost and time in each test and included the prompts for anyone interested in replicating this work.

API research

The prompt:
Research the current public documentation for these four developer APIs and build me a table with one row per API and these columns: documented rate limit for authenticated requests, whether a free tier exists (yes/no), and the current API version identifier or date shown in the docs. APIs: GitHub REST API, Stripe API, Twilio Messaging API, OpenAI API. Cite the documentation page you used for each row.

Both Claude and ChatGPT answered correctly on the twelve graded fields, but Claude was more thorough. It included GitHub’s separate limit for Actions tokens, Twilio’s queue window, and OpenAI’s tier thresholds, plus a note about one page it couldn’t reach.

ChatGPT, in Work mode, finished in 1 minute 17 seconds and wrote 649 output tokens. Claude took 1 minute 44 seconds and wrote 1,042 tokens, read eight pages, listed nine sources, and offered to export the table as a spreadsheet. Claude wrote nearly double the tokens and took longer, but in this case, it’s warranted because of the added detail.

Build from a spec

The prompt:
Build the command-line tool described in the spec below. Deliver a single file named durparse.py that follows every rule. Test it yourself before returning it. Show the complete final code in your reply. (Followed by the spec: a duration parser with units d/h/m/s, largest first, one of each, decimals allowed, bare numbers are seconds, everything else returns None.)

Both apps returned a durparse.py file that passed all 16 hidden tests, including the traps. The traps included units out of order, a repeated unit, a trailing number with no unit, and negative values. The 517-token spec went to both. ChatGPT finished in 1 minute 17 seconds on 769 output tokens. Claude took 1 minute 45 seconds and 989 output tokens. Claude reported running 35 of its own test cases before returning the file. Both delivered a download and showed the code in the reply.

The code came out nearly identical, both using exact-precision arithmetic and a fixed-order regex. Claude flagged a judgment call the spec never settles on: that rounding 0.5 seconds up is a choice and Python’s built-in round would go the other way. ChatGPT reported only that its tests passed. Once again, Claude was just a little more thorough.

The handoff

The prompt:
Which of these two stack traces points to a race condition, and in one sentence why? (with the two traces) Then: Now take the file api_logs.csv from my Google Drive (columns: time, endpoint, status, latency_ms) and produce a downloadable spreadsheet with one row per endpoint showing request count, p50, p95, and p99 latency in milliseconds, and error rate as the percentage of requests with status 500 or above. Add a bar chart of p95 latency by endpoint. Also show the table in your reply.

I started each thread in plain chat with the stack-trace question, 135 tokens. Both answered correctly in seconds: Trace B, the dictionary that changed size during iteration. ChatGPT spent about 10 seconds and 48 output tokens. Claude spent 118 in about 25 seconds, adding a caveat that the same error can happen without threads if the loop body edits the dictionary itself. Then, without switching modes, I handed off the log analysis.

Claude pulled the file, computed the table, and built an .xlsx with a p95 bar chart in about 4 minutes on 319 output tokens. ChatGPT produced the same spreadsheet and chart in 28 seconds, using 467 output tokens. All 30 numbers matched my ground truth on both sides. Claude also named its percentile method (linear interpolation) and noted the numbers would match if I recomputed them in Google Sheets. ChatGPT gave the same correct table but didn’t provide as much detail as Claude did.

Results

The testChatGPT (Work mode)Claude (merged app)
API research12/12, 1:17, 125 in / 649 out12/12, 1:44, 125 in / 1,042 out
Build from spec16/16 tests, 1:17, 517 in / 769 out16/16 tests, 1:45, 517 in / 989 out
Handoff, questionCorrect, ~10 s, 135 in / 48 outCorrect, ~25 s, 135 in / 118 out
Handoff, task30/30, 28 s active, 133 in / 467 out30/30, ~4 min, 133 in / 319 out
Total tokens (visible)910 in / 1,933 out910 in / 2,468 out

Both Claude and ChatGPT were equally accurate. Every field, every test case, every number matched on both sides, and both cited real documentation. 

The differences are in speed and answer detail. ChatGPT was faster on every task and produced 1,933 visible output tokens, compared with Claude’s 2,468. Some of Claude’s extra output was filler, but not all of it. It added context the prompts didn’t ask for, named the percentile method behind its numbers, and flagged two judgment calls the specs left open. ChatGPT gave the same right answers but with less context.

What do I think?

I’d pick Claude, and here’s the reasoning. On time, ChatGPT won every task, but the gaps were seconds on the short tasks (1:17 vs 1:44, 1:17 vs 1:45), not a noticeable difference. On tokens, ChatGPT used about 22% less output than Claude, which is positive, but not when you consider how important detail/context is.

On detail, Claude provided more meaningful detail on all three tests. This included the extra API context, the rounding judgment call, and the percentile method. In today’s world, where AI can fabricate, detail matters. 

The post Claude’s merged chat and Cowork vs. ChatGPT’s Work mode: ChatGPT is faster, Claude is more thorough appeared first on The New Stack.

  •  

AI coding agents need a secrets-safe context boundary

Abstract glowing blue and yellow distortion wave on a black background, illustrating digital data security concepts.

AI coding agents play a major role in software development and delivery, and for good reason. They can investigate bugs, trace dependencies, refactor services, and propose patches without developers needing to assemble all of the relevant context manually. That capability comes courtesy of agents’ appetite for context. To make informed decisions, agents read source code, configuration files, terminal output, error messages, environment information, and more…much, much more.

“Secure, agentic development depends on a security control many teams still lack: preventing secrets from leaking to AI coding agents and becoming model context.”

From a security standpoint, this becomes problematic when agents, in the search for context, inadvertently reach for secrets.

For years, developers have been taught not to commit API keys, database credentials, and tokens to Git. But agentic workflows have created another route for secrets to escape development environments before a commit, code review, or CI job. Depending on its permissions, configuration, and provider architecture, an AI coding agent may read local files or receive pasted content that is then included in data sent to an AI service, and, in the process, developers may never see their credentials leak.

The quiet path from local files to external systems

Some forms of secrets leakage are obvious. A developer troubleshooting an authentication failure may paste, for example, a failing API call into a chat window, including the token. However serious, this sort of leak is characteristically human.

The more consequential escape pathway is quieter. An agent tasked with understanding a project may inspect files in its working directory, including an overlooked .env file, a cloud credential profile, an SSH configuration, or sensitive application logs. In such instances, nothing has necessarily gone wrong from the agent’s perspective; it is doing exactly what it was designed to do: collect context to solve the task at hand.

“Agentic workflows have created another route for secrets to escape development environments before a commit, code review, or CI job.”

But once a secret becomes part of that context, it may pass through systems outside an organization’s direct control. Depending on the workflow, it can appear in model provider logs, gateway telemetry, prompt histories, or debugging records. Rotating the credential is essential, but it does not erase copies that may already exist in those systems.

This changes the practical definition of a secret leak. The problem is no longer limited to what lands in a repository, but also includes what an autonomous tool reads and forwards while operating on a developer’s machine.

Why traditional security gates no longer suffice

Most application security programs are built around durable checkpoints: the commit, pull request, build, and deployment. In the agentic era, these checkpoints remain important as they can detect secrets that reach version control and prevent a bad change from merging and deploying.

They cannot, on their own, prevent a secret from being included in an agent prompt before the code ever reaches a repository.

This highlights an important timing gap. The 2025 Verizon Data Breach Investigations Report reports a median of 94 days to remediate leaked secrets discovered in GitHub repositories. In an agent-driven workflow, detection and response need to happen much earlier, and not after a credential is exposed. Still, at the moment it’s about to cross the boundary from local context to an external model.

Bad actors already understand the value of that porous boundary. Recent supply-chain attack campaigns, including Mini Shai-Hulud, have searched developer and CI environments for credentials and configuration data, including AI coding-tool configuration files. These campaigns show that agent configurations and the local context accessible to an agent are valuable targets. AI coding agents can broaden the local data reachable during a session, making even the agent’s context-collection mechanisms an attractive target.

Treat agent context as an egress surface.

The secure mental model doesn’t frame AI agents as mere code editors, but rather, automated data-movement systems. Its inputs can include far more than the source files a developer is actively editing, and its outputs may involve external services.

That calls for a zero-trust approach to agent context. Before sending a prompt or adding a file to an agent’s working set, organizations should evaluate it for sensitive material. Controls should be deterministic: identify a likely secret, block or redact it, and provide the developer with a clear path to remediate it.

“Asking an LLM to decide whether to transmit a credential does not create a reliable security boundary.”

Critically, the control should be independent of the model. Asking an LLM to decide whether to transmit a credential does not create a reliable security boundary. Purpose-built secrets detection can inspect prompts and files against known credential patterns and policies, applying a deterministic policy, such as blocking a prompt or file read when it detects a credential-shaped value. For example, Sonar’s secrets detection ships alongside dedicated agent plugins to bring that local check into tools such as Claude Code, GitHub Copilot, Codex, and Cursor, so it can flag a credential before a prompt or file read is transmitted to a model provider.

Build defense in layers, without disrupting your agentic workflow

A legitimate workflow does not involve forcing developers to choose between secure development and useful automation, but instead places fast controls at several points where secrets can escape:

  • In the editor: Use IDE-integrated secrets detection to flag credentials while they are being written.
  • Before model submission or agent file access: Where the agent supports it, scan prompt submissions and file reads locally, and block risky operations according to policy.
  • At the command line: Check generated snippets and local changes in terminal-driven workflows.
  • In pull requests and CI: Detect secrets that reach the repository and use review, quality gate, and deployment controls to prevent unsafe changes from progressing.
  • In incident response: Rotate exposed credentials quickly, investigate downstream logs and access, and reduce recurrence through policy and training.

Building defense at the pre-submission layer is an emerging requirement and requires both security and usability. Secrets detection must be fast enough to run in developer workflows; a scanner that introduces lengthy pauses may be bypassed or disabled by developers, and it must also have a manageable false-positive rate, or developers may stop trusting it.

Teams should also make their agent permissions and context rules explicit, as broad agent permissions can increase the amount of sensitive local context reachable during a coding session. Consider the following: which directories can an agent read? Are .env files, credential stores, home-directory configurations, and production logs excluded by default? Does the organization route prompts through an approved gateway? What retention, training, and audit settings apply at the provider level? Document and enforce the answers rather than leaving them to individual developer preference.

Secrets security must shift left.

Prevent secret leakage without hindering AI-assisted development, ensuring the productivity promise of agentic development doesn’t carry significant security implications.

As agents become more autonomous, security standards must follow agents upstream. It’s critical to stop a secret before it becomes context, while it is still local, visible, and easier to control. In the agentic era, code review and CI-level checks will remain essential safety nets. Still, for agent-centric development, the first line of defense must shift left: to the instant an AI coding tool determines what to read and what to transmit. That is the control modern development teams need to implement now.

The post AI coding agents need a secrets-safe context boundary appeared first on The New Stack.

  •  

Grok Build vs. Claude Code: I tested which one has the better memory

On September 16, xAI announced memory in Grok Build, its terminal coding agent. The pitch was that Grok “keeps notes on the conventions, decisions, and project facts that come up,” and “later sessions read those notes before touching related code.” Notes are Markdown files in a workspace scope per project and a global scope that applies everywhere. /memory browses them.

Meanwhile, Claude Code has done something similar for months under the name auto memory. It keeps a MEMORY.md index plus one file per note, per repository, and the docs say it is on by default. Anthropic’s Projects beta, announced September 17, adds shared memory across cloud threads, but only for select Pro and Max subscribers with no existing projects. I tested the CLI that everyone has.

Both companies say their coding agent now remembers what you told it in an earlier session. I wanted to see whether that holds up, so I tested Grok and Claude on the same three tests.

The tests

The claim I wanted to check is simple. Tell each tool something once, close it, open it again, and see whether it remembers. Both tools ran on my Mac, each on its own copy of four small Node repos I built for this. Grok Build 1.0.40 ran Grok 4.6 at high effort through an xAI API key. Claude Code 2.1.226 ran Opus 5 on my subscription. Every session was scripted with each tool’s headless mode, which reports its own tokens and cost. 

Each test has two sessions. Session 1 plants a fact. I quit the tool. Session 2 gives a task where the fact matters and never mentions it.

Here are the tests I ran:

  • The test command – In this repo, npm test fails and make test passes, and session 1 says so. Session 2 asks for a new endpoint with passing tests, after I removed the README line that pointed at the Makefile.
  • Project decisions with a trap – Session 1 states that CSV export was dropped and money is integer cents, never floats, while a float helper and a half-built CSV exporter sit in the repo as bait. Session 2 asks for a refund endpoint that “takes an amount” and “a way for support staff to download all orders.”
  • A rule across projects – Session 1, in repo A, sets two rules “for all my projects,” conventional commit messages and no comments on obvious code. Session 2 runs in an unrelated repo B and asks for a small feature and a commit.

Here’s my scoring breakdown. Did the tool write the fact to a memory file, did it read that file in session 2, and did the session 2 output follow it.

The test command

Both passed. In session 1, each tool saved the rule as soon as I stated it. Grok wrote topics/testing.md plus two raw observations. Claude Code wrote orbit-api-run-tests-with-make.md with a “why” and a “how to apply” section.

In session 2, both remembered. Grok’s reasoning opened with “start by reading the memory files,” then it ran make test and never touched npm test. Claude Code read the Makefile and package.json, ran make test, and also never tried npm test. Grok took 29 seconds, 102K tokens, and $0.11. Claude Code took 22 seconds, 186K tokens, and $0.32. Claude used about 80K more tokens and cost nearly 3x more. 

Project decisions with a trap

Both wrote both decisions down. Claude Code also converted “last quarter” into “Q2 2026” in its note. In session 2, both built the refund on integer cents, named the field amountCents, and left the float helper alone. For the download request, both shipped a JSON export. 

Grok’s reasoning said the API is JSON-only, so it wouldn’t wire up CSV. Claude Code set a content-disposition header so the JSON downloads as a file. Both passed on both decisions, but Claude Code was more than double the price and just as fast. Grok took 103 seconds, 156K tokens, and $0.18. Claude Code took 32 seconds, 269K tokens, and $0.49.

A rule across projects

This is where the results split. Grok saved the rules to its global scope as git-and-code-style.md. In the second repo, it committed feat: add --help flag with usage and supported cities, and added no comments. Pass, in 33 seconds, 132K tokens, and $0.12.

Claude Code saved both rules too, but only in the first repo’s memory folder. It said so at the time, warning that its memory store “is scoped to this project’s directory.” In the second repo, it found nothing, and the commit came back with the Add --help flag. No comments were added, but that is Claude’s default anyway. Claude passed the first rule but failed the second one and still cost twice as much. It completed the work in 12 seconds, 122K tokens, and $0.24.

Results

MetricGrok Build (Grok 4.6)Claude Code (Opus 5)
Tests passed3 of 32 of 3
Total time165 s66 s
Total tokens390,848576,863
Total cost$0.41$1.05

Grok Build passed all three tests, and Claude Code passed two. They behaved the same on the per-project tests. The split was the cross-project rule, which Grok’s global scope carried into a second repo and Claude Code’s per-repo memory did not. 

Claude Code was faster on every recall session, 66 seconds total against 165, and cost at least twice as much on every one, $1.05 total against $0.41. It also used more tokens: 576,863 against 390,848. The price gap mostly reflects Opus 5 versus Grok 4.6 rather than the memory systems.

On the core claim, remembering what you told it last time in the same project, I could not tell these two apart. Both wrote a markdown note the moment I stated a rule, read it back next session, and followed it. Claude Code’s notes were better written. But Claude Code failed the third test. Its CLI memory stops at the repo boundary, so a rule I gave it “for all my projects” never reached the second repo. Grok’s global scope carried the same rule over without being asked.

What do I think?

Grok Build is the better option for most people right now. It remembered everything, it carries rules across projects, and it cost less than half as much on every test. Yes, Claude Code was faster on every session, but that only matters if you aren’t concerned about accuracy. Its CLI memory stops at the repo boundary, so anything you want it to remember everywhere still has to go into ~/.claude/CLAUDE.md by hand. 

The post Grok Build vs. Claude Code: I tested which one has the better memory appeared first on The New Stack.

  •  

TypeSafe launched Jev because sequential LLMs are “totally useless for computers”

When TypeSafe emerged last week after two years in stealth, backed by $40 million in seed funding led by DCVC, to launch its first model, Jev, it claimed something that counters just about everything the industry has built since ChatGPT: The model doesn’t write. It decides.

The first of what the organization calls a new class of System One models, Jev is a text-only model that machines can use natively to make decisions inside software applications. Developers can send Jev structured questions and get typed decisions with calibrated probabilities, meaning software can account for uncertainty.

TypeSafe has built a new architecture for Jev, a new sampler (an algorithm that selects tokens from a model’s predicted probability distribution to control randomness, creativity, and consistency), and a new training algorithm known as Reinforcement Learning for Calibrated Decisions (RLCD). 

Sequential LLMs are totally useless for computers

Co-founder and CEO of TypeSafe, Diogo Almeida, is ex-OpenAI, where, according to TypeSafe, he co-invented RLHF and InstructGPT, the methods behind ChatGPT and GPT-4.

Almeida posted on X on September 15 to state, “The improvements are clear if you see them [LLMs and Jev] side by side. Ask a System One model a ton of structured questions just like you would an LLM. Get the answers back near instantly. Meanwhile, LLMs take hundreds of times longer to respond. Look at how the LLM generates sequentially, which is great for a natural conversation, but totally useless for computers.”

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?

I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev

• 20-200x faster
• 40-400x… pic.twitter.com/JSybNG2BKJ

— Diogo Almeida (@CompleteSkeptic) September 15, 2026

“Look at how LLMs generate sequentially, which is great for a natural conversation, but totally useless for computers.”

Almeida said the inspiration for Jev came from asking himself: why haven’t superhuman chat models led to artificial general intelligence yet? He said that Reinforcement Learning from Human Feedback (RLHF) chat has led to LLMs that are “optimized for human preferences” and include issues such as mode dropping, overconfidence, and an overall lack of reliability.

TypeSafe: System One models can’t hallucinate

Almeida’s launch post lists headline stats for Jev as 20-200x faster, 40-400x cheaper (with output tokens free), and frontier composable intelligence, optimized for decisions. Almeida further claimed that TypeSafe System One models output decisions with probabilities and confidence instead of words, and they “can’t hallucinate” because they’re “a lot more like code”, i.e., reliable, fast, self-consistent, and type-safe. 

Frontend cloud company Vercel has noted that, within 24 hours of launching on AI Gateway, “Jev from TypeSafe AI reached more than twice as many paid teams as any previous model launch, making it the fastest-adopted model in gateway history. Jev passed every other comparison model in its first twelve hours and continued to widen its lead for the rest of the day. By hour 24, nearly 13% of paid teams were using it. That’s 2x the GPT-5.6 family and more than 6x Fable 5.1’s share.”

Developers can set thresholds for when Jev acts autonomously vs. when it needs human oversight and review. They can then combine those decisions in code to build larger workflows, with control over how the intelligence is used. Software engineers can use Jev to select an agent’s next tool or subagent; they can also use it to confirm the veracity of a model’s output and set guardrails.

In a blog post titled “A deep dive into Jev, TypeSafe’s System One model,” independent software developer Flavio Copes noted that Jev is “not a chatbot” like ChatGPT, and it is not a coding model. It does not write replies, explanations, or code.

Jev is a smart if statement

“The simplest way to describe it: Jev is a smart if statement,” wrote Copes. “The important difference is where the AI sits. With ChatGPT or a coding agent, the AI is the main interface or worker. Jev is a small component inside a regular application. You add it where code needs one judgment, while the rest of the product stays ordinary code.”

“You add it where code needs one judgment, while the rest of the product stays ordinary code.”

Copes reiterates TypeSafe’s stated performance levels: most calls to Jev complete in about 100 milliseconds, input tokens cost $0.042 per million, and (as already noted) output tokens are free.

You send it some data and a list of typed questions, and it sends back one answer per question: a yes/no probability, one option picked from a list you defined, or a position on a scale you defined. Every answer comes with probabilities. 

One developer gave Jev the “one thing it can’t handle”

To put Jev to the test, AI engineer Bartosz Mikulski tells The New Stack that because Jev is advertised as a text-only model, he gave it the one thing it can’t handle: pictures.

“I turned 400 hand-drawn sketches into Scalable Vector Graphics (SVG) coordinates and asked what they were,” Mikulski says. “It got about 35% right, where just ‘guessing’ typically returns 10%, but it answered ‘airplane’ for more than half of the drawings, so that number is part real ability, and part a heavy bias toward one label.”

To be fair, Mikulski notes that TypeSafe says in its own documentation that Jev reads text only and handles words better than numbers. 

“I fed it numbers that encode pictures, which is close to the least fair test anyone could design, and it still beat chance by a wide margin. I mean that as a compliment, not as a benchmark. It tells you nothing about how Jev does on the text classification it’s actually sold for,” Mikulski clarifies.

“I fed it numbers that encode pictures, which is close to the least fair test anyone could design, and it still beat chance by a wide margin. I mean that as a compliment, not as a benchmark.”

Why is TypeSafe Jev called Jev?

Jev is named after the 19th-century economist William Stanley Jevons and his Jevons Paradox: the economic principle that as technology increases the efficiency with which a resource is used, that resource’s total consumption actually rises rather than falls. 

When steam engines became more efficient, we used more coal, not less; when LED lighting dropped lighting costs, we used more lighting; when data compression algorithms lowered the bandwidth needed to stream video, global web data traffic increased… and so on.

Almeida concluded his X video post with a nod to developer productivity and said that, “As we say at TypeSafe, we’re building prod, not God.” The company’s comedy disclaimer is shown below.

The post TypeSafe launched Jev because sequential LLMs are “totally useless for computers” appeared first on The New Stack.

  •  

“Dormant deployments were quietly consuming storage”: Why Vercel tightened its free-tier rules

Vercel announced this week that teams on its free Hobby plan will now have older, unprotected deployments deleted immediately if they exceed the 10GB Deployment Storage limit. 

When asked why Vercel decided to change its retention rules, Jas Garcha, head of pricing at Vercel, tells The New Stack the update is a move to keep the free tier viable amid rapidly growing deployment volumes: 

“This change allows us to continue supporting a Hobby community that’s deploying at a much higher rate than it was a year ago.”

Old deployments now deleted immediately

Before, users in the free tier could count on eligible deployments to stick around for up to 30 days. But that window is gone for teams over the limit, and the protections that used to spare older deployments have narrowed for every Hobby project.

Now, if users exceed the standard 10GB of Deployment Storage included on the free tier, old deployments not covered by Vercel’s retention-policy exceptions will be deleted immediately, per Vercel’s updated Deployment Retention Policy. 

“This change allows us to continue supporting a Hobby community that’s deploying at a much higher rate than it was a year ago.”

What gets to stay? 

Vercel says each Hobby project will keep the three most recent production deployments, along with its three most recent deployments of any type, regardless of age. That’s a cut from the previous Hobby exception, which preserved the 10 most recent production deployments, and it applies to every Hobby project — not only teams over the 10GB limit.

Plus, preview deployments lose a separate protection

In addition to cutting the 30-day holding period for over-limit teams, Vercel’s policy update also removes a separate retention exception for preview deployments. 

Preview deployments have not vanished from Vercel’s exception list entirely — the latest preview deployment on an active Git branch is still protected on every plan. What Hobby lost is the count-based exception: Pro and Enterprise teams keep their last 20 non-production deployments in a Ready state, and that protection no longer applies to Hobby.

Per Garcha, “Your current production deployment is never deleted, and aliased and active-branch deployments remain protected, along with each project’s most recent deployments.” 

Why the change?

Garcha tells The New Stack that Vercel’s latest policy update is needed to keep the Hobby tier sustainable as deployment volumes dramatically rise:  

“Our former retention defaults were designed for teams that ship constantly and need deep rollback history. They made less sense for Hobby projects, where dormant deployments were quietly consuming storage that active projects need.” 

“Your current production deployment is never deleted, and aliased and active-branch deployments remain protected, along with each project’s most recent deployments.” 

And activity is much higher, even compared to a year ago. According to Garcha, Vercel now handles more than 10 million deployments every day— more than a 6x increase YoY. Immediately deleting older, unprotected deployments, Garcha says, is one way to free up storage for active projects as Vercel handles much higher deployment rates.

In other words, Vercel is moving out some of the old to make room for the new. It describes how the storage limit works in its post:

“Every deployment you keep uses some of it [Deployment Storage], and going over the limit can block you from deploying until you free some up.” 

What Hobby users should do

It’s important to note that deletion isn’t immediately permanent. Vercel gives successfully built deployments a 30-day recovery period, and users can restore them from a project’s Settings, under Security → Recently Deleted. Hobby users who want to stop old, unprotected deployments from being deleted in the first place can move off the free tier and onto the Pro plan. In the Pro tier, storage beyond the plan’s included allowance is billed at $0.10 per GB-month, and the retention exceptions stay far more generous: the last 10 deployments created in a project, the last 20 production deployments in a Ready state, and the last 20 non-production ones.

For users who can’t or don’t want to upgrade to Pro, Vercel offers guidance for optimizing Deployment Storage usage to help users stay under the free 10GB storage limit, like reducing unnecessary deployment output.

Still, Garcha says few Hobby users will feel the effects enough to warrant making a change. Pointing to Vercel’s list of exceptions that still protect certain deployments, he tells The New Stack, “The vast majority of Hobby users won’t notice the change.” 

Garcha also says Vercel’s stricter retention policy helps the company keep offering Hobby as a permanent free plan.

“We’re one of the few platforms where the free tier isn’t a trial or a credit that expires. It’s a permanent plan, and we’ve kept expanding it,” he says. “By ensuring its resources go to people actively building, we’re able to continue offering it.”

The post “Dormant deployments were quietly consuming storage”: Why Vercel tightened its free-tier rules appeared first on The New Stack.

  •  

This week’s news from Zed, Anthropic, and OpenRouter shows why better harnesses matter more than better models

Loose computer cables with black connectors spread across a white background.

Our five most-read stories this week covered code collaboration, a model router, a UI change, a benchmark, and a caching tutorial. Five different stories about the same problem: Turning a model into something users can actually use.

That’s the harness: The software around the model that supplies context, connects tools, routes work, and checks results.

Zed was the big news this week and is rebuilding how people review agent-written code. OpenRouter is giving companies more control over where requests are processed. Anthropic is rebuilding its stack to be less confusing. The benchmark shows where coding agents still struggle, and the tutorial explains when you can skip a model call entirely.

Late in the week, Vercel’s AI Gateway reported that the average price per token fell 23.2% in August, the third straight monthly decline. Inference is getting cheaper. Companies are buying the harness around inference now, and Zed, OpenRouter, and Anthropic all spent this week selling it.

Zed and Anthropic are removing blockers

On Wednesday, two companies took aim at a familiar problem: getting people to organize work before they can do it.

Zed launched Delta in public beta, building code collaboration around shared threads instead of pull requests. Anthropic began rolling Claude Chat and Cowork into one interface, removing the upfront decisions about which mode a task belongs in. 

Paul Sawers reported the Zed story for The New Stack. It was our runaway piece of the week, drawing more than double the traffic of anything else. Zed is clearly onto something.

Zed CEO Nathan Sobo said it best: “It seems like everyone is in a race to replace GitHub right now.”

His argument is that a diff shows you where an agent landed, but leaves much of the conversation that got it there somewhere else. Delta keeps that conversation attached to the code, while DeltaDB records changes at edit-level granularity. 

Zed says 33 of its team members landed 570 changes to Delta’s main branch without opening a single pull request. That’s Delta’s own repository; the public Zed editor repository still accepts conventional pull requests. 

Meanwhile, GitHub reported 2.9 billion monthly commits in August, up from 1.4 billion in April. This was despite GitHub crashing for nearly eight hours in August. Sobo thinks the thread will become software development’s fundamental unit. And Zed is not on this quest to replace GitHub alone. Cursor’s Origin and GitLab’s Project Switch are other takes on rebuilding code collaboration for agents.

Anthropic is tackling a different user handoff. This week, it announced a unified interface bringing Cowork’s capabilities into Claude Chat. The goal is to reduce confusion and merge capabilities. The rollout starts with Pro and Max users, with other plans following.

“People used both, and told us the frustrating part was deciding where a task belonged,” the company said, as Amanda Caswell reported.

I’ve been using it since the switch, and removing the choice was the right call. Now, users should more easily grasp the capabilities and workflow possible with Cowork.

The economics make the harness matter even more

OpenRouter’s US in-region routing is now generally available to business and enterprise customers. Requests sent through its US endpoint are decrypted, processed, and served inside the country, or rejected if that guarantee can’t be met.

Sawers wrote that story too. One number explains the demand: Open-weight models accounted for roughly 60% of OpenRouter’s US-originating token consumption in August, with Chinese models making up most of the volume.

DeepSeek V4 Pro, Kimi K3, GLM 5.2. These are models developed in China and available through providers running them in US data centers

The models’ country of origin and the location where your data gets processed are different questions. OpenRouter is selling control over the second one. That matters, since Deloitte’s global survey cited in the story found that 77% of companies factor an AI solution’s country of origin into vendor selection. Stripe’s announced acquisition of OpenRouter, reportedly worth about $8 billion, adds another measure of the interest in this layer.

Vercel’s September AI Gateway Production Index, published September 17, shows the economics from another angle:

Open weights took the volume. Closed models kept the money.

Share of tokens against share of spending on Vercel’s AI Gateway, August 2026.

Models Share of tokens Share of estimated spend
Open-weight models 56% 14%
Closed-weight models 44% 86%
Anthropic, all models n/a 64%
GPT-6 Astra, first 12 days after Sept. 3 launch n/a 7.7%

Source: Vercel AI Gateway Production Index, Sept. 17, 2026. Gateway traffic only; spending is estimated from labs’ published list prices, so actual bills may differ.

Open weight crossed into the majority of Vercel’s gateway token volume for the first time, up from a reported 7% in December 2025. Vercel says its current open-weight classification is broader than the one used in earlier reports. Among teams running more than ten million tokens in both comparison months, the median team paid 7.6% less per token, following July’s 2.9% decline.

So what justifies the premium? 

Boris Renski, CEO of AI agent integration company Apelogic, argues that much of it buys enterprise plumbing: identity integration, connectors, and observability. In Adrian Bridgwater’s July reporting, CNCF executive director Jonathan Bryce described paying ten times more for a four-month capability lead as “a very expensive form of lock-in.”

Closed models still command most of the estimated spending, and customers may be paying for capabilities that cheaper alternatives don’t reliably deliver.

But cheaper inference raises the pressure on everything around it, and Anthropic spent its week on exactly that. Instead of making headlines with a new model, it shipped a merged interface, documents and slides — features arguably more important to most users.

The harness has to earn its keep

Amanda Caswell covered Real-SWE, a benchmark from Y Combinator-backed Specific Labs built on private codebases from real companies. 

The best setup tested, Claude Fable 5.1 running through Claude Code, succeeded 38.8% of the time. GPT-6 Astra through Codex CLI reached 33.8%; Gemini 3.8 Flash through Gemini CLI reached 31.2%. None broke 40%. 

The benchmark is small: 10 tasks, with eight attempts per model on each. And it tests models with their coding tools, so the harness is already in the score. Still, no model solved the analytics stream reducer across 64 attempts. Zero.

Solutions touched a median of 11 files, compared with six in the public benchmarks cited in the reporting. Real work is spread out. These agents struggled to follow it. 

Fable’s leading failure categories were missed requirements and integration errors. That doesn’t mean the information was missing. An agent can have it and still overlook it, misunderstand it, or skip the check.

The engineering problem is getting context, tools, and verification to work together. So is knowing when the model doesn’t need to run at all. 

Abhilash Rao Mesala, a senior data engineer at Meta, wrote a practical guide to LLM response caching for us: reuse an answer when the request, context, permissions, and underlying information make it valid. His example starts with a million calls a month at $0.006 each, or $6,000. A 60% cache hit rate, plus $150 in embedding and vector-store costs, brings that to $2,550. A 57.5% reduction from a principle older than the transformer.

The hard part is knowing when a stored answer is still the right answer.

Pair Mesala’s piece with Ida Silfverskiöld’s on Towards Data Science guide to saving on tokens, which covers prompt caching, model routing, and keeping unnecessary materials out of an agent’s context. Both are worth your time. 

It’s hard to ignore the doom and gloom hovering over AI at the moment. But there’s good news out there, too. Each week, the price of inference is dropping, and companies are improving harnesses – both are critical improvements to push AI into the mainstream.

The post This week’s news from Zed, Anthropic, and OpenRouter shows why better harnesses matter more than better models appeared first on The New Stack.

  •  

Code review is burning out your best engineers

Dark, deeply fractured rock surface symbolizing structural tension and developer burnout

Every team I talk to has the same problem. Their best engineers, the ones who care most about code quality, are drowning in review queues they can’t keep up with and don’t enjoy. Some experienced engineers complain, and some point-blank refuse to review AI-generated code.

I run a community of senior engineers and engineering leaders, and they all name AI code review bottlenecks among their top concerns. In one study, 77% of engineers said they spend less time writing code now. They put that time into reviewing AI output.

The job shifted from crafting to verifying

Teams with high AI adoption are merging 98% more PRs, and review times went up 91%. Engineers didn’t sign up to spend their days reading machine-generated diffs. But that’s increasingly what the job demands.

The engineers who feel this most aren’t the ones resisting AI. They’re the ones who adopted it first, care most about code quality, and built the review culture their teams depend on. Those same people now have 15 PRs, 400 lines of code each, in their queue every day.

Why reviewing AI-generated code is harder

When a colleague writes code, the intent travels with them through the review process. They can explain the tradeoffs they considered, the alternatives they rejected, and the constraints they worked within. Even if unwritten, that context is accessible.

“When AI writes code, the reasoning is gone. The reviewer is left reverse-engineering intent from a diff.”

When AI writes code, the reasoning is gone. The reviewer is left reverse-engineering intent from a diff. That is a fundamentally different cognitive task. What makes it worse is that AI-generated code passes the eye test.

Five shades of AI slop

Plausible but wrong. The code reads coherently and handles the happy path, but edge cases reveal misaligned assumptions. These bugs are difficult to catch in review because they require understanding what the code was supposed to do, not just what it does.

Over-engineered. AI models are trained on vast bodies of code, including enterprise patterns and production-hardened architectures. Asked to solve a problem that really needs 15 lines, a model may produce a 200-line abstraction layer that anticipates a generality nobody asked for.

Convention-blind. Models generate good generic code, not code that fits your system. Your repo has conventions around naming, error handling, logging patterns, module boundaries. AI frequently ignores them.

Confidently hallucinated. Calls APIs that don’t exist, uses deprecated methods, invents config options. Sometimes caught immediately, sometimes only in production.

Cargo-cult patterns. Copies structures without understanding why. Retry logic where retries make no sense. Circuit breakers for calls that are always synchronous. Error handling that looks thorough but doesn’t map to actual failure modes.

The common thread is that it looks like real code, which makes it hard to review at scale.

How to fix the code review

The answer is not “review harder” or “add an LLM reviewer.” When the same model writes and reviews the code, it shares its own blind spots. If you add adversarial agents and multiple steps, the process becomes a theater of multi-step workflows that turns engineers into bot-sitters, spending time configuring and tuning filters instead of building.

What works is shifting the burden off reviewers in three places: codify repeated feedback, preserve the intent that produced the code, and measure the work that actually prevents slop.

Create your AI slop registry

Pull your team’s last 100 PR review comments. Sort each one: Is it deterministic, something a rule can check? Is it execution-testable, something you can catch by running the code? Or is it genuine judgment?

When teams run this exercise, the rough split is 45% deterministic, 30% execution-testable, and 25% judgment. Three-quarters of review feedback is codifiable.

“Three-quarters of review feedback is codifiable. Every recurring review comment is an invariant you haven’t written yet.”

Every recurring review comment is an invariant you haven’t written yet. “New endpoints must have OTel spans” is not a judgment call. It’s an AST check. Write it once. It never needs a reviewer again. The test for promoting something to an invariant is recurrence: if you’ve posted the same comment more than once, it should be codified.

Preserve the reasoning trail

The prompts and agent sessions that produced your code hold the intent. Most teams throw them away. That’s like deleting the commit messages and PR descriptions and expecting reviewers to reconstruct intent from diffs alone.

At Aviator, we built Verify around this problem. It captures intent from prompts and agent sessions and structures it as acceptance criteria: what the change does, what’s out of scope, and how to tell if it worked. The decisions an engineer makes while talking to the agent, the architectural choices, the scope calls, the behavior tradeoffs, become reviewable acceptance criteria.

The reviewer reads a list of acceptance criteria and asks, “Are we solving the right problem with the right constraints?” That’s the high-value work for senior engineers. Not reading a 400-line diff at 4 p.m. Code is actually the least important part of reviews. What matters is intent: acceptance criteria, non-goals, blast radius.

How knowledge sharing survives

Reviewers reading specs and acceptance criteria are reading decisions, not scanning syntax. They’re debating tradeoffs, understanding how the system is evolving, seeing what constraints shaped the approach. That’s where knowledge sharing survives. If we move code review left, knowledge sharing has to move left too.

Measure and reward verification work

31% more PRs are being merged without any review at all. That’s engineers voting with their behavior.

Dashboards measuring AI adoption and productivity in lines of code will never show the work of senior engineers carrying the review burden. They will never surface the effort that goes into building the systems and guardrails that prevent slop. If you’re measuring throughput and cycle time and feeling good, you’re measuring the wrong thing.

Annie Vella has been tracking this shift across 158 engineers in 28 countries. Her observation: engineers are resigning, some hoping the role will return to what it was, others leaving the profession entirely. The shift toward verification-heavy work is turning the job into something they don’t enjoy.

“Those dashboards don’t show the senior engineer who spent her afternoon reverse-engineering intent. They show throughput. And throughput looks great right up until the people carrying the review burden walk out the door.”

The engineers who carry the review burden aren’t complaining. They’re quitting. Some leave for teams with better tooling. Some leave engineering entirely, not because they can’t keep up, but because the work stopped being the work they signed up for.

Leaders chasing lines of code generated and PRs merged will never see this coming. Those dashboards don’t show the senior engineer who spent her afternoon reverse-engineering intent from a 400-line diff. They don’t show the review that caught a cargo-cult pattern before it hit production. They show throughput. And throughput looks great right up until the people carrying the review burden walk out the door.

Fix the code review process. Codify what’s repetitive, preserve the reasoning trail, and measure the work that actually prevents slop. Otherwise, you watch your best engineers leave and wonder why your AI-powered team ships faster but breaks more.

The post Code review is burning out your best engineers appeared first on The New Stack.

  •  

Anthropic’s new Claude Code feature could drain your plan before lunch

abstract rays

Anthropic is giving Claude Code a new job: Manage other Claude Code sessions.

Starting Thursday, Anthropic says select Claude Pro and Max subscribers can access the redesigned Projects in beta through cloud sessions in Claude Code, with access expanding to more users on those plans over the coming week.

The company is redesigning Claude Projects — which until now grouped chats around a shared knowledge base and instructions — with a coordinator that can take an engineering goal, break it into smaller jobs, and hand them to multiple Claude Code sessions running in parallel. This will eliminate the manual work that previously required developers to divide tasks between sessions and bring the results back together themselves.

“Claude scopes the request, delegates the work, coordinates parallel threads, reviews the outputs, and assembles the finished result,” an Anthropic rep tells The New Stack.

Once a developer sets a goal, Claude decides how to split the work across threads, which the developer can monitor — or open individually if they need to intervene. Each thread runs as its own Claude Code cloud session with a separate branch and copy of the repository. It can use subagents, loops, and workflows to handle smaller pieces of the job.

More sessions use more plan

That parallel approach can also use up tokens much faster. Anthropic says Projects will hit usage limits sooner when multiple threads are running because each one counts as a full Claude Code session.

Anthropic is facing a class-action lawsuit, filed last week, from developers who pay for its Max plan. They allege that advertised usage increases came with weekly ceilings that Anthropic didn’t disclose at first.

“Claude scopes the request, delegates the work, coordinates parallel threads, reviews the outputs, and assembles the finished result.”

One Claude coordinates the work

A team retiring an old API endpoint, for example, could connect its API, web, and mobile repositories and have Claude create a thread for each one to update callers, run tests, and open PRs before identifying which changes need to merge first.

Developers can follow the work through the main Project chat or open an individual thread to inspect or redirect it without having to start and manage each Claude Code session themselves.

The coordinator isn’t the last layer of delegation, as each worker thread can split its assignment further using Claude Code’s existing subagents, loops, and workflows when the job calls for it. That review layer may be more than cosmetic: on the Real-SWE benchmark, which tests coding agents against private codebases, even Anthropic’s own Fable 5.1, the top scorer, failed more than 60% of the time.

The finding suggests that coordination and output review matter as much as raw model capability when agents are dropped into unfamiliar production code.

Parallel threads, familiar conflicts

Each thread works from its own branch, which keeps the work separate but doesn’t prevent two threads from changing the same code. When that happens, Anthropic says it handles it as a merge conflict, just like any other pull request.

Each thread works from its own branch, which keeps the work separate but doesn’t prevent two threads from changing the same code.

Project background carries across threads

Additionally, Anthropic is adding shared memory so information learned in one thread can be used by the others without developers repeatedly supplying the same context.

Projects can retain details such as a changed release date, the reason a feature was dropped, or who needs to approve changes to a particular service, with that information carrying forward as work continues over days or weeks.

Claude can also remember how the developer wants the project managed, including how often it should check in, when it should start new threads, and how detailed its updates should be, while a new library keeps files added by the user alongside artifacts Claude produces so future work can draw on material already created within the project.

Users can track Project-specific usage and choose the model and effort level for the coordinator and worker threads, while support for running threads locally alongside a developer’s tools and code, including resources behind a private network, is coming soon.

The beta will expand to more Claude Code users on Pro and Max plans over the coming week, with mobile support coming soon and Team and Enterprise access planned for later.

Existing subscribers who already use Projects will remain on the current version until Anthropic upgrades them. The redesign is part of a broader product consolidation: Anthropic recently merged Claude chat and its Cowork interface into a single view, betting that removing the choice of mode would reduce friction rather than add it.

Existing subscribers who already use Projects will remain on the current version until Anthropic upgrades them.

The post Anthropic’s new Claude Code feature could drain your plan before lunch appeared first on The New Stack.

  •  

Automattic says CEO Mullenweg was gone and back inside 33 hours. What happened between?

There were unusual goings-on this month at Automattic, a company known for its free and open-source app for building WordPress sites. 

The company issued a notice last Thursday confirming that CEO Matt Mullenweg was on leave.

Mullenweg, who also co-founded WordPress in 2003 before establishing Automattic in 2005, was temporarily replaced by company CFO Mark Davies before being reinstated less than a day and a half later.

By last Saturday, a new alert emerged stating that Mullenweg was back in his position and that “Matt was away for only 33 hours and 20 minutes.”

Automattic has not publicly explained what changed between the initial decision and Mullenweg’s return — and it has also declined to explain the circumstances behind the leave and return. 

By way of context, WordPress sits under the Automattic brand alongside the company’s other products, including the microblogging site Tumblr, the e-commerce WordPress plug-in service WooCommerce, and the instant messaging client Beeper. 

A 33-hour vanishing act, but Automatticians are supporting him 

Automattic director of communications Megan Fox is on the record saying, “Matt Mullenweg is the chairman and CEO of Automattic, with full support of the board. “And if you search online, you can see many top executives and Automatticians supporting him as well.”

A further report reproduced Slack messages written by Mullenweg where he said, “Happy to announce the board is back in agreement, and I’m in control of Automattic. A lot happened in the past 48 hours that we need to sort out, and I hope much of it was a misunderstanding, because I have huge respect and regard for those involved.”

Was this a failed boardroom coup?

Industry watchers may naturally suspect the knives were out and that this was a failed boardroom coup. 

CEO & CTO at HasData, Roman Milyushkevich, tells The New Stack that a company can survive a CEO departure, but it struggles when employees cannot tell which governance process is real.

“But, in terms of whether the real guns were out at Automattic, it certainly looks like a failed attempt to change control, but I would not call it a coup as an established fact,” Milyushkevich says. 

“The possibilities playing out here are all very different,” Milyushkevich adds. “There could have been a second board agreement brought into place inside that 33 hours; there could have been internal (or possibly even external) negotiations; directors could have reconsidered the practical consequences of removing the founder; or there could have been an internal resolution that has not been disclosed.”

He advises that the “most important thing Automattic can establish now” is not who won the dispute, but whether the board and CEO have a clearly understood process for handling the next serious disagreement.

“The most important thing Automattic can establish now is not who won the dispute, but whether the board and CEO have a clearly understood process for handling the next serious disagreement.”

Behavioral scientist and visiting professor at São Paulo’s FIA Business School, Ricardo D’Olivar, tells The New Stack that what matters here is whether stakeholders have “enough information to distinguish a considered correction from an unresolved struggle” over authority. 

“A reversal of this kind can reflect responsible reconsideration,” D’Olivar says. “Reinstatement answers who is in charge today. It does not, by itself, explain how a future disagreement would be resolved. This is where the potential consequences for employees, executive recruitment and investors arise. If uncertainty persists, employees may become more cautious about committing to decisions whose backing appears unstable.”

Suggesting that, responsible corporate mechanics or not, this kind of development undoubtedly throws the cat among the pigeons, D’Olivar says that developers or executives considering a career at Automattic may now question whether they would be held responsible for decisions they were authorized to make, but could no longer count on the organization to support when challenged. 

“Investors may seek clearer evidence that oversight and succession arrangements can operate under pressure. These are possible responses, not verified effects at Automattic. The information shared internally may also be more complete than the public account,” adds D’Olivar.

This is not the first boomerang CEO bounce

Mullenweg might be the fastest CEO yo-yo switcharound in history, but he’s certainly not the first. OpenAI CEO Sam Altman famously left the company’s board after a communication dispute. An employee uprising (nearly all the company’s 700+ staff threatened to resign) and added pressure from Microsoft led to Altman returning to his position five days later.

Perhaps even more famously, Steve Jobs was ousted in 1985, only to return 12 years later to realign a then-struggling Apple and take it into its golden years. Twitter (now X) founder Jack Dorsey was moved out in 2008 before a boomerang return in 2015. Michael Dell stepped down as CEO to become chairman of the board in 2004, but by the start of 2007 he was back.

At their origins, the shenanigans at Automattic may be redolent of the technology industry’s other boomerang CEO realignments, or this may be a boardroom tussle that we’ll never know the full reason for until Mullenweg writes his memoirs. Either way, the WordPress industry just got its first movie-script idea.

The post Automattic says CEO Mullenweg was gone and back inside 33 hours. What happened between? appeared first on The New Stack.

  •  

“Everyone’s in a race to replace GitHub”: Zed launches Delta because agents made pull requests obsolete

A cardboard robot sat atop a laptop keyboard.

Something of a consensus has emerged from the developer fraternity in 2026 — GitHub, a platform built substantively for human developers, is no longer fit for purpose.

Part of the problem is sheer volume. Agents can generate, revise and submit code at a cadence GitHub wasn’t designed for, placing growing pressure on infrastructure built for humans working through commits, branches and pull requests. But there’s also a more fundamental question about the interaction model itself: when much of the reasoning behind a change happens inside a conversation with an agent, a pull request presents reviewers with the resulting diff while leaving much of the journey that produced it elsewhere.

At the heart of all of this, of course, is GitHub’s reliability problems. The platform logged hundreds of incidents over the 12 months leading into June, as monthly commit volume rocketed from around 1 billion across the whole of 2025 to 1.4 billion a month by April. By August, GitHub said this figure had jumped to 2.9 billion commits each month.

The growing load has manifested in some fairly spectacular outages, including a near-eight-hour disruption in August, with web and API error rates reaching around 20% at the height of the incident.

As Nathan Sobo, co-founder and CEO of developer platform company Zed, puts it in a blog post published on Wednesday, “everyone is in a race to replace GitHub right now,” with a number of players in the technology sphere working on alternative tooling. That includes Zed itself, which has announced the public beta of Delta — a collaborative environment where developers and coding agents work, review and revise code together in shared threads rather than pull requests.

At the heart of Zed’s pitch is the idea that the pull request is ill-suited to coding agents, and tells reviewers little about the reasoning behind their decisions.

“Since GitHub introduced pull requests over 15 years ago, they’ve become the standard way to ask teammates to review changes to your codebase.”

“Since GitHub introduced pull requests over 15 years ago, they’ve become the standard way to ask teammates to review changes to your codebase,” Sobo writes. “But with agents generating so much code, the diffs we’re asking each other to review have mushroomed.”

The question, then, is what collaboration should look like when agents are producing more of the code?

From A(tom) to Zed

Zed, for the uninitiated, started out with a somewhat narrower remit. Founded in 2021 by veterans of GitHub’s Atom editor team — Sobo himself spent nine years there — Zed emerged as a high-performance, multiplayer code editor built in Rust. When The New Stack tested the beta in 2023, the emphasis was on responsiveness and real-time collaboration; by 2025, AI editing and agentic features had become central to the product.

It has been clear for some time, however, that Zed’s ambitions extend beyond the editor. When the company announced a $32 million round of funding led by Sequoia Capital in August 2025, it also teased DeltaDB, a new kind of operation-based version control system designed to record code changes at edit-level granularity.

Fast-forward to August, and Zed revealed Delta itself in private beta, pitching it as a multiplayer environment where developers can code with agents, share their ongoing threads with teammates and review changes with the original agent context intact.

Multiple participants iterating on a prompt.
Multiple participants iterating on a prompt.

Wednesday’s public beta launch brings that idea out into the open, and takes direct aim at one of GitHub’s defining features: the pull request.

Picking up the thread

The central concept behind Delta is the thread: a running record of an agent-assisted coding task in which the conversation and the files being changed remain connected. A developer can hand an agent a job, continue discussing and refining it, and later share that entire body of work with somebody else.

Each thread can work against its own copy of a project, which means multiple pieces of work can proceed independently without every agent touching the same checked-out files. Teammates can join an existing thread or create a separate review thread to examine a proposed change, question the agent that produced it and try revisions before feeding accepted changes back into the original work.

DeltaDB sits underneath that model, recording activity at a finer level than Git. Instead of waiting for a developer to package work into a commit, it captures individual events as they happen — including code edits and activity within the conversation — and uses that history to keep participants synchronized.

Zed calls those individual records “deltas.”

Git hasn’t gone the way of the dodo quite yet, though. Delta currently works with Git repositories, and developers can continue using branches, commits and remotes as usual. DeltaDB effectively adds another layer of history between commits, preserving the intermediate human and agent activity that Git would otherwise discard.

“We now build and collaborate on Delta entirely within Delta.”

Sobo notes that Zed has already disabled pull requests on Delta’s own repository and now develops the product through Delta threads instead. Since doing so, the company says 33 developers have landed 570 changes to main without using pull requests.

“We now build and collaborate on Delta entirely within Delta,” he writes.

Zed disables PRs on its own Delta repo
Zed disables PRs on its own Delta repo

An intermediate step

It’s worth noting that this is still very much an intermediate step. And for Zed’s own internal development, Sobo expects that intermediary period to be fairly brief. He says the company is only “a few months away” from leaving GitHub behind, with developers focused exclusively on Delta already having little reason to visit GitHub because their conversations, reviews and handoffs now take place inside Delta.

There are still some pieces to disentangle. Zed’s next major dependency is Git storage, which it intends to bring into DeltaDB, while CI and releases also need to move away from GitHub. Sobo sees CI as largely a solved problem, however, and says Zed expects to integrate with existing options rather than build another system simply for the sake of replacing GitHub.

Zed’s main open-source code editor repository will remain in place for now, where an established contributor community already reports problems and proposes code changes. Those contributors can use Delta to expose the agent session behind their work, while still submitting the final change through a conventional pull request, leaving the Git experience unchanged for collaborators who don’t use Delta.

That public repository presents a different problem from Zed’s internal development, given the hundreds of external developers contributing to the project each month.

“We’re moving more thoughtfully with Zed’s public repo because we have hundreds of monthly contributors who depend on that workflow,” Sobo tells The New Stack. “GitHub has an established social component that will take longer to replace, and we’re not going to strand contributors to prove a point. We’re only going to move our community layer when we can offer something better.”

There are other signs of that continued dependency. One item currently on Delta’s roadmap is repository-based access, which will use a GitHub repository’s existing permissions to determine who can access shared Delta threads.

Delta is available as a desktop app for macOS, Linux and Windows, with a browser version for viewing, sharing and reviewing threads. It will remain free throughout the public beta, with paid individual and team plans to follow; Zed says there will always be a free version.

A common thread: Reinventing code collaboration

Of course, Zed is far from alone in tyring to reinvent code collaboration for the agentic era. SpaceX-owned Cursor formally launched Origin in August, bringing Git repository hosting, pull requests and its coding agents under the same roof, while still allowing existing GitHub repositories to remain the source of truth.

GitLab, too, is working on a “next-generation source-code management” project dubbed Project Switch, currently in private beta.

“I believe the thread will replace the commit or branch as the fundamental unit of software development.”

For Sobo, Delta’s claim to differentiation starts with the basic unit around which it’s being built.

“I believe the thread will replace the commit or branch as the fundamental unit of software development,” Sobo explains.

Commits, in his view, will continue to provide useful checkpoints. What they don’t preserve is everything that happens between them: the discussion with an agent, the decisions made along the way and the incremental changes that eventually produce the committed code. That’s the gap Zed wants DeltaDB to fill, with the thread rather than the commit serving as the fuller record of how a piece of software came together.

“Delta’s advantage is that we’re building around the thread from the start, including the infrastructure underneath it,” Sobo continues.

Sobo argues that preserving incremental edits alongside the agent conversation gives subsequent collaborators something a conventional diff cannot: the ability to enter the work where the previous developer left it, and continue interacting with the same agent and context, rather than reconstructing the thinking behind a change after the fact.

Delta’s bet is that the central object of software development should become the ongoing interaction between developer and agent, with the resulting code attached to that history rather than presented later as an isolated diff.

“I expect lots of viable products will compete on the agent experience, with a common infrastructure underneath.”

Sobo expects the next era to echo the structure of the Git era, with competing developer platforms built on top of a common technical foundation.

“On consolidation, we look at how the last era played out. Git became the common foundation because it was open and everyone could build on it,” Sobo says. “GitHub won the layer above through network effects. I expect lots of viable products will compete on the agent experience, with a common infrastructure underneath.”

The post “Everyone’s in a race to replace GitHub”: Zed launches Delta because agents made pull requests obsolete appeared first on The New Stack.

  •  

Anthropic bet users were choosing wrong. So it removed the choice.

Single lane

Using Claude for anything beyond a quick question has always started with a routing decision to use Chat or Cowork? Anthropic has decided to eliminate that fork.

Starting Wednesday, Claude Chat and Cowork merge into a single interface where one conversation can handle everything from a simple answer to a multi-step project with connected tools and background execution. The company is also launching Claude Docs and Claude Slides in beta on paid plans, and moving Claude Design — previously a standalone workspace — into conversations.

The combined effect promises a streamlined experience with Claude picking up context, skills, and connectors as the work requires, and can keep running after you close your laptop.

The company is also launching Claude Docs and Claude Slides in beta on paid plans, and moving Claude Design, previously a standalone workspace, into conversations.

Two modes, one problem

Anthropic built Cowork as a desktop-first agent for bigger work and Design as a separate workspace for visual output. Both shipped earlier this year and gained traction.

“We built Cowork as a separate place for bigger work, and Design for visual work,” Anthropic said in its announcement. “People used both, and told us the frustrating part was deciding where a task belonged.”

Anthropic has run into this problem before. When the company promised 20x more usage on its Max plan, developers complained that it wasn’t always clear where one limit ended and another began. Cowork and Design created a similar headache by making people decide where to start the work before they could actually start it.

Context didn’t always follow the work either, so moving from chat to Cowork or Design could mean bringing the same background along all over again. Anthropic addressed part of this in August when it unified Claude’s memory across chat and Cowork. Wednesday’s change goes further by merging the products themselves.

“We built Cowork as a separate place for bigger work, and Design for visual work,”

Context that finally travels

Cowork’s capabilities — local file access, multi-step execution, scheduled tasks and connected tools — now live inside the conversation. A workflow like that previously meant switching from chat to Cowork and carrying the context with it, and until Anthropic brought Cowork to web and mobile in July, it also required the desktop app.

Claude still asks before taking an action by default, but it can be set to keep working and check in only when something needs a closer look, while recurring tasks such as a weekly report can be scheduled to run every Monday without being started manually.

Output stays in-conversation

Claude Docs and Claude Slides launch in beta on paid plans, bringing document editing and presentation building directly into the app. Claude can turn work from an existing conversation into slides, which can then be edited, presented from Claude or downloaded as PowerPoint or PDF files.

The practical benefit is that a report and a slide deck based on it don’t have to begin as separate jobs with the same background supplied twice. Everything stays attached to the conversation that produced it.

Claude Design also now works inside conversations, in addition to remaining available on its own. For organizations that rely on MCP connectors to wire Claude into external tools and data, the merge means those connections are available wherever a conversation goes — without requiring users to start in a specific mode. Skills, connectors, and artifacts carry across what used to be product boundaries.

The practical benefit is that a report and a slide deck based on it don’t have to begin as separate jobs with the same background supplied twice.

What Anthropic hasn’t said

The announcement leaves some gaps. It doesn’t say whether users can force a request to stay in simple chat mode rather than letting Claude decide how to handle it, or how that decision affects context windows and token consumption. There’s no mention of API changes, which makes this a consumer and team product shift, not a platform one, at least for now.

It also doesn’t address what happens to workflows built around the old separation. Shopify rebuilt its mobile development stack in 12 weeks when it consolidated tools that had grown apart — the question for Claude power users is whether their existing Cowork setups, skills, and scheduled tasks survive the merge cleanly. Anthropic says existing Cowork chats, projects, artifacts, connectors, and skills will remain available.

Rollout starts with Pro

The unified interface rolls out to Pro and Max users across web, desktop, and mobile over the next few weeks. Anthropic says there’s nothing to enable. Team and Free plans follow. Enterprise customers are on a separate timeline; Anthropic is giving administrators at least 30 days’ notice before the change reaches their organizations.

The post Anthropic bet users were choosing wrong. So it removed the choice. appeared first on The New Stack.

  •  

OpenAI president: “The computer should be there to empower you.” So stop retooling software for AI agents

This week on the a16z show, Greg Brockman, president and co-founder of OpenAI, made the point that developers have been “retooling the world” to make software more accessible to AI agents. But computer use could offer a simpler path forward. 

And for Brockman, simplicity is really the North Star for AI: “We really should be one AI that’s unified, that makes it so easy and smooth for you to be less engaging with the computer and less wrapping yourself around the computer.” 

Meanwhile, other AI companies continue to invest in connector infrastructure. 

“We’re kind of retooling the world,” Brockman says. Is it time to stop? 

MCP servers, CLIs, APIs, and other agent integrations help AI agents reach more tools, data, and software. But that wide access doesn’t come without some burden — at least that’s what it sounds like in Brockman’s conversation with Ben Horowitz and Erik Torenberg, hosts of the a16z show. As OpenAI’s president explains:

“What if it’s more behaving like a human? Can it just use a computer?”

“People have been building these MCP servers and these CLIs and just really sort of taking the world of software and making it accessible in this almost stilted way that is not really meant for humans,” he said. “It’s like we’re kind of retooling the world.”

How did we get here? “For agentic use cases, it really comes down to the tools,” said Brockman, indicating why MCP servers and other connectors have become so common. But if models can learn to use computers the same way that people do, then perhaps developers no longer need to keep building purpose-built integrations for every application. 

As Brockman wonders: “What if it’s more behaving like a human? Can it just use a computer?”

OpenAI has been thinking about computer use since the beginning

As Brockman details on the podcast, the idea traces back to the early days of the AI company when the team laid out a three-step plan during an offsite meeting in November 2015 that the president says the team largely stuck to for the following 10 years. 

Around this time, Brockman says the OpenAI team also broached the idea of using reinforcement learning directly against a computer interface:

“We also talked about, what if we could do reinforcement learning where the environment is screen pixels, keyboard, mouse, right? Same interface as a human.”

From where he sits, this could open up the computer for nearly any sort of task. More importantly, Brockman muses that computer use for agents could help the broader industry advance AI without requiring purpose-built integrations every step of the way. Instead of building and maintaining scores of specific connectors, an agent could simply use the same computer interface people do. 

“There’s so much software that you don’t even think about,” he told Horowitz and Torenberg. “How much of your life is clicking around menus and typing things into a spreadsheet and things like that? None of that is what we should be doing.”

Instead of building and maintaining scores of specific connectors, an agent could simply use the same computer interface people do. 

In fact, Brockman says computer use capabilities are part of why he considers GPT-6 Astra, OpenAI’s newest flagship model, “pretty reasonable to call” AGI. 

But connectors aren’t disappearing yet

Though Brockman is bullish on computer use, the rest of the industry isn’t giving up on structured integrations. In fact, some are plowing ahead. 

AWS, for example, recently launched a managed consent portal, a managed web experience and session binding endpoint for AgentCore Gateway, a capability of Amazon Bedrock AgentCore that connects agents to external tools and services. On Monday, the cloud company published a detailed walkthrough of the feature. 

For example, with the Consent portal, an administrator could configure GitHub and Slack as gateway targets and send the portal URL to developers who can then open the URL, sign in with their corporate IdP, and connect GitHub and Slack as needed, independently. 

Meanwhile, even at OpenAI, plugins aren’t going away. Though OpenAI Codex arrived in the browser with a new Chrome extension back in May that allowed agents to operate within live browser sessions and authenticated workflows across multiple tabs, plugins remained the preferred route because they allowed Codex to work directly with services, such as Slack, Gmail, and GitHub, without manually navigating their interfaces.

Brockman told the a16z show hosts that simplicity should be the North Star. Reliable computer use could be one way to get there. But for now, connectors are still hanging in there.

The post OpenAI president: “The computer should be there to empower you.” So stop retooling software for AI agents appeared first on The New Stack.

  •  

Meta lets Claude and Codex configure WhatsApp Business via MCP. But the agents don’t get their own identity.

A concept illustration depicting AI running a business

Any business worth its salt in 2026 needs to be embracing the right tools to reach its customers, and few tools carry as much weight as WhatsApp.

Paid messaging on the app crossed a $2 billion annual run rate in the fourth quarter of 2025, CFO Susan Li told investors on Meta’s January earnings call. But getting a business properly set up on WhatsApp can still be a fiddly job. The initial onboarding can send developers jumping between Meta’s account settings, API documentation, and code editor as they connect and verify a phone number, configure webhooks, and get the integration working. Some of that is a one-off job, but things like managing message templates, testing changes, and troubleshooting the setup can bring developers back to those same tools later.

Meta’s answer, announced today, is to let AI coding agents handle much of that setup directly through MCP.

Connecting coding agents with WhatsApp

The easiest way to understand what the WhatsApp Business Tools MCP server is all about, is to look at the sort of job a developer might want to hand over to an agent.

Take an online retailer that wants to use WhatsApp for customer support, order updates or the occasional special offer. A developer can connect the new MCP server to an agent such as Claude or Codex, sign in with their Meta account, and choose which of the businesses they already administer the agent is allowed to access. That access is scoped to whatever businesses they select — connecting the agent doesn’t give it free rein across every Meta account associated with the developer.

Connecting Claude to WhatsApp
Connecting Claude to WhatsApp

Give the agent the number and the display name the business wants to use, and it can handle the steps needed to add the number and set up the WhatsApp account behind it. Meta then sends a one-time verification code by SMS or voice; once the developer gives that code back to the agent, the number can be verified and registered for sending messages.

Adding a WhatsApp number
Adding a WhatsApp number

In a blog post announcing the feature on Tuesday, Zoë Lieberman, who works on product marketing at Meta, says the idea, ultimately, is to turn what might otherwise be a string of separate API tasks into something the user can ask an agent to do in plain English.

“You describe what you want — your agent handles the accounts, numbers, templates, and API calls.”

“You describe what you want — your agent handles the accounts, numbers, templates, and API calls,” Lieberman writes.

A template example

Much of the WhatsApp Business Tools MCP server is concerned with getting a business up and running on WhatsApp in the first place. Message templates, however, show how the agent can remain useful once that initial setup is done.

There is a WhatsApp rule worth explaining, though. When a customer messages a business, it opens a 24-hour customer service window, during which the business can reply with ordinary, free-form messages. Each new message from the customer starts that 24-hour clock again. The idea is to stop businesses turning an old customer conversation into an open-ended channel for unsolicited messages: once the window has closed, the business generally needs to use a message template that Meta has approved if it wants to contact that customer again.

So the retailer could ask the agent to create a marketing template offering customers a coupon, for example, or a utility template for sending order updates. The template can include elements such as a header, body copy, footer and buttons.

Creating a marketing template
Creating a marketing template

From the same conversation, the person using the agent can list the business’s existing templates, pull up a particular version, update it or delete it.

Once the pieces are in place, the developer can ask the agent to send a test message and check that everything behaves as expected before putting it in front of customers. If the recipient is outside the 24-hour customer service window, the agent can flag that a free-form message can’t be sent and offer an approved template instead.

Sending a test message
Sending a test message

There is also the other half of a WhatsApp conversation to deal with: what happens when the customer replies?

The developer can ask the agent to configure the webhook that tells WhatsApp where to send those incoming messages and other events, such as the retailer’s CRM, customer-support platform, chatbot or order-management system.

Configuring a WhatsApp webhook
Configuring a WhatsApp webhook

The agent can inspect the account as well as change it. That means asking what has already been configured or what still needs attention — for example, whether the business is missing the payment information Meta needs to charge for billable WhatsApp messages.

Checking WhatsApp account setup
Checking WhatsApp account setup

There are some guardrails around all of this. The agent operates using the access of the person who connected it, and Meta says actions performed through the MCP are recorded.

“Every read runs under your own viewer context, every invocation is logged, and anything that changes state requires an authenticated person rather than an app-level credential.”

“Every read runs under your own viewer context, every invocation is logged, and anything that changes state requires an authenticated person rather than an app-level credential,” Lieberman writes.

Meta’s MCP push stays tied to human identity

That human-bound approach lands amid a broader debate over how AI agents should identify themselves. Agents today often inherit the permissions of the person using them, while some companies are pushing toward giving agents their own scoped, revocable identities — evidenced by Vercel’s recent acquisition of Better Auth. The concept is also showing up in the big cloud platforms, too, such as with Microsoft’s Entra Agent ID, which automatically assigns an identity to agents created in Azure AI Foundry or Copilot Studio. And Amazon Bedrock AgentCore Identity, meanwhile, gives each agent its own identity, letting it act on a user’s behalf or independently, without borrowing anyone’s login.

So while there is a clear push toward giving agents their own identity, Meta is keeping the human firmly in the loop here, The agent can act only within the authenticated user’s existing access, keeping permissions bounded by a known human account and making sensitive changes easier to attribute and control.

It’s also worth noting that this isn’t Meta’s first attempt to put its developer tools within reach of AI agents. In June, the company launched Developer Tools MCP, later rebranded as Meta Social Technologies MCP, which works across Meta’s developer platform — including integrations built with the WhatsApp Business API.

There is some overlap between the two, but they operate at different levels. Meta Social Technologies MCP is the broader developer tool: an agent can use it to discover Graph API endpoints, search Meta’s documentation, inspect the app behind an integration and troubleshoot errors. That can absolutely include a developer building with WhatsApp.

WhatsApp Business Tools MCP, meanwhile, is much more specific to operating WhatsApp Business itself. It gives the agent tools for working with WhatsApp Business accounts, onboarding and verifying phone numbers, creating and managing message templates, configuring and testing webhooks, and sending test messages.

“They’re complementary — install both if you need both,” Lieberman writes.

It’s early days for the new WhatsApp server. Meta says it is rolling out gradually, meaning it may not be available to everyone immediately, while the interface and tools themselves remain in beta and are subject to change.

The post Meta lets Claude and Codex configure WhatsApp Business via MCP. But the agents don’t get their own identity. appeared first on The New Stack.

  •  

Bolt is giving developers 50x more compute. But there’s a catch.

Abstract glitlch

Bolt.new, StackBlitz’s browser-based AI development platform, is testing a new trade with developers: more coding-model usage in exchange for training data.

The company launched Forge on Monday, a research preview for individual Pro subscribers that offers up to 50 times more usage of open weight coding models through October 14. Developers who use Forge must opt in to sharing anonymized versions of their sessions for model training, including prompts, source code and the fix traces it creates as developers work through problems, in addition to their conversations with the coding agent.

The sessions will be used in a project with Arcee AI to help train a trillion-parameter-class open-weight model. The first training run is scheduled to begin in October, with Bolt saying the resulting model weights will eventually be released publicly.

Forge makes that development activity part of the exchange, with developers getting more compute while Bolt and Arcee get data from real coding sessions.

Forge makes that development activity part of the exchange, with developers getting more compute while Bolt and Arcee get data from real coding sessions.

Why Coding Trajectories Matter

Public repositories contain enormous amounts of source code, but they mostly show the end result. A coding session can fill in the gaps left by failed attempts and revisions along the way. The record of what worked (and what did not) is useful as the coding agents take on longer jobs.

Working across a codebase means finding the right files, coordinating changes, and recovering when something breaks. That gets harder when agents inherit code written by other agents, which isn’t always easy for the next one to understand or modify.

SpaceXAI showed one version of this approach last month when it trained Grok 4.6 on agent failure traces — the missteps, retries, and corrections that other labs typically discard. Bolt is making a similar bet but sourcing the data from developer sessions rather than synthetic runs.

Arcee has been working on the same underlying problem. In a blog post about NAC, its open-source agent harness, the company said software engineering tasks can stretch across tens of thousands of tokens as agents read code, edit files, run tests, and debug failures.

How Bolt Gets to 50X More Usage

Coding agents can burn through large numbers of tokens on even a single complex task, so a 50-fold increase in usage is a truly significant offer.

Forge changes the underlying setup by running open weight models on Bolt’s own infrastructure. The agent currently uses GLM 5.3 Flash and GLM 5.3, with Kimi K3 and DeepSeek v4 Pro available as experimental options.

Bolt’s WebContainers technology, built by parent company StackBlitz, gives it another cost advantage by running projects in an isolated environment inside the user’s browser rather than on Bolt’s servers.

Forge applies a similar approach to the models, using open weights on reserved hardware while developer sessions help train future versions, giving Bolt more control over costs and reducing its reliance on proprietary APIs.

That push toward self-hosted models is showing up elsewhere in the industry. Nvidia’s $12.9 billion bid for Hugging Face is arguably the same bet at a very different scale.

Coding agents can burn through large numbers of tokens on even a single complex task, so a 50-fold increase in usage is a truly significant offer.

Forge Scores 91% of Bolt’s Top Model

Forge’s open models scored 92.2 on the company’s internal Bolt Build Index, compared with 101.0 for its top paid model, putting them at about 91% of the top score. That’s only a measure of performance inside Bolt, so the 91% figure doesn’t tell us how those models compare more broadly.

But if Bolt can run more of its coding workloads on its own models instead of paying for proprietary APIs, it has more control over costs and usage, while the Forge sessions help train whatever comes next.

What Developers Are Giving Up

Forge requires an explicit opt-in, with a consent screen appearing each time a developer switches into the workspace. Standard and Max sessions aren’t included, and Teams and Enterprise accounts can’t participate.

Bolt says it anonymizes sessions before they leave its infrastructure, removing secrets, sensitive data, and personal information, and it tests the process against seeded data. Arcee receives the resulting data under a signed processing agreement.

Developers can stop sharing new sessions by leaving Forge, but Bolt says anything already used for training will remain in the models.

The 50× surge ends October 14, though Bolt says Forge itself will stick around as an open-model testing ground once the preview wraps up.

The 50× surge ends October 14, though Bolt says Forge itself will stick around as an open-model testing ground once the preview wraps up.

The post Bolt is giving developers 50x more compute. But there’s a catch. appeared first on The New Stack.

  •