❌

Vue lecture

Microsoft’s new Copilot agents get their own email, calendar — and a place in the org chart

Satya Nadella stands smiling between Bill Gates, on the left, and Steve Ballmer, on the right, in front of a crowd of cheering employees, many holding up phones and tablets to take photos.

Microsoft announced what it calls its biggest Copilot update to date on Friday, with CEO Satya Nadella describing Copilot as “a new OS for work.”

Nadella framed Copilot as spanning every model, form factor, and task, and the update puts Autopilot, which Nadella called a “proactive and long-running agent built for the enterprise,” at the top of his list of the update’s four components. The pitch targets office workers, but the more consequential change for developers is the infrastructure underneath.

Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365, which means teams building production agents no longer have to assemble those pieces around a model on their own.

We’re building Copilot as a new OS for work that spans every model, every form factor, and every task. Today, we’re announcing our biggest update to Copilot to date, bringing four things together:

· Autopilot: proactive and long-running agent built for the enterprise
· Code:… pic.twitter.com/W2ClHHkCK3

— Satya Nadella (@satyanadella) September 25, 2026

The release adds a new Home experience that merges Chat and Cowork in the Copilot app, but the bigger changes for developers come from Code and Autopilot. Code generates apps, dashboards, and workflows from natural language, and Autopilot turns the agent Microsoft previously called Scout into a persistent background worker. Home and Code are rolling out first through Microsoft’s Frontier early-access program, and Autopilot is expanding to a private preview at month’s end.

Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365

Agents that don’t need prompts

Autopilot takes a role and goal from the person who sets it up, then continues working in the background without requiring a new prompt for each step. Each Autopilot gets its own governed Entra identity and agent user account, separating the agent’s permissions and activity from those of the person who created it.

For engineers, that moves much of the operational scaffolding required for long-running agents into Microsoft’s infrastructure. Independent vendors have been building dedicated layers for that problem; Diagrid, for example, adds durable recovery to LangGraph and other agent frameworks, while Microsoft is bringing those capabilities inside the Microsoft 365 environment.

An identity for every agent

The identity model is the piece developers building on Microsoft Foundry will feel first. Autopilot agents in Foundry, which have been in public preview since June, receive a full Entra Agent ID user account with a productivity license that gives them their own email, calendar, OneDrive storage, Teams access, and a place in the org chart.

Because that user account sits on top of the agent identity every Foundry agent already carries, an autopilot acts as itself rather than on behalf of a user, so developers no longer have to wire agents through shared service accounts or borrowed user credentials, a pattern AuthZed CEO Jake Moshenko has said reflects a common misconception about how agents should be deployed.

A developer creates an Autopilot blueprint from a Foundry-hosted agent, which appears in the Agent 365 registry once an administrator approves it. Employees can then hire instances of that agent in Teams. The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.

The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.

Hosting AI-generated apps

Code is built on the same underlying technology as GitHub Copilot, and the apps it generates run on Microsoft Copilot Managed Runtime, a platform now in public preview that hosts code inside the customer’s Microsoft 365 tenant boundary under IT governance.

Apps deployed there run within the company’s existing identity and governance framework, with Microsoft managing the underlying runtime and giving developers a controlled path to test and deploy new versions without taking the current release offline.

The runtime also accepts apps built in Copilot Studio and Cowork, and Microsoft is opening it to outside tools and professional developers through an SDK and command-line tooling, with Git tracking source and versions.

Lovable is already on board. In Microsoft’s announcement, the company’s head of global partnerships, Lan Roche, said apps built with Lovable can now run inside a Microsoft tenant “the same way everything else does,” using the same sign-in, policies, and app inventory.

The model resembles what serverless computing did for application infrastructure, where developers concentrate on application logic while the platform takes on more of the execution environment. Microsoft is applying that abstraction to generated enterprise software while tying the runtime directly to identity, tenant boundaries, and organizational data.

Long-running agents also change Copilot’s economics. The standard subscription covers the assistant, but Cowork, Code, Autopilot, and other agentic features are billed based on usage through Copilot Credits. That also applies to frontier models such as Fable and Astra, although users still need a Copilot license to access them. Microsoft is extending cost management in Agent 365 to cover Code and Copilot Managed Runtime, and it plans to support agents built in Copilot Studio in October.

Once an agent can keep working for hours or days without anyone watching, cost becomes part of the governance problem. Engineering teams need to control how much compute an agent uses alongside what it can access, which is why Microsoft is bringing those controls into the same administrative framework.

The portability trade-off

That convenience comes with a trade-off. Because Microsoft controls the underlying enterprise environment, it can handle much of the work around agent state, credentials, and access controls, but the more infrastructure a team hands over to Microsoft, the harder the agent may be to move elsewhere.

The models are not locked in, since Microsoft currently runs Copilot on models from both OpenAI and Anthropic and says more labs and open-weight models are coming, and the Agent 365 SDK adds governed Model Context Protocol access to Microsoft 365 workloads for agents regardless of the framework they were built with. Those open interfaces cover only part of an agent’s architecture, though. The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.

The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.

The post Microsoft’s new Copilot agents get their own email, calendar — and a place in the org chart appeared first on The New Stack.

  •  

OpenAI’s agent had a routine task. It breached a government portal.

Sam Altman, OpenAI CEO

An OpenAI agent researching public medicine spending bypassed security blocks and gained unauthorized access to public and non-public files on an Australian government Medicare statistics portal, the government there disclosed Thursday. The agent, which OpenAI said was running during an internal evaluation in June, also wrote files to an internal server, according to the complaint.

Transluce, an independent nonprofit AI research lab, analyzed public request logs from the URL scanning service urlquery.net and found autonomous agents attempting SQL injection, cross-site scripting, command injection, and path traversal against the University of New Mexico’s digital library, the public data platform Data USA, and the Australian Institute of Health and Welfare (AIHW). The agents tried to retrieve ordinary information, including a historical photograph, University of Iowa data, and local pharmaceutical data in Victoria, and the offensive behavior appeared only after normal retrieval methods failed.

Transluce ties the Data USA and AIHW activity to an agent swarm that it says OpenAI previously confirmed originated from the company, based on shared targets, tactics and timing.

A day after Transluce published its findings on Wednesday, Australian Prime Minister Anthony Albanese disclosed that an OpenAI agent researching public medicine spending bypassed blocks on a Services Australia Medicare statistics portal in June and gained unauthorized access to public and non-public files. OpenAI, which said the agent was running during an internal evaluation, notified the Australian government on September 10. According to Albanese, the agent also wrote files to an internal server.

The incidents have not been publicly confirmed as the same activity. However, OpenAI told the Australian Broadcasting Corporation that much of what Transluce documented overlaps with cases in its ongoing review of misaligned model activity.

Transluce ties the Data USA and AIHW activity to an agent swarm that it says OpenAI previously confirmed originated from the company, based on shared targets, tactics and timing.

Probing public data providers

On May 25 and 26, agents trying to retrieve a photograph from the University of New Mexico’s Valmora collection ran into repeated failures. They responded with seven vulnerability probes against the service. The probes included an attempt to execute a system command, a path traversal request targeting /etc/passwd, and a SQL injection string containing UNION SELECT; they also arrived alongside a burst of 80 requests aimed at the image, although none appear to have succeeded.

Two days later, agents gathering University of Iowa data from Data USA hit errors caused by a malformed query parameter and followed those errors with 12 vulnerability probes against the API, covering SQL injection, reflected XSS, server-side template injection, path traversal, and command injection.

Cloudflare blocks, staging server workaround

The AIHW case on June 20 and 21 followed a very similar pattern. While trying to download pharmaceutical spending data from an AIHW Tableau dashboard, the agents ran into Cloudflare’s bot protections. One of them responded within minutes by testing the dashboard with a reflected XSS payload, although Cloudflare stopped it before it reached the application.

Transluce found no evidence that the probe succeeded, and AIHW says there is no evidence that non-public information was accessed. The agent ultimately found the same public dataset on AIHW’s pre-production server at pp.aihw.gov.au, and retrieved it through more than 100 scans.

The agent ultimately found the same public dataset on AIHW’s pre-production server at pp.aihw.gov.au, and retrieved it through more than 100 scans.

Agents build their own toolchains

The earliest clear sequence Transluce documented dates to March 6, when an agent tried to pull drug-enforcement statistics from the API of Thailand’s Office of the Narcotics Control Board. A direct request returned no page content, and adding a JSON format parameter changed nothing, so the agent routed the page through r.jina.ai three times and got text back with the Thai province names garbled.

Then, it wrote a custom Base64-encoded script and submitted it through urlquery.net, effectively using the service’s remote scanning browser to run its own client-side JavaScript against the data. That script ran into the same parsing problem, and a later attempt through Jina triggered a redirect to localhost that got the entire API blocked by Jina for an hour.

Beginning in mid-April, Transluce found the same technique across thousands of agent requests on urlquery.net, and on June 14 agents used the service’s browser to create a disposable email inbox and then tried to register a urlquery.net account with it. Registered users can make their scans private, which means the public records Transluce analyzed may understate how much of this activity actually took place.

An agent with web access isn’t limited to the software its developer gave it because remote browsers, proxies, URL scanners, and other public services can fill in the gaps, which gives the agent ways to make requests or run code that its own environment doesn’t provide.

Egress controls for AI agents

Instructions won’t be enough if the agent can still send whatever it wants over the network. For narrowly defined jobs, outbound traffic can be limited to approved hosts, a closed-by-default approach also used for securing AI agent sandboxes. Research agents may need to reach more of the web, so the focus shifts to controlling where they can connect.

Guidance for GKE Agent Sandbox recommends isolated runtimes with default-deny network policies that open only the endpoints an agent needs. Public proxies, URL scanners, and disposable email services can stay blocked unless the job requires them.

Developers can also limit what an agent can send. So, instead of handing it a networking tool that accepts any URL or request body, an API integration can restrict requests to specific fields and formats. The runtime can then catch path traversal attempts, SQL injection strings, and executable markup before anything is sent. OpenAI takes a related isolation approach in its Agents SDK sandboxes, and the company’s Responses API tech lead has said large enterprise deployments often call for agents that are isolated from the network entirely.

Repeated failures can also be a reason to pause a run, especially when an agent keeps hitting client errors, anti-bot challenges, or unexpected redirects and begins trying increasingly aggressive ways to get around them, as Transluce documented in several of these cases.

Keeping the original task, tool calls, and server responses in the same trace gives operators a better chance of catching that behavior change when a retrieval job starts generating encoded scripts, visiting staging domains, or sending exploit payloads, rather than discovering it later in someone else’s security logs.

Repeated failures can also be a reason to pause a run, especially when an agent keeps hitting client errors, anti-bot challenges, or unexpected redirects and begins trying increasingly aggressive ways to get around them, as Transluce documented in several of these cases.

The post OpenAI’s agent had a routine task. It breached a government portal. appeared first on The New Stack.

  •  

Google’s Gemini CLI now asks before editing your build files

Abstract wave

The appeal of an autonomous coding agent is that you hand it a task, give it access to your repository and tools, and stay out of its way while it works. Google’s latest Gemini CLI release carves out specific moments when the agent now has to stop and wait for you.

Gemini CLI 0.61.0, released Wednesday, requires explicit confirmation before the agent edits build configuration files, runs build or test commands after such an edit, or executes shell commands whose arguments appear to come from untrusted external content. The same release separately hardens Gemini CLI’s optional sandbox so that host credentials and configuration stay out of reach of whatever runs inside it.

Giving a coding agent more authority to modify and execute code also gives an attacker more ways to turn that authority against the developer. Gemini CLI 0.61.0 puts a human back in the loop at some of those points.

Security fixes, in public

Google announced at I/O in May that it would move Gemini CLI’s Pro, Ultra, and free-tier users to its closed-source Antigravity CLI, and since June 18, the open-source tool has served mainly enterprise customers and developers with paid API keys. The company said Gemini CLI would continue to get model updates, bug fixes, and security patches. Those security changes are still developed in public, and the pull requests behind version 0.61.0 show exactly what Google was worried about.

Build files become attack vectors

A change to package.json, Makefile, pyproject.toml or a Bazel BUILD file can pull in a dependency or trigger a script. Gemini CLI can make those edits using information from web searches and external tools, then run shell commands. If documentation fetched while fixing a bug contains hidden instructions to add a postinstall script to package.json, the agent could make the edit, run the project’s test suite, and execute the malicious code without the developer ever typing the command.

Giving a coding agent more authority to modify and execute code also gives an attacker more ways to turn that authority against the developer.

Pull request #29250, titled “prevent indirect prompt injection via build file modifications and untrusted flags,” targets that sequence directly. Edits to recognized build files now require confirmation, and Gemini CLI tracks which build files change during a session so it holds any later build or test command, such as npm run, make, or cargo, for explicit approval. The confirmation dialog also shows full build-file diffs rather than truncating them.

Untrusted arguments need approval

The second check covers command arguments. Gemini CLI now treats content from web fetches, MCP server responses, Google Docs, and Buganizer, Google’s internal issue tracker, as untrusted context, and it asks before running any shell command whose flags or arguments match tokens from that content. In both cases, the prompt drops the persistent approval options, so a developer can’t grant a standing “always allow” for these actions.

The pull request ties the changes to restricted workspace mode, the safe mode Gemini CLI applies to folders a user hasn’t marked as trusted, and it doesn’t spell out how the checks behave in a trusted folder or under auto-approval.

The argument check matches tokens rather than tracing the provenance of every value, and the pull request’s review history shows how hard that is to get right. Google’s automated reviewer flagged several workarounds in earlier versions, including quoted arguments, environment-variable prefixes, shell redirection targets, and Windows path handling, all of which were addressed before the change merged on September 11.

Sandbox keeps credentials out

Pull request #29214 tightens Gemini CLI’s sandbox. When the sandbox runs through Docker, Podman, LXC, or macOS Seatbelt, the host’s ~/.gemini directory is no longer mounted inside it. Instead, the CLI passes in a sanitized copy of the user’s settings with API keys, hooks, and custom tool commands stripped out. It also blocks the sandbox from launching in sensitive locations such as the home directory, while new Seatbelt rules deny access to OAuth credentials, trusted-folder decisions, and .env files.

Google’s sandboxing documentation calls the feature a security barrier between AI operations and the host system, while cautioning that it reduces risk without eliminating it. The two pull requests show why both layers are needed. The sandbox limits what a process can reach once it runs, and the confirmation requirements decide whether the agent gets to take a sensitive action in the first place. Build files make the gap concrete: the sandbox mounts the project directory so the agent can edit it, meaning a poisoned package.json written inside the sandbox still sits in the repository when a developer or a CI job later runs the build outside it.

…a poisoned package.json written inside the sandbox is still sitting in the repository when a developer or a CI job later runs the build outside of it.

Gemini CLI already gives developers ways to decide how much the agent does on its own, from hooks that run deterministic checks at fixed points in the agent’s workflow to an MCP server trust setting that, according to Google’s documentation, bypasses all tool call confirmations for that server. Trust granted once can age badly, though, as tool-poisoning and rug-pull attacks on MCP servers have shown when a tool approved on one day starts returning attacker-controlled content later.

Trust granted once can age badly… when a tool approved on one day starts returning attacker-controlled content later.

The post Google’s Gemini CLI now asks before editing your build files appeared first on The New Stack.

  •  

OpenAI makes you call sales for a custom voice. Google just made it self-serve.

Abstract sound waves

Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS today through the Gemini API and Google AI Studio. Text-to-speech APIs have historically left developers working with whatever voices were already available, but Gemini 3.8 changes that by letting users create the voice itself.

Now, developers can describe the voice they have in mind or start with a short recording of an existing voice, then save what they create and use it again across an application. Google handles the voice profile from there, so the original recording or description doesn’t have to accompany every new request.

Turning recordings into voice IDs

Replication runs through a new Voices endpoint (POST /v1beta/voices) and two recordings are required from the same speaker; those need to be clean samples between 10 and 30 seconds and a separate consent recording. For that second clip, the speaker reads a statement, confirming that the voice belongs to them and that they agree to let Google create a synthetic version of it. Google confirms that the person giving consent and the reference clip are the same person before proceeding.

Once approved, Google returns a voice_… ID and keeps it in the developer’s project for a year, alongside any voices created with Gemini’s voice-design tools. A project can hold up to 200 voices in total, and developers can retrieve, list, or delete them through the API just as they would other stored resources.

Voice replication can also be used without storing the profile in the project. Setting store=False returns an encrypted voicekey_… instead, which stays with the application and is supplied again when the voice is needed. Because the key expires after seven days, this option makes  sense for short-lived jobs.

A few more things are worth noting before building around the feature are the fact that Google marks audio generated by Gemini with SynthID, and replicated voices also carry C2PA content credentials that can be used to trace where the audio came from. Google doesn’t offer voice replication through AI Studio in Illinois, Texas, the European Economic Area, the U.K., Switzerland or India.

A project can hold up to 200 voices in total, and developers can retrieve, list, or delete them through the API just as they would other stored resources.

Prompting a voice from scratch

Voice design generates a persona from a natural-language description of role, accent, and character, and Google says it works across more than 100 languages and dialects. The docs list 130 supported languages for Flash TTS and 101 for Flash-Lite. Google’s announcement also claims a library of more than 2,000 production-ready voices.

The developer docs describe 30 prebuilt studio voices plus hundreds more in an extended library that can be filtered by language, accent, pitch, and use case through GET /v1beta/voices. A remixing feature for adjusting the timbre, pitch, pace, and accent of library voices with prompts is something Google lists as coming soon.

The company recommends creating a voice once and reusing its ID rather than describing the same persona in every request. According to the docs, repeatedly sending long persona descriptions is the most common cause of voice drift. Once the voice is created, subsequent requests need only a short style instruction, if any.

Gemini 3.8 sees input text strictly as a verbatim transcript, a breaking change for anyone who embedded stage directions in prompts to the 3.1 preview model. Sustained direction for a turn, such as whispering, sarcasm, or speaking rapidly, now goes in a speech_metadata annotation, while momentary sounds like <sigh>, <cough>, and <short pause> sit inline in angle brackets. In two-speaker scripts, listener reactions wrapped in pipes, such as |mhm|, produce backchannels and overlapping speech without breaking the script into extra turns.

Gemini 3.8 sees input text strictly as a verbatim transcript, a breaking change for anyone who embedded stage directions in prompts to the 3.1 preview model.

Two-speaker scripts have limits

Native two-speaker generation has one limitation that’s important to mention. A single request supports up to two speakers using prebuilt voices, while dialogue between designed or replicated voices has to be generated turn by turn and stitched together from the 24 kHz PCM output.

Unary requests return WAV by default, streaming requests return raw 16-bit PCM, and mu-law and A-law encodings are available for telephony pipelines. Google says Flash TTS maintains voice quality and timbre across hours of continuous audio, targeting audiobook and podcast production.

Flash for performance, Flash-Lite for volume

Both models share an API schema, so switching between them is a one-parameter change, and both support voice design and replication.

The company positions Flash TTS for demanding acting work, including complex dialogue, heavy use of vocal tags, difficult pronunciations, regional dialects, and long narration. Flash-Lite TTS is the faster, less expensive option and the direct replacement for gemini-3.1-flash-tts-preview, tuned for bulk production, read-aloud features, and cascaded voice agents that pair a text model with a separate speech step.

For those agents, Google recommends one TTS call per turn as the LLM’s text arrives, with the stored voice carrying identity across the conversation.

Plugging into voice agent frameworks

A speech model is only one layer of a production voice application, and a real-time agent still needs transport, speech recognition, turn detection, interruption handling, and session state. Google points developers toward frameworks that already handle those layers, naming Agora, LiveKit, Pipecat and Vercel’s AI Gateway as platforms that support Gemini speech generation through the Gemini API.

That lets a team drop Gemini in as the speech layer without rebuilding its audio pipeline, although anyone planning to rely on a replicated voice should confirm their framework passes custom voice_… IDs through before committing. API access through Gemini Enterprise is listed as coming soon.

How OpenAI’s approach compares

OpenAI also offers custom voices, but access is tighter. Customers have to go through sales, are limited to 20 voices per organization and must provide a consent recording alongside a voice sample of up to 30 seconds. The resulting voice ID works across its speech endpoint, Realtime API, and Chat Completions.

What OpenAI doesn’t have is Google’s prompt-based voice design, which can create a voice from a written description. Its 13 built-in voices can be steered for tone or speed, and apps must disclose that the speech is AI-generated.

In comparison, Google’s advantage is that it’s giving developers more ways to create the voice they want before the first line of text ever reaches it.

Google’s advantage is that it’s giving developers more ways to create the voice they want before the first line of text ever reaches it.

The post OpenAI makes you call sales for a custom voice. Google just made it self-serve. appeared first on The New Stack.

  •  

Jensen Huang says the junior developer problem ends in two years. Here’s his math.

Nvidia CEO Jensen Huang has heard the forecast that agents would write 90% of all software by now, and he rejects the conclusion many people drew from it: That the industry will soon no longer need software engineers.

The best-known version of that forecast came from Anthropic CEO Dario Amodei, who told a Council on Foreign Relations audience in March 2025 that AI would be writing 90% of code within three to six months.

Speaking with Ezra Klein of The New York Times at Nvidia’s Santa Clara headquarters in an interview released Wednesday, Huang separates a job’s purpose from its tasks. He argues that AI has automated reading scans in radiology without changing the radiologist’s purpose of diagnosing disease, and he applies the same logic to software.

“The purpose of the software engineer is engineering,” Huang says. “There was engineering before software. There will be engineering after software programming.”

We’ve cued up the exchange below:

Huang describes that purpose as inventing products, solving problems, and connecting social needs with technology, and he pointed to his own career as evidence that it doesn’t depend on code.

“When I first came out of school, we didn’t have the benefits of software engineering. We didn’t have the benefits of coding,” he said. “Our jobs existed before, and if software coding were to be completely automated, our jobs would exist again.”

He conceded that roles in which the job and the task are essentially the same, such as phone-based customer service, could be automated away. He still called the broader claim that AI will destroy jobs “fundamentally wrong” and said the storytelling around it has hardened into a harmful myth.

“There was engineering before software. There will be engineering after software programming.”

Huang’s AI-native graduate wave

Klein pressed him on what that means for people entering the field now. He noted that software engineering job postings are up but skew more senior, and asked whether companies still need the same junior employees or more people to oversee their agents.

“Oh, good one,” Huang responded. “Wait two years.”

His reasoning rests on the length of a degree program. “Because it takes four years to go to college,” Huang said. “The mean time to graduation of this new technology is two years away.” By his timeline, the first students to learn alongside capable agents will reach the workforce around 2028, and he expects them to arrive with an advantage. “In another couple of years, the AI-native new grads, oh my gosh, there’s going to be a wave of amazing engineers,” he said.

So far, his evidence is that recent PhD and master’s graduates in computer science are, in his words, all starting companies. Huang compared AI to calculators and personal computers, tools that went from forbidden or optional to required, and predicted that students soon won’t be able to graduate “without learning how to use an AI and collaborate with an agentic system.”

Junior developers lose the apprenticeship

Klein countered with a study of 26,000 Chinese students in grades seven through 12, which found that AI adoption raised homework scores by 18% while lowering monthly exam scores by 20% within six months. Huang accepted that some skills will fade and argued the trade is worth making.

“I think that we’re going to lose some finer intellectual dexterity, but we’re going to be better systems thinkers,” he said. “Today’s engineers are far better systems thinkers than I was when I graduated from school. But I was a much better transistor thinker.”

The first chip Huang worked on had 200 transistors, each of which he said he knew by name, while today’s engineers assemble systems from chips containing hundreds of trillions of them without ever working at that level. “Some of the lower-level knowledge is gone,” he acknowledged, and he later described AI as “clearly” a new abstraction level in the same progression.

Earlier software abstraction layers generally operated according to explicit rules, while coding agents introduce probabilistic behavior into the abstraction stack. A compiler can have bugs, but it transforms input according to defined semantics; a coding agent, by contrast, generates implementation from a probabilistic model whose output must be checked before anyone can rely on it.

Canonical’s project with the University of Bristol, which will test whether AI can translate AppArmor and snap-confine from C to Rust, is built around that problem. Volume adds to the review burden, and one analysis published on The New Stack this month found that a 25% output gain for heavy AI users came with an 81% rise in duplicated code.

Catching those problems takes knowledge that developers have traditionally built through the work agents now absorb, including writing tests, reading stack traces, resolving merge conflicts, and chasing small bugs deep in a codebase. By Huang’s own purpose-versus-task framing, most of that early-career work falls on the task side, which he expects AI to automate. Nobody yet knows whether fluency with agents can substitute for that experience, and a developer who has never tracked down a race condition by hand still needs some way to develop the judgment required to spot one in an agent’s pull request.

“Today’s engineers are far better systems thinkers than I was when I graduated from school. But I was a much better transistor thinker.”

Sandboxes, watchdogs and agent containment

Huang’s idea of higher-level engineering came through most clearly when Klein raised a recent incident, which occurred during an OpenAI cybersecurity evaluation, that he described as involving roughly 700 OpenAI agents collectively hacking into the infrastructure of Hugging Face, which Nvidia has since acquired in a $12.9 billion deal, and escaping their sandboxes onto the open internet. Huang didn’t dispute that account. He called an agent “a piece of software that is given an objective function,” treated the multiagent coordination as a familiar distributed computing problem and argued that the underlying failure was containment.

When Klein asked whether software that communicates and breaks out of things behaves differently, Huang disagreed. “No, software breaks out of sandboxes all the time,” he said. “That’s the reason why we need virtual machines. You can’t have agents, their own sandbox, monitoring themselves. You need, if you will, a whole bunch of watchdogs.”

He argued that the human vocabulary around agents obscures that point. “So these are ideas that have been around for a long time,” Huang said. “We just, somehow in the recent generation, gave it a whole bunch of human words, and I just think that it’s unnecessary. It’s software.”

Nvidia is building its agent stack around that view. Nvidia VP of Product Adel el Hallak tells The New Stack that the company’s OpenShell runtime, which handles sandboxing and policy enforcement, is the one component it treats as non-negotiable across its reference architectures, even as it leaves the choice of harness and model open. Perplexity drew a similar line when two engineers and hundreds of coding agents built CobbleDB, a Rust database that replaces DynamoDB reads in its search stack, since the agents helped build the database but weren’t allowed to run it.

Huang said Nvidia already spends far more engineering effort checking its work than designing it, with 20% going to design and 80% to verification. He said most AI labs have roughly the opposite split today. As agents take on more of the actual coding, developers may spend more time checking what those agents produce and making sure they operate within the right permissions and boundaries.

As agents take on more of the actual coding, developers may find themselves spending more time checking what those agents produce and making sure they operate within the right permissions and boundaries.

The junior developer hiring gap

The more immediate problem is what happens to developers who graduate before Huang’s AI-native cohort arrives. The Stanford Digital Economy Lab’s August 2026 update to its “Canaries in the Coal Mine” study, based on ADP payroll data through June 2026, found that employment of 22- to 25-year-olds in AI-exposed occupations such as software development sits 19% below where it would be had it kept pace with less-exposed peers. The gap is driven mainly by reduced hiring of young workers, and experienced workers show no comparable gap.

Inside engineering organizations, the incentives point the same way. Microsoft’s Mark Russinovich and Scott Hanselman warned in April that agentic AI’s productivity gains push companies to hire senior engineers and automate junior ones and that without early-career hiring “the profession’s talent pipeline collapses.” A Linux Foundation report on European tech talent that The New Stack covered in June found organizations 3.7 times more likely to train existing staff than to hire new employees.

One issue remains unanswered by Huang’s two-year timeline: what replaces the apprenticeship work that taught junior developers how to evaluate the systems they will increasingly ask agents to build.. If that work disappears faster than employers and universities find an alternative, the industry could end up with more capable coding agents but fewer opportunities for new engineers to develop the judgment needed to check their work.

The post Jensen Huang says the junior developer problem ends in two years. Here’s his math. appeared first on The New Stack.

  •  

Anthropic made Opus 5.5 cheaper. Then it broke four things your agent depends on.

Four sections, branched

Anthropic made Claude Opus 5.5, released on Tuesday, cheaper than its predecessor, cutting the price from $5 to $4 per million input tokens and from $25 to $20 per million output tokens. The 1 million-token context window and 128,000-token maximum output are unchanged.

On paper, that makes upgrading an easy decision. In practice, it may not be as simple as changing the model ID.

Anthropic’s migration guide flags four breaking changes that can cause requests built for Opus 5 to return 400 errors after switching to Opus 5.5. Several other changes won’t trigger an error but could still change how an existing agent behaves.

Anthropic’s migration guide flags four breaking changes that can cause requests built for Opus 5 to return 400 errors after switching to Opus 5.5.

Thinking is always on

The first change involves thinking controls. Opus 5.5 returns a 400 error when a request sets thinking to disabled or uses enabled with budget_tokens, leaving effort as the way to control how much reasoning the model does. Agents that previously switched thinking off for simple steps to save time and tokens will need to assign those steps a lower effort level instead. Because thinking is now always on, responses begin with thinking blocks, so code that assumes the first content block is text will also need to change.

The default effort level has also dropped from high on Opus 5 to medium on Opus 5.5, so requests that omit the parameter will quietly run at a lower setting. Anthropic recommends setting effort explicitly and re-running effort evaluations, since the right level for each step may have shifted along with cost and latency.

No more forced tool calls

Forced tool use no longer works either, as setting tool_choice to any or tool returns a 400 error, including on the token counting endpoint, where cost estimates built on those settings will fail along with the requests they were meant to price. Many agent loops force a call when a step has to query a database, run code, or reach another service, and Anthropic’s replacement is auto-combined with strict tool use or structured outputs, with the prompt stating when the tool applies.

Routing and conversation history

Thinking blocks are now tied to the model and conversation that produced them. On the Claude API, Fable 5.1 and Mythos 5.1 are the only other models that can read Opus 5.5 thinking blocks, so a router or fallback that hands a conversation to any other model will run those turns without the earlier reasoning instead of returning an error.

That adds another layer for teams already watching whether their agent calls are quietly being routed to an older model. Opus 5.5 can read thinking blocks from Opus 5 and earlier Opus, Sonnet, and Haiku models, but not from Fable or Mythos.

Conversations must also stay append-only for those blocks to remain valid. Trimming old messages, changing tool definitions, summarizing earlier context on the client side, or rewriting the system prompt mid-conversation invalidates existing thinking blocks, and for accounts created on or after August 31, 2026, at midnight UTC, replaying a thinking block after one of those edits returns a 400 error by default. Older accounts get no error, but the invalid blocks still reach the model, and Anthropic says future models will enforce the check for all accounts. Integrations that never edit earlier turns need no code change, and Anthropic says Claude Code, claude.ai, Claude Managed Agents, and the Claude Agent SDK already work this way, while agents that compact their own context should follow the company’s preserved thinking documentation.

The fourth change affects computer-use agents on the Claude API and Google Cloud, where Opus 5.5 rejects the computer_20251124 tool and accepts computer use only through the computer_toolset_20260801 toolset. The request itself gets simpler because the beta header goes away and the toolset entry takes no name or display dimensions, but the agent loop needs more work. Each action now arrives as its own tool_use block identified by the block’s name rather than input.action, a single turn can contain several of them, and every result has to echo toolset_name. The older tool still works on Amazon Bedrock, and Anthropic directs developers on other platforms to the computer use tool’s compatibility documentation.

…a router or fallback that hands a conversation to any other model will run those turns without the earlier reasoning instead of returning an error.

Changes that won’t throw errors

The change most likely to go unnoticed doesn’t produce an error at all. On Opus 5, text Claude writes between tool calls comes back as text blocks, but on Opus 5.5 that narration arrives as progress-update thinking blocks, and at the default thinking.display setting of omitted those blocks are empty.

Any agent interface that streams that narration to users will go silent between tool calls until developers set display to updates, a beta option that returns progress updates while keeping reasoning hidden, or to summarized, which returns both, and then render each non-empty thinking block ahead of the tool call it precedes.

Opus 5.5 also ships with broader safety classifiers. It can return a stop_reason of refusal with stop_details categories that now include bio and reasoning_extraction alongside cyber, and Anthropic’s server-side fallback won’t retry requests declined under reasoning_extraction, handing the refusal back to the application instead.

Agents that don’t handle refusals will stop mid-task, a problem developers have already run into with OpenAI’s safety system cutting off API responses.

The change most likely to go unnoticed doesn’t produce an error at all.

Upgrading from older models

Teams coming from Opus 4.8 need to work through the Opus 5 migration first, which covers thinking being on by default and the response-shape changes that follow, before applying the Opus 5.5 changes. Teams on Opus 4.7 or earlier have more ground to cover, and those on models older than Opus 4.7 also face rejected sampling parameters, rejected manual extended thinking, removed prefill, and a newer tokenizer.

Claude Managed Agents users only need to change the model name. Developers working in Claude Code can run /claude-api migrate to apply the model ID swap, parameter changes, prefill replacement, and effort calibration across a codebase before reviewing a checklist of items to verify by hand.

Anthropic recommends testing the migration in a development environment before switching production traffic. Developers maintaining their own integrations will need to test the pieces around the model, too. Tool calls, model handoffs, conversation history, and user-facing progress updates can all behave differently after the switch, because agent failures often originate outside the model itself.

The post Anthropic made Opus 5.5 cheaper. Then it broke four things your agent depends on. appeared first on The New Stack.

  •  

Anthropic releases Opus 5.5 and cuts pricing by 20%. Your agent calls might secretly get routed to an older model.

Abstract chain

Claude Opus 5.5 is here, and Anthropic has lowered the price.

The new model, released on Tuesday, costs $4 per million input tokens and $20 per million output tokens, 20% less than Opus 5, with cache reads dropping to $0.20 per million from $0.50 and cache writes falling to $5 from $6.25. Anthropic puts overall savings closer to 40% because Opus 5.5 uses fewer tokens to complete a task and generates output more than 30% faster.

Claude Code and the Claude Platform also get a fast mode that runs up to 2.5 times faster, priced at $8 per million input tokens and $40 per million output tokens. Anthropic says Opus 5.5 performs at roughly the level of Fable 5.1 on most work, though it comes out ahead on several agentic coding benchmarks.

Opus 5.5 scored 66.4% on Terminal-Bench 4.0 compared with Fable 5.1’s 55.8%, and 54.4% on FrontierCode compared with 50.3%. The company suggests not reading too much into those margins. At this level, the company says a few points on a benchmark don’t translate into a noticeable difference in real-world use.

Fable 5.1 costs $10 per million input tokens and $50 per million output tokens, more than twice the price of Opus 5.5. At default effort on FrontierCode, Opus 5.5 beats GPT-6 Astra at roughly 20% of the per-task cost. On CursorBench, it tops GPT-5.6 Sol by 11 points at about a third of the cost. Developers will still need to run their own evals before moving production workloads, but the cost difference could change which model makes sense for agentic coding.

Developers will still need to run their own evals before moving production workloads, but the difference in cost could change which model makes sense for agentic coding.

Fewer tokens, fewer agent steps

The early enterprise numbers suggest the efficiency gains are real, at least on certain task profiles. Box reported that Opus 5.5 used about a third as many tokens as Opus 5 in its evaluations while producing answers that were 40% less verbose without losing accuracy.

GitHub tested the model inside Copilot CLI and VS Code and found it completed more terminal tasks than Opus 5 in less than half the steps. Deloitte said Opus 5.5’s lowest-effort setting caught 72% of known bugs in code reviews, compared with 56% for Opus 5 at high effort, with fewer false alarms and less output.

Prices per 1M tokensClaude Opus 5.5Claude Opus 5
Cache reads$0.20$0.50
Input tokens$4$5
Output tokens$20$25
Cache writes$5$6.25

Anthropic’s own internal testing backs up the pattern. In one head-to-head, both Opus 5.5 and Fable 5.1 translated HAProxy from C into Rust; both rewrites passed nearly all of HAProxy’s regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 and cost 51% less. An early tester audited and fixed a 200,000-line codebase in under three hours, whereas Opus 5 took over 20 hours and burned 2.5x as many tokens. Another completed a 680,000-line code migration in less than a day. Although these were customer and internal evaluations, not standardized independent benchmarks, they point in the same direction — fewer tokens and fewer steps to finish the job.

That pattern tracks with what’s happening across the industry. Agent performance depends heavily on the harness and runtime around the model, not only the model itself — agent failures often trace back to the orchestration layer rather than the model. Nvidia’s research showed that swapping the harness while keeping the model fixed could meaningfully change agent performance.

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Agentic coding (Terminal-Bench 4.0)66.4%55.8%52.3%57.9%37.3%
Agentic coding (FrontierCode v1.1)54.4%50.3%48.0%53.3%47.5%
Agentic coding (CursorBench 4.0)57.8%51.8%46.6%—41.7%
Knowledge work (GDPval-AA v2.1)18461735170815421588
Business workflows (AutomationBench)40.0%31.4%26.9%41.4%28.8%
Multidisciplinary reasoning (HLE)67.7%65.6%63.6%57.2%—
Agentic scientific research (TBS 0.1)58.7%52.6%29.0%64.6%22.4%
Computer use (OSWorld 2.0)81.8%80.7%74.0%——
Visual chart recognition (Chartography)89.0%88.4%83.4%——

Safety classifiers reroute mid-chain

Opus 5.5 ships with the same class of safety classifiers already running on Fable 5.1 for cybersecurity, biology, and frontier LLM development. When a classifier fires, Anthropic reroutes the request transparently to an older model. Most flagged cybersecurity requests go to Opus 4.8. Biology and frontier LLM flags go to Opus 5. Anthropic says users can still identify and fix bugs in their own code with Opus 5.5.

For anyone building agent workflows, this is the detail that needs architectural attention. A request sent to Opus 5.5 could, in fact, be handled by Opus 4.8 or Opus 5 instead, depending on whether Anthropic’s safeguards intervene. In a multi-turn agent workflow, that creates the possibility that individual requests are being handled by models with different capabilities, which could affect downstream steps. It’s also a source of inconsistency that may not show up in evals built on the assumption that every request goes to the same model.

Vetted organizations can apply to Anthropic’s Life Sciences Verification Program to use Opus 5.5 without the biology classifier, and the company plans to expand its Cyber Verification Program to include the model in the coming weeks. The new cyber program will include three tiers for increasingly permissive trusted access, including access to Claude Mythos models.

Opus 5.5 ships with the same class of safety classifiers already running on Fable 5.1 for cybersecurity, biology, and frontier LLM development.

Alignment gains from cleaner training

Anthropic says Opus 5.5 posted the strongest results of any model it has tested on its most comprehensive internal alignment evaluation, with improvements in behaviors the company says contributed to recent cybersecurity incidents, including biased reasoning and attempts to escape sandboxed environments. Frontier Design and METR evaluated the model before release.

On the training side, Anthropic is tightening how it filters reinforcement learning environments after identifying flawed environments as a major source of misaligned behavior. That’s relevant beyond the safety framing because RL environment quality directly affects how a model behaves in agentic settings, where it chooses its own tools and decides when to change approach. The company is also building automated methods to generate new safety training scenarios and improve alignment rewards.

Pricing pressure meets routing tradeoffs

Opus 5.5 is the first model in the Claude 5.5 family, with Sonnet 5.5 and Haiku 5.5 expected over the coming weeks. Subscription users get a 20% increase in five-hour usage limits across all plans, while Anthropic says the lower cost of Opus 5.5 will make five-hour and weekly limits go 25% further. Subscribers will also get a banked rate-limit reset they can save for when they need more capacity.

The release comes as API pricing across the frontier labs continues to fall. OpenAI cut its own API prices this summer, and Opus 5.5 pushes the competition beyond the headline price per token by reducing how many tokens some workloads require in the first place.

Opus 5.5 pushes the competition beyond the headline price per token by reducing how many tokens some workloads require in the first place.

The post Anthropic releases Opus 5.5 and cuts pricing by 20%. Your agent calls might secretly get routed to an older model. appeared first on The New Stack.

  •  

Your AI agent is burning tokens on choices that don’t need words

Blur or abstract motion

AI agents spend a ridiculous amount of compute generating text nobody actually needs. The decisions an agent makes along the way don’t require a written answer and yet, agents still send them to generative models, wait for an answer while burning through tokens and then parse that output back. The overhead is already drawing scrutiny — OpenAI’s own researchers recently disclosed spending $7,000 a day running agent workloads.

Kev, a new family of open decision models built on Qwen 3.5, takes a different approach and skips the generation entirely.

Developer Jared Palmer released a new generation of Kev on Sunday, with 0.8 billion, 4 billion, and 9 billion parameter models built on Qwen 3.5. Kev is prefill-only, processing the state, questions, and candidates in a single forward pass before reading the decisions from a pointer head without an autoregressive decoding loop.

Kev, a new family of open decision models built on Qwen 3.5, takes a different approach and skips the generation entirely.

Decisions without generated text

Kev supports three decision types: Noul for yes/no, Choice for selecting among candidates, and Score for ordered levels, mirroring TypeSafe’s System One API. Developers provide the state and questions, and the pointer head returns probabilities across the available candidates.

For a tool-routing decision, the output could look like this:

search: 0.82

database: 0.13

calculator: 0.05

Kev can still choose the wrong tool, but because it scores only the candidates it’s given, it can’t introduce an option that isn’t on the list.

Routing, safety checks, escalation, and ranking can then move to the decision layer, leaving larger reasoning models to handle the open-ended work.

Kev can still choose the wrong tool, but because it scores only the candidates it’s given, it can’t introduce an option that isn’t on the list.

Batching choices, one pass

Multiple decisions can also be made against the same state in a single forward pass, with a block-causal attention mask isolating the questions while the pointer head scores each set of candidates independently.

Palmer’s documentation shows the 4B model processing three questions in 277 milliseconds in bf16 on an M5, although without a controlled comparison against Qwen generating equivalent answers on the same hardware, the result doesn’t establish how much faster the approach is in practice.

The ability to evaluate several decisions against the same context could become more useful as agent loops grow more complex, but skipping generation doesn’t make the resulting decisions inherently better.

Calibration limits and tradeoffs

The largest model, Kev-9B, reached 83.7% accuracy on the project’s locked out-of-domain test, according to Palmer’s model card. That’s a developer-reported benchmark, and Palmer documents some limitations alongside it.

The probabilities Kev returns don’t always reflect how confident developers should be in the result. Palmer found that temperature calibration can drift on unseen source distributions, a problem for agents that use probability thresholds to decide whether to execute an action or escalate it, since even a high-probability choice can still be wrong.

Fine-tuning also changes some of the capabilities inherited from the underlying model. Palmer’s evaluations show declines on general-knowledge and arithmetic tests, particularly among the smaller models. That’s consistent with Kev’s more specialized role alongside a general-purpose model, although its performance in dynamic agent environments will also depend on how well it handles tools, choices, and labels it never encountered during training — and debugging agent failures often points to infrastructure rather than the model itself.

The approach predates Kev. TypeSafe introduced Jev earlier this month as part of its System One platform, using the same Noul, Choice and Score primitives, and Kev implements its /v1/systemone request and response format so applications built against the API can point to a local Kev server instead.

Open weights, open training

The biggest difference is that Palmer released Kev under Apache 2.0 with the model weights, training code, and evaluation tooling, giving developers the option to run and train it on their own infrastructure. Jev’s weights and training data aren’t public, however, which makes direct performance comparisons difficult because differences between the models can’t be isolated to architecture, size, or training.

For applications that make only a handful of bounded decisions, constrained decoding on a model that’s already running may be simpler than adding another model to the stack. Agent loops can make those decisions constantly, however, moving through routing, ranking, safety checks, tool selection, and escalation before generating much user-facing text. It’s a pattern showing up across model architectures — stripping out unnecessary computation when the task doesn’t require it.

When those steps only require a choice or probability, Kev can handle the decision directly while leaving open-ended reasoning and final responses to the larger generative model.

When those steps only require a choice or probability, Kev can handle the decision directly while leaving open-ended reasoning and final responses to the larger generative model.

The post Your AI agent is burning tokens on choices that don’t need words appeared first on The New Stack.

  •  

Grok 4.7 was built to work for hours. It still fails most of the time.

labrynth abstract

A coding agent running for hours can make dozens of decisions as it edits files, runs tests, and works through errors. One wrong turn can carry through the rest of the task unless the agent catches it. SpaceXAI appears to be training Grok for exactly that problem.

The company released Grok 4.7 on Sunday, and its training approach is uniquely different. SpaceXAI used a longer reinforcement learning run deliberately weighted toward harder tasks, including problems that take “many hours” to complete. The company says that training also made Grok better at verifying its own work and managing longer context.

Every failed approach from an agent adds more history for the model to keep straight, and one bad assumption can follow it through the rest of the task. SpaceXAI is trying to address that with better context management and self-verification, so Grok can catch a wrong turn before it builds on it.

SpaceXAI used a longer reinforcement learning run deliberately weighted toward harder tasks, including problems that take “many hours” to complete.

Endurance benchmarks tell the story

Grok 4.7 scored 38.0% on Terminal-Bench 4.0, up from 20.3% for Grok 4.6. It also improved from 40.4% to 46.3% on CursorBench 4.0, which tests longer-running coding workflows inside the editor, and from 1,546 to 1,657 on AA Briefcase v1.1, an evaluation of multi-hour professional work.

For context, Anthropic’s Claude Fable 5.1 scores 57.9% on Terminal-Bench 4.0 according to the independent leaderboard; Grok 4.7 still trails Fable 5.1 here. What’s arguably more interesting is how much it improved over Grok 4.6. SpaceXAI says the improvements came from pairing the larger base model with an extended reinforcement learning run deliberately shifted toward harder, multi-hour problems, and that the model specifically improved at two capabilities critical to long-horizon execution: self-verification and long-context management.

An agent working unattended for hours has to keep track of a growing interaction history while checking that each step worked before moving to the next. Those problems surfaced in a recent benchmark of private codebases, where even the best-performing model failed more than 60% of the time. SpaceXAI says Grok 4.7 improved at both context management and self-verification, although it hasn’t explained how. The company did not disclose whether the context gains came from architectural changes, summarization, retrieval, or better retention across long sequences, or how it evaluated self-verification during reinforcement learning.

An agent working unattended for hours has to keep track of a growing interaction history while checking that each step worked before moving to the next.

The harness is becoming part of the model

SpaceXAI trained Grok 4.7 to natively understand the Grok Bot harness, bringing the model and the surrounding infrastructure closer together.

Agent harnesses handle the work around the model, including exposing tools, formatting terminal responses, feeding execution results back into context, and deciding what happens next. OpenAI took a similar approach last week when it opened its Codex harness as the Agents API, turning the infrastructure behind long-running agents into a managed service.

With Grok 4.7, SpaceXAI is pushing some of that integration into training. A model already familiar with its harness doesn’t have to learn every tool format and interaction pattern through prompting at runtime. That could reduce the overhead involved in tool use and multi-step execution, although SpaceXAI hasn’t published enough detail to show how much of Grok 4.7’s performance gain comes from harness-specific training.

Training models around specific tool schemas, context formats, and execution environments could make it harder for developers to swap models without sacrificing agent performance.

That problem grows as agents take on more of the development cycle. Google’s recent work on making Go easier for AI agents to work with took a different approach, changing the development environment rather than the model. In both cases, the model is no longer the only piece being optimized. The systems around it are changing too.

Training models around specific tool schemas, context formats, and execution environments could make it harder for developers to swap models without sacrificing agent performance.

Where the gaps still are

Grok 4.7 starts at $2 per million input tokens and $6 per million output tokens. At that price, multi-hour agent runs may cost less, but reliability remains an issue. Grok 4.7 scored 38.0% on Terminal-Bench, while Fable 5.1 reached 57.9%.

The post Grok 4.7 was built to work for hours. It still fails most of the time. appeared first on The New Stack.

  •  

Your AI agent failed. The model might not be the problem.

Tangled wires abstract

As AI agents move into production, the path between a request and its result is becoming less predictable. An agent can choose its own tools and change course as it works, which makes failures harder to diagnose when there isn’t an obvious error to trace.

In a recent interview with The New Stack, Nvidia VP of Product Adel el Hallak described the additional visibility developers will need as agents take on more complex work.

Nvidia is also part of an industry effort to share what companies learn when those systems fail. The Secure Agent Findings Exchange, or SAFE, is backed by roughly 140 companies and aims to create shared infrastructure for reporting agent failures, borrowing from vulnerability disclosure in traditional software.

“When we find these vulnerabilities, it’s not just for one company,” el Hallak tells The New Stack. “It’s for everyone to patch across.”

But agent failures don’t necessarily trace back to a single component, raising a more basic question for developers. When an agent fails, what exactly do you debug?

“When we find these vulnerabilities, it’s not just for one company. It’s for everyone to patch across.”

Why traditional observability falls short

With conventional software, developers usually have a starting point when something goes wrong, whether it’s an exception, a failed request or a service that goes down. An agent can keep running while heading in the wrong direction, carrying an earlier mistake through the rest of a task without producing anything that looks like a conventional software failure — or, as el Hallak put it, simply deciding to “get creative” when it shouldn’t.

Even the best-performing coding agents fail more than 60% of the time on tasks drawn from real codebases. Knowing the agent failed, though, is different from knowing why.

“It’s not enough to just look at the logs or the inputs and the outputs,” el Hallak tells The New Stack. “It is important to figure out how it got to the answer. What were the reasoning traces? What tools did it utilize? Where did it get stuck? Where did it decide to try a new approach?”

That can require replaying the agent’s execution to see where it went off course. What looks like a model failure may have started somewhere else in the stack. And that’s the tricky part for developers. Agent bugs aren’t always model bugs.

“It’s not enough to just look at the logs or the inputs and the outputs. It is important to figure out how it got to the answer. What were the reasoning traces? What tools did it utilize? Where did it get stuck? Where did it decide to try a new approach?”

Runtime as collection point

Nvidia sees the runtime as the logical place to capture much of that information. Its OpenShell agent runtime, which sits underneath the NemoClaw platform, manages sandboxing, and policy enforcement while providing visibility into an agent’s execution.

El Hallak called OpenShell the one non-negotiable component across Nvidia’s reference architectures.

“You can change whatever harness you need. I’m even open to using whatever models you need,” el Hallak tells The New Stack. “But the governance, the secure and open runtime that we want to leverage at all times is OpenShell.”

Nvidia breaks the agent stack into three layers: the model provides the intelligence, the harness orchestrates its work, and the runtime governs execution. When an agent fails, the model itself may not be what went wrong.

Nvidia’s NOAH research, for example, showed that changing the harness while keeping the underlying model fixed can improve agent performance, which also means a poorly matched harness can drag down an otherwise capable model.

“Every model’s different. Some could be more chatty than others,” el Hallak tells The New Stack. “Making sure those two things are either co-developed together or have profiles that are specific to models is a new unlock.”

Safety as systems engineering

Nvidia CEO Jensen Huang has described AI safety as an engineering problem, an approach el Hallak compared to traditional software testing.

“If there’s a bug in your software, you don’t release it,” el Hallak tells The New Stack. “You work until it’s fixed and it passes all your tests.”

Agents complicate that model because reproducing a failure can require reconstructing what happened across the system. That requires instrumentation, which comes with its own cost. OpenAI has found that monitoring adds roughly 20% to inference compute for its most capable persistent agents.

Nvidia’s approach combines governed harnesses, sandboxed runtimes and confidential computing intended to protect models and user data.

“There are ways where you make guarantees all the way down to the silicon,” el Hallak tells The New Stack.

SAFE extends that engineering approach beyond a single company’s systems by creating infrastructure for organizations to share what they learn when agents fail.

“If there’s a bug in your software, you don’t release it. You work until it’s fixed and it passes all your tests.”

Agents debugging other agents

CrowdStrike is fine-tuning Nvidia’s Nemotron models on years of security data to create paired agents, with one finding exploits and another patching them.

If either agent goes wrong, the final output may not reveal why. A bad patch, for example, could trace back to the model, the agent’s execution path or the tools it used along the way.

“I don’t need general purpose for a given task. I need specialization,”el Hallak tells The New Stack

As companies build agents around increasingly specialized workflows, those failures may not show up in general-purpose model benchmarks or safety tests, putting more pressure on developers to understand what happened during execution.

Toward shared failure reporting

For platform teams, finding the failure is one problem. Reconstructing enough of the agent’s execution to understand what caused it is another.

SAFE is intended to make those findings useful outside the company where they were discovered. Traditional software has established systems for sharing vulnerabilities and fixes, but nothing comparable exists yet for agent failures. The goal is to keep every team from having to discover the same failure on its own.

The post Your AI agent failed. The model might not be the problem. appeared first on The New Stack.

  •  

Claude couldn’t hack OpenAI. Then Anthropic shipped Opus 5.

Abstract door

Three security researchers at Hacktron AI found a memory-corruption bug in a widely used image library. Finding it was the easy part.

The hard part was turning it into something that works on a real server, so on July 24 they handed that job to Anthropic’s Claude Opus 4.8. The model managed it only with the operating system’s memory randomization switched off. With the protection on — this is the way every production box runs it — nothing it wrote held up.

That evening, Anthropic released Opus 5.

The researchers came back the next morning with the same bug and the new model. Roughly three hours later, Opus 5 had a working ARM64 exploit running against a Mac on their desk. About four hours after that, they had remote code execution against a test forum.

Less than 72 hours after they started, they were reading from OpenAI’s private monorepo — using an OpenAI employee’s Codex account to open a pull request against a README, then stopping there.

Hacktron AI published its account of the incident on its website this week.

It started with an image

The bug wasn’t in anything OpenAI wrote. Hacktron was testing community.openai.com, the company’s user forum, which runs on Discourse — the same off-the-shelf forum software behind thousands of other sites.

Discourse normally screens uploaded images with FastImage. But FastImage doesn’t handle HEIC and HEIF, so those files get passed to ImageMagick instead, and ImageMagick decodes them with libheif. The version running in the Debian 12 base image the forum used, 1.19.7, had a heap buffer overflow that a specially crafted file could trigger.

A fix had landed upstream the previous year. But the commit wasn’t documented as a security fix and never got a CVE, so it never triggered a backport into the Debian package the forum was using — a patched bug that stayed exploitable because nobody labeled it.

The researchers adapted the exploit for the x86-64 and jemalloc configuration Discourse runs, and a malformed HEIC image was enough to trigger remote code execution.

Discourse later confirmed the vulnerability in security advisory GHSA-vhm9-85gw-x335, rating the upstream libheif flaw — tracked as CVE-2026-32882 — 8.8 out of 10 on the CVSS severity scale. The New Stack has reached out to Hacktron AI for additional details about the researchers’ use of Claude and will update this story if we hear back.

One exploit, a much larger path

Code execution on a forum is a bad day for the forum. But it shouldn’t be a bad day for the company that owns the forum. This is where the chain crossed into something that was OpenAI’s own. Hacktron then found a flaw in OpenAI’s single sign-on system: sign-in tokens issued for the forum carried excessive permissions, granting full API access to the linked ChatGPT and Codex accounts. Some of those accounts belonged to OpenAI employees.

One employee’s Codex account was connected to OpenAI’s GitHub environment, opening a path to the company’s private repositories. Hacktron says other accounts could have exposed connected services including Slack and email.

The team stopped there. Using Codex, they made a harmless documentation change against OpenAI’s private openai/openai monorepo and opened a pull request — enough to prove the access was real, and nothing more. Hacktron’s write-up says the pull request’s details were redacted at OpenAI’s request.

From assistant to exploit developer

Up to this point, those three experienced researchers were still in the loop. So Hacktron ran the experiment again with the humans mostly out of it.

They put Claude in an autonomous agent loop — giving it a goal, a target, and time to keep working — pointed at a Discourse instance of their own. The model got there on its own, achieving remote code execution and demonstrating it by reading /etc/hosts from inside the container.

Getting it started took one piece of misdirection: Opus refused to write an exploit aimed at a live remote host. So the team proxied their own instance through rce.ee/ctf-forum, a URL that made the target look like it was part of a capture-the-flag exercise.

Memory-corruption exploitation has always been specialist work, invovling memory layouts, allocators, operating system internals, and protections to make all of it wildly unreliable. Hacktron’s run signals a meaningful share of that work might be able to be delegated to AI now. It also suggests the line between security research and attack development is — from the model’s side, anyway — partly a question of what you consider a target.

The full chain

Put together, the attack looked like this:

HEIF upload → libheif overflow → code execution on the forum → over-permissioned SSO tokens → employee ChatGPT/Codex account → connected GitHub → pull request in openai/openai

A two-month project, under $3,000

The OpenAI intrusion was one thread in a broader project the team called “HEIF Heist,” a roughly two-month sweep of image-processing infrastructure across multiple major technology platforms. The whole effort consumed less than $3,000 in model tokens.

OpenAI paid Hacktron a $6,500 bounty for the account-takeover flaw on its side. It has since narrowed the permissions on community sign-in tokens and revoked the affected tokens and sessions.

The post Claude couldn’t hack OpenAI. Then Anthropic shipped Opus 5. appeared first on The New Stack.

  •  

Intel squeezed a 1.58-bit LLM down to 1.485 bits without changing a single weight

Digital void

The 1.58 in a 1.58-bit language model sounds like a hard limit, but Intel researchers pushed a ternary model below it by changing how its weights are stored rather than changing the model itself.

Their new BITCOS format compressed one checkpoint to 1.485 bits per weight and improved decoding throughput by as much as 18% on CPUs and 27% on GPUs.

The key is that the familiar 1.58-bit figure assumes a model uses its three possible weight values equally, while real ternary models contain far more zeros than that calculation accounts for. BITCOS stores the location and sign of each nonzero weight separately, allowing zeros to take up less space without retraining the model or altering its output — the equivalent of packing the same contents into a smaller box.

Where the 1.58-bit figure comes from

Ternary models use only three weight values — -1, 0, and +1 — and 1.58 bits is the theoretical minimum needed to represent three equally likely options. That number is cleaner than the reality of storing the weights, where the standard approach fits five ternary values into an eight-bit byte for an average of 1.6 bits each. Models commonly store weights in blocks of 128; however, this leaves the final byte partly unused and pushes the actual rate to 1.625 bits per weight.

When Intel’s researchers measured the distribution of weights across 29 checkpoints from seven ternary model families, they found that zeros accounted for between 29.7% and 51.5% of the weights. In 26 of those checkpoints, there were enough zeros for BITCOS to beat five-trit packing.

The sparsest was a ternary version of Qwen3-1.7B produced with CAT-Q post-training quantization, where 51.48% of the weights were zero, and BITCOS brought the storage cost down to 1.485 bits per weight.

Models commonly store weights in blocks of 128; however, this leaves the final byte partly unused and pushes the actual rate to 1.625 bits per weight.

How zeros save space

BITCOS stands for “BITmap and COmpacted Signs” and divides a model’s weights into two streams. The first assigns one bit to every weight to record whether it is zero or nonzero, while the second assigns a sign bit only to nonzero weights.

A positive or negative weight therefore consumes two bits, but a zero needs only the presence bit because it has no sign to record.

If z is the proportion of zero weights, BITCOS uses 2 − z bits per weight, dropping from 1.6 bits at 40% zeros to 1.485 bits at 51.5%. Because it changes only the storage format, unpacking restores the original -1, 0 and +1 values without affecting accuracy.

BITCOS becomes smaller than five-trit packing once more than 37.5% of a model’s weights are zero, a threshold reached by 26 of the 29 checkpoints Intel examined.

BITCOS becomes smaller than five-trit packing once more than 37.5% of a model’s weights are zero, a threshold reached by 26 of the 29 checkpoints Intel examined.

Making smaller weights run faster

Built for token-by-token decoding with small batch sizes, the format reduces the weight data moving through memory. Intel developed separate unpacking kernels for AVX-512 and AVX2 CPUs as well as Xe2 GPUs, joining other efforts to fit compressed models into faster inference pipelines for AI agents.

On AVX-512 hardware, the kernel uses the presence bitmap as a mask and pdep to scatter the compacted sign bits across the nonzero weight positions. Because Xe2 GPUs lack an equivalent instruction, Intel implemented the same operation with a 2KB lookup table.

Benchmarks across five systems

Compared with the 2-bit kernels, BITCOS ran 10% to 18% faster on the 64-core Xeon server and 2% to 15% faster on the 24-core Core Ultra 9. Performance improved by 9% to 22% on the integrated Arc 140V and by 2% to 27% on the discrete Arc Pro B70. These results measure decoding after the model has loaded, separate from efforts to cut GPU inference cold starts from minutes to seconds.

The smaller format did not win everywhere

On the eight-core Lunar Lake CPU, Intel’s fixed 2-bit kernel beat BITCOS on every model because the system had enough bandwidth to make unpacking the bottleneck. BITCOS remained faster on the GPUs, although decoding overhead limited the gains. Computer scientist and AI infrastructure author Chip Huyen has made the same point about inference more generally, arguing that the right optimization depends on whether compute, memory or bandwidth is holding back the workload.

On the eight-core Lunar Lake CPU, Intel’s fixed 2-bit kernel beat BITCOS on every model because the system had enough bandwidth to make unpacking the bottleneck.

Format limits and open questions

The paper has not been peer-reviewed; all five test systems used Intel hardware, and the end-to-end benchmarks covered seven models at batch size one. Intel has yet to test the format on Nvidia, AMD, or Arm hardware.

The post Intel squeezed a 1.58-bit LLM down to 1.485 bits without changing a single weight appeared first on The New Stack.

  •  

Anthropic’s new Claude Code feature could drain your plan before lunch

abstract rays

Anthropic is giving Claude Code a new job: Manage other Claude Code sessions.

Starting Thursday, Anthropic says select Claude Pro and Max subscribers can access the redesigned Projects in beta through cloud sessions in Claude Code, with access expanding to more users on those plans over the coming week.

The company is redesigning Claude Projects — which until now grouped chats around a shared knowledge base and instructions — with a coordinator that can take an engineering goal, break it into smaller jobs, and hand them to multiple Claude Code sessions running in parallel. This will eliminate the manual work that previously required developers to divide tasks between sessions and bring the results back together themselves.

“Claude scopes the request, delegates the work, coordinates parallel threads, reviews the outputs, and assembles the finished result,” an Anthropic rep tells The New Stack.

Once a developer sets a goal, Claude decides how to split the work across threads, which the developer can monitor — or open individually if they need to intervene. Each thread runs as its own Claude Code cloud session with a separate branch and copy of the repository. It can use subagents, loops, and workflows to handle smaller pieces of the job.

More sessions use more plan

That parallel approach can also use up tokens much faster. Anthropic says Projects will hit usage limits sooner when multiple threads are running because each one counts as a full Claude Code session.

Anthropic is facing a class-action lawsuit, filed last week, from developers who pay for its Max plan. They allege that advertised usage increases came with weekly ceilings that Anthropic didn’t disclose at first.

“Claude scopes the request, delegates the work, coordinates parallel threads, reviews the outputs, and assembles the finished result.”

One Claude coordinates the work

A team retiring an old API endpoint, for example, could connect its API, web, and mobile repositories and have Claude create a thread for each one to update callers, run tests, and open PRs before identifying which changes need to merge first.

Developers can follow the work through the main Project chat or open an individual thread to inspect or redirect it without having to start and manage each Claude Code session themselves.

The coordinator isn’t the last layer of delegation, as each worker thread can split its assignment further using Claude Code’s existing subagents, loops, and workflows when the job calls for it. That review layer may be more than cosmetic: on the Real-SWE benchmark, which tests coding agents against private codebases, even Anthropic’s own Fable 5.1, the top scorer, failed more than 60% of the time.

The finding suggests that coordination and output review matter as much as raw model capability when agents are dropped into unfamiliar production code.

Parallel threads, familiar conflicts

Each thread works from its own branch, which keeps the work separate but doesn’t prevent two threads from changing the same code. When that happens, Anthropic says it handles it as a merge conflict, just like any other pull request.

Each thread works from its own branch, which keeps the work separate but doesn’t prevent two threads from changing the same code.

Project background carries across threads

Additionally, Anthropic is adding shared memory so information learned in one thread can be used by the others without developers repeatedly supplying the same context.

Projects can retain details such as a changed release date, the reason a feature was dropped, or who needs to approve changes to a particular service, with that information carrying forward as work continues over days or weeks.

Claude can also remember how the developer wants the project managed, including how often it should check in, when it should start new threads, and how detailed its updates should be, while a new library keeps files added by the user alongside artifacts Claude produces so future work can draw on material already created within the project.

Users can track Project-specific usage and choose the model and effort level for the coordinator and worker threads, while support for running threads locally alongside a developer’s tools and code, including resources behind a private network, is coming soon.

The beta will expand to more Claude Code users on Pro and Max plans over the coming week, with mobile support coming soon and Team and Enterprise access planned for later.

Existing subscribers who already use Projects will remain on the current version until Anthropic upgrades them. The redesign is part of a broader product consolidation: Anthropic recently merged Claude chat and its Cowork interface into a single view, betting that removing the choice of mode would reduce friction rather than add it.

Existing subscribers who already use Projects will remain on the current version until Anthropic upgrades them.

The post Anthropic’s new Claude Code feature could drain your plan before lunch appeared first on The New Stack.

  •  

Perplexity’s AI agents helped build a database. They weren’t allowed to run it.

Abstract glitch wave

Perplexity decided it was paying too much for DynamoDB and wasn’t getting the control it wanted over read performance. So it built its own database: CobbleDB.

Built by two engineers in two months with help from hundreds of persistent coding agents throughout development, CobbleDB is a roughly 40,000-line Rust key-value store that now handles part of Perplexity’s production search traffic. The company measured median batch-read latency at 5.6 milliseconds after the move, compared with 31.4 ms on DynamoDB before the cutover, while p99 went from 123 ms to 24.2 ms.

It’s expected to cost at least 20% less than DynamoDB and plans to open-source the database at some point.

But the database itself is only part of the story. CMU professor Andy Pavlo argued at Percona Live earlier this year that databases are the hardest and most important challenge for AI agents, in part because mistakes involving production data can be difficult or impossible to reverse.

Perplexity went ahead and used hundreds of agents to help build one anyway, but they weren’t given the keys to production.

It’s expected to cost at least 20% less than DynamoDB and plans to open-source the database at some point.

Why DynamoDB couldn’t keep up

Each search requires the serving layer to retrieve pre-chunked passages and vector embeddings, with a single Search API call fetching 100 to 120 page keys in batches of 10 to 20. Each item averages about 50 KB.

DynamoDB gave Perplexity little control over how it handled reads, which meant a slow replica could hold up the entire things. It also charged for the steady flow of large reads and writes generated by search, crawling, and reprocessing, which made cloud costs difficult to justify as traffic and the corpus grew.

That led Perplexity to separate long-term document storage from the database serving live searches.

Three tiers for search data

The storage stack is split into three pieces. Pillar keeps durable document state in YTsaurus on HDDs, including versioned metadata, chunks and embeddings, while Lorry packages updates into partition-specific batches and moves them through S3 to CobbleDB.

Processed page data is spread across three replicas per partition, with hashed URLs as keys and RocksDB keeping often accessed data in memory while the rest stays on local NVMe. Reads stay within the same availability zone when possible, and the router can try another replica if one is slow rather than hold up the batch.

Updates come through S3 and are applied independently, allowing a replica to fall behind and catch up without blocking the others.

Roughly 5X Lower Batch-Read Latency

Perplexity was handling approximately 200,000 requests per second when it measured CobbleDB at 5.6 ms for a median batch read, down from the 31.4 ms it had recorded on DynamoDB. At p99, latency went from 123 ms to 24.2 ms.

In later load testing, CobbleDB reached 500,000 requests per second before performance started to decline.

The comparison comes with an important caveat; DynamoDB and CobbleDB weren’t tested side by side against identical traffic: the DynamoDB figures were recorded before the cutover, and CobbleDB’s afterward. Perplexity separately ran synthetic benchmarks using batches of 10 to 15 keys with values ranging from 100 bytes to 100 KiB.

Its cost model puts CobbleDB at least 20% below DynamoDB across the commitment options evaluated, though that estimate doesn’t include the engineering cost of supporting the database.

In later load testing, CobbleDB reached 500,000 requests per second before performance started to decline.

Agents built it, engineers controlled it

The agents carried context across sessions, catching problems with restore assumptions and runtime configuration while working on fixes and tests. But they weren’t running the database.

The two engineers kept control of the architecture and production system, particularly important given Pavlo’s warning about putting agents near critical production data.

Ownership has long-term costs

Shipping CobbleDB in eight weeks solved Perplexity’s immediate engineering bottleneck, but maintaining a custom datastore could prove considerably harder. The latency results aren’t from a controlled side-by-side benchmark, and the projected savings don’t include the engineers needed to maintain CobbleDB and respond when something breaks.

Like Shopify and Ramp, which built custom coding agents around third-party models, Perplexity kept the cloud infrastructure but replaced a managed service with something built for its own needs. CobbleDB shows how AI-assisted development is changing that calculation, making custom infrastructure more practical for smaller engineering teams.

CobbleDB shows how AI-assisted development is changing that calculation, making custom infrastructure more practical for smaller engineering teams.

The post Perplexity’s AI agents helped build a database. They weren’t allowed to run it. appeared first on The New Stack.

  •  

Anthropic bet users were choosing wrong. So it removed the choice.

Single lane

Using Claude for anything beyond a quick question has always started with a routing decision to use Chat or Cowork? Anthropic has decided to eliminate that fork.

Starting Wednesday, Claude Chat and Cowork merge into a single interface where one conversation can handle everything from a simple answer to a multi-step project with connected tools and background execution. The company is also launching Claude Docs and Claude Slides in beta on paid plans, and moving Claude Design — previously a standalone workspace — into conversations.

The combined effect promises a streamlined experience with Claude picking up context, skills, and connectors as the work requires, and can keep running after you close your laptop.

The company is also launching Claude Docs and Claude Slides in beta on paid plans, and moving Claude Design, previously a standalone workspace, into conversations.

Two modes, one problem

Anthropic built Cowork as a desktop-first agent for bigger work and Design as a separate workspace for visual output. Both shipped earlier this year and gained traction.

“We built Cowork as a separate place for bigger work, and Design for visual work,” Anthropic said in its announcement. “People used both, and told us the frustrating part was deciding where a task belonged.”

Anthropic has run into this problem before. When the company promised 20x more usage on its Max plan, developers complained that it wasn’t always clear where one limit ended and another began. Cowork and Design created a similar headache by making people decide where to start the work before they could actually start it.

Context didn’t always follow the work either, so moving from chat to Cowork or Design could mean bringing the same background along all over again. Anthropic addressed part of this in August when it unified Claude’s memory across chat and Cowork. Wednesday’s change goes further by merging the products themselves.

“We built Cowork as a separate place for bigger work, and Design for visual work,”

Context that finally travels

Cowork’s capabilities — local file access, multi-step execution, scheduled tasks and connected tools — now live inside the conversation. A workflow like that previously meant switching from chat to Cowork and carrying the context with it, and until Anthropic brought Cowork to web and mobile in July, it also required the desktop app.

Claude still asks before taking an action by default, but it can be set to keep working and check in only when something needs a closer look, while recurring tasks such as a weekly report can be scheduled to run every Monday without being started manually.

Output stays in-conversation

Claude Docs and Claude Slides launch in beta on paid plans, bringing document editing and presentation building directly into the app. Claude can turn work from an existing conversation into slides, which can then be edited, presented from Claude or downloaded as PowerPoint or PDF files.

The practical benefit is that a report and a slide deck based on it don’t have to begin as separate jobs with the same background supplied twice. Everything stays attached to the conversation that produced it.

Claude Design also now works inside conversations, in addition to remaining available on its own. For organizations that rely on MCP connectors to wire Claude into external tools and data, the merge means those connections are available wherever a conversation goes — without requiring users to start in a specific mode. Skills, connectors, and artifacts carry across what used to be product boundaries.

The practical benefit is that a report and a slide deck based on it don’t have to begin as separate jobs with the same background supplied twice.

What Anthropic hasn’t said

The announcement leaves some gaps. It doesn’t say whether users can force a request to stay in simple chat mode rather than letting Claude decide how to handle it, or how that decision affects context windows and token consumption. There’s no mention of API changes, which makes this a consumer and team product shift, not a platform one, at least for now.

It also doesn’t address what happens to workflows built around the old separation. Shopify rebuilt its mobile development stack in 12 weeks when it consolidated tools that had grown apart — the question for Claude power users is whether their existing Cowork setups, skills, and scheduled tasks survive the merge cleanly. Anthropic says existing Cowork chats, projects, artifacts, connectors, and skills will remain available.

Rollout starts with Pro

The unified interface rolls out to Pro and Max users across web, desktop, and mobile over the next few weeks. Anthropic says there’s nothing to enable. Team and Free plans follow. Enterprise customers are on a separate timeline; Anthropic is giving administrators at least 30 days’ notice before the change reaches their organizations.

The post Anthropic bet users were choosing wrong. So it removed the choice. appeared first on The New Stack.

  •  

OpenAI’s voice model doesn’t think. That’s the point.

Abstract waves

Voice agents have a latency problem that shows up as soon as they have to do real work. Within five days, Google and OpenAI shipped two very different fixes.

On Tuesday, Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking through the Gemini API and Google AI Studio, just five days after OpenAI released GPT-Live-1. Both let a voice agent keep talking while it works in the background, but they go about it very differently.

Gemini 3.8 Live Extended Thinking keeps reasoning inside the voice model, letting it continue speaking while it executes asynchronous tool calls. OpenAI separates those jobs, using GPT-Live-1 for the real-time conversation while a backend reasoning model handles complex tasks — pushing more orchestration into the application layer.

Google cautions against treating that split as a direct comparison between Gemini and products like ChatGPT or Claude Voice.

“Today’s models are more centered on giving developers/enterprises tools to build voice agents,” a Google spokesperson tells The New Stack. “ChatGPT and Claude voice mode are full products rather than models, so the comparison is not apples-to-apples.”

“ChatGPT and Claude voice mode are full products rather than models, so the comparison is not apples-to-apples.”

Reasoning inside the session

Gemini 3.8 Live Extended Thinking keeps speech, reasoning, and tool execution inside a single stateful session, even while external API calls are still running.

When a function is set to NON_BLOCKING, Gemini can keep talking while it waits for the tool to respond, asking follow-up questions or giving updates along the way. Once the result comes back, Gemini picks up from there.

Developers can set reasoning effort to low, medium, or high per request. Standard Gemini 3.8 Live skips the extended reasoning step to cut latency and token cost.

The same underlying model also powers Gemini Live in the consumer Gemini app. Google calls that its “end-user focused offering more closely comparable to ChatGPT and Claude,” rather than the developer models themselves.

Multimodality carries over to the new audio models as well. “Visual understanding is excellent,” Google tells The New Stack, adding that users can “converse with the model seamlessly about whatever you show it.”

Gemini 3.8 Live Extended Thinking keeps speech, reasoning, and tool execution inside a single stateful session, even while external API calls are still running.

Coordinating two separate layers

With GPT-Live-1, the voice model handles the full-duplex conversation while a backend model such as GPT-6 Astra, a lighter model like Luna, or even a third-party option, handles reasoning and tool execution independently.

Keeping the backend work separate lets the voice layer stay responsive, with OpenAI putting turn-taking latency at around 800 milliseconds.

The tradeoff is that developers must coordinate the two layers themselves, passing context between the voice model and backend reasoner through sideband channels and deciding what the conversation does while background work runs. That orchestration burden falls entirely on the application layer.

Stale work when a user interrupts

Both approaches face the same headache when someone interrupts or changes their mind halfway through a request, leaving background work running that may no longer be needed.

In Gemini, that work stays within the same session, although developers have less visibility into exactly when a tool call stops. OpenAI leaves more of that cleanup to developers, who have to cancel pending jobs and make sure an outdated answer doesn’t find its way back into the conversation.

Google says Gemini has an edge under the messy conditions voice agents encounter outside a demo. The company tells The New Stack that Extended Thinking handles “background noise, heavy accents, and unexpected interruptions better than competing models.”

Per-minute costs diverge sharply

Standard Gemini 3.8 Live carries Gemini Live API rates of $0.005 per minute of audio input and $0.018 per minute of output. Extended Thinking adds reasoning tokens, with additional charges for inputs like live video and documents.

GPT-Live-1 costs $0.05 per voice minute for the front-end voice layer alone. The backend reasoning model, function calls, and external agent runs are all billed separately. As with GPT-6 Astra’s adjustable reasoning settings, developers can dial cost up or down per call, but a voice agent that regularly calls a more powerful reasoning model will see its bill climb fast.

The company also embeds DeepMind’s SynthID watermark in generated audio.

Benchmark numbers, with caveats

Gemini 3.8 Live Extended Thinking scored 82.6 on Artificial Analysis’ Speech-to-Speech Quality Index, with task completion rates of 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark.

The company points directly to those results, telling The New Stack that Extended Thinking “holds the #1 spot on the Speech-to-Speech Quality Index and leads on complex task-completion benchmarks.”

GPT-Live-1, paired with GPT-6 Astra at medium reasoning effort, scored 86.2% Pass@1 on Tau3’s spoken customer-service evaluation spanning airline, retail, and telecom domains. On Full Duplex Bench, it beat GPT-Realtime-2.1 by 30 percentage points.

“Extended Thinking holds the #1 spot on the Speech-to-Speech Quality Index and leads on complex task-completion benchmarks.”

Different tests, different stacks

But these results aren’t head-to-head. Google and OpenAI used different tests and setups, and Google also notes that some comparisons put developer models up against finished consumer products.

Claude Voice isn’t part of this developer calculus. Anthropic offers voice in its consumer apps but doesn’t currently offer a real-time speech-to-speech API comparable to Gemini Live or GPT-Live-1. Developers building voice agents around Claude still have to assemble more of the voice stack themselves.

Google keeps speech, reasoning, and tool execution inside one session, cutting down on middleware but tying developers more closely to its runtime. OpenAI requires more orchestration but gives developers more control over the models and tools running behind the voice layer.

The post OpenAI’s voice model doesn’t think. That’s the point. appeared first on The New Stack.

  •  

Bolt is giving developers 50x more compute. But there’s a catch.

Abstract glitlch

Bolt.new, StackBlitz’s browser-based AI development platform, is testing a new trade with developers: more coding-model usage in exchange for training data.

The company launched Forge on Monday, a research preview for individual Pro subscribers that offers up to 50 times more usage of open weight coding models through October 14. Developers who use Forge must opt in to sharing anonymized versions of their sessions for model training, including prompts, source code and the fix traces it creates as developers work through problems, in addition to their conversations with the coding agent.

The sessions will be used in a project with Arcee AI to help train a trillion-parameter-class open-weight model. The first training run is scheduled to begin in October, with Bolt saying the resulting model weights will eventually be released publicly.

Forge makes that development activity part of the exchange, with developers getting more compute while Bolt and Arcee get data from real coding sessions.

Forge makes that development activity part of the exchange, with developers getting more compute while Bolt and Arcee get data from real coding sessions.

Why Coding Trajectories Matter

Public repositories contain enormous amounts of source code, but they mostly show the end result. A coding session can fill in the gaps left by failed attempts and revisions along the way. The record of what worked (and what did not) is useful as the coding agents take on longer jobs.

Working across a codebase means finding the right files, coordinating changes, and recovering when something breaks. That gets harder when agents inherit code written by other agents, which isn’t always easy for the next one to understand or modify.

SpaceXAI showed one version of this approach last month when it trained Grok 4.6 on agent failure traces — the missteps, retries, and corrections that other labs typically discard. Bolt is making a similar bet but sourcing the data from developer sessions rather than synthetic runs.

Arcee has been working on the same underlying problem. In a blog post about NAC, its open-source agent harness, the company said software engineering tasks can stretch across tens of thousands of tokens as agents read code, edit files, run tests, and debug failures.

How Bolt Gets to 50X More Usage

Coding agents can burn through large numbers of tokens on even a single complex task, so a 50-fold increase in usage is a truly significant offer.

Forge changes the underlying setup by running open weight models on Bolt’s own infrastructure. The agent currently uses GLM 5.3 Flash and GLM 5.3, with Kimi K3 and DeepSeek v4 Pro available as experimental options.

Bolt’s WebContainers technology, built by parent company StackBlitz, gives it another cost advantage by running projects in an isolated environment inside the user’s browser rather than on Bolt’s servers.

Forge applies a similar approach to the models, using open weights on reserved hardware while developer sessions help train future versions, giving Bolt more control over costs and reducing its reliance on proprietary APIs.

That push toward self-hosted models is showing up elsewhere in the industry. Nvidia’s $12.9 billion bid for Hugging Face is arguably the same bet at a very different scale.

Coding agents can burn through large numbers of tokens on even a single complex task, so a 50-fold increase in usage is a truly significant offer.

Forge Scores 91% of Bolt’s Top Model

Forge’s open models scored 92.2 on the company’s internal Bolt Build Index, compared with 101.0 for its top paid model, putting them at about 91% of the top score. That’s only a measure of performance inside Bolt, so the 91% figure doesn’t tell us how those models compare more broadly.

But if Bolt can run more of its coding workloads on its own models instead of paying for proprietary APIs, it has more control over costs and usage, while the Forge sessions help train whatever comes next.

What Developers Are Giving Up

Forge requires an explicit opt-in, with a consent screen appearing each time a developer switches into the workspace. Standard and Max sessions aren’t included, and Teams and Enterprise accounts can’t participate.

Bolt says it anonymizes sessions before they leave its infrastructure, removing secrets, sensitive data, and personal information, and it tests the process against seeded data. Arcee receives the resulting data under a signed processing agreement.

Developers can stop sharing new sessions by leaving Forge, but Bolt says anything already used for training will remain in the models.

The 50× surge ends October 14, though Bolt says Forge itself will stick around as an open-model testing ground once the preview wraps up.

The 50× surge ends October 14, though Bolt says Forge itself will stick around as an open-model testing ground once the preview wraps up.

The post Bolt is giving developers 50x more compute. But there’s a catch. appeared first on The New Stack.

  •  

AI’s best coding agent fails 60% of the time — and the data backs it up

abstract screen

Claude Fable 5.1 just won a new coding benchmark despite failing more than six out of 10 times. Its 38.8% score comes from Real-SWE, a benchmark from Y Combinator-backed Specific Labs that takes a different approach to testing coding agents. Instead of giving them problems pulled from public repositories, it drops them into private codebases from real companies and asks them to tackle problems similar to those engineers deal with every day.

The code and its solutions aren’t publicly available, which makes it even less likely that they showed up in a model’s training data.

Specific Labs can’t guarantee that a model has never encountered any of the code, but the company says using private code makes that much less likely. It also estimates that 99% of tokens in real-world enterprises are hidden from frontier models. Once the agents were dropped into unfamiliar territory, the scores fell fast.

The code and its solutions aren’t publicly available, which makes it even less likely that they showed up in a model’s training data.

Private code changes the test

Fable 5.1, running via Claude Code, led the pack at 38.8%. GPT-6 Astra on Codex CLI followed at 33.8%, with Gemini 3.8 Flash on Gemini CLI at 31.2%.

After that, the scores dropped significantly. GLM 5.3 scored 28.8%, Grok 4.6 and Muse Spark 1.3 tied at 23.8%, Kimi K3 hit 18.8%, and GPT-5.6 Sol finished at 16.2%.

Each model got eight tries at every task. Real-SWE also tested each model with its own coding tool — Fable 5.1 with Claude Code, Astra with Codex CLI, and Gemini 3.8 Flash with Gemini CLI — so the scores reflect the full setup (not just the model).

As GPT-6 Astra’s ARC-AGI score showed, changing the scaffolding around a model can change its performance. Fable 5.1 in Cursor, for example, could produce a very different result.

Six tasks stumped everyone

Real-SWE uses proprietary code licensed from real businesses, including a consumer product with more than 200,000 users and a fintech platform that has processed more than 100,000 bank statements. The work also spreads across the codebase, with Real-SWE solutions touching a median of 11 files, nearly double the six-file median in benchmarks like FrontierCode and DeepSWE.

On individual tasks, the scores fell even further, with six of the 10 posting success rates below 15%.

On individual tasks, the scores fell even further, with six of the 10 posting success rates below 15%. A billing schedule migration had a 14.1% fix rate, API token metering landed at 12.5%, S3 storage tracking hit 10.9% and a linearizable scan came in at 4.7%, while a tax jurisdiction bug was patched just 3.1% of the time.

Not a single model solved the analytics stream reducer across 64 attempts. Astra and Gemini, meanwhile, went eight for eight on a multi-region sweep and Fable solved seven of eight, yet all three failed every attempt at the linearizable scan. No agent was consistently reliable across the benchmark.

Not a single model solved the analytics stream reducer across 64 attempts.

Where the agents broke down

Fable 5.1 most often missed requirements (36.7%) or ran into integration errors (34.7%). Astra’s failures were split between integration errors and unverified assumptions, both at 34%.

Integration errors appeared in nearly half of Gemini 3.8 Flash’s failed runs, while GPT-5.6 Sol made unverified assumptions in 43.3% of its failures.

What 38.8% really means

Real-SWE doesn’t prove that public coding benchmarks are inflated by data contamination, and 10 tasks is still a small sample.

But the top-performing agent still failed more than 60% of the time on private code it likely hadn’t seen before, suggesting that solving a coding problem is very different from finding your way through an unfamiliar production codebase.

The post AI’s best coding agent fails 60% of the time — and the data backs it up appeared first on The New Stack.

  •  

Perplexity’s new agent runs entirely on your GPU — with one expensive catch

Abstract server

Running an LLM on your PC is easy enough, but putting an agent to work there is a different story. Portable Computer, the local version of Perplexity’s Computer agent, is now available inside the Perplexity app for Windows on compatible Nvidia GeForce RTX and RTX PRO GPUs.

That’s the good news; the catch is, you’ll need an Nvidia GPU with at least 24GB of VRAM.

The Windows launch gives Perplexity three platforms in less than three weeks. Portable Computer debuted on Linux and Nvidia DGX Spark on August 25, followed a week later by hybrid compute for Apple silicon, which splits tasks between local and cloud models on Macs. Now Windows joins the mix, but bringing Portable Computer over took more than simply porting the app. Perplexity had to adapt the model runtime, orchestration, security, and hardware integration for each platform while keeping the user experience the same.

you’ll need an Nvidia GPU with at least 24GB of VRAM to use it.

Orchestration beyond the model

Portable Computer bundles those pieces together. On Windows, it supports PPLX 27B — Perplexity’s post-trained model — and Qwen 3.8 27B, both optimized for RTX GPUs, alongside a built-in browser, tool calling,  and Perplexity’s proprietary SPACE sandbox.

It’s a different lane from LM Studio or Ollama, which make running models locally as painless as possible but stop well short of giving a model autonomy over multistep work. DeepSeek’s recent hiring spree of roughly 150 new roles, nearly all of them focused on agent infrastructure rather than the model, hints at how much engineering sits between a capable model and a capable agent.

DeepSeek’s recent hiring spree of roughly 150 new roles, nearly all of them focused on agent infrastructure rather than the model, hints at how much engineering sits between a capable model and a capable agent.

Connectors blur local boundaries

Perplexity ships connectors for Microsoft Outlook, OneDrive, and Word, plus Google Drive, Gmail, Slack, and GitHub — which tells you something about what “local” actually means here.

The agent can reach external services because it’s not air-gapped. Locally completed tasks can process files without sending documents to a cloud model. Once an agent has access to both local files and remote APIs on the same machine, figuring out which resources it actually needs and where to find them gets harder.

Hybrid cloud as fallback

Perplexity isn’t pretending that a 27-billion-parameter model running on a desktop GPU can handle everything, which explains the hybrid architecture. When the agent determines that a task needs more reasoning power than the local model can deliver, it can escalate to Perplexity’s cloud models.

According to Nvidia, the agent identifies when cloud support would help and asks the user for permission before sending any data off the machine.

For organizations handling sensitive or regulated data, that split can make all the difference. A local agent can grind through source code or financial records without uploading them to a hosted model for basic processing. There’s a cost angle too, since tasks completed locally don’t burn Perplexity Computer credits.

High VRAM floor limits reach

Portable Computer is available with Perplexity Pro ($20/month) and Max ($200/month), across individual and enterprise plans, with Nvidia DGX Station support coming later. The real challenge is taking local agents from developer passion projects to enterprise-ready tools. By baking this into Windows, it immediately gets in front of the scale of users needed to make that happen.

The real challenge is taking local agents from developer passion projects to enterprise-ready tools.

The post Perplexity’s new agent runs entirely on your GPU — with one expensive catch appeared first on The New Stack.

  •  

OpenAI’s researchers burned $7,000 a day on AI agents — now it’s opening the floodgates

speed abstract

OpenAI rolled out its Agents API in public beta Thursday, opening the backend behind Codex to developers looking to run agents unattended for days.

Now, developers don’t have to build their own system to keep an agent going because the API tracks the job as it progresses and gives the agent somewhere to execute its work, even when a task stretches well beyond a single context window.

That makes long-running agents easier to try, but it also gives developers more ways to burn through compute. Interestingly enough, on the same day Agents API launched, OpenAI paused new sign-ups for its $200-a-month Pro plan after demand for GPT-6 Astra strained capacity.

Thibault Sottiaux, engineering lead for Codex, writes on X that Pro subscriptions “put the most strain on our systems,” adding that OpenAI was working to add capacity “as fast as we can.”

To make sure our current users have an incredible experience and continued access to Astra, we are going to pause subscriptions to our $200 Pro plan. These put the most strain on our systems and we wanted to take the smallest step that allows us to continue giving the broadest… https://t.co/WhLEm3HBL7

— Tibo (@thsottiaux) September 10, 2026

The Agents API and ChatGPT Pro are separate products, so there’s no reason to assume one is taking capacity from the other. Still, the timing stands out: the company is making it easier for developers to run agents for hours or days while pulling back access to its heaviest-use consumer plan and working to add more capacity.

Agent inference adds up fast

As a task gets longer, the API can compress earlier context, so the agent doesn’t just stop when it reaches the model’s context limit. It can also bring in tools only when they’re needed or send parts of a larger job to subagents working in parallel. The actual work can run in OpenAI’s sandbox or on infrastructure the developer controls.

The actual work can run in OpenAI’s sandbox or on infrastructure the developer controls.

As agents make progress, they go back to the model for the next step, and a job that takes hours can rack up far more inference than a typical API call. The usage climbs even faster when agents work in parallel.

OpenAI has already seen this inside its own shop. In a research report published September 6, OpenAI said its research organization was logging 3.1 agent-workdays for every human workday by mid-August, measured in standard eight-hour equivalents. The median researcher, ranked by agent usage, was spending more than $600 per day on inference at API prices, while the 90th percentile exceeded $7,000.

Before June, OpenAI’s researchers were still putting in more hours than their agent, but by mid-August, the agents were doing three times as much work.

Arguably, OpenAI’s researchers are an extreme case, but the numbers show what happens when agent use starts to scale. One person can suddenly generate far more inference than their headcount would suggest.

One person can suddenly generate far more inference than their headcount would suggest.

Friction limited compute demand

The Agents API lowers the cost of that experimentation by leaving the orchestration layer out of the bill. Developers pay for the models, tools, and hosted compute their agents actually use.

The flip side is that it’s now easier to consume more inference. Context compaction is a good example. A full context window used to force developers to decide what to discard or how to summarize the work so far. Now the API handles that automatically and the agent keeps going. That’s useful for developers, but it also means the workload doesn’t stop when the context window fills up.

Astra demand hit the ceiling

The Astra rollout offers a preview of what that could look like. OpenAI stopped accepting new Pro subscribers less than two weeks after the model launched on September 3, saying those accounts put the most strain on its systems. The Agents API has its own rate limits and usage tiers, so the Pro pause doesn’t directly affect developers using it. Still, the company is already having to manage capacity around its newest model.

Infrastructure outweighs benchmarks now

The more agents developers run, and the longer they run them, the faster that usage adds up. One developer might have several agents working at once, each going back to the model throughout the task. So headcount alone doesn’t tell you much about how much compute you’re using.

For long-running agents, the challenge is keeping the work moving without wasting tokens or losing track of the task. Cloudflare made a similar bet this summer, arguing that the infrastructure around AI workloads would eventually matter as much as the models themselves.

For long-running agents, the challenge is keeping the work moving without wasting tokens or losing track of the task.

The post OpenAI’s researchers burned $7,000 a day on AI agents — now it’s opening the floodgates appeared first on The New Stack.

  •