❌

Vue normale

Reçu aujourd’hui — 28 septembre 2026Infra

Nvidia launches Open Agent Safety Platform to lock down rogue AI agents

28 septembre 2026 à 12:00

OpenAI, Anthropic, Meta, and Google have all recently disclosed that their models broke out of their test environments and reached real systems. Nvidia’s response, announced Monday, is a runtime that locks agents into kernel-enforced sandboxes and a watchdog on its own silicon that can shut them down.

The Nvidia Open Agent Safety Platform combines OpenShell 0.1.0, the Apache 2.0 agent runtime the company announced at GTC in March, with Nvidia Sentry, a watchdog service that runs on the company’s BlueField-4 data processing units (DPUs).

The new OpenShell release adds a policy prover that checks that an agent’s various permissions can’t be combined into something the operator didn’t intend — like hacking HuggingFace.

Since the BlueField DPU is a separate processor with its own trust domain, it can watch the agent’s traffic to the model and keep an eye on all of its actions and reasoning. Then, when things go awry, it can cut the agent off at the network level.

Justin Boitano, Nvidia’s vice president of enterprise AI, said in a press briefing that the recent incidents “have highlighted a fundamental hurdle for AI agents, and that is that model-level safeguards alone can’t govern what agents can access or do.”

“To date, model safety has been about training good behavior into the model. The industry calls that model alignment,” Boitano said. “For probabilistic systems, this approach has obvious limitations. That’s why we’re introducing a deterministic system to mediate and enforce how these agents behave.”

Credit: Nvidia.

The summer of sandbox escapes

OpenAI disclosed on July 21 that GPT-5.6 Sol and a research prototype had exploited a zero-day in the package proxy that was their sandbox’s only network path and gone on to reach Hugging Face’s production database.

Only a few days later, Anthropic reported that three of its models had found unintended internet access at its evaluation partner Irregular and, among other things, accessed a real company’s database and published a live malicious package to PyPI.

Meta followed on Aug. 6 with a pre-release Muse Spark model that read and modified a real website’s database after the same Irregular misconfiguration handed it a real site name.

Most recently, Google announcedthat Gemini had broken into the networks for three companies, again from an Irregular test environment.

Irregular, which runs pre-release cyber evaluations for frontier labs, is a common factor here, and it’s worth noting that the company is on Nvidia’s partner list for this launch.

Nvidia’s own technical blog puts the breakouts down not to a new capability but to “a combination of tools, time, and ambiguous instructions.”

Asked whether the platform would have prevented the Hugging Face breach, Boitano said each incident is unique, but “from what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on.”

Enforcing policy outside the agent

OpenShell is the core of the platform, and it hasn’t changed all that much since Nvidia first showed it at GTC in March.

With OpenShell, which Nvidia originally announced in parallel with its NemoClaw distribution of OpenClaw, each agent runs in a kernel-isolated sandbox with no network access except through a supervisor that sits outside the workload.

“Traditional sandboxes, whether we’re talking micro VMs or containers or VMs, were built for application-level isolation,” Boitano said. “Every agent running within your company needs to run in its own isolated sandbox with security controls that are outside of the agent’s reach.”

A prover, not a judge

OpenShell is now at version 0.1.0, and the important new component added in this update is a policy prover. This Prover checks that the permissions a given policy grants always stay within the boundary the operator actually intended.

“It is deterministic. It is mathematical reasoning. So this is not LLM as a judge,” Ali Golshan, Nvidia’s senior director of AI software, said during the briefing. Because of that, he said, it runs “roughly at two orders of magnitude higher performance and speed.”

In Golshan’s example, a policy can, for example, bar an agent from reading code on GitHub and posting it externally.

“An agent can bypass this by spawning two sub-agents: one that can read from GitHub, that can talk to another one, that could also then post outside,” he said. The prover models the combined access of the entire agent fleet to find that path.

In Nvidia’s own tests, agents running with reduced safeguards spent up to two hours trying to talk an AI reviewer into granting write access to a protected repository. The prover, the company says, gave the reviewer evidence of what the request actually allowed, and no protected writes occurred.

Sentry: the safety island

Nvidia Sentry adds an additional hardware layer to this system. It runs on BlueField-4 in a trust domain separate from the host, and according to Nvidia, it can quarantine an agent in milliseconds.

With a DPU in the system, the agent’s model endpoint gets routed “through a proxy on the DPU, so that you can see all of the reasoning traces of the agents on the host,” Boitano said.

Unlike OpenShell, Sentry isn’t open source, though Boitano said it has open APIs and that OpenShell can work with other network enforcement hardware.

He compared it to autonomous vehicles. “There’s a primary system that might be running the perception system, and then a safety island that ensures the safety of the system.”

“The DPU is really optional in these architectures,” Boitano said. “In a lot of cases, just using OpenShell on CPUs is honestly good enough for providing sort of strict access control for the agents.”

The DPU, he said, is for “frontier use cases of evaluating models or systems where you might have the guardrails off the models, so it could be for red teaming.”

Who’s building on it

Anthropic is integrating OpenShell with Claude Managed Agents, which already keeps the agent loop on Anthropic’s infrastructure and pushes tool execution into customer-controlled sandboxes.

SpaceXAI says it’s using the platform for Cursor coding agents and Grok models, while Salesforce has added OpenShell audit events and permission approvals into Slack.

SAP is embedding the runtime into Joule Studio and is also contributing code.

OpenAI and Google, two of the four labs whose agents went rogue this summer, aren’t on the partner list. Neither is AWS.

Asked whether Anthropic and OpenAI plan to run OpenShell and Nvidia Sentry for their own training runs, Boitano said to look for the partners’ own blog posts.

The post Nvidia launches Open Agent Safety Platform to lock down rogue AI agents appeared first on The New Stack.

Reçu avant avant-hierInfra

OpenAI releases GPT-6 Sol and Luna — and cuts token prices in half

22 septembre 2026 à 20:00

OpenAI on Tuesday released GPT-6 Sol and Luna, which will complement the flagship GPT-6 Astra model in OpenAI’s lineup. As of now, there is no GPT-6 Terra.

The new GPT-6 pricing

The headline news here is that OpenAI cut the price per million input/output tokens by half or more, compared to the previous version. GPT-6 Sol will cost $2/$10 per million input/output tokens (vs. $4/$20 for GPT-5.6 Sol), and GPT-6 Luna will come in at $0.10/$0.50 (vs. $0.20/$1.20).

The GPT-5.6 pricing was always meant to be promotional, but for the new GPT-6 models, this is the default price, an OpenAI spokesperson tells The New Stack.

“Improvements in caching and inference let us serve these models at lower cost, and we’re passing those savings directly on to users and customers,” OpenAI explains in its announcement.

Benchmarks

As you would expect, the new models show clear improvements over the GPT-5.6 predecessors, but for the most part, these are not all that extreme.

On a benchmark like Zapier’s AutomationBench — which checks how well the models work on a set of business workflow tests — GPT-6 Luna improves by 5.4 percentage points over the previous version, for example

Credit: OpenAI

On the DeepSWE v1.1 software engineering benchmark, GPT-6 Sol essentially matches Anthropic’s Fable (68.8% at max effort vs. 69.9% for Fable 5 at xhigh effort), but at only 20% of the cost. Luna, at max effort, hits scores similar to Claude Opus 5 and Fable 5 at medium effort, at a significantly lower cost.

And OpenAI focuses on this cost comparison across its announcement—with a special focus on price per task instead of straight-up token pricing.

Credit: OpenAI

Anthropic resets the comparison

Since Anthropic released Opus 5.5 earlier on Tuesday, OpenAI’s comparisons are already out of date — such is the way of this AI era. Anthropic, too, reduced its per-token pricing for Opus 5.5 to $4/$20, down from $5/$25, but that still leaves Anthropic’s model twice as expensive as the comparable GPT-6 Sol.

In its announcement, when comparing GPT-6 Sol to Opus 5, OpenAI was able to claim significant cost savings when compared to Anthropic’s model — and for the most part that still holds, but Anthropic says Opus 5.5 also uses fewer tokens per task, which, according to the company, works out to 40% lower costs than Opus 5 on typical workloads.

It’s worth noting that no one has run Sol and Opus 5.5 head-to-head yet. Sol likely stays cheaper per task on OpenAI’s AutomationBench numbers, but Opus 5.5 posts higher scores than GPT-5.6 Sol on shared benchmarks in Anthropic’s testing.

Since it’s almost impossible to know how many tokens an agent will use to finish a task, though, these pricing changes still don’t make it any easier for a user to budget.

Prompt caching

For developers building agents, the caching changes may matter more than token prices. OpenAI says it improved prompt caching for GPT-6 to deliver higher cache hit rates by default, with discounts of up to 90% on cached input tokens.

One positive change, too, is that developers can now change the reasoning effort and tool availability without invalidating the cache. With explicit breakpoints, developers can choose where a cached prefix ends, and a new dashboard and diagnostics tool show what’s getting cached and what isn’t.

GitHub says these improvements cut the share of prompt tokens that require fresh processing by more than half over the past several months, across billions of requests to OpenAI models.

Anthropic made a similar move with Opus 5.5, which cuts cache read prices by 60% for token-billed usage, on top of the 20% per-token cut.

Style changes

Models aren’t just about benchmarks, though. With GPT-6 Sol, OpenAI made its models answer more directly, rather than in the previous — already reined-in — more conversational style. “Expect to see more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall without losing substance,” OpenAI says.

Credit: OpenAI

Alignment

Given the HuggingFace incident, it’s no surprise OpenAI is emphasizing its alignment work for GPT-6 Sol and Luna, too.

OpenAI says both models improve on their GPT-5.6 counterparts across its alignment evaluations, including fewer misleading claims about their own coding work. On an internal coding deception test, GPT-6 Sol’s rate fell to 1.3% from 10.4%.

When given a deliberately broken search tool — and graded on whether it disclosed the problem instead of guessing — Sol failed to disclose the problem 4.9% of the time, down from 77.5%.

What is a bit more concerning, though, is that when researchers asked the model to respect an explicit warning like an “access denied” message, GPT-6 Sol still tried to work around those restrictions in 64.4% of runs, down only slightly from 68.2% for its predecessor. Luna improved more, to 42.4% from 76.5%.

OpenAI says these tests cover mostly low-stakes situations and run without full system-level safeguards used in its products.

Credit: OpenAI

On a simulated message board seeded with unauthorized instructions, including requests to disclose private information, Sol took the specified action in 11.3% of runs where it found the board, down from 51.9%. Luna and Astra took none, though OpenAI notes Luna also found the board less often.

Anthropic, by contrast, says Opus 5.5 is the strongest performer on its most comprehensive alignment test and names METR and Frontier Design as pre-release external testers. Opus 5.5 also ships with safeguards that reroute requests, sending most cybersecurity tasks to Opus 4.8 and anything flagged by Anthropic’s biology or frontier LLM development classifiers to Opus 5.

Availability

GPT-6 Sol and Luna are available in ChatGPT Work and Codex starting Tuesday for Plus, Pro, Business, Enterprise, and Edu users.

Free and Go users get Luna in the desktop app.

Neither model is in Chat yet. OpenAI says it plans to roll them out gradually throughout the day to keep service stable, so they may not appear right away.

The post OpenAI releases GPT-6 Sol and Luna — and cuts token prices in half appeared first on The New Stack.

Harness rebuilt its Git repository for nonstop AI agent traffic

9 septembre 2026 à 13:00

As one of Harness’s field CTOs, Martin Reynolds spends much of his time asking engineering leaders one question with no easy answer.

It’s how they’re keeping up with all the pull requests their coding agents now produce. One of them recently answered with two words, “we’re not,” and explained that his team’s threshold for pushing code into production had dropped, Reynolds tells The New Stack.

TNS spoke with Reynolds a few days after Harness launched a rebuilt Code Repository and a new AI Code Review product — and a few weeks after GitHub’s nearly eight-hour platform outage on August 17.

In our conversation, we discussed how the review bottleneck arose, why Harness rebuilt its Git repository for agent traffic, and which parts of the pipeline should remain deterministic.

Drowning in pull requests

Reynolds says he first ran into the bottleneck during Harness’ early trials of GitHub Copilot and Amazon CodeWhisperer.

“We were getting more PRs, but all the PRs were getting stuck,” he says, “and the test team was shouting, saying, we can’t keep up with all of this.”

“Imagine what that test team feels like right now.”
—Martin Reynolds, Harness Field CTO.

The 1.5x to 2x increase in new code pushed the testing teams to the breaking point, he says, and now, he sometimes sees teams at 10x, with some claiming 50x. “Imagine what that test team feels like right now.”

Reynolds notes that during hallway conversations with engineering leaders at the conference, drowning in pull requests was a recurring theme. And what he sees when talking to customers tends to split three ways: Some have raised their risk tolerance, some have a backlog they can’t manage, and most sit in the middle.

“The somewhere in the middle, I think, is the most common,” Reynolds says. “We’re using some kind of another AI tool to help us in that space, but it doesn’t necessarily solve the problem.”

Review what’s changing, not the scaffolding

Ideally, a reviewer opening a pull request should see the most important changes first, Reynolds says, and he suggests reviewers should come from whoever has worked on that part of the codebase before, “not the person who did the prompt or wrote the code.”

“This other stuff is like 30 files because they updated a dependency. That’s less important in terms of getting eyes on,” he says. Reviewers should “actually review what’s changing rather than a bunch of stuff that’s scaffolding around it.”

“It’s not just the model on its own,” he also notes. Harness spent “a good chunk of the last 12 months” building what it calls a software delivery knowledge graph, a map of a customer’s pipelines, deployments, incidents, and policies, so the reviewer can pull context “at speed and not burn lots of tokens.”

The company’s own example is a migration flagged because an earlier incident review found an unindexed CREATE INDEX statement had locked a production table for 14 minutes.

By the company’s own count, its engineers saved more than 10,000 hours of manual review time a month. The day Harness launched, GitHub’s Copilot code review began reviewing pull requests opened by bots, including its own coding agent.

Agents don’t work nine to five

Recently, Harness customers on GitHub “would quite often genuinely send us screenshots of GitHub being down,” Reynolds says.

The reason for GitHub’s struggles, he believes, is that GitHub “was ultimately built for people, teams of maybe up to 10, 15, who are changing code, creating pull requests. Those pull requests will be there for a few hours to maybe a couple of days.” But agents “don’t work nine to five.”

Harness has been selling a repository service since 2023, when it launched Harness Code on top of its open-source Git project, and Reynolds says the company rebuilt it as “a ground-up AI-first repository that works for humans and AI.”

He describes it as Kubernetes-based, running across multiple clouds and regions, tested at thousands of commits per second, and used by about 20 enterprise customers during beta, none of which Harness has published.

GitHub CTO Vlad Fedorov’s postmortem on the August 17 outage said “a critical infrastructure component in our Central US data center failed to scale” as traffic hit a new peak. GitHub now handles 2.9 billion commits a month, a little more than 1,000 a second on average.

That’s not the scale Harness operates at, of course, but for its enterprise users, that may just be an advantage.

Harness is also starting to look beyond the traditional process. The capabilities for an autonomous delivery lifecycle exist today, Reynolds argues, but “are organizations and companies ready for that? I’m not entirely sure.”

Either way, he says deterministic tooling needs to stay, and test results still come from the test runner. “There’s no need to rip those out and replace them. It’s like, where can you enhance them?”

The reviewer is the part of this launch most teams will touch first. It works on pull requests that already live on GitHub, and moving a repository is a long project at most enterprises.

The engineering leader who told Reynolds “we’re not” doesn’t need a new Git host to change that answer. He needs something that tells his reviewers which files in a pull request still need a human and which 30 came with a dependency bump.

That’s a much smaller promise than an autonomous delivery lifecycle, but for now, it’s also likely the more useful one.

The post Harness rebuilt its Git repository for nonstop AI agent traffic appeared first on The New Stack.

OpenAI launches GPT-6 Astra and says welcome to the “AGI era”

3 septembre 2026 à 20:02

OpenAI on Thursday launched GPT-6 Astra, its newest flagship model. The company describes it as “the world’s most intelligent and aligned model,” and as far as the benchmarks go, that seems about right.

During a press briefing ahead of the launch, OpenAI President Greg Brockman took things a bit further than the benchmarks, though. After acknowledging that AGI remains a “gray, fuzzy thing,” he suggested future observers might look back at Astra as the model that marked its arrival.

“I think it’s not unreasonable to feel that we are now in the AGI era,” Brockman said.

Asked whether OpenAI was formally declaring that it had achieved AGI, he said the term was no longer tied to a contractual trigger (referring to its earlier agreement with Microsoft) and instead described it as a “mission concept or spiritual concept.”


More from TNS on OpenAI Astra


“I do leave it up to the reader to decide for themselves if this qualifies for them,” Brockman said. “For me personally, I do think we’re there. I do think there’s a pretty good argument for it. But again, I think this is the beginning of a journey, not the end.”

He closed the briefing with this: “Welcome to the AGI era.”

“Welcome to the AGI era.”
—OpenAI President Greg Brockman.

OpenAI’s biggest training run yet

According to OpenAI’s Aidan Clark. Astra was OpenAI’s largest training run to date. “It’s the first time we’ve pre-trained on more than 100,000 GPUs at our Stargate site in Texas,” he said. Clark also noted that Astra is the first OpenAI model release where earlier models played a significant role in supervising the training process.

Credit: OpenAI.

You likely won’t be able to use Astra just yet, though. The rollout starts with enterprise customers that already have access through OpenAI’s Daybreak program.

OpenAI says Astra will roll out to Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS, “in the coming days.”

Pro, Business, and Enterprise users will also get access to GPT-6 Astra Pro, and eligible API customers will be able to use Astra with Zero Data Retention.

Cost

Once it is available in the API, Astra will cost $10 per million input tokens and $50 per million output tokens. That is 2.5 times Sol’s current promotional price, although it matches Anthropic’s pricing for Fable 5.1.

It is far above Muse’s standard price of $1.25 per million input tokens and $4.25 per million output tokens. Meta also offers a Contributor version for $0.10/$0.20, respectively, but says data from that tier can be used to improve its products. Google’s introductory prices for Gemini 3.8 Flash are $0.75/$3.75.

A higher per-token price does not necessarily mean a higher bill if the model completes a job in fewer steps and needs fewer retries. OpenAI says Astra uses fewer tokens on several evaluations and in partner tests. The launch data is too sparse to show whether those savings offset the price premium.

“The price per task is what matters,” Brockman said.

Unlike with GPT-5.6, OpenAI has not announced Luna, Terra, and Sol variants for GPT-6. For now, the lineup consists of Astra and Astra Pro.

Where Astra leads — and where it doesn’t

OpenAI says that, unless noted otherwise, the models in its evaluations ran at maximum effort. That can improve benchmark results, but it can also increase latency and token use.

Astra’s DeepSWE v1.1 results are a clear improvement over OpenAI’s models. It scored 74.1% on the 113-task agentic coding test, compared with 70.8% for Sol.

But earlier this week, Meta reported 75.4% for Muse Spark 1.3 at its maximum reasoning setting. Muse 1.3 is available, but the maximum setting is under safety review and will not be generally available at launch.

That is a surprising win for Meta. On a benchmark of this size, the 1.3-point difference is roughly equivalent to one or two tasks, but it still shows how far Meta has come.

The public DeepSWE leaderboard currently puts Gemini 3.8 Flash and Claude Opus 5 at 74%, with Sol at 73%. The reported uncertainty ranges overlap, so these results do not establish a clear leader. OpenAI’s chart excludes Muse and uses a 67.4% Fable 5.1 result, making Astra’s advantage appear larger than the broader set of results would suggest.

Astra’s larger gains come outside coding.

The standout is its 98.6% score on ARC-AGI-3. OpenAI ran Astra with a Responses API harness that retains reasoning between turns and uses compaction to manage long contexts. The company previously demonstrated that those system choices can substantially raise ARC-AGI-3 scores without changing the underlying model, and the benchmark measures Astra and OpenAI’s agent systems together.

Astra’s 97.6% FrontierMath Tier 4 result appears to cover the 41 private problems in the 43-problem tier. Epoch AI, which runs the benchmark, says OpenAI funded its development and has exclusive access to part of it, so that’s worth keeping in mind.

Credit: OpenAI.

Astra scored 95.9% on BenchCAD’s 1,000-file Vision2Code subset with Python tools, compared with 84.3% for Fable 5.1 and 83.3% for Sol. BenchCAD asks models to reconstruct CAD programs from rendered views and scores the geometric overlap of the resulting 3D models. OpenAI notes that the Claude results used modified evaluation settings, but the gap relative to Sol remains significant.

On the Terminal-Bench Science task, which asks agents to complete 70 command-line research tasks across five scientific fields, OpenAI reports 64.6% for Astra. Anthropic reports 52.6% for Fable 5.1, while the existing public leaderboard tops out at 30% for Opus 5.

Anthropic also devoted much of its Fable 5.1 announcement to science, including early wet-lab results for protein binders designed by Mythos 5.1, the same underlying model with fewer safeguards. Both companies are trying to move their models beyond answering science questions and into research workflows.

What changes in Codex?

For developers, Astra’s handling of jobs that outgrow the context window may matter more than the benchmark gains.

Codex currently relies on compaction, which summarizes earlier work to free up context. That process can discard exactly the detail an agent may need later: why a previous fix failed, which tests ran or which small requirement the user added at the beginning of the job.

Astra can instead keep notes across context windows and search earlier messages and tool output. The feature is experimental behind a config.toml setting for now. OpenAI says it will become the default for Astra in the coming weeks.

Astra can also ask the user a question without stopping work that does not depend on the answer. This prevents a single unresolved decision from blocking the rest of the job, a common failure mode for coding agents.

OpenAI showed Astra operating applications including KiCad, Excel, Blender, and Power BI, as well as performing browser-based form entry and website QA.

On the OSWorld V2-Offline benchmark, which tests work across desktop applications, OpenAI says Astra scored 72.6%, up from 65.7% for GPT-5.6 Sol. It also cut the average time per task from about 75 minutes to 40.

Anthropic has reported a higher result of 77.9% for Fable 5.1, but says that the test used a different OSWorld release and should not be compared with previously published scores.

OpenAI also changed the Codex harness. On Mind2Web, Astra and the new harness completed tasks 1.9 times faster than the current Sol-based setup.

More capable, but harder to monitor

OpenAI’s claim that Astra is its most aligned model rests partly on an internal test in which Astra went outside an authorized target in 0% of impossible-task scenarios, compared with 48.2% for Sol.

OpenAI describes the older model as running “without production safeguards”; however, it does not make the role of the surrounding safety setup clear enough for a direct comparison.

The company also disclosed that Astra’s written reasoning was harder to monitor than Sol’s in evaluations specifically designed to elicit monitoring evasion. OpenAI attributes the decline partly to Astra having greater control over its written reasoning on simpler tasks and completing problems with fewer written reasoning steps.

“Progress in intelligence does not guarantee progress in alignment,” OpenAI Chief Scientist Jakub Pachocki said. He added that OpenAI “will withhold scaling until we can regain enough confidence” in its ability to monitor future models.

Cyber capabilities come with tighter access

OpenAI says Astra has crossed the Critical cybersecurity threshold in its Preparedness Framework. In company tests, the model developed exploits for hardened browsers and operating systems. It also found two previously unknown vulnerabilities while OpenAI was evaluating it against recent V8 bugs. The company says it is disclosing them to the maintainers. OpenAI says its reported cyber results reflect access to Daybreak Blue, not Astra’s default production configuration.

OpenAI describes ExploitBench and ExploitGym as tests of whether models can turn known software vulnerabilities into working exploits. On ExploitGym, Astra scored 42.4%, up from 30.3% for Sol, but OpenAI removed the usual six-hour time limit for both models.

On ExploitBench, the models scored 100%.

OpenAI notes that the version of Astra that’s available through standard access will refuse some advanced cybersecurity work, including exploit discovery.

OpenAI is giving an initial group of vetted defenders less restricted access through Daybreak and says it will expand Astra access through Daybreak Blue in the coming weeks. Daybreak Blue is an access program for authorized defensive work, not a separate Astra model or reasoning mode.

For developers using the API, a cybersecurity safety check will stop a task outright rather than pause it and wait for approval.

OpenAI’s Mia Glaese also warned that users outside its trusted-access programs may experience slowdowns, pauses, or blocks while performing cybersecurity work — and sometimes while doing unrelated work. That’s been an issue for many recent model launches, of course, but it’s frustrating for users.

“At launch, this is something that people should expect,” she said.

The post OpenAI launches GPT-6 Astra and says welcome to the “AGI era” appeared first on The New Stack.

Nvidia PAIR lets you put your idle Macs and PCs to work for AI agents

3 septembre 2026 à 18:00

Nvidia’s bet on open models and local AI has been taking shape for a while now. Its acquisition of Hugging Face will only accelerate that, but on the product level, too, the company is making some headway in getting more potential users to run AI models on their machines.

On Thursday, Nvidia launched the Nvidia Personal AI Router (PAIR), an open-source software router for your home network that lets you use idle Macs and PCs in your house to run small models on demand and speed up agentic workflows with the help of subagents.

PAIR is very much meant to speed up agents like NemoClaw, OpenClaw or Hermes by allowing them to use more subagents in parallel. The idea is for a lead agent to split up tasks across subagents, whose model requests can then run on these idle machines.

Credit: Nvidia.

Nvidia calls it a “virtual inference router” and is very clear that this is not a new inference engine. Instead, it uses the Ollama or LM Studio installs already running on a given machine and lets them run the models. Once installed on every machine, PAIR can then discover the different systems on your local network (using mDNS) and check if they are able to run a request.

The caveat here is that this does not split up a single inference request across different machines. Nvidia stresses that this doesn’t merge GPUs or pool VRAM into a single accelerator and can’t split a single inference request across machines.

“Agents can send a request through the familiar local interface it expects,” Nvidia explains. “PAIR receives the request through its proxy, identifies its engine and model requirements, and selects one eligible node. That node executes the request from start to finish and sends the response back through PAIR. The agent continues to see one connection while PAIR handles placement behind it.”

Credit: Nvidia.

PAIR supports Windows, macOS, and Linux machines with compatible GPUs. In practical terms, that’s Nvidia GeForce RTX 20 series GPUs and newer (which Nvidia says is the baseline), a Mac with M4 silicon or newer, or an Nvidia DGX Spark (and then, once they are available later this year, RTX Spark PCs and laptops).

It’s nice to see Nvidia supporting Macs here, which have obviously become quite popular for running local models and agents like OpenClaw, even though they don’t use Nvidia GPUs.

Each system can host different models, but PAIR will only route a request to a machine with the required engine enabled and the exact requested model available. Installing the same model on several machines gives the router more options for distributing concurrent requests.

Credit: Nvidia.

PAIR keeps an eye on which computers are available and once a user gets back to work (or gaming) on a given machine and reclaims the GPU, the local inference engine is stopped.

In Nvidia’s example, using PAIR with two PCs running high-end RTX 5090 GPUs with 32 GB of RAM (and a $5,000 price tag right now, despite their initial $2,000 MSRP) and the Qwen3.6 35B A3B model sped up work with five subagents by about 1.6x.

Few people have a bunch of RTX 5090 cards at home (and maybe a few DGX Spark desktops, too), so it remains to be seen what this will look like in a scenario with maybe a Mac Studio, a few Mac minis and maybe a gaming PC, but even that should speed up a local agent workflow as well.

Availability

Nvidia PAIR is now available as a beta. To get started, you just need to install it on every machine you want to be part of the network, have it discover and pair those systems, and make sure you have Ollama or LM Studio installed on them (and the models downloaded).

PAIR can also help with setup by installing Ollama or LM Studio and initiating model downloads on paired machines.

The post Nvidia PAIR lets you put your idle Macs and PCs to work for AI agents appeared first on The New Stack.

Google ships its third Gemini Flash model in six weeks

2 septembre 2026 à 18:56
Google Gemini

Google hasn’t released a Gemini Pro model in a while, but on Wednesday, the company launched yet another set of Gemini Flash and Flash Cyber models.

Even though Gemini 3.7 Flash is only three weeks old (and this is Google’s third Flash release in six weeks), the new Flash model often outperforms it by quite a margin, especially on agentic coding tasks and agentic computer use.

Like before, Flash 3.8 will be available at $0.75/$3.75 per million input/output tokens. That’s the introductory price, though. It will expire on December 31, 2026, and go to $1.50/$7.50 then.

Fairwind gates Flash 3.8 Cyber

Flash 3.8 looks to be a very good model, and Google is right in calling it its “workhorse model,” but Flash 3.8 Cyber is worth a mention, too. With this, Google is essentially copying the Anthropic playbook.

Google describes the Cyber model as its “most capable cybersecurity model with frontier-level performance in vulnerability detection and automated patching.” And because of this, it’s only available to a select number of “trusted defenders” through a new program Google calls Fairwind.

These include about 650 trusted partners like Accenture, CrowdStrike, the Center for Internet Security, Datadog, Palo Alto Networks, Snowflake, and Wiz.

“The defender’s edge comes from shrinking the time between detecting a flaw and patching it,” Four Flynn, Google’s VP for Security and Privacy, writes in the announcement. “Through Google’s Fairwind Program, government and enterprise partners gain autonomous tools to repair systems faster and at scale, keeping them one step ahead of agentic-speed threats.”

With Flash 3.5 Cyber, Google had launched a limited-access program, but that didn’t have a name and seemed a bit ad hoc at the time.

Gemini 3.8 Flash: long-horizon coding and agents

As for Flash 3.8, Google notes that the new model often outperforms models like GPT-5.6 Sol from OpenAI and Claude Sonnet 5 and Opus 5 from Anthropic when it comes to working on complex engineering tasks. On benchmarks like DeepSWE, for example, it matches Opus 5 and beats GPT-5.6 Sol and Sonnet 5.

There’s a caveat here, though, in that Google also stresses that “3.8 Flash works harder.” But working harder means it will also use more tokens, and Google, to its credit, is open about that and notes that, “On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels.”

To counteract this, developers can choose between different reasoning modes, though, just like with previous models.

Credit: Google.

Google’s benchmarks do not include Anthropic’s Fable 5.1, which was only released a day earlier. That’s a much stronger model, but also far more expensive.

Nobody would call Fable 5.1 a workhorse model, but it’s worth highlighting that on general agentic benchmarks like Terminal-Bench 4.0, where Flash 3.8 scores 19.1%, Fable 5.1 hits 55.8%, and Opus 5 gets to 51.8%. Meanwhile, on Terminal-Bench 2.1, which focuses on coding, Flash 3.8 does better than its competitors.

Terminal-Bench is the outlier here, though, as the Gemini model actually does quite well on other agentic tasks, even when compared to other flagship models.

Where Google’s model still struggles, despite some impressive gains, is on computer use (59% vs. 75.4% on OSWorld-2.0 for Opus 5) and GDPVal, a benchmark that tests models on knowledge work tasks, where Google, at 1545, remains well behind Opus 5 at 1824 but is finally catching up to Sonnet 5 at 1584.

Google notes that it was able to improve the model in such a short time in part because it’s now using agentic loops to improve the models, too. “Both of today’s releases are powered by the same foundational intelligence, and further accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models,” the team writes.

Chinese models close the gap

One thing worth noting is that Google’s comparison focuses on OpenAI and Anthropic, but it’s hard not to notice that on some benchmarks like DeepSWE 1.1, some of the Chinese models like GLM-5.3 and GLM-5.3 Flash, as well as DeepSeek v4 Pro and Kimi K3, are very much in the same league, too — and they often offer even better price/performance ratios.

Credit: Google.

Cyber benchmarks

As for Flash 3.8 Cyber, Google says it delivers “frontier-level performance in autonomous vulnerability discovery” on CyberGym, where it beats 3.5 Flash Cyber, which was released in July, and “significantly larger frontier models.”

CyberGym only covers C and C++ code, so Google also tested the model on an internal benchmark that spans 20 languages and reports a success rate above 70% there.

On Collinear’s CWE-Bench, the model scores a pass@1 of 47.2%, just behind “a leading frontier model at 47.8%” that Google doesn’t name.

Benchmarks only go so far, though. Chrome’s security team says the model produced 2.6 times more correct patches than “the best commercial models that are much larger,” while Wiz reports 7.5% to 9.7% higher recall on its internal penetration testing benchmark at 2.3 to 5.2 times lower cost.

Safety

On the safety front, Google says that 3.8 Flash is shipping “with safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense, while enabling beneficial use cases, as per our Frontier Safety Framework.”

Gemini Flash 3.8 Cyber has a more permissive set of safeguards, but that’s why it is only available to a select group of users, too.

Google does note that the new model is far more robust when it comes to prompt injections, too, with (almost) class-leading results in the Gray Swan IPI benchmark, for example.

Gemini Pro?

The release of the next Gemini Pro model will be interesting. These Flash models are improving rapidly, and while Google has fumbled the Pro launch a bit, that model may have been worth the wait. Google’s strategy to bet on its Flash models in the meantime also means that it has been able to push a price/performance narrative that has been harder to tell for other U.S.-based frontier labs.

At this rate, though, we may see a Gemini 3.9 Flash before Gemini 4 Pro arrives.

Availability

Gemini 3.8 Flash is now available in the usual Google products like Antigravity, Google AI Studio, Android Studio, and Stitch, as well as Gemini Enterprise.

Consumers with AI Pro and Ultra subscriptions can also use it in the Gemini App, AI Mode in Google Search, and Gemini in Google Sheets.

The post Google ships its third Gemini Flash model in six weeks appeared first on The New Stack.

Anthropic’s Fable 5.1 is a bit cheaper, a bit smarter, and refuses a lot less

1 septembre 2026 à 21:38

On Tuesday, Anthropic launched the latest versions of its flagship Fable and Mythos models. Anthropic promises that the updated models will offer stronger performance at a lower cost, largely because the company is reducing cache read pricing.

Like before, Fable 5.1 and Mythos 5.1 are the same model under the hood, but while Fable 5.1 is now generally available, Mythos 5.1 remains restricted to Anthropic’s trusted access program because it features safeguards that Anthropic says are “specifically designed to support work in cybersecurity and the life sciences.”

Pricing remains unchanged at $10/$50 per million input/output tokens, but cache reads are now only $0.25 per million, down 75%.

Fable 5.1, like Fable 5, remains restricted to API usage and users with Max, Team Premium, and Enterprise Premium plans (at 50% of their weekly usage limits). Pro subscribers can use it via usage credits.

Fable 5.1 benchmark results

In the benchmarks Anthropic shared, Fable 5.1 easily outperforms the older model, Anthropic’s Opus 5, and OpenAI’s GPT-5.6 Sol. Nobody is bothering with comparisons to Gemini 3.1 Pro anymore.

Credit: Anthropic

In general, Anthropic says the new model outperforms Fable 5 in virtually all aspects at a lower cost. This depends a bit on the benchmark, but for the most part, it comes down to being able to step down the reasoning mode by one or two notches and still get the same results as Fable 5 at higher settings, all with Fable 5.1 using fewer tokens overall.

In general, Anthropic says the new model outperforms Fable 5 in virtually all aspects at a lower cost.

Many of the improvements aren’t dramatic, except for a major step-up in Fable 5.1’s performance on the Terminal-Bench-Science 0.1 benchmarks, which test the model’s capabilities around using agents for scientific research. Here, Fable 5.1 scored 52.6% versus Fable 5’s 24.7%.

In most other areas, the improvements are within 2-4% of either Fable 5 or Opus 5.

Because of this, Fable 5.1 in Claude Code defaults to the high reasoning mode setting, but medium on Claude Cowork and Claude.ai.

Credit: Anthropic

Claude Code safeguards loosen up

One issue with Fable 5 was that it was often too cautious and deferred to the Opus models when it sensed that the user was pushing its safeguards. Now, Anthropic says, it has made these safeguards more precise, “ensuring that they’re less likely to flag benign content (like queries about medical issues or cyber defenders using the model to make their systems safer), but still ensuring they provide robust protection against genuine threats.”

For coders who use Claude Code, the most immediate change is that the model’s safeguards will hopefully get out of the way more often.

For coders who use Claude Code, the most immediate change is that the model’s safeguards will hopefully get out of the way more often.

Anthropic says its cyber safeguards now intervene about 60% less per session than they did with Fable 5, and the model is now allowed to identify vulnerabilities in code.

Penetration testing, exploit generation, and scanning binaries for vulnerabilities remain off the table, though, and Anthropic says it will continue to redirect those requests to its Opus models.

The biology safeguards got a similar update and now intervene 85% less often on benign requests about basic biology and medical questions.

What about those watermarks?

Anthropic recently shared its technique for watermarking text generated by its models (for models released after August 2, 2026.

The watermark is an invisible numerical marker that signals how likely it is that Claude wrote a given piece of text. Anthropic says it doesn’t affect output quality and contains no data about the user, organization, or conversation. Until now, though, nobody outside Anthropic had a way actually to check for it.

With this release, Anthropic is opening a detection API in private preview “to eligible organizations as required under EU law (such as regulators, law enforcement, media, fact-checkers, independent researchers, educational organizations, and EU civil society groups).”

Enterprises that are obligated to verify watermarking to comply with the EU’s AI Act will also get access, with a wider rollout planned for the future.

Data retention

One point of contention for enterprises that wanted to use Fable 5 was Anthropic’s retention policy. While Anthropic offers zero data retention to qualifying customers on its other models, Fable 5 users had to opt into a 30-day retention period.

Now, Anthropic is rolling out what it calls Enterprise Frontier Safeguards (EFS), which it says delivers the privacy of zero data retention while preserving its ability to monitor for misuse.

“EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic,” the company writes in its announcement. “It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention.”

Credit: Anthropic

Anthropic says it designed this new system based on customer feedback. Businesses in regulated industries, after all, were essentially unable to use the model under those data retention rules.

“We therefore sat down with customers to design a solution that could provide the best of both worlds: the privacy of ZDR and the safety allowed by monitoring across time and accounts,” Anthropic explains.

In practice, this means EFS moves that data into the customer’s own cloud storage, such as Amazon S3, Azure Blob Storage, or Google Cloud Storage, where it remains encrypted with the user’s encryption keys. It then runs Anthropic’s automated detection against it with no human review on Anthropic’s side, and routes any alerts back to the customer.

As of now, this tooling covers Claude Code, Claude Enterprise, the Claude Platform, and Claude on Amazon Bedrock, Microsoft Foundry, and Google’s Agent Platform. The company says there will be no additional cost beyond the customer’s own cloud storage bill.

EFS will roll out in phases starting this fall. Until then, Anthropic will not retain any data from usage of Fable 5 and Fable 5.1.

Anti-distillation

Fable 5.1 also ships with what Anthropic calls strengthened mechanisms to avoid distillation. The company, of course, has been quite vocal about its suspicions that Chinese labs have been distilling its models at an industrial scale.

The first of them is a change to the API itself: New accounts can no longer “manually edit Claude’s prior context in a multi-turn conversation while preserving the transcript of Claude’s prior thinking.”

This, Anthropic says, closes off “a common, publicly documented distillation technique.” However, it’ll also trip up agent harnesses that rewrite or compact conversation history, and the company says the restriction will extend to existing accounts with future model releases.

The post Anthropic’s Fable 5.1 is a bit cheaper, a bit smarter, and refuses a lot less appeared first on The New Stack.

Meta just beat OpenAI and Google at real-time transcription

1 septembre 2026 à 19:06

Meta’s Superintelligence Labs on Tuesday launched Muse Voice Transcribe, a new real-time speech recognition model that, on some benchmarks, outperforms virtually every other comparable model for real-time speech processing.

Meta’s lab describes the model as its first “real-time audio perception model.” With Muse Spark, the company also recently shipped another speech-to-text capable model, though not one that specializes in this use case.

The model can distinguish among more than 20 speakers, Meta says, and has been trained on over 70 languages (with 25 of them “extensively verified”), including cases where multilingual speakers switch languages in the middle of a conversation. It also supports long conversations of over an hour.

Credit: Meta.

It’s now available via the Meta Model API, Meta AI for Mac, and in Muse Code. The API pricing seems reasonable at $3.00 per 1,000 audio minutes (or $0.18 per hour).

Unlike with its Muse Glimmer models, Meta will not make the open weights of this model available, a Meta spokesperson told The New Stack.

Benchmark lead, in English

On Artificial Analysis’s AA-WER Streaming speech-to-text accuracy benchmark, the model’s word error rate is 3.1%, ahead of competitors like Cartesia Ink-2 (3.4%), ElevenLabs’ Scribe v2 Real-time (3.6%), GPT Live Transcribe (3.9%), and Gemini 3.5 Transcribe Live (4%). This benchmark only applies to English speech, though.

When it comes to recognizing distinct speakers, all models still struggle more than most users would like, but here, too, in these real-time use cases, Muse Voice Transcribe leads the pack with a 17.5% error rate across several standard benchmarks.

Credit: Meta.

How it all works

Under the hood, Muse Voice Transcribe is an autoregressive multimodal model from the Muse Spark family, Meta says, and the interesting part is how it decides when to talk.

Audio comes in as 80-millisecond chunks (12.5 per second), each compressed into a single soft token. At every chunk, the model makes a choice. It either emits a text token or a special “next audio” placeholder, which the system then replaces with the next chunk of audio.

When the audio stops, an “empty audio” token signals to the model that no more audio is coming, and it flushes any text it’s still holding.

Because the model controls how much audio it hears before committing to a word, it also controls its own latency. Meta calls this “adaptive delay.” The idea here is that difficult words get more context, while easy words get transcribed almost immediately.

That tradeoff is learned during the models’ reinforcement learning phase, where the word error rate reward and the delay reward are multiplied rather than added.

The system uses a similar mechanism for detecting speakers.

Real-time transcription has quietly become one of the most crowded corners of the AI market this summer, with OpenAI, Google, xAI, and Alibaba all shipping streaming models within weeks of each other, on top of the specialists that were already there. A 0.3-point lead on a benchmark won’t hold for long in a competitive field like that.

What Meta has, however, is a built-in reason to keep pushing. Every product it really cares about, from its glasses to the Mac app, needs this to work as well as possible.

The post Meta just beat OpenAI and Google at real-time transcription appeared first on The New Stack.

Google’s new forecasting model beats everyone. You can’t use it at work (yet).

31 août 2026 à 21:41

On Monday, Google launched TimesFM-3, a 330-million-parameter time-series forecasting model trained on over a trillion real-world and synthetic data time points.

The new model is now available on Hugging Face under a non-commercial license.

Large language models are great at predicting the next word. For businesses, time-series forecasting models essentially try to do the same thing, but for data. Over the last few years, there’s been a lot of work in building better forecasting models. Last year saw the launch of models like Chronos-2 from Amazon and Moirai 2.0 from Salesforce, while more recently, Datadog launched its Toto 2.0 model.

These models represent a relatively new breed of forecasting models, as they can ingest multiple time series. As Google research scientists Ayush Jain and Rajat Sen explain in the announcement, “most real-world forecasting problems are inherently multivariate: where multiple time series and auxiliary external features jointly impact the future forecast of a time series.”

Past sales, they explain, only tell part of the story. “A good forecast should also draw on sales of related products (e.g., ice cream cones, syrups), historical foot traffic, and known future events like weather forecasts, promotions, and holidays.”

TimesFM-3 is Google’s first model that was natively pre-trained to handle multiple time series and to do so with zero-shot generalization. This also allows it to forecast multiple related time series in parallel and to include historical data, such as past foot traffic.

In the benchmarks Google shared, TimesFM-3 outperforms all of these, often by a significant margin. The team looked at Salesforce’s Gift-Eval, Amazon/AutoGluon’s FEV-Bench, and Time.

What’s maybe the most surprising here is that TimesFM-2.5, which was state-of-the-art when it launched in September 2025, is now at the bottom of the benchmarks. That’s how fast this field is developing.

The architecture

Like its predecessors, TimesFM-3 is a decoder-only transformer that chops each time series into patches of 32 data points and treats them roughly the way a language model treats tokens.

What’s new is that these tokens now flow through two alternating kinds of attention layers. The first one looks backward across time within a single series and keeps things strictly causal, so the model can’t see values it shouldn’t know yet. The other looks sideways across all series at a given moment, which is how a promotion in one product line, for example, can inform the forecast for another.

Decoding changed, too. Earlier versions generated forecasts one patch at a time, Google’s researchers explain, which added latency and compounded errors along the way.

Instead, TimesFM-3 appends masked placeholder tokens for the entire forecast horizon and then fills them all in with a single forward pass.

The non-commercial license

Google decided to launch the new model under a non-commercial license. That’s becoming a bit of a trend in the world of model builders.

TimesFM-2.5 still shipped with the Apache 2.0 license — as do Toto 2.0 and Chronos-2.

The TimesFM-3 source code is still under the Apache license. Still, Google notes that “for the time being, TimesFM 3.0 pretrained weights are distributed under the separate timesfm-non-commercial-license-v1.0 license and are restricted to non-commercial, non-production use. Commercial or production use of the default pretrained weights is not permitted.”

Google will soon replace TimesFM-2.5 as the model that powers BitQuery’s AI.FORECAST command, so the company is actively monetizing these models.

That’s not unusual, of course. Every player in this market already integrates its forecasting models into its own platform, but Google restricting the state-of-the-art weights while also opening a paid path through its data warehouse is a pretty clear signal of where these labs think the money is in the long run.

The post Google’s new forecasting model beats everyone. You can’t use it at work (yet). appeared first on The New Stack.

Z.ai’s GLM-5.3 goes open weight, but its new license aims at hyperscalers

28 août 2026 à 19:42

Earlier in August, Z.ai, the Chinese AI lab behind the viral ox-alpha model that turned out to be GLM-5.3-Flash, launched its flagship GLM-5.3 model. On Friday, the company made the model’s weights available on Hugging Face, which Nvidia may soon own. Several third-party inference services already host it and make it available on services like OpenRouter, which Stripe will soon own.

As The New Stack’s Amanda Caswell reported when the model originally launched, Z.ai’s focus on training GLM-5.3 was on post-training. The result isn’t simply a large jump in benchmark performance over its predecessor, GLM-5.2, but a model that is often ahead of other Chinese open-weight models and can keep pace with current models from the large U.S. frontier labs.

A new license only hyperscalers won’t love

One thing Z.ai definitely changed is the model’s license. While GLM-5.2 shipped under the permissive MIT license, the new model ships under what the company calls the GLM-5.3 license and it has one major caveat: companies that want to host the model (not just route it like OpenRouter or embed it into a product), and have an aggregate revenue of more than $10 billion over any 12 consecutive months, “must pass Z.AI’s security review before using the Software or its derivative works for any commercial purpose.”

For individual users, nothing really changes — the license is more specific to models and includes the rights to run, deploy, and fine-tune. But for hyperscalers (and the neoclouds — once they hit these revenue numbers), there are now hoops to jump through.

The company says it held the GLM-5.3 open weights back for two weeks for safety evaluation and hardening. GLM-5.3 hits 84.5 percent on CyberGym, a vulnerability discovery benchmark, which Z.ai says is the best published result. That number is self-reported, of course, and nobody outside the company has reproduced it yet.

Z.ai says it used the model to find 2,436 vulnerabilities across 269 open-source projects, including the Linux kernel, though only a few dozen of those findings are publicly inspectable so far.

GLM-5.3 is now available for download, local deployment, fine-tuning, and commercial use under the GLM-5.3 License.

Given the model’s advanced cybersecurity capabilities, we conducted two additional weeks of comprehensive safety evaluations before releasing the weights.

Under… https://t.co/imV0F2O6bA

— Zixuan Li (@ZixuanLi_) August 28, 2026

Despite the safety framing, it’s worth noting that the license itself contains no acceptable-use section and also says nothing about cyber or offensive security.

Whether Z.ai changed its license for security reasons or to better monetize its own models is a question worth asking, of course. With the launch of GLM-5.3-Flash, the company stressed that inference was running on Chinese chips, so Z.ai has definitely shown interest in owning the inference layer for these models, after all. And now that these open-weight models are getting so close to the performance of what U.S. frontier labs are producing, there’s more money to be made there, too.

Back in 2023 and 2024, Z.ai also used a custom license for models like ChatGLM3-6B. Under that license, registration was required for commercial use. From then on, though, all new Z.ai models were licensed under the MIT license.

Z.ai’s license is also more restrictive than those used by other Chinese labs. Moonshot, for example, says that if a model-as-a-service provider offers access to Kimi K3 and has more than 100 million active users or more than $20 million in monthly revenue, “‘Kimi K3’ must be prominently displayed on the user interface of such product or service.”

DeepSeek still uses the basic MIT license for its flagship models.

Running GLM-5.3

Under the hood, GLM-5.3 is the same 753-billion-parameter mixture-of-experts architecture as GLM-5.2, with a 1 million-token context window and a maximum output of 128,000 tokens.

The weights ship in BF16 and FP8 and run on vLLM, SGLang, KTransformers, and Hugging Face’s own Transformers library.

Credit: Z.ai.

The two-week gap between the API launch and the open release is new for Z.ai. GLM-5.2’s weights were available on launch day.

Even though the model is now open-weight, you’re not all that likely to run it locally — unless you have a very powerful machine. Even the 2-bit quants, which Unsloth says will still achieve about 86 percent top-1 accuracy, need 245GB of memory. That just fits on a Mac with 256GB of unified memory. The 8-bit quants need 810GB.

For $1.40/$4.40 per million input/output tokens, it’s also significantly cheaper to use than virtually all of its most direct competitors — though the GLM-5.3-Flash model at $0.15/$0.47 has a pretty unbeatable price/performance ratio right now.

What’s next

At first glance, the license is a small change in absolute terms, but it comes in the same week that Nvidia moved to acquire Hugging Face for a reported $12.9 billion and Stripe agreed to acquire OpenRouter. If both deals close, the repository where developers download open-weight models and the marketplace where they rent them will be owned by American companies, while the models themselves increasingly come out of Chinese labs.

Z.ai kept MIT for Flash, so the company hasn’t completely abandoned permissive licensing, but it has stopped applying it to its best model. If GLM-5.4 ships the same way, the company’s MIT years will look like the customer acquisition phase, and open weights will start to look less like a gift to the ecosystem and more like a distribution channel with terms attached.

The post Z.ai’s GLM-5.3 goes open weight, but its new license aims at hyperscalers appeared first on The New Stack.

This duck will teach you reinforcement learning — and pick up your socks

27 août 2026 à 18:20

You could have a mechanical duck waddling through your home before Christmas. 

Hugging Face‘s Pollen Robotics on Thursday opened pre-orders for the Microduck, a $399 (introductory) bipedal duck-adjacent robot that can walk, waddle, and use roller-skates(!), with a beak to pick up objects. And when it falls, it can get back up, too.

The company expects to make the first deliveries of the Microduck before Christmas. It’s available in North America and Europe.

Microduck is the follow-up to Reachy Mini, the desktop robot Hugging Face and Pollen launched last year, which was stationary and focused on interacting with humans. Pollen also still sells Reachy 2, a far larger and pricier humanoid aimed at research labs. Microduck, the team writes in its announcement, is meant to focus on action.

Credit: Pollen Robotics.

“How do you teach a robot to move? How do you train a behavior in simulation, transfer it to real hardware, see what went wrong, and try again? What changes when the robot can leave the desk, carry something, fall over, and recover? It is an ideal platform for developers who want to train physical behaviors, experiment with reinforcement learning, and test how AI moves from simulation into the real world,” the team writes.

Not just a toy

And indeed, Microduck is not just a 25cm-tall toy. It’s an open-source platform with an SDK, virtual training environment, and reinforcement learning scripts to help developers train the robot to perform new tasks. The code is Apache 2.0, though the hardware design files are licensed non-commercially, so nobody is building and selling a clone.

There’s also a full simulator for those of us who just want to play with a duck robot, but if you do buy one, you’ll also get a game controller to control the robot on the fly, too.

Credit: Pollen Robotics.

But if you want to go deep, you can use the physics simulator to teach the robot new movements. On the project’s GitHub page, the team notes how to train the duck’s walking policy across 4,096 virtual ducks in parallel, for example, which results in a usable gait in one to two hours. The repo registers 13 task families in all, including a forward roll and six built around a set of passive wheels that go under the feet.

To train the robot, you’ll need an Nvidia GPU, or you can train it on Hugging Face’s own infrastructure. That’s not incidental. The simulator the ducks run in is MuJoCo Warp, built on Nvidia’s Warp framework, and mjlab, the training framework underneath, reimplements the API of Nvidia’s own Isaac Lab.

Out of the box, the robot comes with seven trained moves, including walking, sitting and standing, kicking, grabbing objects with its beak, roller skating, and getting back up off the ground.

Inside the hardware

As for the hardware, the robot will weigh in at about 800 grams and will be powered by a Rockchip RK3566 with AI accelerator. That’s basically a quad-core Arm Cortex-A55 with a Mali GPU. But now that Nvidia is reportedly in talks to acquire Hugging Face, in a deal first reported by The Information, I would expect a future version to use a slightly more powerful Nvidia-made chip.

Credit: Pollen Robotics.

It features 1GB of on-board memory and 32 GB of storage.

What’s more important, though, is its set of sensors. There’s a single front camera, a small LiDAR sensor with an 8×8 time-of-flight matrix, and 2 inertial measurement units. The camera and LiDAR let it see and place objects around it, while the IMUs keep track of its orientation and balance.

The team is still working out what the final camera resolution and LiDAR range will be.

Pollen Robotics will sell a few accessories as well, including, for example, a Charger Pack for $39 with two batteries and a charger, and a Dev Pack for $119 with three spare motors, five motor cables, two batteries, a dual charger, ten NFC tags, Hugging Face credit, screws, and a screwdriver.

Credit: Pollen Robotics.

The robot also comes with microphones and a speaker. One interesting note here: when you first turn the robot on, it generates its own signature sound that is different from any other Microduck. The robot doesn’t speak, though, as the team notes, the Microduck “communicates through weird little sounds, closer to a creature than an assistant.”

Everything is better with more ducks

The team says having several of them together is what really makes the robots come alive. “Races, football, or simply robots reacting to one another immediately make the experience feel more alive. For developers, it also creates a practical way to explore multi-robot behaviors without a room full of expensive hardware,” they write.

At the end of the day, this is also just a fun project, and in this depressing world, we all deserve some ducking fun every now and then.

The post This duck will teach you reinforcement learning — and pick up your socks appeared first on The New Stack.

Anthropic’s Claude now has a browser of its own

27 août 2026 à 00:05

Anthropic is giving Claude its own browser.

As the company announced on Wednesday, Claude on the desktop (Mac, Windows, and Linux) now has access to a built-in Chromium-based browser in Cowork that allows it to surf the web right where you’re already doing your work.

This finally puts Anthropic’s Claude desktop app in line with OpenAI, which recently gave up its ambitions to build Atlas, its stand-alone browser, and instead added it to its ChatGPT desktop app.

This new feature is now rolling out to paying Pro, Max, and Team plan subscribers.

A browser, not your browser

Previously, for Claude to browse the web, you had to install a Chrome extension. That always felt a bit like a hack, though it worked reasonably well. But having used the ChatGPT app with the built-in browser — and the ability to quickly take screenshots in the browser and discuss them with the model — having a dedicated browser in these apps makes many workflows much easier.

Fun fact: the original Claude Chrome extension launched exactly a year ago, on August 26, 2025. It’s also only been two weeks since Anthropic launched a major update to the Chrome Extension, essentially turning it into a Cowork session. But the AI development cycle moves quickly, and two weeks is a very long time.

Credit: Anthropic.

“Until now, giving Claude the ability to use the web in Cowork meant giving it access to your browser through the Claude in Chrome extension,” Anthropic explains. “When the work is on a page you already have open, that’s still the right choice. But a lot of web tasks don’t need your browser, just a browser, and now Claude has one.”

There is still a use-case for the Claude Chrome extension, though. It’s still useful to have Claude in the sidebar to talk about a page you have open in your browser.

“Claude in Chrome is for the page you already have open, with the accounts you’re already signed in to, such as updating your CRM, working through your inbox, or editing the doc in front of you,” Anthropic explains.

Also, if you have the extension installed already, it will remain the default for Claude to surf the web, unless you explicitly change your preferred browser in the Claude Desktop settings.

Bringing your logins across

The team notes this is very much “Claude’s browser, not yours.” But since it’s separate from your day-to-day browser, it doesn’t have your login info, for example. To work around that, Anthropic lets you bring your logins from Chrome, Edge, and Firefox on macOS — and Firefox on Windows and Linux.

The company very explicitly notes that logins from your banking and email providers, as well as any single sign-on sites, are explicitly excluded, unless you really insist on giving Claude access to your Schwab accounts.

The reason this list is inconsistent based on the operating system is likely that Chrome and Edge on Windows ensure that third-party applications can’t read your local cookies as an infostealer countermeasure. Firefox still stores these in a plain SQLite database. On MacOS, Chrome’s key is in Keychain, and other apps can request it from the user.

The risk is not zero

In this context, Anthropic also stresses that there is always a risk of prompt injections when letting an agent loose on the web. The company built safeguards into the Chrome extension and the built-in browser, but also notes that despite its best efforts, “the risk is not zero.”

Since the browser isn’t likely to be logged into many sensitive sites, if something goes wrong, the blast radius here is hopefully small.

The post Anthropic’s Claude now has a browser of its own appeared first on The New Stack.

Z.ai’s GLM-5.3-Flash is cheap, good, and served on Chinese chips

26 août 2026 à 18:26

Ox-alpha, the stealth model that quickly became the most popular model on OpenRouter in the last few days, is actually Z.ai’s GLM-5.3-Flash, a 320 billion-parameter hybrid model (with 18 billion active parameters) that the team specifically trained for ultra-low-cost inference.

On Wednesday, Z.ai unmasked the stealth model and made its weights available on Hugging Face under the MIT license.

It’s also already available on a number of inference platforms, including OpenRouter, where it’s currently available at $0.075 per million input tokens and $0.25 per million output tokens (though those prices reflect a 50% discount).

Benchmarks: It’s good, but not Fable

While some of the early hype put the model at the level of Anthropic’s Claude Fable 5, the benchmarks don’t bear this out. But the model can, for the most part, keep up with a Claude Opus 4.8 and OpenAI’s GPT-5.6 Terra when set to its max-effort reasoning mode.

On the Artificial Analysis Intelligence Index, GLM-5.3-Flash sits at 57 points, in line with GPT-5.6 Terra, Google’s Gemini 3.7 Flash, Meta’s Muse Spark 1.2, and Qwen 3.8 2.4T A95B.

When it comes to its performance in driving AI agents, which may be a better indication of how it will perform in real-world use cases, it’s doing even better than its competitors.

The model is able to understand multimodal inputs, including images, videos, and files. Z.ai also stresses that it trained the model to do better at visual tasks like building presentations and websites, as well as at standard knowledge work tasks like working with documents, spreadsheets, and dashboards, all of which also benefit from the model’s improved visual capabilities.

Credit: Z.ai.

One caveat: the model is very chatty and burns a lot of tokens, which isn’t unusual for smaller models that need more reasoning steps to achieve this kind of performance. Because the inference costs are so low, though, that’s not too much of an issue.

Cheap inference on Chinese chips

None of these competing models comes close to matching Z.ai’s pricing, and that may just be the most important takeaway from this launch. These Chinese open-weight models — including those from Alibaba, Deepseek, and Moonshot (Kimi) — are getting closer and closer to the performance of what American frontier labs can achieve — and they are making them available at a very low cost (and for free for those who have the hardware to run them).

Credit: Z.ai.

One fact that tripped many early ox-alpha users up was the fact that the lab behind the model was able to serve 100 trillion free tokens per day (according to OpenCode). Very few infrastructure providers can handle that. But as it turns out, Z.ai did all of this on Chinese AI chips, something the company heavily stresses in its announcement.

“Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs,” Z.ai writes. “This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.”

Credit: Z.ai.

Linear and sparse attention

The company doesn’t go into details, but notes that the individual chips are limited in compute and memory capacity. But Z.ai built its own inference engine based on SGLang for this serving stack — with the help of its flagship GLM-5.3 model powering an infrastructure agent, which the company says created “a feedback loop in which the model helped optimize the system serving the model itself.”

Z.ai doesn’t say what chips the model was trained on, but the company does note that it was trained on a 30 trillion multimodal pre-training corpus and that the team used a hybrid architecture that combines linear and sparse attention.

Credit: Z.ai.

“Linear attention captures local dependencies through state modeling, while sparse attention retrieves relevant global context through a lightweight indexer,” the team explains, and also notes that compared to the full GLM 5.3 model, GLM 5.3-Flash’s architecture reduces the compute by 3x and KV cache size by 4.4x.

The model has a context window of 1 million tokens, so those reduced KV cache sizes make a substantial difference.

For developers, the calculus here is pretty straightforward. A model that matches Opus 4.8 on the benchmarks that matter for agents, at a tenth of the price, is hard to ignore. For the U.S. labs, the harder question is what happens now that Chinese labs can serve at this scale.

The post Z.ai’s GLM-5.3-Flash is cheap, good, and served on Chinese chips appeared first on The New Stack.

IBM’s new Granite 4.2 models add reasoning and stay dense

25 août 2026 à 21:13

On Tuesday, IBM launched the latest family of its open-weight Granite large language models (LLMs). Weighing in at 3 billion, 8 billion, and 30 billion parameters, IBM is taking a very different approach to model building from some of the competition here, with dense, decoder-only reasoning models that it pre-trained from scratch.

Many recent models have moved from all-attention Transformers toward hybrid Mamba/attention architectures, including Nvidia’s Nemotron 3 family. But IBM tried that with its Granite 4.0 models. That generation included conventional dense, dense-hybrid, and hybrid MoE models.

With Granite 4.1, IBM returned the main family to an all-attention, dense Transformer architecture. At the time, IBM said that these new models outperformed the older generation, “while using a simpler — and therefore more flexible — architecture for fine-tuning for downstream tasks.”

Reasoning with Granite

IBM describes the 4.2 family as a “reasoning-focused release.” Models can run in thinking and non-thinking modes, but there is also a low-effort mode that only spends a low number of reasoning tokens for answering easy questions.

When the company launched the Granite 4.1 models, IBM still argued that reasoning models weren’t efficient enough, so “turning to less expensive, non-reasoning models with similar benchmark performance for select tasks like instruction following and tool calling makes sense for enterprise users.” At this point, though, the team clearly believes that reasoning is a necessary feature — even as it remains optional for the 4.2 models.

Unlike some other models in its class, including Qwen-3.8 27B, Muse Glimmer 30B, and Google’s Gemma 4 31B, it’s a text-only model, though. The others are all multi-modal, though IBM does offer Granite Vision 4.1 4B, for example, and there’s always a chance we’ll get a 4.2 version of this model, too.

It’s worth noting that IBM also launched two new speech recognition models in the Granite Speech family on Tuesday as well.

Training Granite

The Apache 2.0-licensed models were pre-trained on 15 trillion tokens and in five phases, including a long-context training phase that now brings the family’s context window to 512,000 tokens (although the released configuration natively supports 128K).

IBM notes that the model’s training set also included 1 trillion tokens of synthetic code, generated by IBM’s CodeAlchemy pipeline (though the models’ overall coding performance remains average).

For the most part, all of the models share this same pipeline, but the 8B and 30B models also went through an additional agentic reinforcement learning step to allow them to call tools, edit and run code, work in the terminal, and search the web. “Combined with reinforcement learning from human feedback (RLHF) alignment, this approach produces models better equipped for complex, multi-step agentic work,” IBM writes.

The 3B model also supports tool calling. Just don’t expect too much from it.

Benchmarks? It’s fine.

In terms of benchmarks, the Granite 4.2 models aren’t breaking any new ground, but it’s worth noting that the small 8B model often gets very close to the results of the larger 30B model — and it’s easy to run on virtually any modern Mac and even some relatively lower-end Nvidia RTX GPUs.

Qwen 3.8 27B, especially, beats the Granite models across the board, especially when it comes to coding, where the IBM models deliver inconsistent results overall.

As always, though, benchmarks only tell part of the story. For the right kind of usecase, frontier models are often overkill and IBM argues that the main argument for Granite is its performance in high-throughput agentic tasks, for example.

“This release extends the Granite family with a clear goal: helping enterprises build agents that can reason, act, and adapt during real-life workflows,” IBM writes in its announcement. “Now that AI systems are being asked to carry out tasks in the real world, our expectations have risen. It’s no longer enough to answer clearly and concisely. An AI system must be able to plan, call applications, and execute complex tasks in a reliable and consistent way — while staying light enough to actually use without breaking the bank.”

The post IBM’s new Granite 4.2 models add reasoning and stay dense appeared first on The New Stack.

Anthropic gives chat and Cowork one memory

25 août 2026 à 19:00

On Tuesday, Anthropic launched a major update to how Claude remembers things. The new system combines Claude’s memory in Cowork and chat into a single memory that now bridges both services. Previously, Cowork’s memory was bound to individual projects and had nothing to do with what Claude knew about you from chat.

It’s been a bit of an ongoing theme for Anthropic to bring Chat and Cowork closer together, and with this, Anthropic notes, Claude will remember what it knows about you no matter where you chat with it. That’s especially useful in Cowork, where you are more likely to do in-depth work. With the unified memory, Cowork can now remember important facts and, Anthropic argues, you won’t have to re-explain yourself as often. 

Cowork gets memory

“Ask Cowork to draft an update for your manager, and it already knows who that is and how she likes updates written,” Anthropic explains. “Brainstorm the agenda in chat for a conference you’re organizing; when Cowork builds the budget and logistics doc, it knows the headcount, the city, and the speakers.” 

Credit: Anthropic.

With this update, Claude now also updates the memory feature while you work with it. Previously, it would only do so after the chat ended. Since Cowork sessions could go on for hours or days, it makes sense to update these memories on the fly now.

“The context you’ve built up across months of conversations — for instance, your Q3 priorities and the status of your projects — is there the moment you hand Cowork a task, and what comes up in Cowork carries back to chat,” Anthropic says.

What Claude won’t remember

One thing Anthropic notes is that Claude, by default, won’t store anything it considers sensitive in its memory. This includes health data or information about a user’s religious beliefs, race, ethnicity, and gender identity, for example. 

Anthropic acknowledges that for some users, this may be exactly the kind of information they want Claude to remember, so there is an option in the settings menu to include what would otherwise be sensitive topics. There does not seem to be a way to tune this to only include certain topics, though. It’s all or nothing right now.

Credit: Anthropic.

 Some things never get stored in memory, though: sensitive identification numbers (SSN, government ID numbers, etc), criminal history, immigration status, or anything that violates Anthropic’s Acceptable Use Policy (AUP).

In a bit of a sideswipe at OpenAI, an Anthropic spokesperson also notes that “Claude also remains ad-free, so nothing in memory is used for ad targeting.” OpenAI, of course, also recently launched Computer History, a feature that can optionally keep tabs on everything you do on your Mac — including in Claude — and make that part of its recallable history.

Credit: Anthropic.

In practice, these memory files take the form of small files that are organized by topic. You delete them but you can’t change them directly. You can, however, chat with Claude to make changes right from the settings menu.

The unified memory feature is on by default for Free, Pro, and Max plans across web, desktop, and mobile. By default, the system won’t store any sensitive information, though. For users on Enterprise and Teams plans, all memory features are off by default. 

The post Anthropic gives chat and Cowork one memory appeared first on The New Stack.

Perplexity’s Computer agent can now run locally — if you can afford it

25 août 2026 à 17:34

Perplexity, in partnership with Nvidia, has taken Computer, its agentic AI assistant, and brought it to the desktop as Portable Computer.

While there is a lot of interest in local AI right now, getting started can still be difficult — and expensive. Portable Computer definitely makes it easier to get started, but to run it, you need the right kind of machine.

Local AI’s hardware bill

As of now, there are two main options. The easiest way is to get a DGX Spark desktop from Nvidia running DGX OS. But if you have a more traditional PC that’s running Ubuntu on ARM or x64 hardware, and you have an Nvidia RTX card with at least 24GB of GPU VRAM, you’re also in business.

A DGX Spark will set you back $4,800, and even an older RTX 3090 card with 24GB of VRAM currently costs well over $1,500.

Thanks to RAMmageddon, this won’t likely get cheaper anytime soon.

Bringing Computer to the desktop took more than replacing a cloud model with a smaller one. It needs to be able to read and edit files, run shell commands, and process PDFs locally — but it also needs to connect to external services. All of this, too, needs a local sandbox for the agent to run in.

To account for these smaller models, Nate Kupp, Perplexity’s vice president of Computer Enterprise and Infrastructure, tells The New Stack that the team had to “revisit almost everything throughout the stack.” The company reused many of Computer’s capabilities, he says, but changed the harness and model configuration for local hardware.

Indeed, the harness accounted for most of the engineering work, says Kupp. The model still has to plan tasks, call tools, manage files, and execute multi-step tasks, all with fewer parameters than the much larger models that Perplexity typically uses in its cloud service.

Perplexity says the orchestrator that maintains the agent loop is deterministic code and not yet another AI model. In this system, the local model proposes an action, while the orchestrator assembles the context, enforces policy, and runs approved tool calls in an OS-level sandbox.

Inside the local harness

According to the company, that sandbox restricts processes, filesystem paths, and network access. If the sandbox isn’t available, the harness disables itself before making any tool calls instead of running them outside the sandbox.

The harness is also designed around the model’s practical context limit. Perplexity says Qwen3.8-27B supports a 260,000-token context window but begins to struggle beyond 100,000 tokens. Because of this, portable Computer keeps the core prompt and tool set small and only loads additional skills as needed.

In addition, Perplexity also turned commonly used connectors into command-line tools rather than exposing their larger Model Context Protocol (MCP) definitions directly to the model.

In recent weeks, it has become increasingly clear how important the harness is for agentic performance, something Kupp also noted in the briefing.

With the same Qwen3.8-27B base model and DGX Spark hardware, Computer scored 82.6% on Perplexity’s internal 53-task Local Knowledge Work Bench, compared with 77.6% for Pi and 74% for Hermes. On ParseBench-100, a 100-task subset covering charts, layouts, tables, text, and formatting, Computer scored 65.1%, while Hermes scored only 34.6% and Pi 13.9%.

Perplexity says it plans to open-source the internal benchmark.

When a local task leaves the machine

With Portable Computer, each task starts on the local device. If the local model can’t complete a step, it can ask a cloud model for advice. Perplexity says the harness selects the relevant context, flags potentially sensitive information, shows the user what would be sent from the machine, and asks for approval before making that call.

The cloud model returns text guidance but does not get direct access to the device’s files or tools. The local orchestrator retains control of execution and incorporates the advice into the same local run.

In one demo, Portable Computer reviewed a folder of tax documents on an Nvidia DGX Spark, while in a second demo, the local agent analyzed a CSV file and then posted its findings to Slack — demonstrating the agent’s ability to reach beyond the local machine.

Portable Computer includes connectors for Google Drive, Gmail, Slack, and GitHub. When used, web searches and connector calls leave the device, while Perplexity says model inference and private-document processing remain local.

Perplexity’s argument centers on convenience and privacy. Developers who want to go through all of the necessary steps can already combine a local model server with tools and an execution environment. Still, Kupp argues that Portable Computer comes ready to run. “You don’t have to fiddle around with inference,” he says.

Despite the similar name, Portable Computer isn’t Personal Computer, Perplexity’s Mac and Windows application for working with local files and native apps. Portable Computer runs the model and harness locally, but “we’re not doing computer use at this point,” Kupp says.

Starting with Nvidia

At launch, users can choose between Qwen3.8-27B and PPLX 27B, Perplexity’s post-trained version of that model.

Nvidia’s Nemotron 3.5 Lightning is coming later. Users can also bring their own models and inference servers.

Sadly, even if you have a very powerful Mac (or are pre-ordering one of the new ones), Portable Computer isn’t available for you yet. Windows support, however, is scheduled to follow in September.

“For now,” Kupp says, Perplexity is “very focused on Nvidia across DGX and RTX,” though he also says that the company is considering other hardware.

Portable Computer will be available to subscribers of Perplexity Pro, Max, Enterprise Pro, and Enterprise Max. Tasks completed locally don’t consume usage-based credits unless the system has to reach out to Perplexity’s cloud models.

The post Perplexity’s Computer agent can now run locally — if you can afford it appeared first on The New Stack.

❌