❌

Vue lecture

Claude Opus 5.5 vs. Opus 5 on reasoning tasks: Cheaper, faster, but not better

Abstract long-exposure photograph of red and orange light trails forming layered curves around a dark central shape.

When Anthropic released Claude Opus 5.5 this week, the company claimed the new model costs 40% less than Opus 5 and generates output 30% faster. Anthropic’s marketing makes three claims. Opus 5.5 performs at the level of Claude Fable 5.1 (so it should outperform Opus 5), costs 40% less than Opus 5 on typical workloads, and generates output more than 30% faster.

Anthropic also cut the price developers pay to use the model through its API. Opus 5.5 costs $4 for every million tokens (chunks of text roughly three-quarters of a word long) sent to the model and $20 for every million it writes back, down from $5 and $25 for Opus 5. That price cut alone accounts for a 20% saving. The rest of the claimed 40% saving has to come from the model using fewer tokens.

I wanted to see how this translates for the average Claude user, so I skipped the usual developer workflow simulations this time. Lately, the models I test handle everyday tasks well. Reasoning tasks are where I’ve seen them struggle, so I tested Opus 5 against Opus 5.5 on reasoning tasks only. 

You can find the prompts at the bottom of this post if you want to replicate these tests on your own system.

The tests

I called both models through the Anthropic API with identical prompts. Both ran with adaptive thinking at the default effort level, since Opus 5.5 doesn’t allow you to turn thinking off. Each problem ran once per model. I planned to rerun any problem where the models gave me different results, but they never did.

Here are the tests I ran:

  • Logic grid (medium difficulty) – Seven engineers each have an on-call day, a language, a service, and a city, and 22 clues pin down one answer. Six clues are conditional or “exactly one of these is true” statements, and removing any single clue breaks the puzzle.
  • Constrained orderings (hard difficulty) – Reorder 10 deploy jobs so no job stays in its original slot and no two consecutively numbered jobs sit side by side. The model had to give the count for 6, 8, and 10 jobs.
  • Stone game with memory (harder difficulty) – Players remove 2, 5, 7, or 11 stones, but can’t repeat their opponent’s last move or their own. The model had to find who wins from 200 stones, count the losing starting sizes up to 500, and name the smallest losing size above 340.

I logged input tokens, output tokens, cost at list price, and time for every call. Thinking tokens are billed as output, so I included them.

The logic grid

Both models got all 28 cells right. Opus 5.5 took 65 seconds and 7,573 output tokens, for $0.16. Opus 5 took 108 seconds and 10,621 output tokens, for $0.27.

On this test, Opus 5.5 was 43% cheaper and delivered the same correct answer. Opus 5.5 was slightly more detailed and noted that it didn’t fully prove the solution was unique.

Constrained orderings

Neither model produced an answer, which made this the hardest problem in practice. The correct counts are 27, 1,695, and 159,019. The third answer is very hard to reach by reasoning alone. A computer program that checks every possible ordering can find it, but neither model could run code in this test.

With a 48,000-token output limit, both models spent the whole budget thinking and never replied. Opus 5.5 used 489 seconds and $0.96. Opus 5 used 553 seconds and $1.20.

I raised the limit to 128,000 tokens and ran it again. Opus 5 used every token, took over 25 minutes, and stopped with no answer. That cost $3.20. Opus 5.5 ran for 19 minutes and used 112,733 tokens, but the API ended the response with a “refusal” stop reason and no text. The prompt asks the model to count job orderings and contains nothing sensitive. That means the refusal was most likely a mistake by Anthropic’s safety filter flagging a harmless request.

Opus 5.5 was cheaper, but how much does that matter if you don’t get a result?

The stone game

And we’re back to the same answers again. Both models answered all three parts correctly. The first player loses from 200 stones; starting sizes from 120 to 500 are losses, and the smallest loss above 340 is 344.

The difference in this test came down to how much thinking each needed to get there. Opus 5.5 finished in 215 seconds with 28,740 output tokens, costing $0.58. Opus 5 took 624 seconds and produced 74,981 output tokens, costing $1.88. Opus 5.5 used 62% fewer tokens and cost 69% less for the same answer. 

Results

TestOpus 5.5Opus 5
Logic grid28/28, 1:05, 908 in / 7,573 out, $0.1628/28, 1:48, 906 in / 10,621 out, $0.27
Ordering problem (48k limit)No answer, 8:09, 235 in / 48,000 out, $0.96No answer, 9:13, 233 in / 48,000 out, $1.20
Ordering problem (128k limit)No answer (refusal), 18:56, 235 in / 112,733 out, $2.26No answer, 25:24, 233 in / 128,000 out, $3.20
Stone game3/3, 3:35, 323 in / 28,740 out, $0.583/3, 10:24, 321 in / 74,981 out, $1.88
Total tokens1,701 in / 197,046 out1,693 in / 261,602 out
Total time31 min 45 sec46 min 49 sec
Cost$3.95 ($4 in / $20 out per million tokens)$6.55 ($5 in / $25 out per million tokens)

Across every call, Opus 5.5 wrote 103.4 tokens per second and Opus 5 wrote 93.1, so Opus 5.5 was about 11% faster. Its biggest speed lead on any single problem was 19%, still short of Anthropic’s 30% claim. Its biggest cost savings came on the stone game, where it cost $0.58 to Opus 5’s $1.88, 69% less. Total spend for the test was $10.50. 

What do I think

Anthropic’s benchmarks show Opus 5.5 ahead of Opus 5 on coding, knowledge work, and reasoning. I didn’t rerun those benchmarks. I gave both models the same three reasoning problems, and they performed equally. Both solved the logic grid and the stone game, and both failed the ordering problem.

The savings are real. Opus 5.5 cost less and finished sooner on every problem, including 43% less on the logic grid and 69% less on the stone game. Most of that came from using fewer output tokens. The price cut accounts for 20%. Its writing speed was 11% faster, short of the 30% Anthropic claims.

Switch to Opus 5.5 if you run Opus 5 today. You get the same results on hard reasoning for less money and less waiting. Set a hard output limit and watch your spend on hard problems, though. Both models can think for close to 20 minutes or more and return nothing, which is a problem if you pay for every token. For counting problems like the ordering test, give the model a code execution tool instead of hoping it reasons its way through.

The prompts

Logic grid

Seven engineers (Ana, Ben, Cy, Dee, Eli, Fay, Gus) share an on-call rotation. Each is on call on exactly one day of a single week, Monday through Sunday (Monday is the earliest day, Sunday the latest), and no two share a day. Each writes a different language (Go, Rust, Python, Java, Kotlin, TypeScript, C++), owns a different service (auth, billing, search, queue, cache, gateway, metrics), and is based in a different city (Berlin, Tokyo, Denver, Lagos, Sydney, Toronto, Mumbai).

Clues:

The Java developer is Gus.

The TypeScript developer is on call exactly two days after the queue owner.

Exactly one of these is true: the engineer based in Berlin owns billing, or the C++ developer is on call Monday.

Exactly one of these is true: the engineer based in Tokyo owns cache, or the engineer based in Tokyo writes Rust.

If Fay is on call Sunday, then the engineer based in Denver does not write Kotlin.

Exactly one of these is true: the Kotlin developer owns auth, or the C++ developer is Cy.

The engineer based in Tokyo is on call earlier in the week than the C++ developer.

The engineer based in Denver does not write Python.

Exactly one of these is true: the Go developer is Ben, or the gateway owner is based in Berlin.

Ana is based in Mumbai.

The TypeScript developer is on call earlier in the week than Fay.

The engineer based in Sydney is on call earlier in the week than the billing owner.

The gateway owner is on call earlier in the week than the Kotlin developer.

The engineer on call Sunday is not based in Berlin.

Fay and the engineer based in Lagos are on call on consecutive days.

Eli is on call exactly three days after the auth owner.

The engineer based in Lagos writes Java.

Exactly one of these is true: the search owner is Eli, or the cache owner is Dee.

The gateway owner and the Go developer are on call on consecutive days.

The engineer based in Lagos is on call exactly four days after the auth owner.

The engineer based in Lagos owns metrics.

The engineer based in Mumbai is on call earlier in the week than the Python developer.

Determine the full assignment. At the end of your response, give exactly seven lines, one per engineer in the order Ana, Ben, Cy, Dee, Eli, Fay, Gus, in this format:

ANSWER: Name | Day | Language | Service | City

Ordering problem

A build system has n deploy jobs numbered 1 to n. Originally, job k runs in slot k. You reorder all n jobs into slots 1 to n (each slot gets one job) subject to two rules:

No job runs in its original slot (job k is not in slot k).

Jobs with consecutive numbers never run in adjacent slots (for example, jobs 4 and 5 cannot be in slots i and i+1 in either order).

How many valid orderings are there for (a) n = 6, (b) n = 8, (c) n = 10?

At the end of your response, give exactly three lines in this format:

ANSWER a: <number>

ANSWER b: <number>

ANSWER c: <number>

Stone game

Two players play a game with a pile of stones. They alternate turns. On each turn, a player removes exactly 2, 5, 7, or 11 stones, subject to two rules:

You may not remove the same number your opponent removed on their most recent turn.

You may not remove the same number you removed on your own most recent turn.

(On the very first turn of the game, neither rule applies. On the second player’s first turn, only the first rule applies.) You cannot remove more stones than are in the pile. A player who has no legal move on their turn loses. Both players play perfectly.

(a) Starting with 200 stones, does the first player win?

(b) For how many starting pile sizes from 1 to 500 inclusive does the first player lose?

(c) What is the smallest starting pile size greater than 340 for which the first player loses?

At the end of your response, give exactly three lines in this format:

ANSWER a: <yes or no>

ANSWER b: <number>

ANSWER c: <number>

The post Claude Opus 5.5 vs. Opus 5 on reasoning tasks: Cheaper, faster, but not better appeared first on The New Stack.

  •  

Query decomposition doesn’t fix context starvation — it just moves it

Abstract digital art showing warped light lines surrounding a void, illustrating data compression and AI context starvation.

I have a small example that would best communicate the message I am trying to convey: say you built a chat widget for GitLab’s public documentation (the corpus we are experimenting with in this article) and one of the developers sends this kind of message:

We got an email saying our card was declined for something called “quarterly reconciliation” and I need to know what actually happens now. On top of that, I think we’ve gone over our seat count; there are more people in the group than seats we bought. Our CI has been queuing all week and I want to know whether the compute minutes we purchased last month rolled over or if we lose them. Our finance lead also needs to be the one who gets the invoices from now on, not me. And last thing, is the REST API rate limited? We’re building an internal dashboard and would rather find out now than after it breaks.

Five separate asks: the declined payment, the seat overage, compute-minute rollover, changing who receives invoices, and API rate limits. Each one is answered by a specific passage in GitLab’s public documentation, and you labeled which passage answers which before running anything, so you knew in advance exactly what a correct system needed to find.

Then you run the message through a pipeline that follows current best practice. It splits the query into five clean sub-queries, retrieves for each one independently, merges the results, drops near-duplicates, reranks the merged pool against the original message, and packs the highest-scoring passages into a 2,000-token context.

The pipeline retrieved all five correct passages, but only one of them survived into the packed context; that is one of five asks, not one of five sentences. The packer found, scored, and threw away the other four before the model ever saw them. The same message with no decomposition at all managed three out of five.

The failure has a name, and it isn’t the one you’re thinking of

I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.

The definition is deliberately about allocation rather than retrieval, because allocation is the part nobody watches. If the passage answering the fifth question was found, scored, and then squeezed out by three passages about the first question, the fifth sub-intent is starved, and every recall metric you have will report that the system worked perfectly.

“I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.”

Two failure modes already in circulation describe something different, and it’s worth separating them cleanly:

Semantic dilution happens at retrieval time: when you embed a five-part question as a single vector, you get a centroid that sits somewhere between five topics and lands close to none of them, so the evidence is never found. Decomposition fixes this, which is why it spread.

Context poisoning is about what is present, not what is missing. Wrong, stale, or adversarial content enters the window and corrupts what the model generates downstream. Poisoning is a contamination problem. Starvation is an absence problem, and policy, not accident, produces the absence.

I borrowed the word from operating systems. In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work. Every ingredient of that situation is present in a retrieval pipeline: a fixed resource, competing demands, and a policy that decides who gets served. A relevance-greedy packer is priority scheduling with no aging term, and under priority scheduling without aging, valid low-priority work waits forever.

Decomposition is the right fix to the wrong half of the problem

Split that message into five single-intent queries, and each one embeds cleanly, so per-sub-query recall climbs sharply. This is well-trodden ground. LlamaIndex ships a SubQuestionQueryEngine that breaks a complex query into sub-questions and synthesizes the responses. LangChain’s MultiQueryRetriever generates query variants and returns the unique union of what they retrieve. RAG-Fusion applies reciprocal rank fusion across the per-query result lists. The technique works, and it isn’t mine.

“In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work.”

The context window did not grow. Let me explain: after decomposition, you have n result sets competing for one fixed token budget, and something downstream has to decide the split. In most production pipelines, that something is a short, unremarkable sequence: merge the pools, drop near-duplicates, rerank the merged pool against the original query, then greedily fill until the budget closes.

That sequence is a scheduler. It has a priority function, which is the reranker score, and it has no fairness constraint of any kind. A sub-intent with three strongly-scoring passages takes three slots. A sub-intent whose single correct passage scores mid-pack takes none of them.

So the failure did not go away. It moved from the embedding, where it has a name and people watch for it, into the packer, where it has neither. It also moved somewhere with much worse instrumentation, because recall@k per sub-query is the metric decomposition usually gets validated with, and that number goes up. It goes up at the same time as coverage inside the packed context goes down. You ship on a green dashboard.

The harness

The corpus, GitLab’s public documentation: 10,000 chunks and 2.2M tokens, split on heading boundaries and capped at 480 tokens each. Sixty-one single-intent questions span nine topics, from seat management to rate limits, each labeled with the one passage that answers it. I built multi-intent queries by concatenating those questions while varying n across 2, 3, 5, and 7, randomizing the order so position doesn’t confound topic, and varying topical distance so half the queries draw everything from one topic and half span distinct ones. That produces 100 queries, 25 at each value of n, whose correct decomposition I know exactly.

Two decisions matter more than the rest:

The metric is not recall. Recall tells you what the retriever found. What I need is what survived into the packed context, per sub-intent. So I log each sub-intent twice: once for whether its correct passage reached the candidate pool, and once for whether it reached the packed context. The gap between those two numbers is the entire argument.

Every question has to be retrievable on its own before it’s allowed in. A question enters only if its correct passage ranks in the top 10 for its own isolated query, under both retriever configurations, and both scored 100% recall@10 on that test. Since each sub-query’s candidate pool is exactly its own top 10, passing that gate guarantees the correct passage sits in the pool for every decomposed arm. Any sub-intent that then fails to appear was denied by the packer rather than missed by the retriever, which removes the most obvious objection to everything below.

The core measurement uses no language model. Because queries are composed from known sub-questions, the decomposer is an oracle so that anyone can reproduce the main result with no API key.

That invites an objection, so I tested it. A real LLM decomposer, blind to n, disagreed with my ground truth on 41% of the queries, and on inspection it was right every time. Five of my sixty-one supposedly single-intent questions contain two distinct information needs. What is excess storage usage, and what happens when we go over the free limit? is two questions wearing one question mark. Adjusted for those five, agreement is 100 out of 100. That is not evidence decomposers are reliable, because my queries are joined by fixed connectives and splitting on those alone recovers n perfectly, which real messages never allow. What it caught was an error in my own labels, and that is the best argument I have for the oracle design.

Results

Every arm runs at 2,000, 4,000, and 8,000 tokens against two retriever configurations. The stronger pairs are BAAI/bge-base-en-v1.5 with BAAI/bge-reranker-base; the weaker pairs are a quantized BAAI/bge-small-en-v1.5 with Xenova/ms-marco-MiniLM-L-6-v2. Retrieval is in-memory cosine similarity over a NumPy array because, at 10,000 chunks, a vector database would be slower to write, slower to run, and harder to verify.

Sub-intent coverage at a 2,000-token budget on the stronger configuration:

Table showing coverage (in %) per arm.

The production-default pipeline starves 31.1% of sub-intents whose evidence it had already retrieved. It beats no decomposition by nine points, while a flat B/n split, which is the crudest allocator anyone could write, beats it by fourteen.

Floors work, but not the obvious floor. Reserving one passage per sub-intent before the greedy fill satisfied 99.4% of its reservations and bought only seven points. The mechanism fires correctly and reserves the wrong passage because it picks each sub-intent’s best chunk by score against the original query. The original query asks about all five intents at once. Selecting that same reservation by score against its own sub-query pushes coverage to 89.4% and cuts allocation starvation from 30.6% to 10.1%. That is a change of about four lines.

Two of my own recommendations died here. I expected reranking against the original query to beat reranking against the fragment, and it loses by seventeen points. The incomparable score scales I worried about turn out to help, because each sub-query’s best match ends up at the top of its own scale, producing per-intent fairness for free. I also expected deduplication before allocation to matter, but near-duplicates consume 1.0% of the budget and removing them moves coverage by 0.3 points.

The crossover: starvation by tokens-per-sub-intent, which is simply the budget divided by n:

Table showing tokens per sub-intent.

Below roughly 1,000 tokens per sub-intent, allocation policy dominates. Above it, nothing you do to the allocator matters, because everything fits anyway.

The retriever comparison is the one I’d lead with. Upgrading the retriever moves coverage on the production-default arm from 56.2% to 68.9%, a gain of 12.7 points. Changing the allocation policy on the same retriever moves it from 68.9% to 89.4%, a gain of 20.5 points. In its sharpest form: the weaker retriever with a fragment-scored floor reaches 84.5%, and beats the stronger retriever with a greedy packer at 68.9% by sixteen points. A worse retriever with a better allocator wins.

Position: Held within a fixed n so query difficulty doesn’t contaminate the comparison; starvation across the seven positions of an n=7 query runs 1.3%, 21.3%, 30.7%, 48.0%, 48.0%, 32.0%, and 13.3%. That is a serial-position curve. The packer protects what you asked for first, protects what you asked for last a little less, and drops the middle. The fragment-scored floor flattens it to 4.0%, 9.3%, 6.7%, 8.0%, 22.7%, 5.3%, and 6.7%.

“A worse retriever with a better allocator wins.”

Topical distance: Sub-intents that span distinct topics starve about twice as often as sub-intents drawn from one topic, at 38.1% against 18.3% for n=7. I predicted the opposite. A topically coherent query gives the reranker a coherent target, and it scores all the correct passages similarly. In contrast, a scattered query lets it latch onto some topics and abandon others.

What the user actually sees

Everything above is retrieval-side. What decides whether any of it matters is what reaches the person who wrote the message, so I generated real support replies from 80 packed contexts and had every reply graded per sub-intent, with both the generation and the grading blind to which arm produced which context.

When the correct evidence reached the packed context, the reply addressed that question 100% of the time, across 261 out of 261 cases, in both arms. Coverage predicts the generated outcome exactly, which is the strongest justification I have for measuring it.

When a sub-intent was starved, the reply answered it anyway 48.1% of the time, based on whatever else happened to be in the window. It explicitly flagged the gap 45.6% of the time, with some version of “I’ll follow up on that separately.” It went silent only 6.3% of the time.

“Starvation mostly does not produce silence; it produces unsupported answers.”

I expected silence, and I was wrong. Starvation mostly does not produce silence; it produces unsupported answers. Whether those answers are actually incorrect is the next experiment, because this harness measures whether a question was addressed, not whether the answer was right.

Some limits: composed queries are cleaner than real support messages, which carry pronouns, implicit context, and conditional clauses. This is one corpus and one embedding family. I drafted the gold labels with model assistance and verified them myself. The model writing those replies was strong, so a cheaper production model would plausibly flag fewer gaps and invent more.

What an allocator actually looks like

Give every sub-intent a floor, and choose it by fragment score. Not the naive floor, which satisfies 99% of its reservations and buys seven points. Select a reservation by relevance to the sub-intent it protects, not by relevance to the message as a whole.

Rerank against the fragment rather than the original query. This inverts what I expected and what I have seen recommended. Scores from different fragments are not comparable across sub-intents, and that incomparability is doing useful work.

Don’t spend your effort on deduplication; near-duplicates cost 1.0% of the budget here. Dedup is worth doing, but it isn’t why your fifth question went unanswered, and treating it as the fix will cost you weeks.

Log per-sub-intent coverage: You already computed it to pack, and it predicts the generated outcome perfectly. A sub-intent that received zero passages is the best predictor available that your reply is about to assert something you cannot support.

Where parallel decomposition breaks

“If it’s late can I get a refund” is one clause and two intents, and the second one’s retrieval target depends on the first one’s answer. Parallel decomposition treats them as siblings. It retrieves the late-delivery policy and the general refund policy, packs both, and misses that the passage you actually need covers refunds for late delivery, which may match neither sub-query particularly well.

There are two ways out: You can tag dependencies at decomposition time, or run a deferred second pass that re-retrieves conditional clauses once the first round resolves.

I would take dependency tagging, for three reasons: A second pass costs a full retrieval round trip inside a latency budget a support bot does not have. The tag is reusable, because a dependent sub-intent should not hold a floor reservation. At the same time, its parent is unsatisfied, so it feeds the allocator directly instead of bolting on a separate mechanism. And it fails visibly, since an untagged dependency shows up as a starved sub-intent in the coverage signal. In contrast, a deferred pass that resolves the wrong condition produces a confident wrong answer with nothing to flag it.

The cost is real; dependency tagging pushes work onto the decomposer, which is already the weakest component in the chain, and I have not measured tagged against untagged. That is a design position rather than a result, and it is the one thing here I am asking you to take on argument instead of evidence.

What to measure on Monday

Take your production pipeline and compute one number: your context budget divided by the average count of distinct questions per incoming message. If that number lands below roughly 1,000 tokens, your allocation policy costs more than your retriever does, and the reranker upgrade sitting in your backlog will buy you less than reserving one slot per question.

On my corpus, the retriever upgrade was worth 12.7 points of coverage, and the allocation change was worth 20.5, which is why I think the ordering is wrong in most pipelines I’ve seen. That ordering is the falsifiable part. Run the same two comparisons against your own corpus, and if the retriever wins, I want to see the numbers, because that result would tell me the crossover sits somewhere other than where I measured it.

The cheaper thing to do first takes an afternoon. Log, for every multi-intent request, how many sub-intents ended up with zero passages in the packed context. A support system that cannot tell you which question it dropped will keep answering that question anyway, about half the time, out of whatever else was in the window.

The post Query decomposition doesn’t fix context starvation — it just moves it appeared first on The New Stack.

  •  

“Impressive level of openness”: Xiaomi goes way beyond the usual open-weight playbook with MiMo-V2.6

A picture of an open laptop

New models are coming out thick and fast, almost on a weekly cadence, ranging from the powerful proprietary systems coming out of the major US AI labs to the more open alternatives being released by some of China’s biggest tech companies.

On Tuesday alone, Anthropic debuted Claude Opus 5.5, while OpenAI launched GPT-6 Sol and Luna, each accompanied by their the usual claims about how they outperform their rivals. Amidst all the hullabaloo of the frontier-model frenzy, however, Xiaomi also debuted MiMo-V2.6, another powerful open model from one of China’s growing ranks of AI developers.

All the initial headline numbers look pretty promising, too. The flagship MiMo-V2.6-Pro is a trillion-parameter model, with 42 billion parameters active at a time, a one-million-token context window, and support for text, images, audio and video. Broadly speaking, that puts it in the same frontier territory as the latest models from OpenAI and Anthropic: GPT-6 Sol has a 1.05-million-token context window, while Claude Opus 5.5 has a one-million-token window, though neither company discloses comparable parameter counts.

Xiaomi, for its part, makes broad claims of frontier-level performance across coding, agentic tasks, cybersecurity, multimodal work and research. Independent analysis lends some weight to those claims –Artificial Analysis gives MiMo-V2.6-Pro an Intelligence Index score of 46, ranking it first among the 114 large open-weight models it tracks.

Artificial Analysis  Intelligence Index
Artificial Analysis Intelligence Index



So far, so good. But arguably the bigger story in Xiaomi’s offering is the manner in which it trained the model, how much of that process it showed in public, and what it’s releasing afterward.

A public record

Xiaomi livestreamed its RL training through a public dashboard, exposing metrics from the production reinforcement-learning runs in real time over a five-day period starting on September 15. By the time the runs had finished, the dashboard showed costs of $854,044 for the smaller MiMo-V2.6-Flash model and $2,620,670 for Pro — about $3.5 million combined.

Xiaomi livestreamed its RL runs over a 5-day period.
Xiaomi livestreamed its RL runs over a 5-day period.

It’s worth noting that this figure covers only the RL stage; Xiaomi hasn’t said what pretraining the models cost. Even so, public RL bills are rare. The closest precedents came last year, when MiniMax said the RL phase of its 456-billion-parameter MiniMax-M1 cost $534,700 in GPU rental, and DeepSeek put the RL training of its 671-billion-parameter R1 at $294,000. Both were leading open reasoning models when they launched, though the comparison only goes so far: MiMo-V2.6-Pro is larger, and its RL run targeted longer, agentic tasks.

Shortly after the stream began, Fuli Luo, who leads Xiaomi’s MiMo team after previously working at DeepSeek, took to X to explain the thinking behind the project. The team, she said, had spent almost six months exploring how far RL could be pushed, increasing the amount of training, the variety of environments and agent setups, and the resources used to grade the model’s attempts.

“We’ll open-source the details piece by piece over the coming weeks,” she added.

Nearly half a year of silence. We spent it studying one problem: how far RL can scale.

MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task…

— Fuli Luo (@_LuoFuli) September 16, 2026

Responding on X, Hugging Face co-founder and chief science officer Thomas Wolf called the move an “Impressive level of openness on such a large run.”

However, what Xiaomi’s putting out alongside the finished models is arguably just as interesting. The company has released the model weights under the permissive MIT license, alongside its technical report and a 9-billion-parameter Qwen-based model, intended as a starting point for further agentic RL research.

“Impressive level of openness on such a large run.”

Xiaomi says it has also “fully open-sourced” a broader set of RL resources: more than 7,000 task environments spanning software engineering, vulnerability reproduction, knowledge work and web development; an end-to-end training framework covering everything from environment interaction to reward evaluation and policy optimization; and lightweight agent harnesses for experimenting with different tools, prompts and context setups. At the time of writing, however, Xiaomi’s link to the open-source collection on Hugging Face contain only the three model releases, with the 7,000-plus environments and other supporting resources not surfaced there. Luo had said earlier that Xiaomi would be open-sourcing the various elements “over the coming weeks.”

As the results began arriving this week, attention in the research community quickly moved beyond the benchmark score to what Xiaomi had committed to releasing overall. Elie Bakouch, a former Hugging Face researcher who is now a research engineer at Prime Intellect, singled out the promised RL resources.

“The most insane part, they will release ~7k RL training data and the framework leading to this top 6 model on AA,” Bakouch writes on X. “They also shipped the model + tech report less than 1 week after starting the final RL run.”

Wolf went further, arguing that access to the environments in which models learn may now be especially valuable for open research, as more model development shifts toward RL with verifiable rewards (RLVR). Because RLVR depends on tasks whose outcomes can be automatically checked — whether code passes a test, for example — the environments themselves become a crucial ingredient in training.

“Releasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now.”

“Releasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now,” Wolf writes. “The equivalent of sharing high quality pretraining data, but in the new RLVR paradigm.”

Open-weight vs open-source

So while the benchmarks around Xiaomi’s latest model are notable in their own right, it’s the company’s approach that is generating much of the fanfare so far.

Indeed, MiMo-V2.6 serves as a useful example of a distinction that often gets muddied in the AI sphere: “open-weight” and “open-source” are routinely used as though they mean the same thing, but they don’t. Many “open” models amount largely to downloadable weights — essentially, the vast collection of numerical values a model learned during training, which can then be used to run or fine-tune it — while much of what went into producing them remains closed.

Some companies have gone further in muddying those terms. Meta, for example, has often referred to its Llama models as open-source despite significant restrictions that have led open-source advocates to push back heavily on that description.

And so MiMo-V2.6 goes further than most open-source releases. Its MIT license carries none of the conditions that the likes of Moonshot’s Kimi K3 and Alibaba’s Qwen3.8-Max attach for large commercial users. And if the environments are released as promised, outside researchers will have much more of the post-training process to inspect and build on.

The post “Impressive level of openness”: Xiaomi goes way beyond the usual open-weight playbook with MiMo-V2.6 appeared first on The New Stack.

  •  

A third option is emerging in the fight over AI and your data

Split-screen video interview with The New Stack host Alex Wilhelm and VAST Data cofounder Jeff Denworth.

Not your keys, not your coins. Not your model, not your data?

Over the summer, the tech industry was consumed by a debate about AI use in the enterprise and the need to protect IP. If an enterprise used proprietary models, was data leakage a necessary evil?

Companies seemed to have two options: They could use state-of-the-art, proprietary models and risk losing control of their data, or they could use open-weight models and never kiss the frontier.

Thankfully, a third option is emerging.

Consider the concern: Company A wants to use LLM B from AI Lab C, and they want to avoid training AI Lab C how to eat Company A’s lunch by building its capabilities into LLM B. A good way to resolve the tension would be to let Company A run LLM B on its own infrastructure, so there’s no risk of its information fleeing on the wind.

AI agents are “creating a whole different set of requirements at the data layer.”
–Vast Data co-founder Jeff Denworth

But that raises another problem: AI Lab C doesn’t want to allow Company A to run LLM B on its own GPUs because it doesn’t want to hand over its model weights. It’s the same IP issue the company ran into, in reverse. You have to solve the trust problem in both directions!

Enter VAST Data co-founder Jeff Denworth and a new product called DataEnclave, which aims to let AI labs and enterprise-scale companies deploy proprietary models in secure compute environments without risking data transfer in either direction. (DataEnclave uses Nvidia’s Confidential Computing technology to make the system tick; Vast Data’s core product is AI OS, infrastructure that fits beneath a company’s AI applications.) 

The New Stack had Denworth on the podcast to chat about the confidential computing market. I was curious about timing. Why did Vast build DataEnclave now? Nvidia began rolling out Confidential Computing in a serious way in 2024, after all. Denworth argues that the market needed the core technology, yes, but also demand.

And until late 2025, AI demand was modest compared to today’s token totals. Once agentic coding tools took off, corporate demand for AI products soared. This led to the pricing crisis we saw in early 2026, and the secure AI usage debate we endured over the summer. 

Performance drove demand, demand drove usage, and usage dug up fresh problems to solve. Now the question for the market is whether or not DataEnclave has solved enough concerns on both sides of the proprietary AI-proprietary data equation. The market will sort that out as it moves through early access and into general availability.

Our conversation goes deep into the arc of AI, where companies are in their AI journey today, and how much data remains to be unlocked inside the enterprise. If you want to feel the acceleration, it’s a fun one!

The post A third option is emerging in the fight over AI and your data appeared first on The New Stack.

  •  

Anthropic made Opus 5.5 cheaper. Then it broke four things your agent depends on.

Four sections, branched

Anthropic made Claude Opus 5.5, released on Tuesday, cheaper than its predecessor, cutting the price from $5 to $4 per million input tokens and from $25 to $20 per million output tokens. The 1 million-token context window and 128,000-token maximum output are unchanged.

On paper, that makes upgrading an easy decision. In practice, it may not be as simple as changing the model ID.

Anthropic’s migration guide flags four breaking changes that can cause requests built for Opus 5 to return 400 errors after switching to Opus 5.5. Several other changes won’t trigger an error but could still change how an existing agent behaves.

Anthropic’s migration guide flags four breaking changes that can cause requests built for Opus 5 to return 400 errors after switching to Opus 5.5.

Thinking is always on

The first change involves thinking controls. Opus 5.5 returns a 400 error when a request sets thinking to disabled or uses enabled with budget_tokens, leaving effort as the way to control how much reasoning the model does. Agents that previously switched thinking off for simple steps to save time and tokens will need to assign those steps a lower effort level instead. Because thinking is now always on, responses begin with thinking blocks, so code that assumes the first content block is text will also need to change.

The default effort level has also dropped from high on Opus 5 to medium on Opus 5.5, so requests that omit the parameter will quietly run at a lower setting. Anthropic recommends setting effort explicitly and re-running effort evaluations, since the right level for each step may have shifted along with cost and latency.

No more forced tool calls

Forced tool use no longer works either, as setting tool_choice to any or tool returns a 400 error, including on the token counting endpoint, where cost estimates built on those settings will fail along with the requests they were meant to price. Many agent loops force a call when a step has to query a database, run code, or reach another service, and Anthropic’s replacement is auto-combined with strict tool use or structured outputs, with the prompt stating when the tool applies.

Routing and conversation history

Thinking blocks are now tied to the model and conversation that produced them. On the Claude API, Fable 5.1 and Mythos 5.1 are the only other models that can read Opus 5.5 thinking blocks, so a router or fallback that hands a conversation to any other model will run those turns without the earlier reasoning instead of returning an error.

That adds another layer for teams already watching whether their agent calls are quietly being routed to an older model. Opus 5.5 can read thinking blocks from Opus 5 and earlier Opus, Sonnet, and Haiku models, but not from Fable or Mythos.

Conversations must also stay append-only for those blocks to remain valid. Trimming old messages, changing tool definitions, summarizing earlier context on the client side, or rewriting the system prompt mid-conversation invalidates existing thinking blocks, and for accounts created on or after August 31, 2026, at midnight UTC, replaying a thinking block after one of those edits returns a 400 error by default. Older accounts get no error, but the invalid blocks still reach the model, and Anthropic says future models will enforce the check for all accounts. Integrations that never edit earlier turns need no code change, and Anthropic says Claude Code, claude.ai, Claude Managed Agents, and the Claude Agent SDK already work this way, while agents that compact their own context should follow the company’s preserved thinking documentation.

The fourth change affects computer-use agents on the Claude API and Google Cloud, where Opus 5.5 rejects the computer_20251124 tool and accepts computer use only through the computer_toolset_20260801 toolset. The request itself gets simpler because the beta header goes away and the toolset entry takes no name or display dimensions, but the agent loop needs more work. Each action now arrives as its own tool_use block identified by the block’s name rather than input.action, a single turn can contain several of them, and every result has to echo toolset_name. The older tool still works on Amazon Bedrock, and Anthropic directs developers on other platforms to the computer use tool’s compatibility documentation.

…a router or fallback that hands a conversation to any other model will run those turns without the earlier reasoning instead of returning an error.

Changes that won’t throw errors

The change most likely to go unnoticed doesn’t produce an error at all. On Opus 5, text Claude writes between tool calls comes back as text blocks, but on Opus 5.5 that narration arrives as progress-update thinking blocks, and at the default thinking.display setting of omitted those blocks are empty.

Any agent interface that streams that narration to users will go silent between tool calls until developers set display to updates, a beta option that returns progress updates while keeping reasoning hidden, or to summarized, which returns both, and then render each non-empty thinking block ahead of the tool call it precedes.

Opus 5.5 also ships with broader safety classifiers. It can return a stop_reason of refusal with stop_details categories that now include bio and reasoning_extraction alongside cyber, and Anthropic’s server-side fallback won’t retry requests declined under reasoning_extraction, handing the refusal back to the application instead.

Agents that don’t handle refusals will stop mid-task, a problem developers have already run into with OpenAI’s safety system cutting off API responses.

The change most likely to go unnoticed doesn’t produce an error at all.

Upgrading from older models

Teams coming from Opus 4.8 need to work through the Opus 5 migration first, which covers thinking being on by default and the response-shape changes that follow, before applying the Opus 5.5 changes. Teams on Opus 4.7 or earlier have more ground to cover, and those on models older than Opus 4.7 also face rejected sampling parameters, rejected manual extended thinking, removed prefill, and a newer tokenizer.

Claude Managed Agents users only need to change the model name. Developers working in Claude Code can run /claude-api migrate to apply the model ID swap, parameter changes, prefill replacement, and effort calibration across a codebase before reviewing a checklist of items to verify by hand.

Anthropic recommends testing the migration in a development environment before switching production traffic. Developers maintaining their own integrations will need to test the pieces around the model, too. Tool calls, model handoffs, conversation history, and user-facing progress updates can all behave differently after the switch, because agent failures often originate outside the model itself.

The post Anthropic made Opus 5.5 cheaper. Then it broke four things your agent depends on. appeared first on The New Stack.

  •  

OpenAI releases GPT-6 Sol and Luna — and cuts token prices in half

OpenAI on Tuesday released GPT-6 Sol and Luna, which will complement the flagship GPT-6 Astra model in OpenAI’s lineup. As of now, there is no GPT-6 Terra.

The new GPT-6 pricing

The headline news here is that OpenAI cut the price per million input/output tokens by half or more, compared to the previous version. GPT-6 Sol will cost $2/$10 per million input/output tokens (vs. $4/$20 for GPT-5.6 Sol), and GPT-6 Luna will come in at $0.10/$0.50 (vs. $0.20/$1.20).

The GPT-5.6 pricing was always meant to be promotional, but for the new GPT-6 models, this is the default price, an OpenAI spokesperson tells The New Stack.

“Improvements in caching and inference let us serve these models at lower cost, and we’re passing those savings directly on to users and customers,” OpenAI explains in its announcement.

Benchmarks

As you would expect, the new models show clear improvements over the GPT-5.6 predecessors, but for the most part, these are not all that extreme.

On a benchmark like Zapier’s AutomationBench — which checks how well the models work on a set of business workflow tests — GPT-6 Luna improves by 5.4 percentage points over the previous version, for example

Credit: OpenAI

On the DeepSWE v1.1 software engineering benchmark, GPT-6 Sol essentially matches Anthropic’s Fable (68.8% at max effort vs. 69.9% for Fable 5 at xhigh effort), but at only 20% of the cost. Luna, at max effort, hits scores similar to Claude Opus 5 and Fable 5 at medium effort, at a significantly lower cost.

And OpenAI focuses on this cost comparison across its announcement—with a special focus on price per task instead of straight-up token pricing.

Credit: OpenAI

Anthropic resets the comparison

Since Anthropic released Opus 5.5 earlier on Tuesday, OpenAI’s comparisons are already out of date — such is the way of this AI era. Anthropic, too, reduced its per-token pricing for Opus 5.5 to $4/$20, down from $5/$25, but that still leaves Anthropic’s model twice as expensive as the comparable GPT-6 Sol.

In its announcement, when comparing GPT-6 Sol to Opus 5, OpenAI was able to claim significant cost savings when compared to Anthropic’s model — and for the most part that still holds, but Anthropic says Opus 5.5 also uses fewer tokens per task, which, according to the company, works out to 40% lower costs than Opus 5 on typical workloads.

It’s worth noting that no one has run Sol and Opus 5.5 head-to-head yet. Sol likely stays cheaper per task on OpenAI’s AutomationBench numbers, but Opus 5.5 posts higher scores than GPT-5.6 Sol on shared benchmarks in Anthropic’s testing.

Since it’s almost impossible to know how many tokens an agent will use to finish a task, though, these pricing changes still don’t make it any easier for a user to budget.

Prompt caching

For developers building agents, the caching changes may matter more than token prices. OpenAI says it improved prompt caching for GPT-6 to deliver higher cache hit rates by default, with discounts of up to 90% on cached input tokens.

One positive change, too, is that developers can now change the reasoning effort and tool availability without invalidating the cache. With explicit breakpoints, developers can choose where a cached prefix ends, and a new dashboard and diagnostics tool show what’s getting cached and what isn’t.

GitHub says these improvements cut the share of prompt tokens that require fresh processing by more than half over the past several months, across billions of requests to OpenAI models.

Anthropic made a similar move with Opus 5.5, which cuts cache read prices by 60% for token-billed usage, on top of the 20% per-token cut.

Style changes

Models aren’t just about benchmarks, though. With GPT-6 Sol, OpenAI made its models answer more directly, rather than in the previous — already reined-in — more conversational style. “Expect to see more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall without losing substance,” OpenAI says.

Credit: OpenAI

Alignment

Given the HuggingFace incident, it’s no surprise OpenAI is emphasizing its alignment work for GPT-6 Sol and Luna, too.

OpenAI says both models improve on their GPT-5.6 counterparts across its alignment evaluations, including fewer misleading claims about their own coding work. On an internal coding deception test, GPT-6 Sol’s rate fell to 1.3% from 10.4%.

When given a deliberately broken search tool — and graded on whether it disclosed the problem instead of guessing — Sol failed to disclose the problem 4.9% of the time, down from 77.5%.

What is a bit more concerning, though, is that when researchers asked the model to respect an explicit warning like an “access denied” message, GPT-6 Sol still tried to work around those restrictions in 64.4% of runs, down only slightly from 68.2% for its predecessor. Luna improved more, to 42.4% from 76.5%.

OpenAI says these tests cover mostly low-stakes situations and run without full system-level safeguards used in its products.

Credit: OpenAI

On a simulated message board seeded with unauthorized instructions, including requests to disclose private information, Sol took the specified action in 11.3% of runs where it found the board, down from 51.9%. Luna and Astra took none, though OpenAI notes Luna also found the board less often.

Anthropic, by contrast, says Opus 5.5 is the strongest performer on its most comprehensive alignment test and names METR and Frontier Design as pre-release external testers. Opus 5.5 also ships with safeguards that reroute requests, sending most cybersecurity tasks to Opus 4.8 and anything flagged by Anthropic’s biology or frontier LLM development classifiers to Opus 5.

Availability

GPT-6 Sol and Luna are available in ChatGPT Work and Codex starting Tuesday for Plus, Pro, Business, Enterprise, and Edu users.

Free and Go users get Luna in the desktop app.

Neither model is in Chat yet. OpenAI says it plans to roll them out gradually throughout the day to keep service stable, so they may not appear right away.

The post OpenAI releases GPT-6 Sol and Luna — and cuts token prices in half appeared first on The New Stack.

  •  

TypeSafe launched Jev because sequential LLMs are “totally useless for computers”

When TypeSafe emerged last week after two years in stealth, backed by $40 million in seed funding led by DCVC, to launch its first model, Jev, it claimed something that counters just about everything the industry has built since ChatGPT: The model doesn’t write. It decides.

The first of what the organization calls a new class of System One models, Jev is a text-only model that machines can use natively to make decisions inside software applications. Developers can send Jev structured questions and get typed decisions with calibrated probabilities, meaning software can account for uncertainty.

TypeSafe has built a new architecture for Jev, a new sampler (an algorithm that selects tokens from a model’s predicted probability distribution to control randomness, creativity, and consistency), and a new training algorithm known as Reinforcement Learning for Calibrated Decisions (RLCD). 

Sequential LLMs are totally useless for computers

Co-founder and CEO of TypeSafe, Diogo Almeida, is ex-OpenAI, where, according to TypeSafe, he co-invented RLHF and InstructGPT, the methods behind ChatGPT and GPT-4.

Almeida posted on X on September 15 to state, “The improvements are clear if you see them [LLMs and Jev] side by side. Ask a System One model a ton of structured questions just like you would an LLM. Get the answers back near instantly. Meanwhile, LLMs take hundreds of times longer to respond. Look at how the LLM generates sequentially, which is great for a natural conversation, but totally useless for computers.”

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?

I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev

• 20-200x faster
• 40-400x… pic.twitter.com/JSybNG2BKJ

— Diogo Almeida (@CompleteSkeptic) September 15, 2026

“Look at how LLMs generate sequentially, which is great for a natural conversation, but totally useless for computers.”

Almeida said the inspiration for Jev came from asking himself: why haven’t superhuman chat models led to artificial general intelligence yet? He said that Reinforcement Learning from Human Feedback (RLHF) chat has led to LLMs that are “optimized for human preferences” and include issues such as mode dropping, overconfidence, and an overall lack of reliability.

TypeSafe: System One models can’t hallucinate

Almeida’s launch post lists headline stats for Jev as 20-200x faster, 40-400x cheaper (with output tokens free), and frontier composable intelligence, optimized for decisions. Almeida further claimed that TypeSafe System One models output decisions with probabilities and confidence instead of words, and they “can’t hallucinate” because they’re “a lot more like code”, i.e., reliable, fast, self-consistent, and type-safe. 

Frontend cloud company Vercel has noted that, within 24 hours of launching on AI Gateway, “Jev from TypeSafe AI reached more than twice as many paid teams as any previous model launch, making it the fastest-adopted model in gateway history. Jev passed every other comparison model in its first twelve hours and continued to widen its lead for the rest of the day. By hour 24, nearly 13% of paid teams were using it. That’s 2x the GPT-5.6 family and more than 6x Fable 5.1’s share.”

Developers can set thresholds for when Jev acts autonomously vs. when it needs human oversight and review. They can then combine those decisions in code to build larger workflows, with control over how the intelligence is used. Software engineers can use Jev to select an agent’s next tool or subagent; they can also use it to confirm the veracity of a model’s output and set guardrails.

In a blog post titled “A deep dive into Jev, TypeSafe’s System One model,” independent software developer Flavio Copes noted that Jev is “not a chatbot” like ChatGPT, and it is not a coding model. It does not write replies, explanations, or code.

Jev is a smart if statement

“The simplest way to describe it: Jev is a smart if statement,” wrote Copes. “The important difference is where the AI sits. With ChatGPT or a coding agent, the AI is the main interface or worker. Jev is a small component inside a regular application. You add it where code needs one judgment, while the rest of the product stays ordinary code.”

“You add it where code needs one judgment, while the rest of the product stays ordinary code.”

Copes reiterates TypeSafe’s stated performance levels: most calls to Jev complete in about 100 milliseconds, input tokens cost $0.042 per million, and (as already noted) output tokens are free.

You send it some data and a list of typed questions, and it sends back one answer per question: a yes/no probability, one option picked from a list you defined, or a position on a scale you defined. Every answer comes with probabilities. 

One developer gave Jev the “one thing it can’t handle”

To put Jev to the test, AI engineer Bartosz Mikulski tells The New Stack that because Jev is advertised as a text-only model, he gave it the one thing it can’t handle: pictures.

“I turned 400 hand-drawn sketches into Scalable Vector Graphics (SVG) coordinates and asked what they were,” Mikulski says. “It got about 35% right, where just ‘guessing’ typically returns 10%, but it answered ‘airplane’ for more than half of the drawings, so that number is part real ability, and part a heavy bias toward one label.”

To be fair, Mikulski notes that TypeSafe says in its own documentation that Jev reads text only and handles words better than numbers. 

“I fed it numbers that encode pictures, which is close to the least fair test anyone could design, and it still beat chance by a wide margin. I mean that as a compliment, not as a benchmark. It tells you nothing about how Jev does on the text classification it’s actually sold for,” Mikulski clarifies.

“I fed it numbers that encode pictures, which is close to the least fair test anyone could design, and it still beat chance by a wide margin. I mean that as a compliment, not as a benchmark.”

Why is TypeSafe Jev called Jev?

Jev is named after the 19th-century economist William Stanley Jevons and his Jevons Paradox: the economic principle that as technology increases the efficiency with which a resource is used, that resource’s total consumption actually rises rather than falls. 

When steam engines became more efficient, we used more coal, not less; when LED lighting dropped lighting costs, we used more lighting; when data compression algorithms lowered the bandwidth needed to stream video, global web data traffic increased… and so on.

Almeida concluded his X video post with a nod to developer productivity and said that, “As we say at TypeSafe, we’re building prod, not God.” The company’s comedy disclaimer is shown below.

The post TypeSafe launched Jev because sequential LLMs are “totally useless for computers” appeared first on The New Stack.

  •  

Grok 4.7 was built to work for hours. It still fails most of the time.

labrynth abstract

A coding agent running for hours can make dozens of decisions as it edits files, runs tests, and works through errors. One wrong turn can carry through the rest of the task unless the agent catches it. SpaceXAI appears to be training Grok for exactly that problem.

The company released Grok 4.7 on Sunday, and its training approach is uniquely different. SpaceXAI used a longer reinforcement learning run deliberately weighted toward harder tasks, including problems that take “many hours” to complete. The company says that training also made Grok better at verifying its own work and managing longer context.

Every failed approach from an agent adds more history for the model to keep straight, and one bad assumption can follow it through the rest of the task. SpaceXAI is trying to address that with better context management and self-verification, so Grok can catch a wrong turn before it builds on it.

SpaceXAI used a longer reinforcement learning run deliberately weighted toward harder tasks, including problems that take “many hours” to complete.

Endurance benchmarks tell the story

Grok 4.7 scored 38.0% on Terminal-Bench 4.0, up from 20.3% for Grok 4.6. It also improved from 40.4% to 46.3% on CursorBench 4.0, which tests longer-running coding workflows inside the editor, and from 1,546 to 1,657 on AA Briefcase v1.1, an evaluation of multi-hour professional work.

For context, Anthropic’s Claude Fable 5.1 scores 57.9% on Terminal-Bench 4.0 according to the independent leaderboard; Grok 4.7 still trails Fable 5.1 here. What’s arguably more interesting is how much it improved over Grok 4.6. SpaceXAI says the improvements came from pairing the larger base model with an extended reinforcement learning run deliberately shifted toward harder, multi-hour problems, and that the model specifically improved at two capabilities critical to long-horizon execution: self-verification and long-context management.

An agent working unattended for hours has to keep track of a growing interaction history while checking that each step worked before moving to the next. Those problems surfaced in a recent benchmark of private codebases, where even the best-performing model failed more than 60% of the time. SpaceXAI says Grok 4.7 improved at both context management and self-verification, although it hasn’t explained how. The company did not disclose whether the context gains came from architectural changes, summarization, retrieval, or better retention across long sequences, or how it evaluated self-verification during reinforcement learning.

An agent working unattended for hours has to keep track of a growing interaction history while checking that each step worked before moving to the next.

The harness is becoming part of the model

SpaceXAI trained Grok 4.7 to natively understand the Grok Bot harness, bringing the model and the surrounding infrastructure closer together.

Agent harnesses handle the work around the model, including exposing tools, formatting terminal responses, feeding execution results back into context, and deciding what happens next. OpenAI took a similar approach last week when it opened its Codex harness as the Agents API, turning the infrastructure behind long-running agents into a managed service.

With Grok 4.7, SpaceXAI is pushing some of that integration into training. A model already familiar with its harness doesn’t have to learn every tool format and interaction pattern through prompting at runtime. That could reduce the overhead involved in tool use and multi-step execution, although SpaceXAI hasn’t published enough detail to show how much of Grok 4.7’s performance gain comes from harness-specific training.

Training models around specific tool schemas, context formats, and execution environments could make it harder for developers to swap models without sacrificing agent performance.

That problem grows as agents take on more of the development cycle. Google’s recent work on making Go easier for AI agents to work with took a different approach, changing the development environment rather than the model. In both cases, the model is no longer the only piece being optimized. The systems around it are changing too.

Training models around specific tool schemas, context formats, and execution environments could make it harder for developers to swap models without sacrificing agent performance.

Where the gaps still are

Grok 4.7 starts at $2 per million input tokens and $6 per million output tokens. At that price, multi-hour agent runs may cost less, but reliability remains an issue. Grok 4.7 scored 38.0% on Terminal-Bench, while Fable 5.1 reached 57.9%.

The post Grok 4.7 was built to work for hours. It still fails most of the time. appeared first on The New Stack.

  •  

Open-weight models now handle a majority of tokens on Vercel’s AI Gateway. But Anthropic still takes 64% of the spend.

Isometric illustration of a retro-style computer monitor

The trend is clear: open-weight models are taking an increasingly large bite out of production AI usage.

On Monday, The New Stack reported that open-weight models accounted for 60% of OpenRouter’s US token consumption in August, with Chinese-developed models making up the majority of that volume. The latest data point hails from Vercel, whose AI Gateway routes tens of trillions of tokens each month across the applications running on its infrastructure.

As per Vercel’s September report, published on Thursday and covering activity through August, open-weight models handled 56% of all tokens routed through the gateway, the first time they have accounted for a majority of monthly token volume. In December 2025, their share was just 7%; by April it had reached 13%, and it rose every month thereafter. Vercel’s previous report, published in August, put July’s open-weight share at 36%.

So the pattern was already clear. But last month, Vercel CEO Guillermo Rauch took to social media to declare that August 22 had been a “record day for open weight share of tokens on Vercel AI Gateway,” accounting for 62% of traffic.

Rauch saw the milestone as just an early indication of where usage is heading, with enterprises still early on the adoption front.

“This is very likely just the start, because enterprise adoption is still early.”

“This is very likely just the start, because enterprise adoption is still early, and harnesses, CLIs, IDEs, SDKs, etc need to be adapted to be model agnostic,” Rauch wrote at the time.

Open-weight token share on AI Gateway: December '25 to August '26
Open-weight token share on AI Gateway: December ’25 to August ’26 (Credit: Vercel)

Tokens and dollars: Anthropic dominates spend

For context, Vercel launched AI Gateway last year as a way for developers to access models from multiple providers through a single interface, saving them from having to manage separate API keys, accounts and rate limits. The service sits between applications and the underlying model providers, routing requests while tracking usage and costs — giving Vercel a useful vantage point into which models its customers are actually running in production.

Token volume, in this context, is essentially a measure of how much model inference is flowing through the gateway. Vercel counts input and output tokens, along with reasoning, cached-input and cache-creation tokens.

While it’s a good proxy for the amount of work being handed to different models, it shouldn’t be confused with the amount of dollars being spent. Open-weight models from the likes of DeepSeek, Moonshot AI and Z.ai are generally cheaper to run than the proprietary models offered by US frontier labs — and so handling 56% of Vercel’s token volume doesn’t mean open-weight models are taking 56% of the money passing through its gateway.

Indeed, Vercel’s data shows that open-weight models accounted for just 14 cents of every estimated dollar spent through AI Gateway in August, despite processing 56% of its tokens. Their share of spending remains far behind their share of usage, although Vercel says the open-weight share of gateway spending is on the rise.

Open-weight share of tokens vs spend on AI Gateway
Open-weight share of tokens vs spend on AI Gateway (Credit: Vercel)

Across Vercel’s AI Gateway, the average price per token fell 23.2% in August, marking a third consecutive monthly decline. Among teams that processed more than 10 million tokens in both July and August, the median cost per token fell 7.6%.

Anthropic, meanwhile, has remained remarkably consistent at the spendy end of the market. Its models accounted for 64 cents of every dollar spent through the gateway in August. Vercel says the Claude-creator’s share has never fallen below 61% in any month since December 2025, with its models occupying the top two positions by spend throughout that period — often taking third spot, too.

Top 3 models by spend, by lab.
Top 3 models by spend, by lab. (Credit: Vercel)

Loyalty lies in the model

There has been plenty of movement within that Anthropic share, however. Fable 5 fell from 13.2% of total gateway spend in July to 4.9% in August, while the cheaper Opus 5 climbed to 22.5%. More broadly, Vercel’s data suggests that 90% of teams using Fable reduced their usage, with more moving those workloads to Opus 5 than to any other model.

Opus ultimately gained almost twice as much usage as Fable lost, which Vercel attributes to the newer model handling similar workloads at roughly half the price. Or, in other words, Anthropic kept the dollars even as customers shifted toward a cheaper model within its own lineup.

“Lab loyalty doesn’t follow brand, it follows model profile, and consistency wins.”

“Lab loyalty doesn’t follow brand, it follows model profile, and consistency wins,” Vercel’s report authors note.

Anthropic's share of spend by model
Anthropic’s share of spend by model (Credit: Vercel)

This trend was evidenced elsewhere, too. Within five days of Z.ai launching GLM-5.3-Flash, the new model was processing three times the daily volume of GLM-5.2.

But Vercel’s data also suggests customers are more than prepared to cross lab boundaries when a replacement fails to meet the same needs on capability and price: more than three-quarters of the volume lost by Google’s Gemini 3 Flash moved to models from other providers, including OpenAI and Anthropic. And the consequence for Google wasn’t insignificant: its overall share of token volume on the gateway fell from 30% to 5%, with the decline in Gemini 3 Flash alone accounting for 22 of those 25 percentage points.

“When a new model preserves what users valued in its predecessor, the lab retains its customers,” the authors note. “When it doesn’t, those customers fill the need through other providers.”

The post Open-weight models now handle a majority of tokens on Vercel’s AI Gateway. But Anthropic still takes 64% of the spend. appeared first on The New Stack.

  •  

OpenAI’s voice model doesn’t think. That’s the point.

Abstract waves

Voice agents have a latency problem that shows up as soon as they have to do real work. Within five days, Google and OpenAI shipped two very different fixes.

On Tuesday, Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking through the Gemini API and Google AI Studio, just five days after OpenAI released GPT-Live-1. Both let a voice agent keep talking while it works in the background, but they go about it very differently.

Gemini 3.8 Live Extended Thinking keeps reasoning inside the voice model, letting it continue speaking while it executes asynchronous tool calls. OpenAI separates those jobs, using GPT-Live-1 for the real-time conversation while a backend reasoning model handles complex tasks — pushing more orchestration into the application layer.

Google cautions against treating that split as a direct comparison between Gemini and products like ChatGPT or Claude Voice.

“Today’s models are more centered on giving developers/enterprises tools to build voice agents,” a Google spokesperson tells The New Stack. “ChatGPT and Claude voice mode are full products rather than models, so the comparison is not apples-to-apples.”

“ChatGPT and Claude voice mode are full products rather than models, so the comparison is not apples-to-apples.”

Reasoning inside the session

Gemini 3.8 Live Extended Thinking keeps speech, reasoning, and tool execution inside a single stateful session, even while external API calls are still running.

When a function is set to NON_BLOCKING, Gemini can keep talking while it waits for the tool to respond, asking follow-up questions or giving updates along the way. Once the result comes back, Gemini picks up from there.

Developers can set reasoning effort to low, medium, or high per request. Standard Gemini 3.8 Live skips the extended reasoning step to cut latency and token cost.

The same underlying model also powers Gemini Live in the consumer Gemini app. Google calls that its “end-user focused offering more closely comparable to ChatGPT and Claude,” rather than the developer models themselves.

Multimodality carries over to the new audio models as well. “Visual understanding is excellent,” Google tells The New Stack, adding that users can “converse with the model seamlessly about whatever you show it.”

Gemini 3.8 Live Extended Thinking keeps speech, reasoning, and tool execution inside a single stateful session, even while external API calls are still running.

Coordinating two separate layers

With GPT-Live-1, the voice model handles the full-duplex conversation while a backend model such as GPT-6 Astra, a lighter model like Luna, or even a third-party option, handles reasoning and tool execution independently.

Keeping the backend work separate lets the voice layer stay responsive, with OpenAI putting turn-taking latency at around 800 milliseconds.

The tradeoff is that developers must coordinate the two layers themselves, passing context between the voice model and backend reasoner through sideband channels and deciding what the conversation does while background work runs. That orchestration burden falls entirely on the application layer.

Stale work when a user interrupts

Both approaches face the same headache when someone interrupts or changes their mind halfway through a request, leaving background work running that may no longer be needed.

In Gemini, that work stays within the same session, although developers have less visibility into exactly when a tool call stops. OpenAI leaves more of that cleanup to developers, who have to cancel pending jobs and make sure an outdated answer doesn’t find its way back into the conversation.

Google says Gemini has an edge under the messy conditions voice agents encounter outside a demo. The company tells The New Stack that Extended Thinking handles “background noise, heavy accents, and unexpected interruptions better than competing models.”

Per-minute costs diverge sharply

Standard Gemini 3.8 Live carries Gemini Live API rates of $0.005 per minute of audio input and $0.018 per minute of output. Extended Thinking adds reasoning tokens, with additional charges for inputs like live video and documents.

GPT-Live-1 costs $0.05 per voice minute for the front-end voice layer alone. The backend reasoning model, function calls, and external agent runs are all billed separately. As with GPT-6 Astra’s adjustable reasoning settings, developers can dial cost up or down per call, but a voice agent that regularly calls a more powerful reasoning model will see its bill climb fast.

The company also embeds DeepMind’s SynthID watermark in generated audio.

Benchmark numbers, with caveats

Gemini 3.8 Live Extended Thinking scored 82.6 on Artificial Analysis’ Speech-to-Speech Quality Index, with task completion rates of 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark.

The company points directly to those results, telling The New Stack that Extended Thinking “holds the #1 spot on the Speech-to-Speech Quality Index and leads on complex task-completion benchmarks.”

GPT-Live-1, paired with GPT-6 Astra at medium reasoning effort, scored 86.2% Pass@1 on Tau3’s spoken customer-service evaluation spanning airline, retail, and telecom domains. On Full Duplex Bench, it beat GPT-Realtime-2.1 by 30 percentage points.

“Extended Thinking holds the #1 spot on the Speech-to-Speech Quality Index and leads on complex task-completion benchmarks.”

Different tests, different stacks

But these results aren’t head-to-head. Google and OpenAI used different tests and setups, and Google also notes that some comparisons put developer models up against finished consumer products.

Claude Voice isn’t part of this developer calculus. Anthropic offers voice in its consumer apps but doesn’t currently offer a real-time speech-to-speech API comparable to Gemini Live or GPT-Live-1. Developers building voice agents around Claude still have to assemble more of the voice stack themselves.

Google keeps speech, reasoning, and tool execution inside one session, cutting down on middleware but tying developers more closely to its runtime. OpenAI requires more orchestration but gives developers more control over the models and tools running behind the voice layer.

The post OpenAI’s voice model doesn’t think. That’s the point. appeared first on The New Stack.

  •  

Chinese AI models dominate OpenRouter’s US token consumption. It can now guarantee that traffic stays entirely in the US.

Illustration of data-center servers marked with location pins and connected by routing paths

Everyone knows the open-weight model pitch by now: companies can download the weights, customize them, run them on infrastructure of their choosing, and retain far greater control over where their data is processed — often at a much lower cost than using proprietary models.

Moreover, open-weight models are now thought to trail the leading frontier models by only around four to five months. Nvidia, the world’s most valuable company, is betting heavily on that future. In early September, it agreed to acquire Hugging Face — the sprawling “GitHub for AI” that hosts more than three million models — for $12.9 billion, while pledging to keep the platform open to different models, clouds and computing providers. And on Thursday, Nvidia detailed how Nvidia is using its own open-weight Nemotron model to manage its vast global supply chain in partnership with Palantir.

That power also comes with serious security questions. OpenAI president Greg Brockman recently warned that increasingly capable open-weight models — pointing specifically to China’s GLM-5.3 — could “significantly accelerate the threat landscape” as models with advanced cyber capabilities become freely downloadable and modifiable.

But for businesses accessing those models through third-party services, there is another concern closer to home: where their own data goes when they use those models, particularly when the model originated in China.

China and the open-weight factor

Hugging Face data from February showed models from Chinese developers accounted for 41% of downloads in the preceding 12 months, ahead of the US at 36.5%. Over on OpenRouter, meanwhile, open-weight models now account for around 60% of tokens consumed by US-originating requests, with the company noting that Chinese models constitute the majority.

OpenRouter: Share of monthly tokens (Sept. '25 - Aug. '26)
OpenRouter: Share of monthly tokens (Sept. ’25 – Aug. ’26) — US and EU

And that’s why OpenRouter is now giving companies a way to put a geographic fence around that traffic. The AI model marketplace has officially launched US in-region routing into general availability for business and enterprise customers, promising that requests sent through its US endpoint are decrypted, processed and served entirely inside the country — or rejected if that can’t be done.

The feature itself had been quietly available in some form before now, with OpenRouter updating its documentation in early August to say US in-region routing was available to enterprise customers by request. It’s also worth noting that this is in addition to European in-region routing, which it says has been available since October 2025.

Started in early 2023 by former OpenSea CTO Alex Atallah, OpenRouter serves as an interface to the crowded AI model market, with developers able to switch between hundreds of models from myriad providers via a single API. Payments giant Stripe recently announced plans to acquire the company in a reported $8 billion deal, while a slew of other companies including Cursor, Ramp, and Meta, are also building their own model routers.

The reason why model routers are such hot property right now is largely down to economics. Developers have traditionally hard-coded applications to send everything to the same model, while a model router can instead make that choice request by request, sending easier jobs to cheaper models while reserving the pricier frontier systems for the work that actually needs them.

That intermediary role is also what makes OpenRouter’s new residency controls possible: it already decides which provider serves each request, and can now restrict that choice to provider endpoints operating in the US.

Keeping Chinese models inside the US

In a blog post announcing the new feature on Wednesday, Cailee Moberg, who works on OpenRouter’s product team, notes that while US-developed models from Nvidia and Thinking Machines are contributing to the broader open-weight model boom, Chinese models dominate usage and raise tough questions for companies concerned about their data.

“Models from Chinese labs are still most of the [open-weight model] volume, and procurement approval for those models can be difficult.”

“Models from Chinese labs are still most of the [open-weight model] volume, and procurement approval for those models can be difficult,” Moberg writes.

In its 2026 State of AI in the Enterprise report, Deloitte concluded that sovereign AI was on the rise, noting that 77% of companies “now factor country of origin into their vendor selection,” while nearly 60% construct their AI stacks “primarily with local vendors.”

And this at least partly explains why OpenRouter is now offering in-region routing for US customers. Moberg points to DeepSeek V4 Pro, Kimi K3 and GLM 5.2 as specific examples. All three are available through US In-Region Routing because Baseten, Fireworks and Azure serve them from US data centers. Companies could already keep these models inside the US by self-hosting them or using a US provider directly; OpenRouter’s new routing gives its own customers that residency guarantee without having to manage those deployments themselves.

OpenRouter maintains a live list of models eligible for US in-region routing, ranging from proprietary frontier models from OpenAI and Anthropic to open-weight models from the major Chinese labs.

“In-Region Routing allows teams with data residency requirements to get the price and performance gains from Chinese open-weight models,” Moberg continues. “When a US or EU provider hosts a model, requests go to that provider and the lab is not involved.”

“In-Region Routing allows teams with data residency requirements to get the price and performance gains from Chinese open-weight models.”

The technical change happens at the routing layer. With OpenRouter’s standard global endpoint, a request can be served by an eligible provider operating in any region, so even using a model from a US company does not guarantee that the request itself is processed in the US. With us.openrouter.ai, the request is decrypted on OpenRouter infrastructure inside the US and the pool of providers is filtered to endpoints OpenRouter has approved as operating there.

If no compliant US provider can serve the requested model, OpenRouter returns a 404 error. Companies can also enforce the regional restriction through OpenRouter’s Guardrails at the workspace, team or API-key level, while tools that would send prompt data outside the US are disabled on the regional endpoint.

So while none of this ultimately changes where the DeepSeek, Kimi or GLM models are developed, in-region routing alters which copies of those models its US customers can be routed to, and where their prompts are handled along the way.

The post Chinese AI models dominate OpenRouter’s US token consumption. It can now guarantee that traffic stays entirely in the US. appeared first on The New Stack.

  •  

“Machine translation is still broken for most of the world’s languages”: Cohere builds non-reasoning for a reason

A scattered pile of overlapping alphabet cutouts in bright blue, pink, green, gold, red, and silver.

Enterprise AI company Cohere announced North Small Translate last week, a mixture-of-experts (MOE) open-weight machine translation model that works across 50 languages.

Developers can download the weights for noncommercial use under CC BY-NC 4.0. Cohere offers commercially licensed deployment through Model Vault, which is a Cohere-managed inference environment. Cohere positions the model as part of its sovereign AI strategy, aimed at organizations that want greater control over where their models run and how their data is handled.

North Small Translate builds on Cohere’s multilingual and translation lineage, which includes its Tiny Aya and Command A Translate model families. The company claims North Small Translate outperforms “similarly sized open-weight models” under 1T parameters, as well as API-based translation models in various dimensions of machine translation on average. 

Cohere co-founder Nick Frosst tells The New Stack that the model’s efficiency draws from the fact that it is non-reasoning, i.e., it relies on learned statistical patterns without a step-by-step logic process, which means it uses fewer tokens.

Machine translation is still broken for most of the world’s languages

“We spent nine years scaling an architecture invented to fix translation, and machine translation is still broken for most of the world’s languages,” Frosst says. “General-purpose models get you most of the way and then stop. The next phase of enterprise AI in this space is smaller, more specialized, and runs inside your own walls.”

“…machine translation is still broken for most of the world’s languages.”

In Cohere’s reported evaluation using WMT26 benchmarks, the company states that North Small Translate leads with a WMT26 All Languages benchmark score of 83.60, compared with 81.56 for Qwen 3.5 397B A17B, 76.50 for GLM 5.2 FP8, 81.37 for DeepL NextGen, 79.46 for Gemma 4 31B (on), and 68.20 for Google Translate. 

With its mixture-of-experts architecture and 218 billion total parameters, with 25 billion active. Cohere points to North Small Translate’s smaller compute & memory footprint than other models. Some model-to-model comparisons in this space aren’t fully substantiable, since not every vendor discloses parameter counts.

With current solutions, long documents start to fall apart

“Machine translation allows documents to be translated from one language to another automatically. With current solutions, long documents start to fall apart,” Frosst says. “Google Translate scores 21.3 on our long-context test, Gemma 4 31B 19.4; we score 48.9. That’s [for example] a safety manual that reads fine on page one… and has drifted by page ten. The other risk is where the text goes. Once you push HR policies or regulated documents through a third-party API, that data has left your building, and necessarily that means your control over it is diminished.”

“The risk [in machine translation] is where the text goes. Once you push HR policies or regulated documents through a third-party API, that data has left your building and necessarily that means your control over it is diminished.”

Explaining why the model offers “stronger translation performance” across complex enterprise translation tasks, Frosst says the model can support work spanning “a high volume” of sensitive documents. 

As well as its 50 languages (32 ‘high-resource’ languages + 18 others), the Cohere team explains that the model also supports translation-workflow-focused capabilities, such as structured translations (i.e., Markdown or JSON documents), instruction following (i.e., recommended tone & format), and terminology guides (i.e., providing specific vocabulary to use in the translation), all as part of the model.

“North Small Translate works with a multi-pass workflow,” explains Frosst. “The model translates, reviews its own output, finds errors, and fixes them – and this is the same loop we used in training. We ship both because standard is one pass and built for volume, while the agentic [version] spends more tokens for 84.36 against 83.60 on WMT26. That difference ends up being worth it when the document is a contract or a safety procedure, for instance, but in other cases you’d rather optimize for efficiency.”

“The model translates, reviews its own output, finds errors and fixes them.”

Model ‘steerability’ drives suggesting language tone and formatting

This model uses the same architecture as prior Cohere models but improves performance through post-training advances, including reinforcement learning and new datasets, specifically for machine translation tasks.

Frosst concludes that, across the translation model marketplace, generative machine translation models offer the highest quality and steerability (i.e., suggesting tone, formatting, etc.) but typically cost much more than Neural Machine Translation (NMT) models commonly used in commercial use cases. 

North Small Translate was developed in partnership with RWS, an AI solutions company pioneering in language technology and services. Collaboration with RWS, specifically with its Language Weaver research and science teams along with its language experts, helped shape the model’s real-world translation performance throughout development. 

As noted above, developers can access the weights free of charge for non-commercial use in three quantizations. There is also a Hugging Face Space and an API for those who lack the required hardware. 

The post “Machine translation is still broken for most of the world’s languages”: Cohere builds non-reasoning for a reason appeared first on The New Stack.

  •  

OpenAI split a voice model’s brain. Then one team deleted 23,000 lines of code.

audio waves

Building an AI voice agent has always been clunkier than it seems. Most voice agents are really a chain of systems passing a conversation back and forth. What you say gets turned into text so a model can figure out how to respond, then that answer has to be turned back into speech; it’s easy to see why things can get robotic fast. Now, OpenAI is trying to collapse that stack.

On Wednesday, the company launched GPT-Live-1 in its API, bringing the native, full-duplex voice architecture behind ChatGPT voice mode to outside developers for the first time. Instead of making developers manage the entire chain, it wants one model to handle the conversation while the heavier thinking happens elsewhere.

Instead of making developers manage the entire chain, it wants one model to handle the conversation while the heavier thinking happens elsewhere.

Full-duplex voice delegation

GPT-Live-1 operates as the conversational frontline. Because it’s natively full-duplex, it can keep up with a conversation as it happens, including when someone cuts in mid-sentence, without developers coordinating separate systems. But the voice model doesn’t have to do all the work alone.

When a request needs more time or more processing, GPT-Live-1 can hand it off to another model in the background. That could be GPT-6 Astra, a smaller model like Luna, or something from another provider entirely.

Waiting on a bigger model can make a voice agent painfully awkward. Ask a difficult question, and you can end up sitting in silence while the model works through it. GPT-Live-1 can keep the conversation going instead — filling pauses, acknowledging the speaker — then work the answer in once the backend is finished.

OpenAI says GPT-Live-1 performs 30 percentage points better than GPT-Realtime-2.1 on Full Duplex Bench. Paired with GPT-6 Astra at medium reasoning, it also takes the top spot on the 𝜏³-benchmark.

What the handoff looks like

OpenAI exposes delegation through an event-driven interface. The voice session generates a delegation_id, sends context to whatever backend system is handling the heavier work, and gets the result back through an event called session.commentary.append. The voice model folds that result into the ongoing conversation rather than reading a block of text aloud. Developers can still see what the model hears and says and control when it takes a turn — they just don’t have to build the entire conversation out of separate systems. OpenAI’s API docs walk through the full pattern, including a working example with the Codex SDK.

Early customers cut code

One early customer deleted 23,000 lines of code after switching to GPT-Live-1.

Tony Stoyanov, co-founder and CTO of EliseAI, a healthcare company testing the API, said the move shrank his codebase by 80%. His team could spend that time on the patient experience instead — making it easier to book appointments and navigate care.

The language-learning company Speak saw the difference in the conversations themselves. In early tests of its Live Tutor Lessons, GPT-Live-1 was nearly 80% less likely to interrupt someone who had simply paused to think. For someone learning a new language, those extra few seconds can be the difference between getting the answer out and having the AI cut them off.

Yelp is already using GPT-Live-1 in Yelp Host and Hatch. CTO Alex Levy said the company is seeing more calls successfully handled by AI, and callers are speaking in fuller, more natural sentences — a sign, Levy said, that the experience on the other end of the phone feels different. A demo released with the announcement shows exactly why, when a restaurant reservation kept moving even with background noise, and people spoke over one another.

Pricing the voice layer

GPT-Live-1 costs $0.05 per minute, or about $3 an hour. Then there’s whatever developers choose to run behind it. If GPT-Live-1 hands a request to GPT-6 Astra, the developer pays for that call too. The more often an agent reaches for a reasoning model, the faster the bill climbs.

OpenAI has been cutting API prices as competition from Anthropic, Google, and Chinese labs heats up, but frontier reasoning still isn’t free.

The more often an agent reaches for a reasoning model, the faster the bill climbs.

The tradeoff is that developers can now be selective about where they spend that money. Something simple, like scheduling an appointment, could go to Luna. A harder question, like one that actually needs multi-step reasoning or tool calls, could go to Astra. OpenAI has already shown how Astra’s adjustable reasoning settings let developers dial cost up or down per call, and GPT-Live-1 gives them a place to apply that same logic to voice.

Platform control tradeoffs

With the older cascaded approach, teams can choose a different provider for each part of the voice stack and swap pieces out when they want. GPT-Live-1 takes over more of the conversation, which also means handing more of it to OpenAI.

The bet is that developers will give up some of that control if it means voice agents can finally keep up with the people talking to them.

The bet is that developers will give up some of that control if it means voice agents can finally keep up with the people talking to them.

The post OpenAI split a voice model’s brain. Then one team deleted 23,000 lines of code. appeared first on The New Stack.

  •  

Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet

Macro view of overlapping textured paper sheets in white, pink, blue, orange and teal.

When Anthropic launched Claude Fable 5.1 this month, it centered the announcement around one benchmark result: its Terminal-Bench-Science score.

In this benchmark, a model gets a terminal and a real scientific research problem to solve independently. Fable 5.1 scores 52.6%, and Fable 5 scores 24.7%. By Anthropic’s scoring, the new model more than doubles the old one.

Anthropic’s published score was produced under conditions most users don’t have access to. The benchmark allows each model up to eight hours per task, and Anthropic has not said what harness or budget it used to get its numbers. When the benchmark’s leaderboard tested Fable 5, it ran the model through Claude Code at maximum effort and spent $14,180 across 210 attempts, about $67 each.

Most people don’t use Fable in a lab, so I wanted to know what an average user would get on these tasks. I recently tested Fable 5 and Fable 5.1 on everyday work and found them far closer than the benchmark suggests. That made me want to run the benchmark’s own tasks the way a home user would and see where the models actually differ.

The benchmark’s 70 tasks are public, so I pulled five of them, one from each science field, and ran both models myself. 

The tests

Terminal-Bench-Science has five categories, each with multiple tests. I chose one test per category that could run in a Python environment. Here’s what I picked:

  • Symbolic regression (mathematics) – A dataset with 100 variables and a hidden formula behind a yes-or-no label. The model must find a predictor that works on data it has never seen.
  • Lorenz-96 assimilation (Earth sciences) – Reconstruct a chaotic atmospheric model from a few uncalibrated sensors with unknown clock offsets. Grading is all-or-nothing on five criteria.
  • Reactor safety control (engineering) – Write a controller for a chemical reactor that finishes every batch as fast as possible without ever exceeding the temperature limit, across public and hidden fault scenarios.
  • Foraging cognitive model (life sciences) – Predict, trial by trial, which lever each of 20 mice will press, graded on sessions the model never saw.
  • Nanoindentation (physical sciences) – Extract material properties from raw indentation curves that include drift, adhesion, defects, and an unknown tip shape.

Each run got a plain terminal, and I set a $12 limit and 60 turns for each test. The full set of ten runs took about 12 hours.

Symbolic regression

This was the only test where a model passed the benchmark’s hidden test. Fable 5.1 worked for 27 turns, found the hidden structure, wrote a predictor, and stopped on its own after 11.8 minutes, 27,088 output tokens, and cost $1.96 to pass this one test.

Fable 5 used all 60 turns over 53.5 minutes, generated 39,461 output tokens, cost $4.20, and failed. I ran Fable 5 a second time to rule out bad luck. It used all 60 turns again, took 60 minutes, generated 60,608 output tokens, cost $6.38, and failed again.

Lorenz-96 assimilation

This was the most expensive pair of runs. Fable 5 hit the $12 cost limit at 45 turns after 97.7 minutes and 92,091 output tokens, ending at $12.63. Fable 5.1 used all 60 turns over 126 minutes, generated 89,789 output tokens, and cost $10.70. On the public leaderboard, Earth sciences is also the field where Fable 5 scores close to zero, and both models failed it here.

Reactor safety control

Neither wrote a controller that passed the grader’s scenarios. This run produced the most output tokens, 157,710 generated by Fable 5.1. It hit the 60-turn limit after 40.9 minutes and cost $11.53. Fable 5 hit the cost limit at 49 turns after 63.8 minutes, 121,978 output tokens, and $12.04. 

Foraging cognitive model

This was the longest run of the testing series. Fable 5.1 was the only model that declared itself finished. It built a model, tested it against its own scoring loop, and declared it done at 43 turns after 53.5 minutes. It created 65,518 output tokens and cost $5.65. But the official grader rejected it. 

Fable 5 never declared anything. It hit the $12 limit at 60 turns, after 139.3 minutes and 62,587 output tokens, ending at $12.13. 

Nanoindentation

Both failed. Both spent most of the run reading raw curves and writing code to segment them. Neither produced a results file the grader accepted. Fable 5.1 ran out of turns at 29.6 minutes, 115,687 output tokens, and $10.91. Fable 5 ran out of money at 48 turns after 34.1 minutes and 114,239 output tokens, ending at $12.59. 

Results

Here are the results by the numbers.

MetricFable 5 scoreFable 5.1 score
Tasks solved0 of 51 of 5
Output tokens430,356455,792
Total cost$53.59$40.75
Total time388 min262 min
Runs ended by cost limit40

The benchmark scores models on all 70 tasks with three trials each, and Anthropic’s 24.7% and 52.6% come from that full suite. The independent leaderboard puts Fable 5 at 21.4%, close to Anthropic’s figure. Fable 5.1 is not on the independent leaderboard yet, so its 52.6% is Anthropic’s number alone. My results, 0% and 20%, are below both. Five tasks are a small sample. 

Getting these results by chance is plausible even if the published scores are exactly right, so this run neither confirms nor contradicts the doubling claim. The direction matched, since the new model did better. The one task Fable 5.1 solved was in mathematics, which is also the field where the leaderboard shows Fable 5 performing best.

What I think

I don’t think a regular user will see much difference between Fable 5 and Fable 5.1. I only ran a small sample of tests, so I can’t prove or reject Anthropic’s benchmark results. But what I saw suggests that gap won’t reach the average user. The one difference that did show up was on the bill. Fable 5.1 failed faster and cheaper, and it never hit my cost limit, whereas Fable 5 hit it four times.

Suppose your work looks more like the benchmark tasks; the harness and the budget matter as much as the model. With a purpose-built harness, hours per task, and a much bigger budget, you may get closer to Anthropic’s numbers.

The post Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet appeared first on The New Stack.

  •  

Claude performed best on a new benchmark for ‘agents that build agents’. But it passed fewer than a quarter of the tests.

Illustration of a developer coding at a monitor, surrounded by floating UI windows, code snippets, and a password field.

AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems that answer questions, process refunds, change bookings, and interact with company systems. But while AI is increasingly central to building these systems, humans still play a major steering role: setting goals, supplying context, choosing architectures, reviewing decisions, and testing the result.

Which raises a more interesting question: what happens when an agent is asked to build another agent entirely on its own?

This is the key question Hyper-𝜏-bench is designed to probe.

Hyper-𝜏-bench asks: How well can AI agents build other agents?

Created and open-sourced in early September by Sierra, the enterprise AI agent company co-founded by tech veteran and current OpenAI board chairman Bret Taylor, Hyper-𝜏-bench builds on the original 𝜏-bench benchmark it introduced back in 2024. But while 𝜏-bench focused on measuring how well a finished agent could interact with users, use tools, and follow company policies, Hyper-𝜏-bench takes it a level up: it evaluates how well an AI developer agent can build that agent in the first place.

Today, most agents are built with the help of other agents like Sierra's Ghostwriter. Yesterday, Sierra open-sourced hyper-𝜏-bench (published as 𝜏^𝜏-bench), a new long horizon agent evaluation that measures how well models can not only act as an agent, but construct one.…

— Bret Taylor (@btaylor) September 9, 2026

In a research paper published on September 4, Sierra researchers tested six combinations of AI model and coding harness. Those included Anthropic models running in Claude Code, OpenAI models in Codex, and Moonshot AI’s Kimi K3 running in both Kimi Code and the open-source OpenCode.

Hyper-𝜏-bench gives a developer agent key materials from a simulated business, such as documents, transcripts, an API, and a codebase, and asks it to build a customer service agent under model and cost constraints. Sierra then tests the finished agent on unseen customer conversations across airline, retail, telecom, and banking, with tasks such as canceling a flight or disputing a fee. It passes when the agent gives the right information and makes the correct changes in the business’s underlying systems. A score of 50%, for example, would mean that the agents succeeded in half of those simulations.

As the leaderboard shows, the best-performing combination was Claude Opus 5 running in Claude Code, at 23.9%, narrowly ahead of GPT-5.6 Sol in Codex at 22%. None of the six autonomous configurations broke 25%.

Hyper-τ-bench pass rates, build time, token spend and serving spend
Hyper-τ-bench pass rates, build time, token spend and serving spend (Credit: Sierra)

Low scores alone aren’t necessarily a problem for a benchmark. Tests aimed at frontier AI systems need to be difficult enough to leave room for improvement and to expose meaningful differences between systems; once the best models routinely ace a benchmark, it becomes much less useful as a measure of progress.

What stands out in Hyper-τ-bench’s initial results, however, is the 82.2% “Human + AI reference” bar on the right, which sits far above every autonomous developer. But that comparison comes with an important caveat.

In an accompanying blog post published on Tuesday by Sierra researchers Ben Shi and Keshav Dhandhania, they describe the result as a model being paired with “an engineer with deep context.” The research paper explains this further: the reference agents were hand-built by a benchmark author working with a frontier model and, crucially, with access to the ground-truth requirements that the autonomous developer agents had to discover for themselves.

So while it’s tempting to read the gap as evidence that human support more than tripled performance, Sierra cautions against that interpretation, saying the 82.2% figure is merely an “oracle reference” rather than a measure of average human performance.

Those overall numbers also hide some dramatic differences by task. Claude Opus 5, for example, reached 72.8% on retail, 55.9% on airline, and 48.2% on telecom, before falling to just 5.9% on banking. GPT-5.6 Sol did somewhat better on banking, at 9%. Banking accounts for 35 of the benchmark’s 53 construction tasks. It is by far the most information-heavy domain: its corpus contains 2,969 individual policy facts, and a single task can depend on as many as 580 of them.

Where the agents lose ground

The overall scores only show whether the finished agents worked. Sierra also examined what the developer agents actually did while building them, and found several recurring problems: they often stopped researching the business too soon, asked too few questions when information was missing, made poor decisions about how much computing power the finished agent should use, and showed little appetite for trying different technical approaches.

“The failures mirror ones human agent developers see.”

Importantly, the researchers argue that these weren’t uniquely machine-like mistakes. “The failures mirror ones human agent developers see,” they write in the paper. That makes the results more interesting: the agents could write code and assemble functioning systems, but they were losing ground on familiar engineering problems such as gathering enough information before building, knowing when to ask questions, and exploring alternatives rather than settling quickly on an answer.

The information-gathering problem was particularly stark in banking. Developer agents opened fewer than 80 of roughly 1,700 available files, instead relying heavily on searches to find documents that appeared relevant. That meant they could start building without having uncovered all the business rules the finished agent needed to follow.

Nor did they make much use of the opportunity to ask the business for information that wasn’t in those files. Across the recorded runs, such interactions accounted for just 0.3% of the developer agents’ tool calls. On some tasks, agents could discover 20 to 25 requirements only by asking questions, yet they asked no more than four. Sierra found that asking really did matter: on tasks where its expert-built reference scored between 95% and 100%, builds that asked no questions scored just 5%, rising to 15% after one question and 25% after two.

Cost was another problem. Hyper-τ-bench limits how much the finished customer-service agent can spend on AI model calls while handling a conversation. Two builds exceeded that allowance — by 3x and 1.3x, respectively —and received a score of zero after penalties. Most went too far the other way: among agents that stayed within the limit, average spending was just 45% of the amount available.

Other weaknesses appeared in the technical choices the developer agents made: what kind of agent they built, which model they chose to run it, and, in some cases, whether they tried to uncover parts of the benchmark that were deliberately hidden from them.

Constructed-agent architectures, serving-model choices and “cheating-adjacent” attempts
Constructed-agent architectures, serving-model choices and “cheating-adjacent” attempts (Credit: Sierra)

There was remarkably little experimentation with different designs, too. Ninety-two percent of builds used a “single LLM tool loop” — essentially one AI model repeatedly deciding whether to respond or call a tool. This also mattered: in one telecom experiment, giving the developer agent a single sentence suggesting a different architecture lifted its score from 31% to 67%.

“Because the system being built is an AI itself, the only way to know if a design works is to run it and read what it says to real users, who the developer never sees while building.”

That creates a blind spot for the developer agent: it has to make design choices without seeing some of the strongest evidence of whether those choices actually improve the customer experience.

“Because the system being built is an AI itself, the only way to know if a design works is to run it and read what it says to real users, who the developer never sees while building,” Shi and Dhandhania write.

And Sierra’s results suggest the developer agents often failed to compensate for that blind spot through enough testing and iteration, instead “shipping the first design that runs.”

On top of that, the agents also tended to favor familiar models. Ninety-six percent of Codex builds chose an OpenAI model to power the finished agent, compared with 13% of builds produced by Kimi. Sierra argues that the pattern suggests developer agents often defaulted to familiar model families rather than testing which option worked best for the job.

Finally, the researchers recorded what they call “cheating-adjacent” behavior in between 17% and 42% of runs, depending on the developer setup. This didn’t mean looking for business requirements they were expected to find; rather, the agents tried things like searching for the benchmark’s hidden test data or probing the grading system—information kept secret so a system cannot simply build to the answers. None of those attempts succeeded, according to Sierra.

Agents build agents

AI is playing a growing role in building agents. Tools such as Microsoft’s Copilot Studio and Salesforce’s Agentforce Builder let people describe an agent in natural language and have AI generate much of its underlying logic.

Sierra itself goes further with Ghostwriter, dubbed the “agent-building agent.” Users can give it instructions, standard operating procedures, transcripts, or recordings and have it build or modify an agent, generate tests, run simulations, and fix problems it finds. Sierra still gives humans the final say: Ghostwriter shows what it has built before anything goes live so it can be reviewed and approved.

Developers can also hand coding agents such as Claude Code or Codex a much broader “build me an agent” task. In all of these cases, though, humans still tend to provide much of the business context, decide what good looks like, and check the result.

Sierra’s own description of agent development helps explain why. Shi and Dhandhania argue that building an enterprise agent, even for people, is often “less like implementing a spec, and more like doing research.”

“Requirements are scattered across handbooks, support, spreadsheets, and the minds of your best frontline reps — so you form a hypothesis, dig up evidence, and build and test to identify which levers actually move performance,” they write.

That is also where Hyper-𝜏-bench’s results become interesting. The autonomous developer agents could write code and produce working systems. Still, they often failed to gather enough information, rarely asked questions when requirements were missing, and experimented little before settling on a design.

And this, perhaps, is why Hyper-𝜏-bench could prove a notable addition to the burgeoning benchmark brigade. AI is already taking on more of the work involved in building agents. The benchmark asks what happens when you remove much of the human guidance that still surrounds that process today—and, at least for now, its results suggest autonomous developer agents still struggle with some of the judgment-heavy parts of the job.

The post Claude performed best on a new benchmark for ‘agents that build agents’. But it passed fewer than a quarter of the tests. appeared first on The New Stack.

  •  

Building trust in agentic RAG starts with evidence

High contrast abstract digital texture evoking complex AI data traces and execution paths.

Basic retrieval-augmented generation (RAG) follows a straightforward pattern. A user asks a question, the system finds relevant content in a knowledge base, and the model uses it to ground its answer. This works for simple lookups, but many real-world retrieval systems need more control over how and where they search.

Agentic RAG lets an agent rewrite the question and choose where and how to search. It may query a knowledge base or an account system, combine lexical, semantic, and graph search, fuse the resulting scores, rerank candidates, discard weak results, and try again. This can find evidence that a single semantic search would miss. It also adds more decision points that should be supported by evidence, and a confident answer may not reveal the retrieval path that led to it. 

“The opportunity comes with a responsibility: more decisions require a clear evidence trail.”

The opportunity comes with a responsibility: more decisions require a clear evidence trail. More control can improve coverage, but control alone cannot create trust. The system earns that trust by showing what it searched and why it accepted a source. It must also disclose what it couldn’t verify. Without that record, it can be harder to understand the basis for even a good answer.

Retrieval is a series of decisions

A retrieval turn may look like a single operation in the application, but the agent is making a chain of choices. It interprets the user’s intent and creates a query. Then it chooses data sources, applies the required filters, and inspects the results. Only then can it decide whether the evidence is sufficient and connect claims to citations.

Each choice deserves care. An agent might search the support index for a billing question, remove a product name while rewriting a query, or find the right policy in the wrong customer’s account. The final answer could sound convincing yet be incomplete, fall outside the intended scope, or be unsuitable to share.

Workflow diagram comparing basic RAG against agentic RAG.
A one-shot retriever makes a single retrieval pass. An agentic retriever may make several, and each can fail with no visible change in the final answer.

A final list of top-k chunks can’t reconstruct this process. By then, the agent may have issued multiple queries and rejected several sources. It may have switched tools or rewritten the query. Each step must record structured data while it happens. This becomes your flight recorder for retrieval:

request      "Can I cancel this contract early?"
query        "early termination enterprise agreement"
source       approved_contracts (tenant=acme, region=US)
accepted     contract_884 §12, effective=2026-01-01, score=0.81
rejected     policy_119, reason="expired 2025-12-31"
decision     evidence sufficient for contract terms; fee amount unverified

Keep the query and its filters. Add source IDs, ranking data, timestamps, and the reason for each branch. There isn’t one correct retrieval method for every request. Lexical keyword search may be best for an exact contract number. Vector search may be best for a paraphrased policy question, while a plain SQL query pulls back an account balance. Graph traversal may be used to connect related documents or entities. The record should say which one the agent chose and why.

Give users and operators visible evidence

Users and operators need different views of the same evidence. A user needs citations that identify the source and the relevant passage or record. Each citation should show the source’s effective date or last updated date, along with the date it was retrieved. Users need plain language when the evidence has limits: “I found the cancellation terms, but I couldn’t verify the current fee for your account.”

Operators need enough detail to produce and improve an answer. Preserve the rewritten queries and search attempts, while protecting those traces with appropriate access controls, redaction, and retention rules. Keep the rejected results, tool calls, applied filters, and any instructions that affected source selection. A citation alone does not establish claim support. An agent can cite a legitimate document that contains related language but doesn’t support the claim it wrote. It can also attach a citation after the answer is generated, leaving it unclear whether the source informed the answer. 

Preserve citation provenance during generation, then verify that each claim is supported before releasing the answer. Store the source IDs passed to the model and map each supported claim back to the excerpt or record that supplied it. If a claim has no source, the application can remove or qualify it. For a high-risk claim, it can hold the answer for review before it reaches the user.

Use a practical replay test. Give an engineer the request and the trace, then ask, “Why this source?” Why was it valid at that time? Why did the system reject the alternative? If the trace can’t answer those questions, it isn’t detailed enough.

Make currency and authority part of retrieval

Semantic similarity measures resemblance, not authority. A policy from last year can match a question perfectly and still not be the appropriate result in the index. A current policy with different wording may be the only one the agent should use.

“A similarity score is an opinion; a scope filter is a rule the system can enforce.”

Treat source metadata as part of retrieval. Start with the effective date and owner, then record the access scope and document type. Approval status and jurisdiction matter for controlled material. Tenant identity is a hard boundary. A similarity score is an opinion; a scope filter is a rule the system can enforce. All of these fields should affect filtering and ranking. A regulatory question may require an approved primary source. A product question may prefer the latest published manual. A customer question must stay inside that customer’s scope.

Workflow diagram of an example where the closest match isn't the most appropriate one.
The closest match is not always the most appropriate source. Metadata rules decide what a similarity score can’t: whether a candidate is current, approved, and inside the caller’s scope.

These rules can run before or after similarity ranking—or at both stages. The placement depends on the data and the risk, but either way, an unauthorized or expired record should be excluded, even when its wording appears to be a closer match. When tenant and scope boundaries are properly implemented, unauthorized records can be excluded before they become candidates.

Conflicting sources need their own path. If two approved policies overlap, the agent should not default to the most convenient paragraph. It should report the conflict and narrow the answer to what both sources support. If that isn’t possible, it should send the request for review. Someone must also own each source throughout its active life and retire it upon expiration. Retrieval cannot establish currency from a document library that is no longer maintained.

Define a retrieval policy for the agent

“Be accurate” is an important goal, but too vague to serve as a retrieval policy. The application needs enforceable rules for when the agent searches and which source types it can use. Separate rules should govern when it may broaden a query and when it must admit the evidence is incomplete.

Apply those rules before the model writes. Customer data stays within the verified customer scope. Regulatory answers use approved sources for the correct jurisdiction and effective date. A missing primary source results in a qualified answer or a request for review. These constraints should be implemented in tool permissions, query filters, and application code rather than relying on the model to remember a sentence in its instructions.

“The agent decides what to ask; the retrieval layer decides what may be returned.”

Tool access needs the same treatment. Searching a public knowledge base carries a different risk than searching contracts, case notes, or a company-wide file store. Give the agent access only to the systems required for the task, and pass verified identity and scope to each search tool. Do not rely on the model to supply them as query arguments. The agent decides what to ask; the retrieval layer decides what may be returned.

Build these controls into the retrieval path before exceptions reach users.

Treat retrieved content as data, not policy

Every document an agentic retriever reads should be treated as untrusted model input, even when the application controls the source. Some of those documents will contain instructions. A wiki page can contain language that asks the agent to disregard its source restrictions. An ingested PDF can carry a line telling the model to prefer it over newer material. In basic RAG, a planted instruction can corrupt the answer. In agentic RAG, it can also steer subsequent searches, right down to the citations the agent presents as evidence.

Workflow diagram showing how an embedded instruction can influence a basic RAG answer.
An embedded instruction can influence a basic RAG answer. In agentic RAG, it can redirect later searches, and memory can carry that influence into future requests.

The governing rule is that retrieved content is data, never policy. Scope and permissions come from the application, and nothing in a document body should be allowed to change authorization or policy. Identity and scope filters belong in tool code and, where possible, in the database itself. 

“The governing rule is that retrieved content is data, never policy.”

The prompt alone is not a sufficient enforcement layer. Query rewrites and tool calls still need to be validated against the retrieval policy. The trace can reveal that a document influenced the next search, making it useful as a tool for detection and investigation. By the time it tells you anything, though, the search has already run. If retrieved material can be promoted into memory, the instruction can outlive the retrieval that introduced it and influence unrelated future requests.

Keep retrieval near the data when it helps

Many RAG systems copy documents into one service, embeddings into another, metadata into a third, and permissions into application code. Each copy can update on a different schedule. That makes answer freshness harder to diagnose and access decisions more difficult to demonstrate.

Keeping more of that work near the operational data can shorten the path. Oracle AI Vector Search stores vector embeddings alongside business data, and its SQL queries can combine similarity search with relational filters and lexical search. A team using Oracle AI Database can keep operational records, vectors, and access rules in a data platform it already controls. Database-enforced access controls can apply row- and column-level policies inside the database, allowing access restrictions to be enforced independently of the retrieval service. 

This arrangement can reduce data copies and make lineage easier to inspect. It doesn’t decide which policy is authoritative, detect a conflict, or prove that a citation supports a claim. The retrieval policy and evaluations still have to do that work. Additional products do not resolve an undefined evidence path.

Test decisions as well as answers

An evaluation that scores only the final prose does not assess much of the decision-making in agentic RAG. Build a small set of requests that exercise those decisions. Include a current-policy question and a case with two tenants holding similar records. Add conflicting documents, an unusual but valid source, and a document that carries embedded instructions to the model. The set should also include a request with the correct result “I can’t verify this.”

Score retrieval separately from generation using measures such as corpus-selection accuracy, recall at k, tenant-isolation violation rate, citation coverage, and claim-support accuracy. Check whether the agent selected the correct corpus and applied all required filters. Inspect the selected sources and their citation-to-claim links. Confirm that the agent appropriately declined or escalated when it lacked evidence. The answer can sound awkward and still retrieve correctly. It can also sound convincing while using an expired policy.

Run these cases after a change to the embedding model, chunking method, index, prompt, ranking rules, or search tool. A higher relevance score means very little if the new index starts to prefer older documents or crosses a tenant boundary. Save production issues as new evaluation cases so the same issue is less likely to recur.

Each answer needs an evidence path

Agentic RAG adds decisions, and confidence grows when the system can account for them. Make the evidence path and retrieval policy visible outputs instead of details buried in logs. When an answer needs review, that record is what lets a person decide whether it deserves their trust.

Trying to implement agentic RAG? Working examples of these patterns, such as agentic RAG with hybrid search, are available in Oracle’s AI Developer Hub.

The post Building trust in agentic RAG starts with evidence appeared first on The New Stack.

  •  

Microsoft built a prompt injection detector. Then it caught a phishing campaign instead.

dice

Microsoft flagged a phishing campaign last week that exploits a gap in how machines read text. Attackers are slipping invisible Unicode tag characters into email bodies; they don’t render on screen but change the underlying string that software processes.

The company says attackers are already using the technique at scale to bypass spam filters and ML-based classifiers, and the same approach could cause problems for AI systems that regularly ingest text from external sources.

Tag characters split keywords

Security researchers have documented a nearly identical technique targeting LLMs, commonly called “ASCII Smuggling,” which uses Unicode tag characters in the U+E0000 to U+E007F range — code points that exist in the character stream but aren’t displayed by most interfaces.

That gives you two versions of the same text: what a person reads and what software receives.

For example:

Human view:     funding
Under the hood: fun⟨U+E0020⟩ding

In the campaign tracked by Microsoft Defender for Office 365, attackers weren’t using tag characters to smuggle hidden instructions into an AI model. They placed them inside high-signal words associated with financial phishing, such as “funding,” “loan,” and “credit,” so that filters scanning for those terms would no longer find an exact match.

A hunting signature for ASCII Smuggling fired on roughly 21,000 messages the day before the campaign started and the next day, it fired on more than 1.3 million. Then, just two days later, the count passed 2.3 million. The whole time, recipients saw ordinary-looking offers for business loans and credit lines.

Just two days later, the count passed 2.3 million.

Tokenizers parse them differently

NLP systems break text into tokens before processing it, and slipping an unexpected Unicode character into a word can change how those tokens are formed. Researchers have already shown that encoding techniques can hide adversarial content from AI systems, although Microsoft’s campaign uses the trick for a different purpose.

NLP systems break text into tokens before processing it, and slipping an unexpected Unicode character into a word can change how those tokens are formed.

Exactly what happens depends on the tokenizer. Some may ignore the tag character while others split the surrounding text differently, so developers have to test the models they’re actually using rather than assume they’ll all behave the same way.

Running the text through standard Unicode normalization won’t necessarily remove the tags, either. NFC and NFD can clean up different representations of the same character, but they weren’t designed to strip Unicode tag characters, which means those tags can still make it through to the next step.

Agents lack email’s defenses

Email providers have other ways to spot a suspicious message beyond the words it contains, but an AI pipeline may be working with far less information.

That then becomes a problem when agents are pulling in outside text and using it to decide what to do next because those invisible characters buried in the text can change how it gets processed along the way, while also making a hidden prompt injection much harder for someone looking at the original to catch.

Normalize before the model

For applications that have no reason to accept characters in the U+E0000 to U+E007F range, the simplest approach is to remove them before the text reaches the model, although that gets trickier when an application has a legitimate reason to keep them.

In those cases, developers can compare the original text with a version that has the tags removed and look for anything that changed, while also testing the tokenizer their application actually uses to see how it handles the same characters. Whatever gets cleaned should stay that way through the rest of the pipeline, rather than checking one version of the text and then sending the untouched original to the LLM.

The subdivision flag edge case

Stripping every Unicode tag character isn’t always safe because some serve a legitimate purpose. The subdivision flag emojis for England, Scotland and Wales rely on invisible tag-character sequences to render, and Microsoft’s initial hunting signature was broad enough to trip on those flags before the team carved out an explicit exception.

Stripping every Unicode tag character isn’t always safe because some serve a legitimate purpose.

The post Microsoft built a prompt injection detector. Then it caught a phishing campaign instead. appeared first on The New Stack.

  •  

“Sorry for the messy rollout”: OpenAI launches GPT-6 Astra to most paying users a day after its unveiling

ominous door

Update: As of 6:30 p.m. Eastern on Friday, September 4, GPT-6 Astra was available on ChatGPT for all paying users with plans that include access to it.

Thibault Sottiaux, a leading member of the technical staff at OpenAI, posted on his X account, “OK nevermind, the team and Astra did a good job and our systems are more scalable than we anticipated. Astra is now rolled out to all Plus and Business users too. Hope you have a blast and let us know how it goes!.”

Earlier in the day Friday, OpenAI CEO Sam Altman posted on X, “GPT-6 Astra is now available to all Pro, Enterprise, and Business Premium users in Work/Codex, and is available in the API. We will start rollout to Plus and Business users next. Thank you for the patience.”

Users who pay for the $8/month ChatGPT Go plan do not and will not have access to GPT-6 Astra or GPT-5.6, according to the company’s pricing tier documentation.

Sottiaux posted on Thursday evening after OpenAI announced the debut of Astra but before it released it to users: “We are starting to release GPT-6 Astra, and we are doing it as carefully and quickly as possible.”

OK nevermind, the team and Astra did a good job and our systems are more scalable than we anticipated.

Astra is now rolled out to all Plus and Business users too. Hope you have a blast and let us know how it goes!

— Tibo (@thsottiaux) September 4, 2026

Our original story, published at 12:24 p.m. Eastern on Friday, continues below:

OpenAI launched GPT-6 Astra on Thursday, but many developers hoping to give it a test drive are still waiting for access.

Hours after the announcement, OpenAI CEO Sam Altman apologized for what he called a “messy rollout,” acknowledging that broad access to Astra had not yet begun for either API customers or ChatGPT subscribers.

“First, sorry for the messy rollout,”

“First, sorry for the messy rollout,” Altman wrote in a post on X. He added that OpenAI expected to begin the broader rollout “in the near future,” starting with ChatGPT Pro subscribers.

first, sorry for the messy rollout.

second, when we screw up, we try to make it right.

third, we should be able to begin broad rollout to API customers and chatgpt subscribers in the near future. as usual we will start with pro subscribers. https://t.co/nKOhW18CDK

— Sam Altman (@sama) September 4, 2026

For developers, that creates an unusual bind. OpenAI has already published Astra’s API documentation and pricing, including its 1.05 million-token context window, support for up to 128,000 output tokens, and standard API rates of $10 per million input tokens and $50 per million output tokens. But the endpoint itself is still rolling out, and OpenAI has yet to explain specifically what went wrong with the rollout. So while Astra’s benchmark results were impressive, developers and independent reviewers who want to evaluate those claims against their own workloads are still left waiting.

Why developers are waiting

OpenAI’s engineering lead for Codex, Thibault Sottiaux, offered a little more detail in an X post of his own about what is happening behind the scenes. He said the rollout will take “a few days” to complete and that OpenAI is bringing new systems and additional compute online as it expands access.

“We are starting to release GPT-6 Astra, and we are doing it as carefully and quickly as possible,” Sottiaux wrote. He added that “many novel systems will operate at scale for the first time” during the rollout and that OpenAI is “bringing a lot of compute up.”

“We are starting to release GPT-6 Astra, and we are doing it as carefully and quickly as possible,”

While this does not establish that compute capacity caused the delay, it does offer some insight into what OpenAI is dealing with as it expands access. The company’s original announcement said Astra would initially be available to a limited group of organizations, with ChatGPT Plus, Pro, Business, and Enterprise users, the OpenAI API, Microsoft Azure, and AWS Bedrock expected to follow “over the coming days.”

Developers who have secured early API access are already encountering surprises beyond the rollout itself. As The New Stack reported this week, Astra’s API introduces a new class of safety-triggered interruptions that look like timeouts but aren’t — which could make a real difference for developers building production workflows around the model.

Banked resets explained briefly

While users wait, OpenAI is trying to make up for at least part of the delay with something it calls a “banked reset.”

Sottiaux said paid ChatGPT subscribers will receive one banked reset for every day they remain without Astra access, beginning September 3, adding that the “team is moving mountains to give access as fast as we can.”

“Team is moving mountains to give access as fast as we can.”

OpenAI has used banked resets before with ChatGPT Work and Codex, giving users a way to replenish their usage after hitting a limit. They aren’t additional API credits or a permanent increase in usage limits but resets users can save until they need them.

What’s different with Astra is how OpenAI is handing them out: one for every day a paying ChatGPT subscriber remains without access. So far, OpenAI hasn’t announced anything similar for developers waiting to use Astra through the API.

The gesture fits a broader pattern of OpenAI experimenting with how it charges for AI: the company recently announced an outcome-based pricing model that would bill only when the model produces a correct result; another sign that its pricing strategy is still very much in flux.

Compute capacity complicates rollout

OpenAI says Astra will roll out to Plus, Pro, Business and Enterprise users, along with the OpenAI API, Microsoft Azure and AWS Bedrock, “over the coming days,” with Pro subscribers first. Sottiaux said the rollout should take a few days to complete.

For developers, the biggest unanswered question is when broad API access will arrive and whether it will come with tighter usage limits. Astra’s persistent-agent capabilities also make the wait more significant for developers.

OpenAI did not immediately respond to The New Stack’s questions about the rollout problems, banked resets, or API timeline.

The post “Sorry for the messy rollout”: OpenAI launches GPT-6 Astra to most paying users a day after its unveiling appeared first on The New Stack.

  •  

GPT-6 Astra’s score of 98.6% looked like AGI. Then researchers read the fine print.

abstract glass

There was no mistaking the divide in March with the release of ARC-AGI-3. While frontier AI models could do little more than register a sub-1% score, humans were able to navigate their new interactive settings.

OpenAI reports a different story for GPT-6 Astra six months on: 98.6%. Put that against the GPT-5.6 Sol it has superseded, which OpenAI puts at 7.8%, and the improvement is hard to miss.

Then again, one has to consider what ARC-AGI is designed for. The whole point is to put models in uncharted interactive territory where they cannot simply rely on training data to find an answer but must work out the mechanics of the environment themselves. Given what ARC-AGI-3 was built to test, a jump to 98.6% is enormous. But the number comes with an important caveat.

Given what ARC-AGI-3 was built to test, a jump to 98.6% is enormous. But the number comes with an important caveat.

The asterisk on Astra’s 98.6%

Astra was evaluated through the company’s Responses API harness, with two settings changed to better reflect how the model performs in real-world use. OpenAI says those changes weren’t made specifically for ARC-AGI-3, but the other models in its comparison were evaluated using different setups.

ARC-AGI-3 requires a model to find its way through an unfamiliar environment, which means the setup it runs in can affect how well it performs.

The model is only part of the story

The gains aren’t limited to ARC-AGI-3. On FrontierMath Tier 4, the model scored 97.6%, followed by 100% on ExploitBench and 99.2% on SRE-Bench with four attempts. Terminal-Bench Science saw one of the biggest jumps, from 22.4% for GPT-5.6 Sol to 64.6%.

OpenAI warns against rolling those results into a single measure of performance, but the range shows how much more the model can take on. In the company’s demonstrations, it works directly inside software such as KiCad, Power BI and Unity, while an experimental Codex feature lets it keep notes and search earlier context when a job runs longer than a single context window.

On offline OSWorld 2.0, Astra scored 72.6% while taking about 40 minutes per task, compared with Sol’s 65.7% and roughly 75 minutes.

On offline OSWorld 2.0, Astra scored 72.6% while taking about 40 minutes per task, compared with Sol’s 65.7% and roughly 75 minutes.

Beyond benchmarks into discovery

The math is where things get more interesting. OpenAI says Astra was involved in two new findings about gaps between prime numbers. Mathematician Julia Stadlmann had already pushed one bound from 246 to 240. With Astra involved, it fell again, this time to 186. The company points to another case where the model helped improve part of a bound that hadn’t budged in more than 80 years.

There’s an important gap in OpenAI’s account, though. It doesn’t spell out what Astra came up with on its own, what researchers suggested or how the work moved between them. So while this goes beyond solving a benchmark with a known answer, it’s not enough to call the math evidence of AGI.

Alignment gains, oversight gaps

Astra is pushing past many of AI’s familiar limits, from unfamiliar problems to longer tasks. Yet even a 98.6% score on ARC-AGI-3 doesn’t settle the AGI debate. Part of the problem is that performance more often depends on the system around the model. Intelligence itself doesn’t improve evenly, either.

In OpenAI’s internal tests involving difficult or impossible tasks without production safeguards, GPT-5.6 Sol went beyond what it was authorized to do 48.2% of the time. Astra didn’t do so even once. Yet when researchers explicitly asked the models to evade monitoring, Astra’s written reasoning was harder to follow than Sol’s. OpenAI says that’s partly because Astra can solve simpler problems in fewer written steps, although it still struggles to conceal its reasoning on more complicated tasks.

If AGI means doing useful intellectual work across different fields, Astra is getting remarkably close to what many people once had in mind. If it means matching human judgment across the board, ARC-AGI-3 can’t establish that.

Epoch AI’s Greg Burnham described Astra as the “end of one era, start of another.”

“…end of one era, start of another.”

[Editor’s note: This article’s headline has been updated to clarify that a 98.6% score on ARC-AGI-3 does not mean the benchmark was “aced.” ARC-AGI-3 scores performance against a human baseline, with 100% representing performance at or above the median human baseline.]

The post GPT-6 Astra’s score of 98.6% looked like AGI. Then researchers read the fine print. appeared first on The New Stack.

  •  

Cut GPU inference cold start from 8 minutes to less than a minute

We instrumented the full path from pod creation to first inference response on a GPU node running a 70B-class model. Eight minutes. Six sequential phases. We expected one bottleneck. We found six, and which one dominates depends on model size.

For a 64 GB model, 65% of the startup time is spent recompiling CUDA kernels that produce identical output every time. For a 203 GB model, 92% of the time is spent downloading weights from S3 through a calling pattern that leaves 98% of available bandwidth idle. Both are fixable with configuration changes. Neither is fixed by default.

“Eight minutes. Six sequential phases. We expected one bottleneck. We found six.”

We define time to first token served (TTFTS) as the wall-clock duration from pod creation to the first inference response leaving the GPU. Not time to first token (TTFT), which measures per-request latency once the model is warm. TTFTS is the one-time startup tax. TTFT begins where TTFTS ends.

Here’s what we achieved:

ScenarioDescriptionBeforeAfterReduction
Pod restart on warm nodeWeights loading + compilation on existing node1.5-8 minunder 30s80-93%
New node from scratchFresh node provisioned, nothing cached8-15 min~5 min40-65%

The warm-node row is what you pay on every pod restart: scale-up events, rolling updates, OOM recoveries. That’s the 80-93% win, and it requires only configuration changes. The cold-node row includes ~2 minutes of fixed infrastructure cost (node provisioning and framework initialization) that no application-layer optimization can remove. The rest is avoidable waste that we eliminated through platform and configuration fixes. The warm-node optimizations are environment variables and a volume mount that work on any Kubernetes cluster. The cold-node optimizations require EKS Auto Mode, which comes pre-configured with pre-compiled NVIDIA drivers, SOCI (Seekable OCI) parallel image pull, and NVMe instance store mounting.

All model startup measurements were taken on p5.48xlarge instances running Amazon EKS Auto Mode, with S3 traffic routed directly (bypassing the NAT Gateway) and container images in a private Amazon ECR repository (same region as compute). Model startup improvement ratios (80-93%) hold consistently across instance types (validated on P-family and G-family). Cold-node times vary with network bandwidth and CPU count. For the weights loading and compilation cache configuration, see Accelerate model loading on Amazon EKS.

The Kubernetes ecosystem has made real progress on the inference stack in 2026. OCI image volumes are now stable for model delivery. Dynamic Resource Allocation (DRA) gives GPUs structured attributes instead of opaque integer counts and provides flexibility in allocating GPUs to workloads. Gateway API has inference-aware routing extensions. But none of these primitives address the full cold-start stack: the six layers between “pod pending” and “first token served,” each with its own bottleneck and its own fix.

The six layers of cold start

When a new inference pod starts on a freshly provisioned GPU node, it passes through six distinct phases before serving its first request:

  1. Node provisioning. Karpenter launches an EC2 instance, boots it, and registers it with the Kubernetes API server (~60-90s).
  2. GPU driver initialization. The driver kernel module must load and expose accelerator devices.
  3. Container image pull. The inference engine image (8-12 GB compressed) must be transferred to the node and extracted.
  4. Model weights download. The model files must stream from object storage into GPU memory.
  5. GPU kernel compilation. torch.compile traces the model graph and generates optimized CUDA kernels.
  6. Engine initialization. CUDA graph capture, KV cache profiling, and HTTP server startup (30-120s depending on whether compilation is cached).

Each layer has a different bottleneck, a different fix, and a different owner.

Which layer dominates depends on model size

Before diving into each layer, one finding shaped every decision we made: the bottleneck is not fixed.

We instrumented the model startup path (layers 4 and 5) and measured each phase independently for two model sizes:

64 GB model (Qwen3.6-35B-A3B):

  • Weights loading: ~29s (35% of model startup)
  • torch.compile: ~53s (65% of model startup)

203 GB model (Llama-4-Scout, TP=4 where TP is tensor parallelism, splitting the model across GPUs):

  • Weights loading: ~423s (92% of model startup)
  • torch.compile: ~34s (8% of model startup)

For models under ~100 GB, compilation dominates. For larger models, network transfer dominates. torch.compile time stays roughly constant (it depends on graph complexity, not parameter count). Weights loading scales linearly with file size.

“For models under ~100 GB, compilation dominates. For larger models, network transfer dominates.”

This means any single-layer optimization has a ceiling.

Layer 1: Node provisioning

On EKS Auto Mode and Karpenter-managed clusters, node provisioning takes approximately 60-90 seconds for accelerated instances from pod pending to node Ready. Karpenter calls the EC2 Fleet API directly and reacts to pending pods within seconds, keeping provisioning at the EC2 launch floor.

Layer 2: GPU driver initialization

The NVIDIA GPU Operator in its default configuration adds 2-3 minutes to node boot while it compiles the driver kernel module from source. This cost repeats on every new node.

When the platform controls the full stack (OS image, kernel version, driver version, boot sequence) it can pre-compile driver kernel modules at image build time. The node boots, runs modprobe to load an already-compiled .ko file, and the GPU is ready in seconds.

This matters more now than it used to. Blackwell-architecture GPUs (G7, G7e instances) require NVIDIA’s open-source kernel modules exclusively. Older Maxwell/Pascal/Volta GPUs can only run proprietary modules. A cluster with both legacy and next-gen GPU nodes needs different drivers, different AMIs, different upgrade cycles. A managed platform that pre-compiles the correct module per instance family eliminates this complexity.

On EKS Auto Mode, the GPU driver loads in seconds (pre-compiled at image build time), compared to the 2-3 minutes a runtime-compilation approach requires.

Layer 3: Container image pull

A production vLLM or SGLang inference image is typically 8-12 GB compressed. Standard containerd pulls layers sequentially, decompresses them one by one in memory, and writes them to disk. At this size, sequential pull takes 2-4 minutes on a cold node depending on instance type and available CPU cores. For larger custom images (30-50 GB compressed), containerd can run out of memory entirely during decompression.

EKS Auto Mode uses SOCI’s parallel pull mode, which replaces containerd’s default snapshotter. The SOCI snapshotter downloads layer chunks concurrently via HTTP range requests and writes each chunk directly to its target byte position on disk (no in-memory ordering buffer). Decompression runs in parallel across all available CPU cores.

Pull time is bottlenecked by CPU-bound decompression, not network bandwidth. We confirmed this directly: a p4d.24xlarge with 400 Gbps networking achieved only ~1 Gbps effective pull throughput because CPU decompression was the constraint. On instances with capable, current-generation CPUs, SOCI parallel pull reduces image pull time from 2-4 minutes to 30-60 seconds. The dominant factor is per-core decompression throughput, which depends on CPU generation and instruction-set support, more than raw core count. A newer CPU with fewer cores can outperform an older one with more.

For a deeper look at how bounded-memory parallel pull handles images exceeding 30 GB without OOM, see Bounded-Memory Parallel Image Pulling for Large Container Images.

Layer 4: Model weights download

The obvious optimization for weights loading: more parallel connections. Split the model files into small chunks, download them concurrently, saturate the network pipe.

We tested it on p5.48xlarge with the 64 GB model streaming from same-region S3. The results were counterintuitive:

Chunk sizeConnections neededWeights load time
256 MB25613.98s
512 MB12814.20s
2 GB3413.62s
4 GB1713.35s
8 GB921.80s (+56%)

256 parallel connections provided no benefit over 17. The only failure mode was 8 GB chunks (exceeding shard file size), which caused a 56% regression.

Why? Because the open-source Run:ai Model Streamer (integrated into vLLM and SGLang) processes S3 range requests sequentially within each worker thread. A worker assigned to a 3.9 GB shard file downloads its byte-range requests one after another on a single connection. The parallelism comes from running multiple workers on different files, not from splitting one file into more pieces.

We settled on 4 GB chunks matching typical SafeTensors shard size (3-5 GB per file) with an aggressive timeout-and-retry for slow requests. S3 GET latency has a measurable long tail: in our testing, a meaningful fraction of requests took 2-3x longer than median, and a single stalled connection holds up the entire model load. Rather than wait, we kill stalled connections after a few seconds below a speed threshold and retry on a fresh connection. This follows S3’s own performance guidance.

For the 203 GB model, these config-only changes reduced weights loading from 423 seconds to 25 seconds (94% improvement). For the 64 GB model, from 29 seconds to 12 seconds. No code modifications, just environment variables. The tuning consists of three settings: chunk size aligned to shard file boundaries (eliminating the serial sub-request problem), a minimum-speed threshold that kills and retries stalled S3 connections, and explicit concurrency matching the number of shard files per tensor-parallel rank.

Layer 5: GPU kernel compilation

Every time a vLLM or SGLang pod starts, PyTorch traces the model’s computation graph and compiles it to optimized CUDA kernels. This takes 34-53 seconds depending on model architecture. The output is identical every time for the same model, GPU type, and tensor-parallel configuration.

And Kubernetes throws it away on every pod restart. Pods use ephemeral storage by default. When a pod terminates, its local filesystem is destroyed. The next pod recompiles from scratch.

“The output is identical every time for the same model, GPU type, and tensor-parallel configuration. And Kubernetes throws it away on every pod restart.”

Point the torch.compile cache directory at local NVMe instance store. GPU instances ship with NVMe that EKS Auto Mode mounts automatically. First pod compiles and writes ~15-30 MB of cached kernels. The second pod on the same node loads pre-compiled binaries in 4-6 seconds. One volume mount and environment variables.

The cache is safe because the compiled artifacts are deterministic: same model architecture + GPU architecture + tensor-parallel degree + PyTorch version equals valid cache. An image update or hardware change triggers exactly one recompilation.

torch.compile time is hardware independent. The same model compiles in ~52 seconds whether running on H100 or A100. The cache hit (4-6 seconds) is equally consistent across GPU types. This means the optimization works identically regardless of instance type.

Layer 6: Engine initialization

After weights are loaded and kernels compiled, the inference engine must capture CUDA execution graphs and profile KV cache memory. With compiled kernels cached, this completes in 30-45 seconds. Without cache, graph capture triggers additional JIT compilation and takes 60-120 seconds.

This is why the torch.compile cache has an outsized impact: it accelerates not just layer 5 but also layer 6. Cached compilation reduces a 2-3-minute combined phase to a 35-50-second combined phase.

Framework initialization (Python interpreter startup and PyTorch import) adds tens of seconds of fixed overhead that cannot be reduced through configuration.

The compounding effect

The six layers compound. Platform fixes (layers 1-3) eliminate 4-8 minutes of overhead: pre-compiled drivers replace 2-3 minutes of runtime compilation, parallel pull reduces image transfer time from 2-4 minutes to 30-60 seconds, and Karpenter keeps node provisioning to its hardware minimum. Configuration changes (layers 4-5) cut the remaining model startup by 80-93%. Engine initialization (layer 6) drops from 60-120 seconds to 30-45 seconds once the compile cache is warm. Together, cold-node TTFTS drops from 8-15 minutes to approximately 5 minutes.

64 GB model (Qwen3.6-35B-A3B), TP=2:

ConfigurationFirst podSubsequent pod (warm node)
Baseline (no tuning)82s82s
+ S3 chunk tuning65s65s
+ torch.compile cache65s16s
Improvement-21%-80%

203 GB model (Llama-4-Scout), TP=4:

ConfigurationFirst podSubsequent pod (warm node)
Baseline (no tuning)457s457s
+ S3 chunk tuning59s59s
+ torch.compile cache59s32s
Improvement-87%-93%

The warm-node subsequent pod number is what matters most for production. It’s what you pay on every pod restart. The 80-93% reduction is consistent across instance types because the optimizations target software bottlenecks (calling patterns, redundant compilation), not hardware limits.

The cost of cold starts at scale

Why does any of this matter? Because GPU nodes are expensive and inference traffic is bursty.

A single p5.48xlarge costs $55/hour on-demand. Even G-family instances commonly used for inference cost $10-20/hour. Every minute of cold start is GPU time you’re paying for but not using. If your autoscaler needs 8+ minutes to bring up new capacity, you must over-provision (burn money on idle GPUs) or accept latency spikes during traffic surges.

“Every minute of cold start is GPU time you’re paying for but not using.”

When model startup drops to 16-32 seconds on warm nodes, the calculus changes. You can scale more aggressively, keep fewer buffer nodes, and respond to traffic spikes without multi-minute startup delays.

What we learned

  1. Decompose before optimizing. For 64 GB models, torch.compile dominates (65%). For 203 GB models, S3 loading dominates (92%). Without measuring each phase independently, we would have optimized the wrong layer.
  2. The bottleneck flips with model size. torch.compile time is roughly constant across model sizes. Weights loading scales linearly. Every team running inference should know which regime they’re in.
  3. “More parallelism” requires understanding the execution model. 256 connections performing sequential work inside each thread is no faster than 17. The bottleneck was the calling pattern, not the concurrency limit.
  4. 15-30 MB can save 53 seconds. The most impactful optimization for smaller models was persisting a tiny cache file. Always check whether an expensive computation produces deterministic output before trying to make it faster.
  5. Platform-level control enables optimizations that configuration alone cannot achieve. Pre-compiled drivers, default-on parallel image pull, and NVMe auto-mounting are infrastructure-layer decisions that compound upward. Together with the config-only changes at the application layer, these changes reduce cold start time from minutes to seconds.
  6. The ecosystem is building the right primitives, but cold start lives between them. OCI image volumes, DRA, inference-aware routing, and local model caches are all real progress. But the compilation bottleneck and S3 tuning gaps sit in spaces that no upstream Kubernetes primitive addresses. Sometimes the highest-impact optimization is a volume mount and two environment variables, not a new API.

For the complete configuration guide, including environment variables, YAML manifests, and instance-specific recommendations, see “Accelerate model loading on Amazon EKS” in the Amazon EKS User Guide.

The post Cut GPU inference cold start from 8 minutes to less than a minute appeared first on The New Stack.

  •