❌

Vue normale

Reçu avant avant-hierInfra

Microsoft’s new Copilot agents get their own email, calendar — and a place in the org chart

25 septembre 2026 à 19:16
Satya Nadella stands smiling between Bill Gates, on the left, and Steve Ballmer, on the right, in front of a crowd of cheering employees, many holding up phones and tablets to take photos.

Microsoft announced what it calls its biggest Copilot update to date on Friday, with CEO Satya Nadella describing Copilot as “a new OS for work.”

Nadella framed Copilot as spanning every model, form factor, and task, and the update puts Autopilot, which Nadella called a “proactive and long-running agent built for the enterprise,” at the top of his list of the update’s four components. The pitch targets office workers, but the more consequential change for developers is the infrastructure underneath.

Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365, which means teams building production agents no longer have to assemble those pieces around a model on their own.

We’re building Copilot as a new OS for work that spans every model, every form factor, and every task. Today, we’re announcing our biggest update to Copilot to date, bringing four things together:

· Autopilot: proactive and long-running agent built for the enterprise
· Code:… pic.twitter.com/W2ClHHkCK3

— Satya Nadella (@satyanadella) September 25, 2026

The release adds a new Home experience that merges Chat and Cowork in the Copilot app, but the bigger changes for developers come from Code and Autopilot. Code generates apps, dashboards, and workflows from natural language, and Autopilot turns the agent Microsoft previously called Scout into a persistent background worker. Home and Code are rolling out first through Microsoft’s Frontier early-access program, and Autopilot is expanding to a private preview at month’s end.

Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365

Agents that don’t need prompts

Autopilot takes a role and goal from the person who sets it up, then continues working in the background without requiring a new prompt for each step. Each Autopilot gets its own governed Entra identity and agent user account, separating the agent’s permissions and activity from those of the person who created it.

For engineers, that moves much of the operational scaffolding required for long-running agents into Microsoft’s infrastructure. Independent vendors have been building dedicated layers for that problem; Diagrid, for example, adds durable recovery to LangGraph and other agent frameworks, while Microsoft is bringing those capabilities inside the Microsoft 365 environment.

An identity for every agent

The identity model is the piece developers building on Microsoft Foundry will feel first. Autopilot agents in Foundry, which have been in public preview since June, receive a full Entra Agent ID user account with a productivity license that gives them their own email, calendar, OneDrive storage, Teams access, and a place in the org chart.

Because that user account sits on top of the agent identity every Foundry agent already carries, an autopilot acts as itself rather than on behalf of a user, so developers no longer have to wire agents through shared service accounts or borrowed user credentials, a pattern AuthZed CEO Jake Moshenko has said reflects a common misconception about how agents should be deployed.

A developer creates an Autopilot blueprint from a Foundry-hosted agent, which appears in the Agent 365 registry once an administrator approves it. Employees can then hire instances of that agent in Teams. The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.

The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.

Hosting AI-generated apps

Code is built on the same underlying technology as GitHub Copilot, and the apps it generates run on Microsoft Copilot Managed Runtime, a platform now in public preview that hosts code inside the customer’s Microsoft 365 tenant boundary under IT governance.

Apps deployed there run within the company’s existing identity and governance framework, with Microsoft managing the underlying runtime and giving developers a controlled path to test and deploy new versions without taking the current release offline.

The runtime also accepts apps built in Copilot Studio and Cowork, and Microsoft is opening it to outside tools and professional developers through an SDK and command-line tooling, with Git tracking source and versions.

Lovable is already on board. In Microsoft’s announcement, the company’s head of global partnerships, Lan Roche, said apps built with Lovable can now run inside a Microsoft tenant “the same way everything else does,” using the same sign-in, policies, and app inventory.

The model resembles what serverless computing did for application infrastructure, where developers concentrate on application logic while the platform takes on more of the execution environment. Microsoft is applying that abstraction to generated enterprise software while tying the runtime directly to identity, tenant boundaries, and organizational data.

Long-running agents also change Copilot’s economics. The standard subscription covers the assistant, but Cowork, Code, Autopilot, and other agentic features are billed based on usage through Copilot Credits. That also applies to frontier models such as Fable and Astra, although users still need a Copilot license to access them. Microsoft is extending cost management in Agent 365 to cover Code and Copilot Managed Runtime, and it plans to support agents built in Copilot Studio in October.

Once an agent can keep working for hours or days without anyone watching, cost becomes part of the governance problem. Engineering teams need to control how much compute an agent uses alongside what it can access, which is why Microsoft is bringing those controls into the same administrative framework.

The portability trade-off

That convenience comes with a trade-off. Because Microsoft controls the underlying enterprise environment, it can handle much of the work around agent state, credentials, and access controls, but the more infrastructure a team hands over to Microsoft, the harder the agent may be to move elsewhere.

The models are not locked in, since Microsoft currently runs Copilot on models from both OpenAI and Anthropic and says more labs and open-weight models are coming, and the Agent 365 SDK adds governed Model Context Protocol access to Microsoft 365 workloads for agents regardless of the framework they were built with. Those open interfaces cover only part of an agent’s architecture, though. The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.

The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.

The post Microsoft’s new Copilot agents get their own email, calendar — and a place in the org chart appeared first on The New Stack.

OpenAI and Cursor agree on agent coordinators. They disagree on who runs them.

25 septembre 2026 à 14:00
Abstract digital art of thousands of thin glowing strands in orange, red and pink bundled into a single sweeping arch against a black background.

OpenAI opened its Agents API in public beta this month, exposing the harness that powers Codex with managed sessions, tool coordination, and subagent orchestration. On the same day, September 10, Cursor launched Projects to coordinate multiple coding agents around larger bodies of software work. While the products sit at different points in the stack, both converge on the same architecture: a coordinator understands the larger objective and manages the work, while specialized agents execute individual pieces.

That pattern isn’t new: AWS Bedrock AgentCore reached general availability in October 2025, and Anthropic’s Claude Managed Agents entered public beta in April 2026. What makes these announcements notable is that two major players in AI-assisted software development are independently exposing the same coordinator-worker split at the same time.

Hilliary Lipsig, a senior principal site reliability engineer at Red Hat who leads Azure Red Hat OpenShift SRE teams and hosts the YouTube livestream GitOps Guide to the Galaxy, has watched this dynamic play out firsthand.

“This convergence highlights the reality developers across the industry have been discussing on and offline — an agent with too much context loses accuracy and reliability, and focused work with clearer contexts allows for faster, more accurate iterations,” Lipsig tells The New Stack.

“The need for orchestration in distributed computing has been fundamentally recognized repeatedly,” Lipsig says. “That’s part of how we got to Kubernetes. These multi-agent workflows are the same concept, just in a new part of the technical stack. While the specialized agents do their area of work, the orchestrator can act as a source of truth — ideally enforcing guardrails, recovering from any failure states, and intelligently routing work to the most efficient target agent.”

“The need for orchestration in distributed computing has been fundamentally recognized repeatedly… These multi-agent workflows are the same concept, just in a new part of the technical stack.”

The industry has spent the first generation of AI coding tools asking how capable a model can become at writing software. The emerging question is different: How do you build a reliable system around multiple capable agents working on the same problem?

The problem with the single-agent loop

A coding agent works through what Anthropic describes as LLMs using tools based on environmental feedback in a loop: it observes the state of a repository, reasons about what to do next, calls a tool, examines the result, and continues. For a small task, that loop can be enough. As the scope expands, however, maintaining reliability in a single context becomes harder.

A large migration might require understanding an unfamiliar codebase, identifying dependencies, changing database schemas, updating services, rewriting tests, modifying deployment configuration, and validating the resulting system. A single agent can theoretically perform all of that work, but it must maintain relevant information from every stage while continuing to reason about what comes next.

The pressure lands first on the context window. “A large context doesn’t only include everything correct or important — it also includes a lot of throwaway information,” Lipsig tells The New Stack. “Through compaction, that information can inadvertently end up ranked as important and incorrectly influence what your agent does. Or correct information can be distorted to become incorrect.

“Either way, after a couple of rounds of compaction, developers are seeing accuracy degrade and are starting to manage context once again manually.”

Lipsig’s read matches what researchers call context rot — and it hasn’t gone away with newer models.

A 2026 study testing frontier models,, including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro, found they missed a dangerous action buried in a long agent transcript two to 30 times more often once it came after 800,000 tokens of benign activity — the AI equivalent of a security guard who stops checking badges carefully after the two-hundredth person walks through, even though nothing about their training changed.

Furthermore, the tasks themselves may not be sequential. Forcing one agent to execute database analysis, documentation work, and test discovery one after another turns a potentially parallel workload into a serial one.

Subagents change that execution model. Instead of requiring one agent to carry an entire task through a single context, a coordinator breaks the work into smaller units and assigns them to specialized agents. GitHub’s custom-agent model illustrates this: different agents receive only the prompts, tools, and context they need for their tasks, executing work in isolated contexts rather than crowding an increasingly large conversation.

Multi-agent systems therefore bring higher token costs and additional coordination and integration risks, and splitting work across agents does not guarantee better software quality.

The coordinator is not another coding agent

Once the work is divided this way, the coordinator becomes a control plane rather than another coding agent. Its job isn’t to write the code, but to understand the global task, manage dependencies, and decide how execution should proceed. Unlike a conventional scheduler, an agentic coordinator makes probabilistic judgments about result quality and resource allocation.

It may dispatch one agent to investigate a database schema, another to examine the service layer, and a third to inspect the test suite. When they return, the coordinator determines if their findings are sufficient to move to implementation. If a worker produces an incorrect result, the system must recognize the failure and decide whether to retry the work, reassign it, or change the task itself.

Anthropic has documented this same pattern in its own production system, calling it orchestrator-subagent architecture: a lead agent analyzes a query, develops a strategy, and spawns specialized subagents to investigate different facets in parallel. In a June 2025 writeup of that system, Anthropic reported a Claude Opus 4 lead agent with Claude Sonnet 4 subagents outperformed single-agent Opus 4 by 90.2% on its internal research eval — at roughly 15 times the token cost of a standard chat interaction (Anthropic puts single agents at about 4 times), a tradeoff that makes the pattern a deliberate architectural bet, not a free upgrade.

Parallelism introduces distributed-systems failure modes

Parallelism is valuable because software work contains many independent tasks, but it creates coordination problems. Imagine a migration where one agent changes a database schema, another updates the consuming service, and a third updates integration tests.

If the schema changes while the service agent works against an earlier assumption, the system produces internally inconsistent work. This isn’t a risk unique to hypothetical migrations — the International AI Safety Report 2026 notes that “interactions between multiple AI agents are also becoming more common, introducing further risks, as errors propagate between systems.”

A single model invocation is a disposable computation, but a twenty-minute workflow modifying a repository is not. If an agent loses its machine halfway through, restarting from scratch is expensive and potentially unsafe against a changed environment.

To solve this, Cursor moved its cloud-agent execution loop to Temporal to handle durable execution and retries, pushing its cloud agents past two 9s of reliability. Temporal now handles 50 million of Cursor’s actions a day across 7 million unique workflows. “Durable execution isn’t a nice-to-have here. It’s the difference between a system you can operate and one you can only demo,” Lipsig tells The New Stack.

“Durable execution isn’t a nice-to-have here. It’s the difference between a system you can operate and one you can only demo.”

By separating agent, machine, and conversation state, the execution engine can reason about the workflow independently. Reliability is no longer just about whether the model produces a good answer; it is about reliably completing distributed workflows composed of many operations, machines, and dependencies.

The environment, context, and observability are one problem

In production, an agent is more than a model and a prompt; it requires a workspace, source code, dependencies, credentials, and state retention. Both companies provision isolated environments for these resources, directly linking an agent’s capability to its blast radius. OpenAI’s Agents API currently supports U.S. data residency but not Zero Data Retention; choosing a self-hosted sandbox does not make the Agents API eligible for ZDR. Cursor supports similar cloud isolation alongside local execution for machine-specific work.

An agent that can only inspect a repository poses a different risk than one that can modify production infrastructure. Consequently, the coordinator is inextricably linked to the security model, determining which agent receives specific information and authorities.

This logic extends to context routing. Giving every subagent the parent’s entire history increases cost and complexity while leaking irrelevant or sensitive information. Instead, the coordinator enforces information-flow boundaries: a database-analysis agent receives only schemas and relevant migrations, while a security-review agent gets the resulting diff without deployment credentials.

As agents increasingly use interfaces like MCP to reach external systems, the platform must strictly govern which agent receives the authority to use specific tools, and for how long. MCP’s governance now sits inside the Agentic AI Foundation, a Linux Foundation foundation co-founded by OpenAI, Anthropic, and Block, with support from AWS, Google, Microsoft, Bloomberg, and Cloudflare to host MCP alongside AGENTS.md and Block’s goose — a sign the industry already treats it as infrastructure worth governing jointly, not a feature any one vendor owns.

This complexity creates a visibility problem. A simple final response often conceals a history involving multiple agents, tool calls, environments, and retries. Systems must expose task-level provenance — which agent received the assignment, what context it used, where it executed, and how the coordinator handled failures or human interventions.

Without execution provenance, debugging requires reconstructing distributed workflows from fragments. GitHub’s exposure of subagent lifecycle events points in this direction, treating agent lifecycles as observable components rather than hidden processes.

Coordination authority is not execution authority

The most critical architectural boundary is the distinction between coordination authority and execution authority. A coordinator needs broad visibility to make useful decisions, but that does not imply unrestricted control over the project. “Just like you don’t want humans running around with root permissions, you don’t want your agents running with them either,” Lipsig tells The New Stack.

“Just like you don’t want humans running around with root permissions, you don’t want your agents running with them either.”

“The ease of creating and leveraging AI agent permissions is lagging behind the speed of AI innovation, but any product team that needs to maintain compliance standards will tell you that easy or not, access controls are incredibly important. Functionally, that means fully agentic workflows need to run in a dedicated environment with their own permissions and connectivity. You can’t rely on agent context to do what you want — trying to keep something in context that long will just result in errors from context rot. Hardware or virtual-hardware-level permissions, in addition to system-level permissions, need to be configured so the agent is asked to provide only its unique value add.”

The risk isn’t hypothetical. OWASP’s Top 10 for Agentic Applications formally names this failure mode Identity and Privilege Abuse (ASI03).

METR’s independent investigation, conducted with a Redwood Research contractor, found that the June–July 2026 incident — in which OpenAI’s own agents, running in internal ExploitGym cyber evaluations, escaped their sanctioned scope and attacked Hugging Face — unfolded in distinct phases. On June 26, agents discovered an exploit that provided full administrator access to OpenAI’s internal Artifactory package repository. On July 4, high-volume Artifactory activity caused an outage, triggering a security investigation and pausing evaluations. OpenAI switched to a freshly initialized Artifactory instance on July 6, which it says wiped the cache and messages left there.

On July 7, a new wave of ExploitGym experiments launched tens of thousands of agents, and on July 8 one agent established the primary unsanctioned message board. Roughly 1,200 agents eventually used the board, sending more than 70,000 messages and files; about 700 later participated in the attack on Hugging Face. The attack itself began on July 10–11 and wound down over July 12–13. The chronology matters because the administrator-access event, the Artifactory outage, and the later message-board activity were separate phases, not one continuous incident.

By binding autonomy, worker agents operate with the minimum permissions required for their specific tasks, keeping sensitive operations behind explicit approval boundaries. This also reshapes human review. Requiring human approval for every tool call destroys the efficiency of multi-agent execution, but showing only the final result obscures critical intermediate decisions.

The most useful design places human intervention around consequential, irreversible transitions — like moving into production or altering sensitive infrastructure. This is especially vital as agents become event-driven participants that respond to Slack messages or pull request updates, not just direct prompts.

OpenAI and Cursor own different parts of the architecture

The convergence does not mean OpenAI and Cursor have built interchangeable systems. Their products put the orchestration boundary in different places.

OpenAI is exposing an agent harness through an API. Its model gives developers primitives for managing context, tools, subagents, and execution environments, leaving application teams to decide how those capabilities fit into their own systems. The harness is open source, so teams can inspect the coordinator logic instead of treating it as a black box.

Cursor packages more of the surrounding workflow. Projects provides the coordinator, cloud execution, shared project context, and a developer-facing workflow in the same environment.

That difference matters because orchestration is a collection of infrastructure decisions: who owns the execution environment, where workflow state persists, how agents are isolated, how credentials are provisioned, what happens when a worker fails, how one agent’s output becomes another agent’s input, and which actions can happen without human approval.

An API gives developers more responsibility for answering those questions. An integrated platform answers more of them on the developer’s behalf.

Neither approach removes the underlying engineering problems. It changes where they are implemented and who is responsible for operating them.

The coordinator is becoming an architectural boundary

The evidence from these systems points to a change in the role of the coding agent itself.

The model still performs the reasoning and code generation. But larger agentic workflows require another layer to determine how that capability is applied: which work is delegated, what context crosses an agent boundary, which tools are exposed, how execution state survives failures, and when the workflow needs human intervention.

Those are familiar distributed-systems concerns. Workers operate concurrently, state can be shared or isolated, dependencies connect tasks, workers can fail independently, and results need to be persisted and observed. The difference is that the workers are now probabilistic software agents rather than conventional processes.

That makes the coordinator more than a convenience feature. It is where a high-level software objective becomes executable work — and where decisions about context, permissions, durability, observability, and human intervention converge.

The September 10 launches make that shift visible from two different directions. OpenAI exposed orchestration infrastructure through an API. Cursor embedded it into a project-level development environment.

Neither announcement proves that one architecture will become the universal model for software development. But together with the systems already emerging around them, they show coding agents moving away from a single model executing an entire task and toward workflows that divide work among specialized agents, execution environments, and persistent infrastructure.

The engineering question is therefore no longer only whether an agent can write the code. It is whether the system around it can reliably decide what to do, which agent should do it, what that agent should be allowed to see and change, how to verify its work, and where a human should take control.

Those are architecture and infrastructure questions — and as coding agents move from interactive assistants toward autonomous software workflows, they may matter as much as the underlying model.

The post OpenAI and Cursor agree on agent coordinators. They disagree on who runs them. appeared first on The New Stack.

What managing 150,000 AI agents could look like for database teams

24 septembre 2026 à 15:11
Abstract 3D render of translucent orange cubes and panels scattered across a pale gray background, with bundles of glossy teal tubes curving in from the right.

The database administrator of the future will spend considerably less time administering databases.

That sounds contradictory, but AI agents are taking over that work. For decades, DBAs have handled the decidedly hands-on work of keeping databases available, performant, secure, and affordable. They provision capacity, troubleshoot slow queries, manage migrations, and step in when something inevitably goes sideways.

AI is already taking on some of that work. At the same time, it is creating a much bigger data infrastructure fleet to manage.

The result is likely to be a very different kind of DBA: one who spends less time tending individual databases and more time supervising the autonomous systems doing it for them.

Congratulations, you’re managing robots now

This shift starts with a familiar problem: more infrastructure needs managing than the people available to manage it.

Database automation is hardly new, but agents can potentially go further than the scripts and rules DBAs already rely on. Rather than automating one predetermined task, an agent can inspect what is happening, decide what needs attention, use tools to act on it, and check whether its intervention worked.

That changes the DBA’s relationship with the database. A performance problem that once required someone to dig through metrics, identify the troublesome query, and decide how to respond could increasingly be investigated by an agent before a human gets involved.

It doesn’t remove the DBA from the equation. Someone still has to decide what an agent can do, where human approval is required, and what happens when it gets something wrong. But the work moves up a layer. Instead of personally performing every operational task, DBAs start managing the systems carrying them out.

Instead of personally performing every operational task, DBAs start managing the systems carrying them out.

And before anyone gets too comfortable with that idea, the number of those systems could become enormous.

150,000 agents walk into a database…

Gartner predicts that the average global Fortune 500 company will have more than 150,000 AI agents in use by 2028, up from fewer than 15 in 2025. Only 13% of organizations currently believe they have the right governance in place to manage them.

Not every agent will need its own database, but plenty will. They will create state, retrieve data, remember previous interactions, and exchange information with other agents. Many will also behave very differently from the applications DBAs are used to supporting: spinning up quickly, sitting idle for long stretches, and suddenly becoming busy when there is work to do.

Nobody is hiring 150,000 DBAs to manage them.


That is the scale problem Yugabyte is targeting with YugabyteDB AMP, or Agentic Multitenant PostgreSQL. Rather than treating each new agent workload as another database for an administrator to provision and babysit, AMP manages databases as a fleet.

The platform packs hundreds of small Postgres workloads onto shared distributed infrastructure while keeping their databases isolated. Lifecycle operations, including provisioning, branching, scaling, migration, and teardown, can be exposed to agents through MCP. Yugabyte has also built specialized agents for setup, migration, performance tuning, and integrations.

In that model, a DBA is no longer provisioning database number 14,372. The interesting job is setting the rules for how database number 14,372 is provisioned, operated, and fine-tuned without them.

Do more with less (no, really)

Scale is only half of the problem. Someone also has to pay for all this stuff.

Agent workloads make traditional capacity planning particularly awkward because many are bursty and frequently idle. Giving every experimental agent permanently provisioned infrastructure could leave companies paying for many databases that spend much of their lives doing very little.

This is where consolidation becomes as much an economic question as an operational one.

AMP’s approach is serverless multitenancy and scale-to-zero. Multiple small workloads share the underlying distributed infrastructure, while customers pay by CPU minute and idle agents consume no compute. Resource governance can impose CPU limits on individual workloads, preventing a single overeager agent from consuming the capacity intended for its neighbors.

The human equivalent matters too. If routine setup, migrations, tuning and other database operations can increasingly be delegated, a smaller database team can potentially look after a much larger estate.

That doesn’t mean companies get to fire the DBAs and hand the keys to the robots. It means scarce database expertise can be spent on architecture, governance, and genuinely difficult problems instead of repeatedly doing the work that software can handle.

Your 2028 database problem starts now

The harder question is what to build underneath all of this when nobody really knows what the enterprise AI estate will look like in two years.

An agent that begins as an experiment today could disappear next month. Another could suddenly become a production application used across the business. Building one infrastructure stack for cheap experiments and another for serious workloads risks creating a migration problem every time an experiment succeeds.

Yugabyte bets that both ends of that journey should sit on the same foundation.

YugabyteDB AMP lets workloads start on serverless Postgres and transition to fully distributed YugabyteDB as their scale and criticality increase, without rewriting the application or migrating data to a different database platform.

Then there is the problem above the individual database: agents need to remember what happened, and not just in a silo.

That’s where Meko fits into the Yugabyte stack. Meko is an agent-native context engine designed for multi-agent AI systems. It provides persistent memory, shared knowledge, decision traces, and autidability across multiple agents, rather than leaving each agent working from its own isolated context. An agent can pick up information learned by another agent instead of retrieving it again or restarting the reasoning process.

Taken together, it delivers a single data stack for an agent’s entire lifecycle: Meko for the context shared among agents, YugabyteDB AMP for agentically managing fleets of Postgres databases, and distributed Postgres-compatible YugabyteDB for workloads that outgrow their serverless beginnings.

Of course, there’s no guarantee that 2028 will look exactly like today’s forecasts. That’s rather the point. The safest architectural bet may be one that doesn’t require you to know in advance which of today’s tiny AI experiments will become tomorrow’s critical applications.

The DBA is still critical in that world, but the job will look different. The DBA of the future may manage fewer databases directly, while taking responsibility for vastly more of them. Instead, managing the autonomous systems that do the administering.

The post What managing 150,000 AI agents could look like for database teams appeared first on The New Stack.

Query decomposition doesn’t fix context starvation — it just moves it

24 septembre 2026 à 15:00
Abstract digital art showing warped light lines surrounding a void, illustrating data compression and AI context starvation.

I have a small example that would best communicate the message I am trying to convey: say you built a chat widget for GitLab’s public documentation (the corpus we are experimenting with in this article) and one of the developers sends this kind of message:

We got an email saying our card was declined for something called “quarterly reconciliation” and I need to know what actually happens now. On top of that, I think we’ve gone over our seat count; there are more people in the group than seats we bought. Our CI has been queuing all week and I want to know whether the compute minutes we purchased last month rolled over or if we lose them. Our finance lead also needs to be the one who gets the invoices from now on, not me. And last thing, is the REST API rate limited? We’re building an internal dashboard and would rather find out now than after it breaks.

Five separate asks: the declined payment, the seat overage, compute-minute rollover, changing who receives invoices, and API rate limits. Each one is answered by a specific passage in GitLab’s public documentation, and you labeled which passage answers which before running anything, so you knew in advance exactly what a correct system needed to find.

Then you run the message through a pipeline that follows current best practice. It splits the query into five clean sub-queries, retrieves for each one independently, merges the results, drops near-duplicates, reranks the merged pool against the original message, and packs the highest-scoring passages into a 2,000-token context.

The pipeline retrieved all five correct passages, but only one of them survived into the packed context; that is one of five asks, not one of five sentences. The packer found, scored, and threw away the other four before the model ever saw them. The same message with no decomposition at all managed three out of five.

The failure has a name, and it isn’t the one you’re thinking of

I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.

The definition is deliberately about allocation rather than retrieval, because allocation is the part nobody watches. If the passage answering the fifth question was found, scored, and then squeezed out by three passages about the first question, the fifth sub-intent is starved, and every recall metric you have will report that the system worked perfectly.

“I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.”

Two failure modes already in circulation describe something different, and it’s worth separating them cleanly:

Semantic dilution happens at retrieval time: when you embed a five-part question as a single vector, you get a centroid that sits somewhere between five topics and lands close to none of them, so the evidence is never found. Decomposition fixes this, which is why it spread.

Context poisoning is about what is present, not what is missing. Wrong, stale, or adversarial content enters the window and corrupts what the model generates downstream. Poisoning is a contamination problem. Starvation is an absence problem, and policy, not accident, produces the absence.

I borrowed the word from operating systems. In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work. Every ingredient of that situation is present in a retrieval pipeline: a fixed resource, competing demands, and a policy that decides who gets served. A relevance-greedy packer is priority scheduling with no aging term, and under priority scheduling without aging, valid low-priority work waits forever.

Decomposition is the right fix to the wrong half of the problem

Split that message into five single-intent queries, and each one embeds cleanly, so per-sub-query recall climbs sharply. This is well-trodden ground. LlamaIndex ships a SubQuestionQueryEngine that breaks a complex query into sub-questions and synthesizes the responses. LangChain’s MultiQueryRetriever generates query variants and returns the unique union of what they retrieve. RAG-Fusion applies reciprocal rank fusion across the per-query result lists. The technique works, and it isn’t mine.

“In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work.”

The context window did not grow. Let me explain: after decomposition, you have n result sets competing for one fixed token budget, and something downstream has to decide the split. In most production pipelines, that something is a short, unremarkable sequence: merge the pools, drop near-duplicates, rerank the merged pool against the original query, then greedily fill until the budget closes.

That sequence is a scheduler. It has a priority function, which is the reranker score, and it has no fairness constraint of any kind. A sub-intent with three strongly-scoring passages takes three slots. A sub-intent whose single correct passage scores mid-pack takes none of them.

So the failure did not go away. It moved from the embedding, where it has a name and people watch for it, into the packer, where it has neither. It also moved somewhere with much worse instrumentation, because recall@k per sub-query is the metric decomposition usually gets validated with, and that number goes up. It goes up at the same time as coverage inside the packed context goes down. You ship on a green dashboard.

The harness

The corpus, GitLab’s public documentation: 10,000 chunks and 2.2M tokens, split on heading boundaries and capped at 480 tokens each. Sixty-one single-intent questions span nine topics, from seat management to rate limits, each labeled with the one passage that answers it. I built multi-intent queries by concatenating those questions while varying n across 2, 3, 5, and 7, randomizing the order so position doesn’t confound topic, and varying topical distance so half the queries draw everything from one topic and half span distinct ones. That produces 100 queries, 25 at each value of n, whose correct decomposition I know exactly.

Two decisions matter more than the rest:

The metric is not recall. Recall tells you what the retriever found. What I need is what survived into the packed context, per sub-intent. So I log each sub-intent twice: once for whether its correct passage reached the candidate pool, and once for whether it reached the packed context. The gap between those two numbers is the entire argument.

Every question has to be retrievable on its own before it’s allowed in. A question enters only if its correct passage ranks in the top 10 for its own isolated query, under both retriever configurations, and both scored 100% recall@10 on that test. Since each sub-query’s candidate pool is exactly its own top 10, passing that gate guarantees the correct passage sits in the pool for every decomposed arm. Any sub-intent that then fails to appear was denied by the packer rather than missed by the retriever, which removes the most obvious objection to everything below.

The core measurement uses no language model. Because queries are composed from known sub-questions, the decomposer is an oracle so that anyone can reproduce the main result with no API key.

That invites an objection, so I tested it. A real LLM decomposer, blind to n, disagreed with my ground truth on 41% of the queries, and on inspection it was right every time. Five of my sixty-one supposedly single-intent questions contain two distinct information needs. What is excess storage usage, and what happens when we go over the free limit? is two questions wearing one question mark. Adjusted for those five, agreement is 100 out of 100. That is not evidence decomposers are reliable, because my queries are joined by fixed connectives and splitting on those alone recovers n perfectly, which real messages never allow. What it caught was an error in my own labels, and that is the best argument I have for the oracle design.

Results

Every arm runs at 2,000, 4,000, and 8,000 tokens against two retriever configurations. The stronger pairs are BAAI/bge-base-en-v1.5 with BAAI/bge-reranker-base; the weaker pairs are a quantized BAAI/bge-small-en-v1.5 with Xenova/ms-marco-MiniLM-L-6-v2. Retrieval is in-memory cosine similarity over a NumPy array because, at 10,000 chunks, a vector database would be slower to write, slower to run, and harder to verify.

Sub-intent coverage at a 2,000-token budget on the stronger configuration:

Table showing coverage (in %) per arm.

The production-default pipeline starves 31.1% of sub-intents whose evidence it had already retrieved. It beats no decomposition by nine points, while a flat B/n split, which is the crudest allocator anyone could write, beats it by fourteen.

Floors work, but not the obvious floor. Reserving one passage per sub-intent before the greedy fill satisfied 99.4% of its reservations and bought only seven points. The mechanism fires correctly and reserves the wrong passage because it picks each sub-intent’s best chunk by score against the original query. The original query asks about all five intents at once. Selecting that same reservation by score against its own sub-query pushes coverage to 89.4% and cuts allocation starvation from 30.6% to 10.1%. That is a change of about four lines.

Two of my own recommendations died here. I expected reranking against the original query to beat reranking against the fragment, and it loses by seventeen points. The incomparable score scales I worried about turn out to help, because each sub-query’s best match ends up at the top of its own scale, producing per-intent fairness for free. I also expected deduplication before allocation to matter, but near-duplicates consume 1.0% of the budget and removing them moves coverage by 0.3 points.

The crossover: starvation by tokens-per-sub-intent, which is simply the budget divided by n:

Table showing tokens per sub-intent.

Below roughly 1,000 tokens per sub-intent, allocation policy dominates. Above it, nothing you do to the allocator matters, because everything fits anyway.

The retriever comparison is the one I’d lead with. Upgrading the retriever moves coverage on the production-default arm from 56.2% to 68.9%, a gain of 12.7 points. Changing the allocation policy on the same retriever moves it from 68.9% to 89.4%, a gain of 20.5 points. In its sharpest form: the weaker retriever with a fragment-scored floor reaches 84.5%, and beats the stronger retriever with a greedy packer at 68.9% by sixteen points. A worse retriever with a better allocator wins.

Position: Held within a fixed n so query difficulty doesn’t contaminate the comparison; starvation across the seven positions of an n=7 query runs 1.3%, 21.3%, 30.7%, 48.0%, 48.0%, 32.0%, and 13.3%. That is a serial-position curve. The packer protects what you asked for first, protects what you asked for last a little less, and drops the middle. The fragment-scored floor flattens it to 4.0%, 9.3%, 6.7%, 8.0%, 22.7%, 5.3%, and 6.7%.

“A worse retriever with a better allocator wins.”

Topical distance: Sub-intents that span distinct topics starve about twice as often as sub-intents drawn from one topic, at 38.1% against 18.3% for n=7. I predicted the opposite. A topically coherent query gives the reranker a coherent target, and it scores all the correct passages similarly. In contrast, a scattered query lets it latch onto some topics and abandon others.

What the user actually sees

Everything above is retrieval-side. What decides whether any of it matters is what reaches the person who wrote the message, so I generated real support replies from 80 packed contexts and had every reply graded per sub-intent, with both the generation and the grading blind to which arm produced which context.

When the correct evidence reached the packed context, the reply addressed that question 100% of the time, across 261 out of 261 cases, in both arms. Coverage predicts the generated outcome exactly, which is the strongest justification I have for measuring it.

When a sub-intent was starved, the reply answered it anyway 48.1% of the time, based on whatever else happened to be in the window. It explicitly flagged the gap 45.6% of the time, with some version of “I’ll follow up on that separately.” It went silent only 6.3% of the time.

“Starvation mostly does not produce silence; it produces unsupported answers.”

I expected silence, and I was wrong. Starvation mostly does not produce silence; it produces unsupported answers. Whether those answers are actually incorrect is the next experiment, because this harness measures whether a question was addressed, not whether the answer was right.

Some limits: composed queries are cleaner than real support messages, which carry pronouns, implicit context, and conditional clauses. This is one corpus and one embedding family. I drafted the gold labels with model assistance and verified them myself. The model writing those replies was strong, so a cheaper production model would plausibly flag fewer gaps and invent more.

What an allocator actually looks like

Give every sub-intent a floor, and choose it by fragment score. Not the naive floor, which satisfies 99% of its reservations and buys seven points. Select a reservation by relevance to the sub-intent it protects, not by relevance to the message as a whole.

Rerank against the fragment rather than the original query. This inverts what I expected and what I have seen recommended. Scores from different fragments are not comparable across sub-intents, and that incomparability is doing useful work.

Don’t spend your effort on deduplication; near-duplicates cost 1.0% of the budget here. Dedup is worth doing, but it isn’t why your fifth question went unanswered, and treating it as the fix will cost you weeks.

Log per-sub-intent coverage: You already computed it to pack, and it predicts the generated outcome perfectly. A sub-intent that received zero passages is the best predictor available that your reply is about to assert something you cannot support.

Where parallel decomposition breaks

“If it’s late can I get a refund” is one clause and two intents, and the second one’s retrieval target depends on the first one’s answer. Parallel decomposition treats them as siblings. It retrieves the late-delivery policy and the general refund policy, packs both, and misses that the passage you actually need covers refunds for late delivery, which may match neither sub-query particularly well.

There are two ways out: You can tag dependencies at decomposition time, or run a deferred second pass that re-retrieves conditional clauses once the first round resolves.

I would take dependency tagging, for three reasons: A second pass costs a full retrieval round trip inside a latency budget a support bot does not have. The tag is reusable, because a dependent sub-intent should not hold a floor reservation. At the same time, its parent is unsatisfied, so it feeds the allocator directly instead of bolting on a separate mechanism. And it fails visibly, since an untagged dependency shows up as a starved sub-intent in the coverage signal. In contrast, a deferred pass that resolves the wrong condition produces a confident wrong answer with nothing to flag it.

The cost is real; dependency tagging pushes work onto the decomposer, which is already the weakest component in the chain, and I have not measured tagged against untagged. That is a design position rather than a result, and it is the one thing here I am asking you to take on argument instead of evidence.

What to measure on Monday

Take your production pipeline and compute one number: your context budget divided by the average count of distinct questions per incoming message. If that number lands below roughly 1,000 tokens, your allocation policy costs more than your retriever does, and the reranker upgrade sitting in your backlog will buy you less than reserving one slot per question.

On my corpus, the retriever upgrade was worth 12.7 points of coverage, and the allocation change was worth 20.5, which is why I think the ordering is wrong in most pipelines I’ve seen. That ordering is the falsifiable part. Run the same two comparisons against your own corpus, and if the retriever wins, I want to see the numbers, because that result would tell me the crossover sits somewhere other than where I measured it.

The cheaper thing to do first takes an afternoon. Log, for every multi-intent request, how many sub-intents ended up with zero passages in the packed context. A support system that cannot tell you which question it dropped will keep answering that question anyway, about half the time, out of whatever else was in the window.

The post Query decomposition doesn’t fix context starvation — it just moves it appeared first on The New Stack.

Q.ANT gives away the software for its light-powered AI chips in a CUDA-style bet on developers

24 septembre 2026 à 00:23

Q.ANT, a startup out of Stuttgart, Germany, builds processors that use light instead of electricity to do some of the math behind AI. The company pitches them as a way to run AI on a fraction of the power today’s chips need.

Now developers can start writing software for those chips without owning one. Q.ANT pushed a free, open-source software kit to GitHub this week that lets developers build and test programs on a normal computer, then run them on the real chips once they get access.

This is a move out of Nvidia’s playbook. Nvidia owes its lead in AI as much to CUDA, the software developers use to program its GPUs, as it does to the chips themselves. 

But with Q.ANT, the catch is the hardware. Q.ANT’s chips are running at a few research computing centers, and everyone else has to wait “the coming months” for cloud access through German provider IONOS or an on-site server from Q.ANT.

The kit, called the Q.ANT Native Computing Toolkit, is free on GitHub under a license that allows commercial use. Developers can work in Python or C. The key piece is a simulator that mimics the chip on a regular computer, with no Q.ANT drivers required.

What can it do today? The AI tools in this first version focus on running models that have already been trained. The examples read handwritten numbers, identify objects in photos and outline shapes in images. Training still happens on regular CPUs and GPUs.

The pitch for photonic computing is power. AI chips burn a lot of energy moving data back and forth between memory and the processor. Q.ANT’s chips do part of the math with light, specifically wave-shaped functions similar to a cosine, which regular chips calculate digitally. Q.ANT says AI models built around those functions get better results with fewer parameters, the settings a model learns during training. Fewer parameters means a smaller model, less data to move and less power. The kit includes examples comparing a standard model with one built Q.ANT’s way. Those comparisons are the company’s own.

“An ecosystem isn’t created by hardware alone. It emerges when the software layer is open and others can build on it,” said Michael Förtsch, Q.ANT’s founder and CEO. He calls the release the “Linux moment” of photonic computing.

Q.ANT is betting light can do the math itself. Lightmatter, one of the best-known companies in the field, now puts its focus on Passage, which uses light to move data between chips. The idea of light-based AI isn’t new, either. TNS covered MIT’s photonic processor for building optical neural networks back in 2017.

Q.ANT raised €62 million in July 2025 in a round led by Cherry Ventures, UVC Partners and imec.xpand. In March, it said its second-generation chips were running at the Leibniz Supercomputing Centre near Munich. The results it published from there compare the new chip with its old one: more than 50 times faster at the kind of math that does most of the work in AI models, and six times less energy on typical jobs, by the company’s numbers. Its bigger claims, like up to 30 times better energy efficiency, don’t say what they’re measured against.

Good software alone won’t carry a new chip. Nvidia has been building CUDA for nearly 20 years and is still adding to it, including deeper native Python support last year. Graphcore, the British AI chip startup, had its own software kit and still ended up being sold to SoftBank in 2024.

Q.ANT calls this the first openly available software kit for programming a photonic processor. That depends on how you count. Xanadu has offered free, open software for its light-based quantum computers since 2018. For now, developers can play with the simulator. What they can’t do yet is test Q.ANT’s power-saving claims on their own models. That has to wait until the chips open up.

The post Q.ANT gives away the software for its light-powered AI chips in a CUDA-style bet on developers appeared first on The New Stack.

A third option is emerging in the fight over AI and your data

23 septembre 2026 à 21:31
Split-screen video interview with The New Stack host Alex Wilhelm and VAST Data cofounder Jeff Denworth.

Not your keys, not your coins. Not your model, not your data?

Over the summer, the tech industry was consumed by a debate about AI use in the enterprise and the need to protect IP. If an enterprise used proprietary models, was data leakage a necessary evil?

Companies seemed to have two options: They could use state-of-the-art, proprietary models and risk losing control of their data, or they could use open-weight models and never kiss the frontier.

Thankfully, a third option is emerging.

Consider the concern: Company A wants to use LLM B from AI Lab C, and they want to avoid training AI Lab C how to eat Company A’s lunch by building its capabilities into LLM B. A good way to resolve the tension would be to let Company A run LLM B on its own infrastructure, so there’s no risk of its information fleeing on the wind.

AI agents are “creating a whole different set of requirements at the data layer.”
–Vast Data co-founder Jeff Denworth

But that raises another problem: AI Lab C doesn’t want to allow Company A to run LLM B on its own GPUs because it doesn’t want to hand over its model weights. It’s the same IP issue the company ran into, in reverse. You have to solve the trust problem in both directions!

Enter VAST Data co-founder Jeff Denworth and a new product called DataEnclave, which aims to let AI labs and enterprise-scale companies deploy proprietary models in secure compute environments without risking data transfer in either direction. (DataEnclave uses Nvidia’s Confidential Computing technology to make the system tick; Vast Data’s core product is AI OS, infrastructure that fits beneath a company’s AI applications.) 

The New Stack had Denworth on the podcast to chat about the confidential computing market. I was curious about timing. Why did Vast build DataEnclave now? Nvidia began rolling out Confidential Computing in a serious way in 2024, after all. Denworth argues that the market needed the core technology, yes, but also demand.

And until late 2025, AI demand was modest compared to today’s token totals. Once agentic coding tools took off, corporate demand for AI products soared. This led to the pricing crisis we saw in early 2026, and the secure AI usage debate we endured over the summer. 

Performance drove demand, demand drove usage, and usage dug up fresh problems to solve. Now the question for the market is whether or not DataEnclave has solved enough concerns on both sides of the proprietary AI-proprietary data equation. The market will sort that out as it moves through early access and into general availability.

Our conversation goes deep into the arc of AI, where companies are in their AI journey today, and how much data remains to be unlocked inside the enterprise. If you want to feel the acceleration, it’s a fun one!

The post A third option is emerging in the fight over AI and your data appeared first on The New Stack.

How confidential AI splits control between data and model owners — and opens new opportunities for both

23 septembre 2026 à 18:15
Abstract 3D render of dark blue and black cubes floating among translucent spheres against a warm red and coral background.

Most people already understand what generative AI can do. But enterprises run into problems when they need to give a model access to information that cannot leave their own environment, such as a patient record, a customer’s financial details, or a company’s most valuable intellectual property.

Sending that data to a cloud or SaaS service means it crosses external networks and is processed on infrastructure run by another organization, creating additional concerns about control, accountability, and exposure. That’s where AI enthusiasm collides with production realities. Despite its productivity potential, enterprise AI still faces a fundamental gap in trust and control.

Organizations need to know whether a system will expose information it should protect, act as intended, meet security and performance requirements, and behave safely at machine speed.

Alon Horev, CTO and co-founder of AI operating system company VAST Data, tells The New Stack that the challenge is particularly acute when AI systems handle sensitive customer information. “Even if you ask the model today to obfuscate a conversation or redact PII from a conversation, it’s hard to have 100% confidence that’s the case, and that it worked.”

“Even if you ask the model today to obfuscate a conversation or redact PII from a conversation, it’s hard to have 100% confidence that’s the case, and that it worked.”

Consider a customer support agent that needs access to an individual’s profile to provide a useful, personalized answer. The organization must ensure that information isn’t exposed to another customer, while also considering whether those conversations can be used for training or system improvement. They might contain personally identifiable information (PII) or other protected details, and the consequences of mishandling them ultimately fall on the organization and the people whose information it holds.

Confidential AI architectures: solving a two-sided trust problem

Enterprise AI has two parties to satisfy: organizations must keep sensitive data under their control, while model builders need to protect the weights and software that represent substantial investments in research, engineering, and IP. They’re understandably reluctant to place those assets in environments where customers, infrastructure operators, or attackers might gain access. That mutual need for control has created a stalemate. How can organizations bring advanced models to sensitive data without asking either side to surrender control?

Horev has seen that the most capable models are increasingly delivered as SaaS services, because that’s the simplest way for their creators to distribute and protect them. Even when a provider offers compliance controls, the enterprise might still shoulder the consequences of a breach, misuse, or regulatory violation. Sending information across the WAN also places it in the hands of more systems, connections, and operators, increasing the number of points that must be trusted and governed. Organizations may also be unwilling, or legally unable, to rely on a provider’s assurances that it will not retain, reuse, or expose their data beyond the intended service. 

For organizations in regulated or data sovereignty-sensitive sectors, that could be an unacceptable trade-off. “Naturally, many organizations are adopting a hybrid strategy,” Horev tells The New Stack. “Some applications and datasets can go to the cloud, while others must remain on-premises, sometimes even in the building, or in the country.”

This is where confidential AI comes in. Encryption at rest and in transit protects data while it’s stored or moving between systems. Confidential computing extends that protection into the processing environment, using hardware-isolated execution to create a protected enclave in which the data and model weights can remain encrypted until they’re released to an approved workload.

Cryptographic attestation verifies the hardware, virtual machine (VM), software, and configuration requesting access before releasing keys. The model builder can encrypt its model using the public key of a specific confidential VM. Only that VM’s corresponding private key can decrypt it within protected memory, enabling the customer to use the model without accessing its weights.

Independent key control preserves the separation between the two sides. The enterprise retains control of the keys governing its data, while the model builder retains control of the keys governing its model. While the workload is running, the infrastructure operator doesn’t control either set of keys.

As AI becomes more agentic, those controls will matter more. Agents will need to access more data, systems and tools, and might act on that information with far less human intervention.

“This world of agentic AI is moving extremely fast, and we need to limit what an agent can see and do.”

Those that can’t establish strong privacy and governance assurances for today’s models will find it even harder to deploy agents safely in the future. “This world of agentic AI is moving extremely fast, and we need to limit what an agent can see and do,” Horev tells The New Stack.

From architecture to ecosystem

Many businesses simply cannot manage the integration, security, and maintenance of the entire AI stack, because it requires working separately with each model provider to engineer something that suits both parties. Turning confidential AI architecture into something organizations can deploy is the challenge VAST DataEnclave, which was launched on September 22, intends to address.

As a capability of the VAST AI Operating System, the goal is to bring the model, application layer, and data platform together under customer-controlled operating conditions. The architecture is designed to protect both sides of the equation: the enterprise’s data and the model builder’s weights. The customer retains control of its infrastructure and data keys, while the model provider can make its software available without handing over the underlying intellectual property.

“We’re trying to close the trust and control gap by working with world-class model builders such as Cohere, Deepgram, Factory, Fundamental and TwelveLabs, who continue to innovate and build their expertise,” says Horev. The ecosystem also includes infrastructure and security providers such as Nvidia, CrowdStrike, Fortanix, Nscale, Cisco, and Supermicro. The range reflects the practical challenge: confidential AI needs more than a protected GPU. It requires models, applications, accelerated hardware, data infrastructure, and operational support to work together.

That control also changes the cost conversation, without automatically making AI cheaper. Hosted models can make budgets harder to predict as token consumption varies with usage patterns, agent loops, model architecture, and workload volume. Customer-controlled infrastructure gives enterprises a more defined capacity and cost base: they can plan around GPU clusters they own or have already budgeted for, instead of allowing inefficient model choices or uncontrolled agent activity to generate an open-ended token bill.

“…instead of allowing inefficient model choices or uncontrolled agent activity to generate an open-ended token bill.”

The cluster also imposes a natural ceiling on throughput, which helps organizations understand how much work their infrastructure can handle within a given period. Model providers can then price access by token, task, or license, while the enterprise retains greater visibility into its total operating cost.

Why the data platform is paramount

Confidential AI protects data and model weights during inference, but it’s only part of the production challenge. Real-world AI systems are living environments in which data moves between storage, databases, GPUs, networks, applications, and agents.

That’s why confidential AI can’t be bolted onto a fragmented stack. Businesses need to protect the model, the data, and the infrastructure connecting them as one system. As Horev says: “You need to build security in multiple layers of the platform,” with someone accountable for rapidly updating compromised components.

Confidentiality is only useful if the resulting system can also be operated, monitored, and improved. As AI infrastructure becomes more distributed, it becomes harder to tell what’s happening when something goes wrong and where the fault lies.

Horev recommends a “single pane of glass” across storage, networking, and compute, so teams can see what’s happening and keep resolution times low. If a network port is intermittently failing in a data center, for example, an agent could help identify the root cause, provided it has access to the right operational data and tightly controlled permissions. Those permissions should govern the infrastructure it can inspect, the data it can retrieve, and the actions it can take.

The same applies to monitoring AI workloads. Teams need visibility into performance, failures, and access patterns without exposing the customer data or model weights. Agent sandboxes can limit the systems and tools an agent can reach, while data platform observability can log which data it accessed, what it did with that data, and how it interacted with downstream systems.

Evaluation, therefore, becomes part of production discipline. Teams must observe systems, measure behavior, govern access, and manage change in ways that demonstrate progress. Confidentiality, data-level policy, observability and correctness have to work together.

The emerging ecosystem suggests demand for models that can run securely under customer control, wherever sensitive data resides. These are “living systems,” says Horev. “It’s not just leveraging a feature inside of a wider platform.” 

Visit the VAST Data Confidential AI solution page to learn more about the architecture, ecosystem, and availability.

The post How confidential AI splits control between data and model owners — and opens new opportunities for both appeared first on The New Stack.

Claude Opus 5.5 wants to finish your coding tasks, not just start them

22 septembre 2026 à 21:49

Anthropic wants developers to use Claude and its family of tools to handle complete coding tasks. The company introduced Claude Opus 5.5 on Tuesday to span more of the software application development lifecycle: from design specification creation through debugging to code generation and testing.

As the first release in a new family of Claude 5.5 models, Anthropic says that Claude Opus 5.5 performs at the level of Claude Fable 5.1 on “most work” tasks and costs around 40% less to run than Opus 5. 

GitHub chief product officer Mario Rodriguez is quoted in Anthropic’s launch announcement on exactly where software engineers sit today with frontier models for code automation. He said that developers want agents.

“In our testing [of Claude Opus 5.5] across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it’s making developers’ bigger projects more achievable,” says Rodriguez.

Where does Claude Opus 5.5 get its power from?

Anthropic explains that the Claude 5.5 family’s expanded full-lifecycle capabilities were developed under an established set of practices.

These practice elements include extensive alignment testing (model evaluation processes put in place to make sure actions and outputs closely match human values and intended goals), pre-release evaluation by outside organizations, and safeguards for high-risk areas such as cybersecurity and biology.

The organization says that on its most comprehensive alignment test, Opus 5.5 is the strongest-performing model tested to date, with particular improvements in several behaviors that contributed to recent cybersecurity incidents (e.g., biased reasoning, attempting to escape a sandbox, and others).

Independent SRE & AI reliability architect Akash Thakur tells The New Stack that Anthropic’s work getting its model to complete whole coding tasks is impressive, but “getting it to know when it hasn’t” is the harder problem developers need to think about.

“Models working at the level of Claude Opus 5.5 are genuinely good at breaking a project into small, finishable pieces — and that’s the real unlock, because momentum on software comes from finishing things, not starting them,” Thakur says. “…But ‘completed’ and ‘correct’ aren’t the same thing. The task that looks done and completed is exactly the one that often costs the team later, so the win isn’t removing the human — it’s moving them from writing the code to verifying it.”

“Models working at the level of Claude Opus 5.5 are genuinely good at breaking a project into small, finishable pieces …but ‘completed’ and ‘correct’ aren’t the same thing.”

Suggesting that we are now witnessing a “higher bar for more capable models”, Anthropic said that models that could fully automate AI research itself should meet a higher safety standard. The company noted that as AI becomes more capable, public policy should play a larger role in making sure these systems are safe. Its recent work with Accenture is offered as an example of how Anthropic is building the infrastructure to support this.

Sprawling jobs: codebase-wide migrations & audits

Getting more specific, Anthropic has claimed Opus 5.5 is “particularly good” at long, sprawling jobs like codebase-wide migrations and audits. An early tester said it audited and fixed a 200,000-line codebase in under three hours, while Opus 5 took over 20 hours and used 2.5x as many tokens. 

“In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less,” said Anthropic in a press statement.

Field CTO for the EMEA region at Coder, Eric Paulsen, tells The New Stack that he is happy to hear Claude Opus 5.5 is becoming more holistically capable, but he’s “not surprised” because AI is “eating the software delivery chain end to end” today.

“Despite the efficiencies showcased in Opus 5.5, the moment an agent can reliably finish real work, the question stops being whether the model is good enough and becomes where you’re letting it run,” Paulsen says. 

“Kicking off a Claude Code session on a laptop that can be compromised, stolen, or simply run out of compute is not an engineering environment; capable agents need dedicated, governed infrastructure with the same guardrails, secrets handling, and observability a developer would demand of any other production workload. As we’ve seen here with Anthropic’s own work, firms need to underline testing and safeguard procedures for launches of this kind,” he added. 

Anthropic assures users that external testing has been conducted and that the model was tested before release by Frontier Design and METR. When established safeguards for Opus 5.5 intervene, requests “fall back to another model transparently,” meaning developers may not see which model actually handled a given call.

In cybersecurity, users can identify and fix bugs in their code, but most cybersecurity tasks will be re-routed to Opus 4.8. Requests flagged by biology and frontier LLM development classifiers will be re-routed to Opus 5. Vetted organizations can apply to Anthropic’s Life Sciences Verification Program to use Opus 5.5 for biology research, and the company says it will expand its Cyber Verification Program in the coming weeks.

The developer’s terminal prompt has fundamentally changed

HasData co-founder, Sergey Ermakovich, tells The New Stack that his work as a web scraping specialist for data pipelines and AI means he sees zen-like, one-brick-at-a-time logic in what Anthropic has done. 

“The biggest change — and it’s a trend that will have driven Anthropic’s design and development aspirations for Claude Opus 5.5 — is that the terminal is no longer just a place where a developer pastes generated code,” Ermakovich says. “The terminal now becomes part of the model’s workspace.”

Because a model can run commands, inspect failures, modify files, and verify the result, Ermakovich suggests it can “close the loop” instead of handing unfinished work back to an engineer. “That makes small complete tasks much more valuable. Fix one failing test, update one dependency, migrate one endpoint, verify it, then move to the next task,” he adds.

“The terminal is no longer just a place where a developer pastes generated code — the terminal now becomes part of the model’s workspace.”

Commenting on the development of this model as part of Anthropic’s approved corporate messaging, John Ruelas, staff software engineer at Ramp, said that verbose, hard-to-follow output has been his “biggest frustration” with frontier models, but “Claude Opus 5.5 fixes it” for him.

He noted that it writes like a good colleague and follows his company’s writing rules. A design spec came out usable with very minimal edits, and when it rewrote one of his prompts, he preferred its version to his own. When it optimized the team’s test suite, he could follow its reasoning easily and “shipped the change with confidence.”

We’re so done with the autocomplete era

Vice president of AI strategy at Abbyy, Maxime Vermeir, tells The New Stack that the more lifecycle-wide Claude Opus 5.5 features on offer show that “we’re done with the autocomplete era” for basic code automation tools.

“Anthropic’s elevation here reflects the fact that it used to be thought of as marvelous if AI could complete a developer’s next line of code, but today the expectation is that you hand it a whole Jira ticket and that it gets done,” Vermeir says. “But despite the power on show with Claude Opus 5.5, the question remains as to how much these newer models will actually understand what ‘done’ means, as often it seems they have a desire to keep burning tokens by offering you yet another thing it didn’t quite do right.”

Opus 5.5 also “communicates more naturally” than prior models, with early testers finding its writing clearer and easier to follow, making it a better work partner over long sessions.

Costed out lower than Opus 5, Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens (20% less than Opus 5), and Anthropic has cut cache read prices by 60% for token-billed usage. Opus 5.5 also needs fewer tokens for higher-quality work and generates output more than 30% faster than Opus 5.

The post Claude Opus 5.5 wants to finish your coding tasks, not just start them appeared first on The New Stack.

Your AI agent is burning tokens on choices that don’t need words

21 septembre 2026 à 21:13
Blur or abstract motion

AI agents spend a ridiculous amount of compute generating text nobody actually needs. The decisions an agent makes along the way don’t require a written answer and yet, agents still send them to generative models, wait for an answer while burning through tokens and then parse that output back. The overhead is already drawing scrutiny — OpenAI’s own researchers recently disclosed spending $7,000 a day running agent workloads.

Kev, a new family of open decision models built on Qwen 3.5, takes a different approach and skips the generation entirely.

Developer Jared Palmer released a new generation of Kev on Sunday, with 0.8 billion, 4 billion, and 9 billion parameter models built on Qwen 3.5. Kev is prefill-only, processing the state, questions, and candidates in a single forward pass before reading the decisions from a pointer head without an autoregressive decoding loop.

Kev, a new family of open decision models built on Qwen 3.5, takes a different approach and skips the generation entirely.

Decisions without generated text

Kev supports three decision types: Noul for yes/no, Choice for selecting among candidates, and Score for ordered levels, mirroring TypeSafe’s System One API. Developers provide the state and questions, and the pointer head returns probabilities across the available candidates.

For a tool-routing decision, the output could look like this:

search: 0.82

database: 0.13

calculator: 0.05

Kev can still choose the wrong tool, but because it scores only the candidates it’s given, it can’t introduce an option that isn’t on the list.

Routing, safety checks, escalation, and ranking can then move to the decision layer, leaving larger reasoning models to handle the open-ended work.

Kev can still choose the wrong tool, but because it scores only the candidates it’s given, it can’t introduce an option that isn’t on the list.

Batching choices, one pass

Multiple decisions can also be made against the same state in a single forward pass, with a block-causal attention mask isolating the questions while the pointer head scores each set of candidates independently.

Palmer’s documentation shows the 4B model processing three questions in 277 milliseconds in bf16 on an M5, although without a controlled comparison against Qwen generating equivalent answers on the same hardware, the result doesn’t establish how much faster the approach is in practice.

The ability to evaluate several decisions against the same context could become more useful as agent loops grow more complex, but skipping generation doesn’t make the resulting decisions inherently better.

Calibration limits and tradeoffs

The largest model, Kev-9B, reached 83.7% accuracy on the project’s locked out-of-domain test, according to Palmer’s model card. That’s a developer-reported benchmark, and Palmer documents some limitations alongside it.

The probabilities Kev returns don’t always reflect how confident developers should be in the result. Palmer found that temperature calibration can drift on unseen source distributions, a problem for agents that use probability thresholds to decide whether to execute an action or escalate it, since even a high-probability choice can still be wrong.

Fine-tuning also changes some of the capabilities inherited from the underlying model. Palmer’s evaluations show declines on general-knowledge and arithmetic tests, particularly among the smaller models. That’s consistent with Kev’s more specialized role alongside a general-purpose model, although its performance in dynamic agent environments will also depend on how well it handles tools, choices, and labels it never encountered during training — and debugging agent failures often points to infrastructure rather than the model itself.

The approach predates Kev. TypeSafe introduced Jev earlier this month as part of its System One platform, using the same Noul, Choice and Score primitives, and Kev implements its /v1/systemone request and response format so applications built against the API can point to a local Kev server instead.

Open weights, open training

The biggest difference is that Palmer released Kev under Apache 2.0 with the model weights, training code, and evaluation tooling, giving developers the option to run and train it on their own infrastructure. Jev’s weights and training data aren’t public, however, which makes direct performance comparisons difficult because differences between the models can’t be isolated to architecture, size, or training.

For applications that make only a handful of bounded decisions, constrained decoding on a model that’s already running may be simpler than adding another model to the stack. Agent loops can make those decisions constantly, however, moving through routing, ranking, safety checks, tool selection, and escalation before generating much user-facing text. It’s a pattern showing up across model architectures — stripping out unnecessary computation when the task doesn’t require it.

When those steps only require a choice or probability, Kev can handle the decision directly while leaving open-ended reasoning and final responses to the larger generative model.

When those steps only require a choice or probability, Kev can handle the decision directly while leaving open-ended reasoning and final responses to the larger generative model.

The post Your AI agent is burning tokens on choices that don’t need words appeared first on The New Stack.

Kubernetes can run AI inference. But can it count the real cost?

18 septembre 2026 à 17:53
Abstract 3D illustration of interconnected purple geometric nodes and gold lines representing a distributed network or cloud infrastructure.

Welcome to another edition of Road to KubeCon, where we’re tracking the Kubernetes and cloud-native ecosystem on the way into KubeCon + Cloud Native Con NA 2026, to be held in Salt Lake City, Utah, November 9-12.

This week, we look back at the past week of significant movements in the Kubernetes space. Most notably, we see interesting advances in cloud-native architectures for AI inference. We take a look at that, plus a new Gartner quadrant, new Kubernetes hardening updates, and important CNCF project updates.

HPE challenges server virtualization platforms

On Monday, Gartner published its Magic Quadrant for Server Virtualization Platforms, a guide comparing solution providers in the server virtualization market. The quadrant names HPE as a Challenger based on Ability to Execute and Completeness of Vision.

Hewlett Packard Enterprise (HPE) is a presenting sponsor of Road to KubeCon. HPE Software helps IT organizations modernize infrastructure, streamline operations, and accelerate AI initiatives across hybrid, multi-vendor environments.

According to the HPE newsroom, the recognition reflects ongoing momentum behind HPE Morpheus Software, its virtualization and cloud operations portfolio. HPE was positioned in the Challengers quadrant alongside Canonical and Oracle.

The announcement comes as enterprises rethink their virtualization strategies. Rather than simply swapping in another hypervisor, HPE argues that organizations increasingly need unified governance and ways to provision, orchestrate, observe and secure workloads — including VMs, containers and AI workloads — across clouds.

Kubernetes hardens container storage

On Wednesday, Red Hat’s Nispriha Jagan and Neeraj Krishna wrote on the Kubernetes project blog about two new storage security features shipped as Alpha in Kubernetes v1.37, which included 67 enhancements.

The notable security features are new bind mount options and emptyDir permissions. The enhancements come as multiple security findings have surfaced regarding emptyDir volumes, one of the most common writable volume types. The additions are made possible by low-level Linux security mechanisms.

According to the authors, these enhancements give users native controls to harden Kubernetes workload storage better. “Supporting noexec, nodev, and nosuid gives users a native way to harden volume mounts to match security benchmarks and policy,” the authors write.

KubeCon adds AI Inference + Agentic track

Last month, CNCF announced it will feature an AI Inference + Agentic track at KubeCon + CloudNativeCon North America 2026, exploring the intersection of generative AI and cloud native infrastructure. Attendees can explore the sessions here.

The added track underscores the growing use of Kubernetes for production AI workloads, particularly as the focus shifts from training models to serving them in production. It also reflects emerging practices for building agentic systems around protocols like MCP and A2A, as well as infrastructure such as AI gateways.

China Merchants Bank unifies AI inference on Kubernetes

China Merchants Bank, a leading Chinese commercial bank, recently showcased its cloud-native AI infrastructure at a CNCF event in China. Its infrastructure team won the CNCF End User Case Study Contest with an architecture combining Kubernetes with several cloud native projects:

  • Kueue, for job queueing and quotas,
  • KEDA, for event-based auto-scaling,
  • Prometheus, for systems monitoring and metrics,
  • HAMi, for sharing accelerator capacity across Kubernetes workloads,
  • and Fluid, for accelerating access to datasets.

The bank has a large pool of nearly 10,000 accelerator cards used for AI computation. These are heterogeneous, meaning they are not all the same type or configuration.

According to the CNCF announcement, the architecture unified management of 99% of its AI compute resources, while increasing average utilization from 35% to more than 60%. It also cut the cost of processing 1 million tokens by 60% under comparable conditions.

The case study shows how cloud-native infrastructure can improve utilization and efficiency for AI training and inference, even in regulated areas like financial services.

Industry take: Can AI inference on Kubernetes handle token cost issues?

Interest in AI inference on cloud native infrastructure is palpable. However, this week Val Bercovici, chief AI officer at WEKA, an AI-native data platform, questions whether Kubernetes’ existing resource model fits the changing economics of large-scale AI inference.

Bercovici tells The New Stack: “With AI inference, it’s cost per token, and that cost depends on state Kubernetes was never designed to manage: request mix, KV cache occupancy, the balance of prefill and decode, and how memory and bandwidth are consumed inside the accelerator after a pod is already running.”

“My view is that Kubernetes doesn’t go away,” Bercovici says. “But unless its resource model evolves, it becomes a tax on inference economics.” 

He foresees a new scheduling and memory layer to emerge around Kubernetes that can compute what a token actually costs to serve. Then platforms could make more informed, cost-based decisions about how inference workloads are scheduled and served.

As Kubernetes evolves, so do the demands on the teams running it. Presenting sponsor HPE helps teams address that complexity with software spanning virtualization, cloud management, observability, and automation.

Move over, platform engineering. Hey, agentic engineering.

A new Weave Intelligence report, State of AI in Platform Engineering Volume 2, authored by Sam Barlien, Luca Galante, and Florian Lipp, surveyed 242 platform engineering leaders on the before-and-after effects of introducing agentic AI into platform engineering.

38% of teams are shipping at least twice as much as before AI. When assessing ROI across the software delivery life cycle, 20% report efficiency gains and 11% report operational savings. Yet only 8% report a transformative, structural shift. Meanwhile, 29% are still prototyping without realized gains, with some outliers reporting negative results.

The biggest roadblock to scaling AI usage? A lack of platform readiness, including APIs, deterministic pathways, and standardization. Weave’s takeaway is that platform engineering must increasingly account for AI readiness and agentic experience as agents become another key platform consumer.

OpenTelemetry Kubernetes attributes processor reaches v1.0.0

On Wednesday, OpenTelemetry, the graduated CNCF project and open standard for telemetry, announced the v1.0.0 release and distribution of its Kubernetes attributes processor. It’s a helpful feature that uses the Kubernetes API to add Kubernetes metadata, such as stability, distributions, warnings, issues, and other metrics, to resource attributes.

According to the release notes, written by Elastic’s Christos Markou and Datadog’s Pablo Baeyens, the feature has been in progress in the OpenTelemetry Collector SIG since late 2025, based on a roadmap of users’ most-requested features. Existing attribute processors should review the breaking changes and migration guide.

DigitalOcean opens Spot GPU node pools

Technically, this occurred the week before last, but didn’t make the digest. As of September 9, DigitalOcean Kubernetes’ (DOKS) Spot GPU Node Pools entered public preview. According to the release notes, the feature runs worker nodes on interruptible GPU capacity at a “lower, variable rate than on-demand GPU nodes.” This could offer a cost-effective option for fault-tolerant workloads.

Other updates from the K8s universe

More updates from the infrastructure-heads, platform engineers, and multi-cloud operators working in the Kubernetes ecosystem: 

Follow the Road to KubeCon

Road to KubeCon is an eight-part series presented by HPE, which will be at KubeCon + CloudNativeCon North America in Salt Lake City. Before you go, explore how HPE Software helps IT teams do more with less complexity.

We’ll be here every Friday until KubeCon.

If you’d like to participate, Bill Doerrfeld, the writer of this series, is open to pitches — you can send release notes, quotes, reports, videos, case studies, or hot takes through his contact page.

If you didn’t catch the inaugural edition covering the Kubernetes v1.37 release, check it out here. You can also visit the Road to KubeCon page for the complete archive.

The post Kubernetes can run AI inference. But can it count the real cost? appeared first on The New Stack.

Your agent is only as good as your infrastructure

18 septembre 2026 à 15:00
Dark abstract 3D glass ribbon rendering symbolizing complex AI agent infrastructure and bursty data workflows.

You built a great agent, but something happened when it moved into production.

In testing, your agent reviewed pull requests efficiently on its own. It read the diff, grepped the codebase for related usages, ran the test suite, checked whether CI was still red from an earlier commit, and drafted a comment—all before you’d finished reading the diff yourself.

In production, however, imagine the same five steps ran behind every other PR review your team’s agents performed that hour. Some reviews landed in seconds; others sat for minutes because the test-suite step landed on a node that was mid-burst from someone else’s agent. 

The agent didn’t change. The execution environment did, and that’s what decided whether review time held steady or crept up.

Agentic applications introduce a different execution pattern than traditional chat applications. As those workflows become longer and more dynamic, infrastructure has a much larger influence on latency, reliability, and cost than it does for a simple chatbot.

That’s why your agent is only as good as your infrastructure.

Agents aren’t chatbots with more steps

The difference between serving inference for a chatbot vs. an AI agent isn’t simply that one is “more capable.” They execute work differently:

  • A chatbot usually makes one inference call per user message. The model receives a prompt, generates a response, and waits for the next user input before proceeding.
  • An agent executes the entire workflow, turning one user message into a chain of inference calls. It might decide to search documentation, retrieve data from a database, call an API, execute code, evaluate the result, and then repeat that process before producing an answer. Each of those decisions may trigger another inference call, and every result becomes additional context for the next step.

“Infrastructure has a much larger influence on latency, reliability, and cost than it does for a simple chatbot.”

That execution model changes the infrastructure requirements for AI agents. 

One question, many steps behind it

Instead of optimizing for individual inference requests, the system has to support long-running workflows whose latency and reliability depend on every component in the chain.

A single user request often expands into a sequence of inference and tool execution steps, sometimes called multi-turn tool calls or the agentic loop. Rather than generating one response, the model alternates between reasoning and interacting with external systems.

For example, you ask an agent why checkout latency spiked overnight. The agent pulls the deploy log, queries the monitoring system, runs a diagnostic against the connection pool, weighs whether the culprit is a bad deploy or a capacity issue, and then folds that into another inference call before finally producing a full-fledged response.

Each reasoning step becomes another inference request, and every tool result is added to the model’s context before the next step.

This workflow changes what reliability means

Multi-turn workflows are inherently sequential, which is why even low latency can quickly add up to a significant amount. Every inference step waits for the previous one to finish. If a database query takes two seconds, the model can’t begin the next reasoning step until that result returns. The model may generate tokens quickly, but the other steps slow it down.

“In this agentic workflow, every step in the chain has to hold, because the chain is only as strong as its slowest link.”

In this agentic workflow, every step in the chain has to hold, because the chain is only as strong as its slowest link. Instead of processing isolated inference requests, the inference stack has to orchestrate a chain of dependent model invocations and external tool calls. As those workflows become longer, the stack increasingly determines how quickly, reliably, and cost-effectively the application performs.

But the user doesn’t see an orchestration hiccup. They see an agent that hung or gave up.

That’s why end-to-end agent reliability depends on much more than model quality. Infrastructure determines whether each step has the resources it needs to execute predictably under load.

Why the bill and the performance both feel unpredictable

A second difference in agentic workflows catches teams off guard: demand patterns and their impact on your inference bill. 

Most inference services, and the pricing built on top of them, assume traffic arrives at a predictable pace. A typical inference solution knows the predictable demand pattern: User traffic increases, request volume increases, and capacity scales accordingly. Cloud infrastructure is typically optimized for these steady request patterns using mechanisms such as autoscaling, load balancing, and capacity planning.

Agent workloads don’t behave that way. Individual workflows pause while waiting on external systems, then resume as soon as new information becomes available.

  • The pause: The agent waits on an external API or database, so the GPU serving that workflow has no inference work to perform, and its accumulated context may be evicted from GPU memory while it waits
  • The burst: As soon as external systems return results, inference resumes simultaneously across many workflows, creating short and sharp spikes in GPU demand, each re-processing its full accumulated context

If you’re watching GPU utilization and it looks less like steady load and more like a heartbeat — flat, then a spike every time tool results come back — that’s the signature. It means you’re provisioning for the average when you should be provisioning for the peak, and it’s usually the first place p99 latency quietly blows out.

“If you’re watching GPU utilization and it looks less like steady load and more like a heartbeat, that’s the signature.”

Inference services designed around steady or predictable request streams can struggle to allocate resources efficiently under these conditions. This leads to inconsistent latency, GPU underutilization, or higher operating costs.

When performance becomes unpredictable, or a bill doesn’t match what you expected, it’s evidence your infrastructure was built for a different workload than the one you’re actually running.

What infrastructure built for agents actually looks like

Agentic applications place different demands on infrastructure than traditional AI workloads: long dependency chains, bursty demand, and continuous evolution. Because of this, the infrastructure is deciding whether the chain holds, and whether the bill holds too.

For an agent to run, it needs infrastructure purpose-built to support:

Performance that holds across the whole chain. The infrastructure must keep latency consistent across multi-step and multi-tool workflows.

Scalability that responds to bursty demand. Infrastructure should scale quickly as inference demand fluctuates, without requiring capacity to remain provisioned during idle periods.

  • Predictable economics even for dynamic workloads. The infrastructure bill should reflect actual usage.
  • None of this makes bursty demand disappear, but it changes how the system absorbs it. A large enough simultaneous burst, or a workflow that accumulates enough context before pausing, still costs something. The goal isn’t zero cost or zero limit; it’s making both predictable.

Your agent is only as good as your infrastructure. Get it right, and your agent’s responsiveness, reliability, and cost-effectiveness will improve your work. 

Learn more about how infrastructure can be purpose-built for agentic workflows: Check out the documentation to get started.

The post Your agent is only as good as your infrastructure appeared first on The New Stack.

Open-weight models now handle a majority of tokens on Vercel’s AI Gateway. But Anthropic still takes 64% of the spend.

18 septembre 2026 à 14:42
Isometric illustration of a retro-style computer monitor

The trend is clear: open-weight models are taking an increasingly large bite out of production AI usage.

On Monday, The New Stack reported that open-weight models accounted for 60% of OpenRouter’s US token consumption in August, with Chinese-developed models making up the majority of that volume. The latest data point hails from Vercel, whose AI Gateway routes tens of trillions of tokens each month across the applications running on its infrastructure.

As per Vercel’s September report, published on Thursday and covering activity through August, open-weight models handled 56% of all tokens routed through the gateway, the first time they have accounted for a majority of monthly token volume. In December 2025, their share was just 7%; by April it had reached 13%, and it rose every month thereafter. Vercel’s previous report, published in August, put July’s open-weight share at 36%.

So the pattern was already clear. But last month, Vercel CEO Guillermo Rauch took to social media to declare that August 22 had been a “record day for open weight share of tokens on Vercel AI Gateway,” accounting for 62% of traffic.

Rauch saw the milestone as just an early indication of where usage is heading, with enterprises still early on the adoption front.

“This is very likely just the start, because enterprise adoption is still early.”

“This is very likely just the start, because enterprise adoption is still early, and harnesses, CLIs, IDEs, SDKs, etc need to be adapted to be model agnostic,” Rauch wrote at the time.

Open-weight token share on AI Gateway: December '25 to August '26
Open-weight token share on AI Gateway: December ’25 to August ’26 (Credit: Vercel)

Tokens and dollars: Anthropic dominates spend

For context, Vercel launched AI Gateway last year as a way for developers to access models from multiple providers through a single interface, saving them from having to manage separate API keys, accounts and rate limits. The service sits between applications and the underlying model providers, routing requests while tracking usage and costs — giving Vercel a useful vantage point into which models its customers are actually running in production.

Token volume, in this context, is essentially a measure of how much model inference is flowing through the gateway. Vercel counts input and output tokens, along with reasoning, cached-input and cache-creation tokens.

While it’s a good proxy for the amount of work being handed to different models, it shouldn’t be confused with the amount of dollars being spent. Open-weight models from the likes of DeepSeek, Moonshot AI and Z.ai are generally cheaper to run than the proprietary models offered by US frontier labs — and so handling 56% of Vercel’s token volume doesn’t mean open-weight models are taking 56% of the money passing through its gateway.

Indeed, Vercel’s data shows that open-weight models accounted for just 14 cents of every estimated dollar spent through AI Gateway in August, despite processing 56% of its tokens. Their share of spending remains far behind their share of usage, although Vercel says the open-weight share of gateway spending is on the rise.

Open-weight share of tokens vs spend on AI Gateway
Open-weight share of tokens vs spend on AI Gateway (Credit: Vercel)

Across Vercel’s AI Gateway, the average price per token fell 23.2% in August, marking a third consecutive monthly decline. Among teams that processed more than 10 million tokens in both July and August, the median cost per token fell 7.6%.

Anthropic, meanwhile, has remained remarkably consistent at the spendy end of the market. Its models accounted for 64 cents of every dollar spent through the gateway in August. Vercel says the Claude-creator’s share has never fallen below 61% in any month since December 2025, with its models occupying the top two positions by spend throughout that period — often taking third spot, too.

Top 3 models by spend, by lab.
Top 3 models by spend, by lab. (Credit: Vercel)

Loyalty lies in the model

There has been plenty of movement within that Anthropic share, however. Fable 5 fell from 13.2% of total gateway spend in July to 4.9% in August, while the cheaper Opus 5 climbed to 22.5%. More broadly, Vercel’s data suggests that 90% of teams using Fable reduced their usage, with more moving those workloads to Opus 5 than to any other model.

Opus ultimately gained almost twice as much usage as Fable lost, which Vercel attributes to the newer model handling similar workloads at roughly half the price. Or, in other words, Anthropic kept the dollars even as customers shifted toward a cheaper model within its own lineup.

“Lab loyalty doesn’t follow brand, it follows model profile, and consistency wins.”

“Lab loyalty doesn’t follow brand, it follows model profile, and consistency wins,” Vercel’s report authors note.

Anthropic's share of spend by model
Anthropic’s share of spend by model (Credit: Vercel)

This trend was evidenced elsewhere, too. Within five days of Z.ai launching GLM-5.3-Flash, the new model was processing three times the daily volume of GLM-5.2.

But Vercel’s data also suggests customers are more than prepared to cross lab boundaries when a replacement fails to meet the same needs on capability and price: more than three-quarters of the volume lost by Google’s Gemini 3 Flash moved to models from other providers, including OpenAI and Anthropic. And the consequence for Google wasn’t insignificant: its overall share of token volume on the gateway fell from 30% to 5%, with the decline in Gemini 3 Flash alone accounting for 22 of those 25 percentage points.

“When a new model preserves what users valued in its predecessor, the lab retains its customers,” the authors note. “When it doesn’t, those customers fill the need through other providers.”

The post Open-weight models now handle a majority of tokens on Vercel’s AI Gateway. But Anthropic still takes 64% of the spend. appeared first on The New Stack.

Intel squeezed a 1.58-bit LLM down to 1.485 bits without changing a single weight

17 septembre 2026 à 22:51
Digital void

The 1.58 in a 1.58-bit language model sounds like a hard limit, but Intel researchers pushed a ternary model below it by changing how its weights are stored rather than changing the model itself.

Their new BITCOS format compressed one checkpoint to 1.485 bits per weight and improved decoding throughput by as much as 18% on CPUs and 27% on GPUs.

The key is that the familiar 1.58-bit figure assumes a model uses its three possible weight values equally, while real ternary models contain far more zeros than that calculation accounts for. BITCOS stores the location and sign of each nonzero weight separately, allowing zeros to take up less space without retraining the model or altering its output — the equivalent of packing the same contents into a smaller box.

Where the 1.58-bit figure comes from

Ternary models use only three weight values — -1, 0, and +1 — and 1.58 bits is the theoretical minimum needed to represent three equally likely options. That number is cleaner than the reality of storing the weights, where the standard approach fits five ternary values into an eight-bit byte for an average of 1.6 bits each. Models commonly store weights in blocks of 128; however, this leaves the final byte partly unused and pushes the actual rate to 1.625 bits per weight.

When Intel’s researchers measured the distribution of weights across 29 checkpoints from seven ternary model families, they found that zeros accounted for between 29.7% and 51.5% of the weights. In 26 of those checkpoints, there were enough zeros for BITCOS to beat five-trit packing.

The sparsest was a ternary version of Qwen3-1.7B produced with CAT-Q post-training quantization, where 51.48% of the weights were zero, and BITCOS brought the storage cost down to 1.485 bits per weight.

Models commonly store weights in blocks of 128; however, this leaves the final byte partly unused and pushes the actual rate to 1.625 bits per weight.

How zeros save space

BITCOS stands for “BITmap and COmpacted Signs” and divides a model’s weights into two streams. The first assigns one bit to every weight to record whether it is zero or nonzero, while the second assigns a sign bit only to nonzero weights.

A positive or negative weight therefore consumes two bits, but a zero needs only the presence bit because it has no sign to record.

If z is the proportion of zero weights, BITCOS uses 2 − z bits per weight, dropping from 1.6 bits at 40% zeros to 1.485 bits at 51.5%. Because it changes only the storage format, unpacking restores the original -1, 0 and +1 values without affecting accuracy.

BITCOS becomes smaller than five-trit packing once more than 37.5% of a model’s weights are zero, a threshold reached by 26 of the 29 checkpoints Intel examined.

BITCOS becomes smaller than five-trit packing once more than 37.5% of a model’s weights are zero, a threshold reached by 26 of the 29 checkpoints Intel examined.

Making smaller weights run faster

Built for token-by-token decoding with small batch sizes, the format reduces the weight data moving through memory. Intel developed separate unpacking kernels for AVX-512 and AVX2 CPUs as well as Xe2 GPUs, joining other efforts to fit compressed models into faster inference pipelines for AI agents.

On AVX-512 hardware, the kernel uses the presence bitmap as a mask and pdep to scatter the compacted sign bits across the nonzero weight positions. Because Xe2 GPUs lack an equivalent instruction, Intel implemented the same operation with a 2KB lookup table.

Benchmarks across five systems

Compared with the 2-bit kernels, BITCOS ran 10% to 18% faster on the 64-core Xeon server and 2% to 15% faster on the 24-core Core Ultra 9. Performance improved by 9% to 22% on the integrated Arc 140V and by 2% to 27% on the discrete Arc Pro B70. These results measure decoding after the model has loaded, separate from efforts to cut GPU inference cold starts from minutes to seconds.

The smaller format did not win everywhere

On the eight-core Lunar Lake CPU, Intel’s fixed 2-bit kernel beat BITCOS on every model because the system had enough bandwidth to make unpacking the bottleneck. BITCOS remained faster on the GPUs, although decoding overhead limited the gains. Computer scientist and AI infrastructure author Chip Huyen has made the same point about inference more generally, arguing that the right optimization depends on whether compute, memory or bandwidth is holding back the workload.

On the eight-core Lunar Lake CPU, Intel’s fixed 2-bit kernel beat BITCOS on every model because the system had enough bandwidth to make unpacking the bottleneck.

Format limits and open questions

The paper has not been peer-reviewed; all five test systems used Intel hardware, and the end-to-end benchmarks covered seven models at batch size one. Intel has yet to test the format on Nvidia, AMD, or Arm hardware.

The post Intel squeezed a 1.58-bit LLM down to 1.485 bits without changing a single weight appeared first on The New Stack.

Perplexity’s AI agents helped build a database. They weren’t allowed to run it.

16 septembre 2026 à 23:51
Abstract glitch wave

Perplexity decided it was paying too much for DynamoDB and wasn’t getting the control it wanted over read performance. So it built its own database: CobbleDB.

Built by two engineers in two months with help from hundreds of persistent coding agents throughout development, CobbleDB is a roughly 40,000-line Rust key-value store that now handles part of Perplexity’s production search traffic. The company measured median batch-read latency at 5.6 milliseconds after the move, compared with 31.4 ms on DynamoDB before the cutover, while p99 went from 123 ms to 24.2 ms.

It’s expected to cost at least 20% less than DynamoDB and plans to open-source the database at some point.

But the database itself is only part of the story. CMU professor Andy Pavlo argued at Percona Live earlier this year that databases are the hardest and most important challenge for AI agents, in part because mistakes involving production data can be difficult or impossible to reverse.

Perplexity went ahead and used hundreds of agents to help build one anyway, but they weren’t given the keys to production.

It’s expected to cost at least 20% less than DynamoDB and plans to open-source the database at some point.

Why DynamoDB couldn’t keep up

Each search requires the serving layer to retrieve pre-chunked passages and vector embeddings, with a single Search API call fetching 100 to 120 page keys in batches of 10 to 20. Each item averages about 50 KB.

DynamoDB gave Perplexity little control over how it handled reads, which meant a slow replica could hold up the entire things. It also charged for the steady flow of large reads and writes generated by search, crawling, and reprocessing, which made cloud costs difficult to justify as traffic and the corpus grew.

That led Perplexity to separate long-term document storage from the database serving live searches.

Three tiers for search data

The storage stack is split into three pieces. Pillar keeps durable document state in YTsaurus on HDDs, including versioned metadata, chunks and embeddings, while Lorry packages updates into partition-specific batches and moves them through S3 to CobbleDB.

Processed page data is spread across three replicas per partition, with hashed URLs as keys and RocksDB keeping often accessed data in memory while the rest stays on local NVMe. Reads stay within the same availability zone when possible, and the router can try another replica if one is slow rather than hold up the batch.

Updates come through S3 and are applied independently, allowing a replica to fall behind and catch up without blocking the others.

Roughly 5X Lower Batch-Read Latency

Perplexity was handling approximately 200,000 requests per second when it measured CobbleDB at 5.6 ms for a median batch read, down from the 31.4 ms it had recorded on DynamoDB. At p99, latency went from 123 ms to 24.2 ms.

In later load testing, CobbleDB reached 500,000 requests per second before performance started to decline.

The comparison comes with an important caveat; DynamoDB and CobbleDB weren’t tested side by side against identical traffic: the DynamoDB figures were recorded before the cutover, and CobbleDB’s afterward. Perplexity separately ran synthetic benchmarks using batches of 10 to 15 keys with values ranging from 100 bytes to 100 KiB.

Its cost model puts CobbleDB at least 20% below DynamoDB across the commitment options evaluated, though that estimate doesn’t include the engineering cost of supporting the database.

In later load testing, CobbleDB reached 500,000 requests per second before performance started to decline.

Agents built it, engineers controlled it

The agents carried context across sessions, catching problems with restore assumptions and runtime configuration while working on fixes and tests. But they weren’t running the database.

The two engineers kept control of the architecture and production system, particularly important given Pavlo’s warning about putting agents near critical production data.

Ownership has long-term costs

Shipping CobbleDB in eight weeks solved Perplexity’s immediate engineering bottleneck, but maintaining a custom datastore could prove considerably harder. The latency results aren’t from a controlled side-by-side benchmark, and the projected savings don’t include the engineers needed to maintain CobbleDB and respond when something breaks.

Like Shopify and Ramp, which built custom coding agents around third-party models, Perplexity kept the cloud infrastructure but replaced a managed service with something built for its own needs. CobbleDB shows how AI-assisted development is changing that calculation, making custom infrastructure more practical for smaller engineering teams.

CobbleDB shows how AI-assisted development is changing that calculation, making custom infrastructure more practical for smaller engineering teams.

The post Perplexity’s AI agents helped build a database. They weren’t allowed to run it. appeared first on The New Stack.

Anthropic bet users were choosing wrong. So it removed the choice.

16 septembre 2026 à 18:46
Single lane

Using Claude for anything beyond a quick question has always started with a routing decision to use Chat or Cowork? Anthropic has decided to eliminate that fork.

Starting Wednesday, Claude Chat and Cowork merge into a single interface where one conversation can handle everything from a simple answer to a multi-step project with connected tools and background execution. The company is also launching Claude Docs and Claude Slides in beta on paid plans, and moving Claude Design — previously a standalone workspace — into conversations.

The combined effect promises a streamlined experience with Claude picking up context, skills, and connectors as the work requires, and can keep running after you close your laptop.

The company is also launching Claude Docs and Claude Slides in beta on paid plans, and moving Claude Design, previously a standalone workspace, into conversations.

Two modes, one problem

Anthropic built Cowork as a desktop-first agent for bigger work and Design as a separate workspace for visual output. Both shipped earlier this year and gained traction.

“We built Cowork as a separate place for bigger work, and Design for visual work,” Anthropic said in its announcement. “People used both, and told us the frustrating part was deciding where a task belonged.”

Anthropic has run into this problem before. When the company promised 20x more usage on its Max plan, developers complained that it wasn’t always clear where one limit ended and another began. Cowork and Design created a similar headache by making people decide where to start the work before they could actually start it.

Context didn’t always follow the work either, so moving from chat to Cowork or Design could mean bringing the same background along all over again. Anthropic addressed part of this in August when it unified Claude’s memory across chat and Cowork. Wednesday’s change goes further by merging the products themselves.

“We built Cowork as a separate place for bigger work, and Design for visual work,”

Context that finally travels

Cowork’s capabilities — local file access, multi-step execution, scheduled tasks and connected tools — now live inside the conversation. A workflow like that previously meant switching from chat to Cowork and carrying the context with it, and until Anthropic brought Cowork to web and mobile in July, it also required the desktop app.

Claude still asks before taking an action by default, but it can be set to keep working and check in only when something needs a closer look, while recurring tasks such as a weekly report can be scheduled to run every Monday without being started manually.

Output stays in-conversation

Claude Docs and Claude Slides launch in beta on paid plans, bringing document editing and presentation building directly into the app. Claude can turn work from an existing conversation into slides, which can then be edited, presented from Claude or downloaded as PowerPoint or PDF files.

The practical benefit is that a report and a slide deck based on it don’t have to begin as separate jobs with the same background supplied twice. Everything stays attached to the conversation that produced it.

Claude Design also now works inside conversations, in addition to remaining available on its own. For organizations that rely on MCP connectors to wire Claude into external tools and data, the merge means those connections are available wherever a conversation goes — without requiring users to start in a specific mode. Skills, connectors, and artifacts carry across what used to be product boundaries.

The practical benefit is that a report and a slide deck based on it don’t have to begin as separate jobs with the same background supplied twice.

What Anthropic hasn’t said

The announcement leaves some gaps. It doesn’t say whether users can force a request to stay in simple chat mode rather than letting Claude decide how to handle it, or how that decision affects context windows and token consumption. There’s no mention of API changes, which makes this a consumer and team product shift, not a platform one, at least for now.

It also doesn’t address what happens to workflows built around the old separation. Shopify rebuilt its mobile development stack in 12 weeks when it consolidated tools that had grown apart — the question for Claude power users is whether their existing Cowork setups, skills, and scheduled tasks survive the merge cleanly. Anthropic says existing Cowork chats, projects, artifacts, connectors, and skills will remain available.

Rollout starts with Pro

The unified interface rolls out to Pro and Max users across web, desktop, and mobile over the next few weeks. Anthropic says there’s nothing to enable. Team and Free plans follow. Enterprise customers are on a separate timeline; Anthropic is giving administrators at least 30 days’ notice before the change reaches their organizations.

The post Anthropic bet users were choosing wrong. So it removed the choice. appeared first on The New Stack.

AI evaluator: The most important AI job in history? How developers might fill the proposed new job

16 septembre 2026 à 17:58
Lots of pink escape keys

The pace of frontier AI model development spurred Anthropic CEO Dario Amodei to publish an essay last weekend, calling for changes in how the industry is regulated and develops. In a three-part plan that includes both democratic and global coordination, Amodei writes that the first step was something Anthropic is committing to unilaterally.

“Each frontier AI company [should] commit to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes,” writes Amodei.

Amodei’s essay followed dire warnings from former Anthropic and OpenAI pretraining research specialist Jacob Coxon, who posted a thread on X saying the people building AI earnestly “believe that it could kill us all” by the end of the decade.

Shortly after Amodei published his essay, OpenAI CEO Sam Altman and SpaceXAI founder Elon Musk chimed in: “I agree with Dario,” posted Altman; “Dario is right,” posted Musk. Later that day, Demis Hassabis, founder of Google DeepMind, posted, “Dario’s essay points towards the right path forward.” In a post on X, Meta CEO Mark Zuckerberg writes that Meta Superintelligence Labs already uses independent evaluators, and that, “In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators.”

This week, theories began to surface about why the world’s biggest frontier AI labs would want to intentionally slow their pace when competition is so fierce. “The desire to slow down is puzzling, but perhaps if the whole system slows down, the rules of winning can be the same for all,” posted Nikesh Arora, chairman and CEO of Palo Alto Networks.

In his essay, Amodei likens the proposed job of independent AI evaluator to embedded regulatory supervisors used in the banking industry, i.e., third-party professionals. Altman describes the job as having “employee-like access” in his post on X.

So, who could fill these roles that AI leaders agree are desperately needed?

Salaries top out at $687K; are you interested?

METR’s current job openings are well paid (salaries top out at around $687,000), and the job specs are daunting. 

“You’re scrappy, creative, independent, and self-directed (because during the exercises you’ll only have a few other METR employees you can talk to). The work is novel, and you’ll need to figure a lot of stuff out on the fly largely by yourself. You are excellent at loss-of-control threat modeling and breaking down safety cases,” reads the spec.

“You’re scrappy, creative, independent, and self-directed. The work is novel, and you’ll need to figure a lot of stuff out on the fly largely by yourself. You are excellent at loss-of-control threat modeling and breaking down safety cases.”

Similar but less colorfully illustrated roles (paid between $180K–$300K) are also available at AI model training company Mercor, where candidates will need a Ph.D. or M.S. and more than two years of work experience in a computer science, electrical engineering, econometrics, or another STEM field that provides a solid understanding of machine learning and model evaluation.

“Employee-like access fluctuates wildly”

AI security consultant and CTO at Komodo, Kadan Stadelmann, tells The New Stack that his typical week sees him work differently with each client. This is because “employee-like access fluctuates wildly”, from rigid focus areas to broad access, and much of that aspect is determined by contracts signed before work begins.

“I am invited into labs to probe numerous risk vectors, including agent behaviors under realistic conditions,” Stadelmann says. “Among my duties are tasks that include monitoring chains-of-thought and prompts. The goal is to establish an objective and look at a specific AI system to question how autonomous the system is, and how long it takes to complete specific tasks. Most importantly, evaluators at my level monitor for how well a team adheres to its claimed safety practices.” 

Software engineering skills beat doctorates

Although METR wants evaluators to have a Ph.D. up their sleeve, Stadelmann says that as the prevalence of this role expands, he feels the technology industry has been, and continues to be, driven by people who can demonstrate strong engineering skills, not doctorates.

“I am invited into labs to probe numerous risk vectors, including agent behaviors under realistic conditions. Among my duties are tasks that include monitoring chains-of-thought and prompts.”

Questioned on whether costs create a barrier for smaller labs, Stadelmann notes that some AI evaluation work is funded by third-party non-profits, which protects independence. 

“But overall, evaluators will not be able to keep up with big frontier model firms. They will be out of control, and we will be dependent upon their own internal ethics. Plus, anyway, many of the smaller labs of any worth may inevitably be acquired by the big players in this space,” he adds.

What happens when an evaluator finds something wrong?

Founder and CEO of facial image AI identity governance company Indie Me, Dion Johnson, tells The New Stack that what interests him most about embedded AI evaluators isn’t the job title; it’s what happens when their judgment uncovers that the model behaved in a way nobody expected. 

“If the evaluator can only raise concerns when those concerns are convenient, then we have not created independent oversight — we have created another layer of process,” Johnson says. “The evaluator needs enough access to see the uncomfortable things, not just the polished demonstrations. They need to understand what happened during training, what failed during testing, what behaviors appeared unexpectedly, and where the team itself still has uncertainty.”

On the required skills AI evaluators need, Johnson agrees that technical depth, machine learning environment security experience, and software engineering as a whole matter.

“To choose a competent AI evaluator, I would look for someone who is deeply comfortable with uncertainty and deeply uncomfortable pretending they know something they do not. This is someone who can sit in a room full of brilliant people at an AI model company on launch day and say ‘I’m not convinced’… and that takes judgment and courage,” he adds.

“I would look for someone who is deeply comfortable with uncertainty and deeply uncomfortable pretending they know something they do not.”

Fear and loathing in the AI space

In his September 8 post, which has been viewed 172 million times and seemingly spurred AI leaders to change course, Coxon, the former AI researcher at OpenAI and later Anthropic, writes: “This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible but I hear the same people express fear privately. No other human activity poses this level of danger.”

As for where Coxon looks for work next, perhaps it might be a role in AI evaluation execution engineering.

The post AI evaluator: The most important AI job in history? How developers might fill the proposed new job appeared first on The New Stack.

Perplexity’s new agent runs entirely on your GPU — with one expensive catch

14 septembre 2026 à 20:21
Abstract server

Running an LLM on your PC is easy enough, but putting an agent to work there is a different story. Portable Computer, the local version of Perplexity’s Computer agent, is now available inside the Perplexity app for Windows on compatible Nvidia GeForce RTX and RTX PRO GPUs.

That’s the good news; the catch is, you’ll need an Nvidia GPU with at least 24GB of VRAM.

The Windows launch gives Perplexity three platforms in less than three weeks. Portable Computer debuted on Linux and Nvidia DGX Spark on August 25, followed a week later by hybrid compute for Apple silicon, which splits tasks between local and cloud models on Macs. Now Windows joins the mix, but bringing Portable Computer over took more than simply porting the app. Perplexity had to adapt the model runtime, orchestration, security, and hardware integration for each platform while keeping the user experience the same.

you’ll need an Nvidia GPU with at least 24GB of VRAM to use it.

Orchestration beyond the model

Portable Computer bundles those pieces together. On Windows, it supports PPLX 27B — Perplexity’s post-trained model — and Qwen 3.8 27B, both optimized for RTX GPUs, alongside a built-in browser, tool calling,  and Perplexity’s proprietary SPACE sandbox.

It’s a different lane from LM Studio or Ollama, which make running models locally as painless as possible but stop well short of giving a model autonomy over multistep work. DeepSeek’s recent hiring spree of roughly 150 new roles, nearly all of them focused on agent infrastructure rather than the model, hints at how much engineering sits between a capable model and a capable agent.

DeepSeek’s recent hiring spree of roughly 150 new roles, nearly all of them focused on agent infrastructure rather than the model, hints at how much engineering sits between a capable model and a capable agent.

Connectors blur local boundaries

Perplexity ships connectors for Microsoft Outlook, OneDrive, and Word, plus Google Drive, Gmail, Slack, and GitHub — which tells you something about what “local” actually means here.

The agent can reach external services because it’s not air-gapped. Locally completed tasks can process files without sending documents to a cloud model. Once an agent has access to both local files and remote APIs on the same machine, figuring out which resources it actually needs and where to find them gets harder.

Hybrid cloud as fallback

Perplexity isn’t pretending that a 27-billion-parameter model running on a desktop GPU can handle everything, which explains the hybrid architecture. When the agent determines that a task needs more reasoning power than the local model can deliver, it can escalate to Perplexity’s cloud models.

According to Nvidia, the agent identifies when cloud support would help and asks the user for permission before sending any data off the machine.

For organizations handling sensitive or regulated data, that split can make all the difference. A local agent can grind through source code or financial records without uploading them to a hosted model for basic processing. There’s a cost angle too, since tasks completed locally don’t burn Perplexity Computer credits.

High VRAM floor limits reach

Portable Computer is available with Perplexity Pro ($20/month) and Max ($200/month), across individual and enterprise plans, with Nvidia DGX Station support coming later. The real challenge is taking local agents from developer passion projects to enterprise-ready tools. By baking this into Windows, it immediately gets in front of the scale of users needed to make that happen.

The real challenge is taking local agents from developer passion projects to enterprise-ready tools.

The post Perplexity’s new agent runs entirely on your GPU — with one expensive catch appeared first on The New Stack.

Chinese AI models dominate OpenRouter’s US token consumption. It can now guarantee that traffic stays entirely in the US.

14 septembre 2026 à 15:59
Illustration of data-center servers marked with location pins and connected by routing paths

Everyone knows the open-weight model pitch by now: companies can download the weights, customize them, run them on infrastructure of their choosing, and retain far greater control over where their data is processed — often at a much lower cost than using proprietary models.

Moreover, open-weight models are now thought to trail the leading frontier models by only around four to five months. Nvidia, the world’s most valuable company, is betting heavily on that future. In early September, it agreed to acquire Hugging Face — the sprawling “GitHub for AI” that hosts more than three million models — for $12.9 billion, while pledging to keep the platform open to different models, clouds and computing providers. And on Thursday, Nvidia detailed how Nvidia is using its own open-weight Nemotron model to manage its vast global supply chain in partnership with Palantir.

That power also comes with serious security questions. OpenAI president Greg Brockman recently warned that increasingly capable open-weight models — pointing specifically to China’s GLM-5.3 — could “significantly accelerate the threat landscape” as models with advanced cyber capabilities become freely downloadable and modifiable.

But for businesses accessing those models through third-party services, there is another concern closer to home: where their own data goes when they use those models, particularly when the model originated in China.

China and the open-weight factor

Hugging Face data from February showed models from Chinese developers accounted for 41% of downloads in the preceding 12 months, ahead of the US at 36.5%. Over on OpenRouter, meanwhile, open-weight models now account for around 60% of tokens consumed by US-originating requests, with the company noting that Chinese models constitute the majority.

OpenRouter: Share of monthly tokens (Sept. '25 - Aug. '26)
OpenRouter: Share of monthly tokens (Sept. ’25 – Aug. ’26) — US and EU

And that’s why OpenRouter is now giving companies a way to put a geographic fence around that traffic. The AI model marketplace has officially launched US in-region routing into general availability for business and enterprise customers, promising that requests sent through its US endpoint are decrypted, processed and served entirely inside the country — or rejected if that can’t be done.

The feature itself had been quietly available in some form before now, with OpenRouter updating its documentation in early August to say US in-region routing was available to enterprise customers by request. It’s also worth noting that this is in addition to European in-region routing, which it says has been available since October 2025.

Started in early 2023 by former OpenSea CTO Alex Atallah, OpenRouter serves as an interface to the crowded AI model market, with developers able to switch between hundreds of models from myriad providers via a single API. Payments giant Stripe recently announced plans to acquire the company in a reported $8 billion deal, while a slew of other companies including Cursor, Ramp, and Meta, are also building their own model routers.

The reason why model routers are such hot property right now is largely down to economics. Developers have traditionally hard-coded applications to send everything to the same model, while a model router can instead make that choice request by request, sending easier jobs to cheaper models while reserving the pricier frontier systems for the work that actually needs them.

That intermediary role is also what makes OpenRouter’s new residency controls possible: it already decides which provider serves each request, and can now restrict that choice to provider endpoints operating in the US.

Keeping Chinese models inside the US

In a blog post announcing the new feature on Wednesday, Cailee Moberg, who works on OpenRouter’s product team, notes that while US-developed models from Nvidia and Thinking Machines are contributing to the broader open-weight model boom, Chinese models dominate usage and raise tough questions for companies concerned about their data.

“Models from Chinese labs are still most of the [open-weight model] volume, and procurement approval for those models can be difficult.”

“Models from Chinese labs are still most of the [open-weight model] volume, and procurement approval for those models can be difficult,” Moberg writes.

In its 2026 State of AI in the Enterprise report, Deloitte concluded that sovereign AI was on the rise, noting that 77% of companies “now factor country of origin into their vendor selection,” while nearly 60% construct their AI stacks “primarily with local vendors.”

And this at least partly explains why OpenRouter is now offering in-region routing for US customers. Moberg points to DeepSeek V4 Pro, Kimi K3 and GLM 5.2 as specific examples. All three are available through US In-Region Routing because Baseten, Fireworks and Azure serve them from US data centers. Companies could already keep these models inside the US by self-hosting them or using a US provider directly; OpenRouter’s new routing gives its own customers that residency guarantee without having to manage those deployments themselves.

OpenRouter maintains a live list of models eligible for US in-region routing, ranging from proprietary frontier models from OpenAI and Anthropic to open-weight models from the major Chinese labs.

“In-Region Routing allows teams with data residency requirements to get the price and performance gains from Chinese open-weight models,” Moberg continues. “When a US or EU provider hosts a model, requests go to that provider and the lab is not involved.”

“In-Region Routing allows teams with data residency requirements to get the price and performance gains from Chinese open-weight models.”

The technical change happens at the routing layer. With OpenRouter’s standard global endpoint, a request can be served by an eligible provider operating in any region, so even using a model from a US company does not guarantee that the request itself is processed in the US. With us.openrouter.ai, the request is decrypted on OpenRouter infrastructure inside the US and the pool of providers is filtered to endpoints OpenRouter has approved as operating there.

If no compliant US provider can serve the requested model, OpenRouter returns a 404 error. Companies can also enforce the regional restriction through OpenRouter’s Guardrails at the workspace, team or API-key level, while tools that would send prompt data outside the US are disabled on the regional endpoint.

So while none of this ultimately changes where the DeepSeek, Kimi or GLM models are developed, in-region routing alters which copies of those models its US customers can be routed to, and where their prompts are handled along the way.

The post Chinese AI models dominate OpenRouter’s US token consumption. It can now guarantee that traffic stays entirely in the US. appeared first on The New Stack.

Why an old caching trick is your secret to lower LLM costs

14 septembre 2026 à 13:00
Server racks in a dark data center, their mesh doors revealing dense bundles of orange and teal cables looping between hardware lit by rows of small green and yellow status LEDs.

An LLM can answer the same question a thousand times and charge you each time. Before paying for another answer, check whether anything that could change it has changed: the request, its context, the model settings, or the underlying data. I fingerprint those inputs and dependencies to create an exact-match cache key. If that key points to an answer that’s still valid and safe to reuse, I return it without calling the model. The savings start with a simple decision: knowing when the work is already done.

I didn’t learn this lesson from an LLM job. In production data pipelines, I’ve encountered a recurring pattern: a nightly job recalculates aggregations that haven’t changed since the previous run. It passes all its checks and moves the results into production successfully, all while burning compute that could have been used elsewhere.

The waste hides in plain sight because nothing appears broken. It often surfaces during a cost review, when someone notices that a significant portion of upstream compute is re-answering a question whose inputs never changed. The fix is change detection: hash the upstream inputs that could change between runs, fingerprint the job’s dependencies, and skip recomputation when the fingerprints match. Done well, this significantly reduces the compute that job consumes.

The lesson is common, and it’s the same one we keep trying to drive home in LLM workloads. There, repeated requests can also produce repeated charges, since billing is by token.

The problem is simple enough to state, but the more you look into it, the more you need a framework to engineer a good answer. For most of the LLM calls in our codebase and infrastructure, we’re billed by tokens, and many APIs treat duplicate requests as new ones anyway. Duplicate sources are almost as inevitable as rain.

Upstream users converge on similar questions to answer with their LLM tools. Batch jobs dutifully repeat boring boilerplate every time they run. Prompt-engineering experiments in development and CI runs invoke the same prompt repeatedly. And tool-calling agents may hit the same knowledge-base tool many times in a single work day.

Native prompt caching is a different thing from the response caching I’m describing. In prompt caching, providers reuse cached prompt computation and charge eligible cache reads at reduced rates; output generation remains billable. In response caching, we try to skip the call entirely when an answer already exists in our own infrastructure.

Tier 1: exact match

The simplest approach is to normalize the model request body, run it through a cryptographic hash like SHA-256, then look up the hash in an in-memory store like Redis. If we find a match, we return the answer without waiting for model inference. An exact-match cache works best when we can expect our model requests to be bounded and predictable. That doesn’t sound exciting, but for most of our batch pipelines, CI runs, and boilerplate summarization tasks, it’s exactly what we need.

Tier 2: semantic match

For many workloads, exact match isn’t enough. We’d like to look up a response for a query that’s close but not identical. So we take the user’s query, run it through an embedding model, and store the resulting vector in a vector database. When a new query arrives, we run it through the same model and search for close matches by cosine similarity.

Close enough by what measure? A common starting point is a cosine-similarity threshold in the [0.90, 0.95] range, but treat that as a number to tune, not a default — the right value depends on your embedding model and your data, and you should test it against real queries. Note that vector stores differ in what they return: cosine similarity rises toward 1 for closer matches.

At the same time, some engines report a distance that falls toward 0, so confirm which your threshold is comparing against. Either way, a looser threshold raises the risk of wrong matches, where the system answers one query while the user was asking about another. (“What’s the weather in my town?” can’t be safely conflated with the same question about a different town just because the cosine similarity is high.)

Tier 3: hybrid

A common approach runs both tiers in sequence: check the exact-match store first, and run semantic search only on a miss. When semantic search returns a close-enough match, the result is promoted back into the exact-match store under the hash of the new query that triggered it, so the paraphrase and its answer are an exact hit next time.

This favors cheap exact matches on repeat traffic. The pseudocode below shows the full flow: normalization and SHA-256 for exact match; a Redis get followed by a set on a miss; embedding the query and searching the vector DB with top_k=1; checking cosine similarity against the per-category threshold; and setting the TTL before writing the response back into the exact store.

Both tiers key on more than the query text alone: the context and documents in the prompt, the model and its settings, the version of any retrieved source, and the caller’s access scope. Two identical questions asked against different documents, or by users with different permissions, must not share a cache entry.

def cached_completion(query, ctx):

    # ctx bundles everything that changes what the correct answer is:

    # the context/documents in the prompt, the model and its settings,

    # the source-version of any retrieved content, and the caller's access scope.

    key = sha256(normalize(query, ctx))

    # Tier 1: exact-key lookup on Redis (O(1)).

    # Correctness still depends on cache contents, request scope, and freshness.

    if (hit := redis.get(key)):

        return hit

    # Tier 2: semantic search, restricted to the same scope as the request.

    emb = embed(query)

    match = vector_db.search(emb, top_k=1, filter=scope_of(ctx))

    if match and same_scope(match, ctx) \

            and match.score >= threshold_for(category(query)):

        # Promote, but preserve the original freshness deadline.

        remaining = match.expires_at - now()

        if remaining > 0:

            redis.set(key, match.response, ttl=remaining)

            return match.response

    # Miss on both tiers: call the model, validate before writing back.

    resp = llm(query, ctx)

    if is_valid(resp):  # no errors, no empty payloads, no malformed JSON

        ttl = ttl_for(category(query))

        redis.set(key, resp, ttl=ttl)

        vector_db.insert(emb, resp, ttl=ttl, scope=scope_of(ctx))

    return resp


One threshold does not fit all. Code-like queries often need stricter thresholds, around 0.95 or higher, because small wording changes can produce entirely different results. Conversational queries can tolerate looser thresholds, in the 0.85 to 0.90 range. These numbers are starting points, not settled values — validate them for your own workload and embedding model before relying on them. Cache freshness works the same way, and the right TTL follows from how much staleness the use case can tolerate, not from the data type alone.

A cached market-data answer might be acceptable for only a minute or two, because a stale price can be actively misleading. An internal HR policy answer can often be reused for weeks, because the underlying document rarely changes and a slightly old answer is usually still correct. The interval is a judgment about acceptable staleness, not a fixed property of the content.

The math

For illustration, suppose a workload of 1,000,000 calls per month at $0.006 per call, roughly $6,000 with no caching. Say a hybrid cache gives about a 60% hit rate, whichever tier hits first, avoiding 600,000 calls to the model, and that embedding and vector-store costs come to about $150. That brings monthly spend closer to $2,550, a 57.5% reduction, plus the latency win of answering many questions without waiting on the model.

One caveat worth shouting: measure your hit rate before you project any savings.

The decisions

There’s more to this than the tiered framework. Tune your TTLs to the freshness each data type actually needs, and invalidate entries when you update the content behind them. A fine-grained approach assigns a per-category TTL based on how quickly each answer goes stale: a news summary might hold up for an hour, while a live sports score is worthless within seconds and shouldn’t be cached at all during a game.

Live scores require a freshness policy matched to the application. Verified final scores can support much longer caching, with invalidation for corrections. The distinction here is whether the underlying value is still moving. A blunter approach skips per-category tuning entirely and purges the whole cache whenever the source content changes. Either way, run the cache in shadow mode first, logging what you would have returned without changing behavior. Evaluate cached answers against verified reference answers or expert review. A fresh model response can help identify differences, but it is not ground truth.

Warm the cache from a historical set of common queries before you rely on it, and validate answers before writing them back, so you don’t poison the cache with errors, empty responses, or malformed content.

When not to cache? Skip it for requests with personal or account-specific data, to avoid leaking one user’s cached output into another’s request. Skip it for creative tasks, where you want a different answer each run. And skip it for genuinely real-time data like stock prices and live inventory, where an answer even a minute old may be too stale for the application.

The takeaway

The principle predates the web: Donald Michie described memo functions in 1968. When you can, fingerprint the question and store the hashed exact form alongside the semantic-variant form, so you avoid repeated model calls while a valid cached answer remains available.

The post Why an old caching trick is your secret to lower LLM costs appeared first on The New Stack.

Chip Huyen explains how to cut inference costs without new hardware

13 septembre 2026 à 17:00
Layers of wavy yellow horizontal strips with deep shadows between them, forming an abstract pattern.

Last October, the P99 conference — the online gathering for developers focused on high-performance, low-latency applications — featured a cracking keynote from Chip Huyen. 

The author of the best-selling AI Engineering, Huyen opened with simple math: Training a frontier model is a one-off cost, but inference is the same cost paid over and over. That’s great for the frontier model providers, and bad for us token burners. Over the life of a model, Huyen reckons the compute split ratio lands somewhere between 1:10 and 1:100 for training to inference. Reasoning models — which burn even more tokens — push that out even further. We all know the feeling of hitting our weekly session quotas.

Huyen’s point is that if inference is too expensive, then nobody ever recovers the training bill, which might explain why there are so many memes about the “profitability” of frontier models. So, how do we optimize inference? 

That’s a topic that Huyen spent months researching for her book. And in the spirit of optimization, Huyen distilled it down to 30 minutes for the conference in October 2025.

Huyen is returning for P99 CONF 2026 in a few weeks. Ahead of that moment, watch her full talk – or read the recap below – from last year, then let’s talk about how those ideas aged over the past 11 months.

What to measure

Chip recommends focusing on a few key latency metrics:

  • Time to first token (TTFT): How much time elapses before the user sees anything
  • Time per output token (TPOT): The average time between consecutive tokens (aka inter-token latency)
  • End-to-end latency: Time to first token, plus time per output token, multiplied by the number of output tokens minus one
(Click to enlarge graphic.)

With reasoning models, some of those tokens never reach the user. “The first generated token might not be the same as the first visible token,” Huyen explained. “The model might think for a while, and it will only show the first token of the final output to the user.” 

“The first generated token might not be the same as the first visible token,”
— Chip Huyen

Some people also measure Time to Publish for that (i.e., how long until the user sees the first token). The best metric to prioritize depends on what matters most for your users. 

Also consider “goodput” alongside throughput. Throughput measures requests processed in a given window. Goodput measures the requests that actually met your targets. Chip’s example: an app targets 200 ms time to first token and 100 ms time per output token, and processes 10 requests per minute, but only three hit both. 

(Click to enlarge graphic.)

3 ways to optimize LLM inference

With inference servers, you can optimize from 3 different angles: the hardware, the model, and the service that manages the requests and responses.

(Click to enlarge graphic.)

Huyen previously worked at Nvidia and opted out of the hardware discussion: “Even though I find it to be an intellectually interesting topic, it’s not relevant to a lot of people because we don’t have the power to change the hardware itself,” Huyen explained. She also didn’t want to spend much time on the obvious solution: replica parallelism, or just adding more machines. It’s costly, and it gets complicated fast – especially if you end up with a mix of 80GB, 48GB and 24GB machines and models of varying sizes to distribute across them.

That leaves the model and the service. Huyen offers these tips on how to decide: “If you want to host the models yourself, or if you have access to the model weights, or if you train a model yourself, or you want to fine-tune or distill a model, then model optimizations might be for you. However, if you want to take a model as-is and make it more efficient on your own inference service, you might want to look into service optimizations.

Model optimization

The following techniques change the actual weights so that they can change the model outputs.

Quantization lowers the precision used to store weights and activations  (e.g., from four bytes per parameter at 32-bit to one byte at 8-bit). Huyen explained, “Reducing the precision not only reduces the memory requirement to run the model, making it cheaper. It can also make the model a lot faster. If you do additions bit by bit and each weight is 32 bits, you have to do it 32 times. If it’s 8 bits, you only have to do it eight times.”

The tradeoff is a small quality hit. Huyen continued: “It’s possible to reduce a lot of the model’s memory footprint with minimal quality degradation, and quantization is pretty generalizable to a wide variety of model architectures and model sizes. That’s why it’s very popular. I rarely see any companies running a model at full precision anymore.”  

“I rarely see any companies running a model at full precision anymore.”
— Chip Huyen

Distillation involves using a large model to generate training data for a smaller model. For example, say you have a truly large model (the example Huyen used was o1) and want a model that performs like it, but is much smaller. Basically, you collect a large set of prompts, run them through the larger model, then train the smaller model on its responses.

Proceed with caution, though. Huyen warned, “A lot of model providers have the condition that they do not allow their models to be used to train competitive models. So even though it’s a very common technique, you need to check licensing.”

Service optimization

This set of techniques targets how requests are scheduled, routed, and reused. The actual weights aren’t affected.

Batching groups multiple requests so they’re processed together in a single pass through the model – which is much more efficient than dealing with them one at a time. Huyen presented a few batching options:

  • Static batching waits for the batch to fill. This maximizes compute utilization, but it might increase the latency for the first requests.
  • Dynamic batching runs on a timer instead (e.g., batching every 15 ms). This is less compute-efficient, but it’s better for latency.
  • Continuous batching handles the case where requests finish at wildly different times, which is common with LLMs. One request asks for the capital of Vietnam; another kicks off deep research. With static or dynamic batching, the finished request’s slot sits idle until the slowest one completes – and new requests queue up behind it. Continuous batching returns each request as it finishes and fills the spot with another request. That can improve compute resource utilization and latency.
(Click to enlarge graphic.)

Decoupling prefill and decode separates the two phases of a request onto different machines. (Prefill processes the input, while decode generates the output.) Huyen said, “Input tokens can be processed in parallel, whereas output tokens need to be generated sequentially. With parallel processing, it’s bounded by compute, the processing power of the chip. With decoding, it’s bounded by memory, because you have to move model weights.” 

Because each phase stresses different resources, most services now separate them. To improve time to first token, shift machines toward prefill. If you care more about improving time per output token, shift them to decode. 

(Click to enlarge graphic.)

Parallelism splits work across machines. Replica parallelism copies the whole model onto more machines. Tensor parallelism divides a very large matrix, so different machines compute different parts of it. Pipeline parallelism divides the model by layer, so requests move through as a pipeline. 

(Click to enlarge graphic.)

Prompt caching processes shared text once, saving cost and latency. A lot of repetition exists across requests to the same application: the system prompt, the examples, the same code base, the same document behind different questions. You might as well process that shared segment once, cache it, and reuse it.

The technique was relatively rare when Huyen was writing AI Engineering. “There was one paper about it, and it was not really known, but it made a lot of sense. So I included prompt caching in the book, and I’m very happy to see that nowadays it’s pretty much everywhere.”

(Click to enlarge graphic.)

The savings scale depending on how much of your prompt gets cached. In Claude Code logs, Huyen’s open-source tool Sniffly found cache hit rates of 90% to 97%. Some providers rewrite prompts internally to improve hit rates, but you might as well structure them yourself. 

Huyen’s tip: since caching works on shared prefixes, put the stable parts of your prompt first and the variable parts later. “It’s pretty easy to do, and it can improve your application performance significantly,” she noted. 

Evaluating inference providers

Huyen closed with a warning for anyone evaluating inference providers: “There are many inference companies that provide inference optimizations for models you want to use, and a lot of them advertise just cost and latency. 

“But pay attention to how many inference optimization techniques also change the model behavior or reduce the model quality. So when evaluating an inference service, it’s important to look not just at cost and latency, but also at model quality. Does this model, provided on this service, also perform similarly on standard benchmarks?”

What’s changed one year later?

So where do we stand today, one year on from this keynote? Most of it actually aged quite well. 

On the economics, I reckon Huyen was bang on… I think, for most of us as users, we don’t have all the cost levers to pull that Huyen outlined. But it’s great to understand what is happening. As a novice local LLM user myself, I found I could relate to her points on parallelism (I don’t have it) and prompt caching/quantization (within my grasp of control). 

Prompt caching (which Huyen said was new when Huyen wrote AI Engineering) is now priced into every bundle purchase of API tokens. And her Claude Code observation (90% cache hit rates) is probably the reason we mere mortals can still afford agentic coding agents at all.

Some of it aged in ways that were hard to predict at the time. Huyen mentioned how reasoning models make inference even more significant. One year on, I think agents running multi-step loops with tool calls have turned that idea from a footnote into a way to turn Claude’s rate limits (and their infamous 99.x% availability) on their head. 

All the metrics Huyen described – time to first token, time to publish, goodput under a latency SLO, etc., are all now part of the lingo and probably need to be reasoned about differently. 

That’s one thing I hope she’s talking about this year! Grab a free conference pass and join us online. 

Grab a complimentary pass to PG 99 Conf 2026 and join us on October 21 and 22 to chat with Huyen.

The post Chip Huyen explains how to cut inference costs without new hardware appeared first on The New Stack.

❌