❌

Vue lecture

Performance engineering from kernel analysis to AI: Adrian Cockcroft’s take

Abstract dark digital landscape with glowing contour lines representing multidimensional performance data and response time distributions.

Over its five-year history, P99 CONF has hosted quite a few speakers who’ve offered pointed takedowns of the namesake metric. At last year’s conference, Adrian Cockcroft didn’t explicitly state that P99s are BS… but he did allude to it.   

If you don’t know Cockcroft, he’s spent decades architecting, scaling, and optimizing resilient, high-performance systems at giants like Sun Microsystems, Netflix, eBay, and Amazon. We could probably dedicate an entire day of P99 CONF to discussing the lessons learned from just some of his projects (Solaris kernel performance, multi-processor optimization, Netflix’s on-prem to cloud migration, Chaos Monkey…) 

Fortunately, RedMonk analyst Rachel Stephens proved the perfect host for a conference that’s all about making things fast. She sat down with Adrian and led us on a whirlwind tour of how AI has impacted performance engineering. Here are some highlights from the chat (full video below).

Note: P99 CONF 2026 – a free + virtual conference on all things performance – is going live October 21-22. Grab a complimentary pass and join us!

From kernel analysis to vibe coding perf tools 

As a performance specialist at Sun in its heyday, getting to the root of performance problems involved lots of digging and divination. Cockcroft recalls, “Back in the old days with Sun, people would look at the output of system metrics in vmstat or whatever, and they’d be guessing what the numbers meant. There was a very vague understanding of what these things meant. The manual page wasn’t very clear.” 

Cockcroft ended up going to the source, literally. “I went and read all the kernel source code and figured out exactly where these numbers came from, exactly what they meant, which ones were approximating what, and wrote all that down.” That led to two performance books: Sun Performance and Tuning and Resource Management.

“My speedup is infinite, because this code would never exist without these tools. I wouldn’t have the time to build them.”

Four decades later, there’s now a wealth of helpful tools for end-to-end tracing, but Cockcroft’s curiosity still lies in what the tools are not showing. He continued, “Everything sort of looks okay in the tools – but the system isn’t behaving well. I usually come in and try to find a new way of looking at the data. A new type of analysis, or go a little bit deeper or finer grain, or stop looking at averages and start looking at distributions, and find all kinds of interesting things that nobody knew were happening.”

Currently, he’s vibe coding tools to better analyze the anomalies he finds. Saved from having to brush up on Python or hunt down graphics library fragments on Stack Overflow, Cockcroft can now stand up custom tooling in minutes. “My speedup is infinite, because this code would never exist without these tools. I wouldn’t have the time to build them.”

Peaks not percentiles

One specific vibe coding project: Cockcroft built (and open-sourced) tooling to get a better understanding of response time distributions. 

Response time distributions have been on Cockcroft’s mind for over a decade. While most people obsess over percentiles – yes, P99 CONF included – Cockcroft is most intrigued by the distribution of response time peaks in a histogram. He believes percentiles don’t work when trying to understand the latency and performance of modern web services. A single number like P99 can’t tell you whether the underlying distribution has one peak or several. And when there’s more than one peak (as is often the case in the real world), the mean, the standard deviation, and even the P99 itself lose most of their meaning.

“Percentiles don’t work when trying to understand the latency and performance of modern web services.”

Image showing what people think response time distributions looks like vs what they really look like
(source: A Tale of Two Histograms)

For example, assume you have a histogram with two response time peaks: a fast one from a cache hit and a slow one for misses that require actual work. As the cache hit rate shifts, each peak’s position remains the same (i.e., the latency values of the fast-response mode and the slow-response mode don’t change), but the peak heights rise and fall. “Your averages and your P99 are changing all over the place, but all that’s really happening is your cache hit rate is changing,” Cockcroft said.

So how do you go beyond measuring P99s and averages? Cockcroft did what he’s done for decades: dive in and build a custom tool. But these days, it’s much simpler thanks to LLMs.

“Your averages and your P99 are changing all over the place, but all that’s really happening is your cache hit rate is changing.”

He had already worked out the statistical approach for analyzing the distribution. Once ChatGPT came out, he quickly used it to build a tool that automated it. Instead of collapsing everything into an average, it identifies an arbitrary number of peaks in a distribution and tracks how they fluctuate over time. It’s implemented in R – a language Cockcroft hadn’t used in a while, but ChatGPT knew quite well – and it’s open source. If you’re curious, learn more in his Percentiles Don’t Work article and “A Tale of Two Histograms” talk and deck (“It was the best of response times, it was the worst of response times…”)

Where do we go from here?

To close, Stephens asked Cockcroft what advice he’d share with teams working on high-performance systems today. His top tip was to start with the macro view to find what’s interesting, then keep digging deeper until you’re inspecting individual slow requests end-to-end.

“Remember the microscope that you got when you were a kid,” Cockcroft said. “First, you have to focus it using the lowest resolution, at 10x, and then you can click it to 100x and adjust that, looking at just one speck now. Once you get that in focus, you click it to 1,000x.”

Cockcroft has spent his career building tools that bring obscure performance issues into focus. We look forward to seeing what others have cooked up with agentic tooling to help identify and solve performance problems this year at P99 CONF. 

Learn about the latest performance optimization techniques, tooling, and case studies at P99 CONF – free and virtual, October 21-22. Grab a complimentary pass and join us!

The post Performance engineering from kernel analysis to AI: Adrian Cockcroft’s take appeared first on The New Stack.

  •  

The rise of agentic AI on Kubernetes: unleashing the new infrastructure layer

Abstract 3D render of blue cubes inside gold wireframe boxes, linked by red rods into a dense cluster, with teal lines connecting outer cubes.

AI is changing expectations around infrastructure and operations, including Kubernetes management. When models run close to the data they use, deployment, scaling, and governance responsibilities tend to shift to platform teams. And as clusters, environments, and operational signals continue to multiply, manual operations often strain under the added weight.

AI may simultaneously provide opportunities to lighten this growing load. Agentic software can now observe a system, reason about it, and act within predefined limits. 

Ultimately, these platforms’ value depends on the quality of the context an agent can see and the boundaries you set. Without cluster state, policy, and access rules, an agent can only guess.

Without cluster state, policy, and access rules, an agent can only guess.

For agentic AI to streamline multi-cluster management, you need clear lines between what the system observes, what it recommends, and what it changes. Drawn well, those lines let teams gain notable speed while still maintaining control.

The impact of AI on computing infrastructure

Teams once treated AI as an application concern; models sat on top of existing systems, and the stack underneath stayed mostly unchanged. Today, AI reaches into more and more customer interactions, while data storage needs simultaneously expand and orchestration pressure grows. A recent Forrester report describes the modern AI computing stack as stretching from the models themselves into and across the infrastructure beneath them.

As AI workloads move into production, they place new demands on the infrastructure beneath them. Many lean on specialized compute, with resource needs that rise and fall through bursts of training and inference. Because conditions shift quickly, they can also call into question whether telemetry remains trustworthy. Each of these demands lands at the infrastructure layer, where the workloads run.

The infrastructure layer of the new AI stack

The infrastructure layer covers compute, storage, and networking. It is a foundation that every workload running on the layer depends on. As AI workloads grow, choices about capacity, placement, and control will increasingly shape the performance of the data, intelligence, orchestration, and experience layers atop the infrastructure.

To operate the infrastructure layer efficiently across many machines and locations, a team may rely on orchestration instead of managing servers by hand. In cloud native contexts, Kubernetes has become a control point for scheduling workloads, applying policy, and presenting a consistent interface across environments. Kubernetes is especially well-suited to support organizations this way when teams need consistent control across an estate spanning data centers, clouds, and edge sites. 

Agentic AI and Kubernetes: the future of the infrastructure layer

Agentic AI can extend automation from fixed rules to systems that adapt to real-time conditions. Traditional automation runs the same script whether the environment has changed, while an agentic system observes the environment, reasons about what it finds, and then takes action.

When you apply agentic capabilities to multi-cluster management, the system follows this same sequence. An agent reads cluster state and operational data, proposes a diagnosis or next step, and then carries out actions based on an approved scope, usually after a person signs off. You can further reinforce these boundaries by routing each request to a specialized agent that receives only the metadata it needs.

The signals that an agent receives from the cluster, the context about policy and access, and the definitions of what the agent may change are the key elements that give agentic systems their value. They also separate agentic AI on Kubernetes from a generic assistant. 

Manual Kubernetes management is less efficient at scale

Admittedly, agentic AI fits some settings better than others. On a small single-cluster footprint, the overhead may outweigh the benefit. Manual Kubernetes management often holds up on a handful of clusters, but it can become unreliable in a rapidly growing estate. After all, each new cluster adds lifecycle work across upgrades, patching, configuration, and renewal. Those tasks can quickly multiply and diverge in hybrid environments.

Configuration drift is a high risk in these situations. Settings that started identical can fall out of sync, and policies can apply unevenly from one team to the next. Individually, these gaps may be manageable, but collectively they raise the odds of an outage or a failed rollout.

Visibility can also erode in an unmanageable way. Clusters spread across data centers, clouds, and edge sites often leave teams with no single view of the whole landscape. When DevOps and platform engineers stitch together signals from separate tools, resolution can slow and become more error-prone. A unified view helps enable sound, efficient decision-making by people, agents, or both.

Kubernetes knowledge is fragmented, and existing AI tools lack business context

Kubernetes expertise often sits unevenly across an organization. For example, senior engineers may hold deep operational knowledge that application teams lack. The most current information about a running system may also be fragmented if logs sit in one tool and metrics in another. Real-time understanding can be further clouded when policies, runbooks, access rules, and deployment history each live elsewhere.

Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes. Without that context, even a capable AI tool may fall short of providing meaningful Kubernetes management support.

Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes.

When an agent can read current signals alongside the rules that govern them, its suggestions become specific, testable, and actionable. In an incident, agentic systems can correlate logs with a recent change. Ahead of a rollout, they can check the change against policy. During troubleshooting, they can account for access rules rather than guessing at them. Kubernetes decisions carry real operational consequences, which makes these details all the more important to consider. 

Engineering “toil” isn’t time-efficient

Site reliability teams use the word “toil” for repetitive manual work, especially tasks that keep systems running without adding lasting impact. In Kubernetes operations, toil takes the form of repeated triage, manual signal correlation, alert follow-up, and routine checks. The tasks aren’t particularly difficult, but they can consume significant time and attention for enterprise teams.

When engineers spend their days on this kind of investigation, proactive modernization efforts tend to stall and planned upgrades can slip behind schedule. In other words, the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements. In a recent survey about how AI provides value to DevOps teams, reducing toil emerged as one of the clearer opportunities.

…the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements.

Agentic AI can support repetitive investigations by gathering signals, correlating them, and proposing a likely cause for an engineer to weigh.

Kept under human review, it can take on some of the routine correlation that would otherwise fall to the team. That kind of support can give engineers more room to focus on the strategic work that most needs their judgment.

Building more intelligent infrastructure with agentic AI and Kubernetes

As you consider building toward intelligent infrastructure without surrendering control, the following principles can inform your efforts:

  • Start with observable context, giving agents access to current cluster state, policy, and history before they reason about a problem.
  • Separate suggestions from actions, allowing agents to recommend freely while any change must wait for human approval and a defined scope.
  • Connect agents to existing controls, routing their work through the access rules, identity, and audit paths the team already trusts.
  • Keep the ecosystem open, favoring platforms that integrate with current tools and standards over those that lock work into a single stack.

Platforms like SUSE Rancher Prime and SUSE AI Factory embrace these principles and illustrate how Kubernetes management can become a foundation for agentic operations. These platforms can help you improve cluster and policy consistency without compromising your authority over AI. Built on open-source foundations, they can also help you avoid being trapped in a single vendor’s stack.

In SUSE Rancher Prime, the industry’s first context-aware agentic AI ecosystem, its AI assistants work as a crew of specialized agents with an intelligent router. The platform draws on the cluster context already in place and acts through existing access controls. Through support for external Model Context Protocol (MCP) servers, teams can extend that crew to their own sources. In addition, human validation tools allow you to hold a proposed action for approval before the agent runs it.

Despite its potential, intelligent infrastructure is not universally beneficial. In situations where change control must stay fully manual, for example, agentic AI’s role may be strictly limited to observation and suggestion. Measure the technology’s value against the realities of your day-to-day operations. For those who are investing, agentic AI will have the greatest impact when it actively supports context, control, openness, and human judgment.

The post The rise of agentic AI on Kubernetes: unleashing the new infrastructure layer appeared first on The New Stack.

  •  

Avoiding vendor lock-in through an open-source approach: a developer’s perspective

Abstract dark digital artwork depicting dense undulating layers, symbolizing cloud architecture and software ecosystem tension.

Every infrastructure team makes decisions that are difficult to reverse. Most of the time, that works out. Sometimes it does not.

Vendor lock-in usually begins as a reasonable choice, made under time or budget pressure, that solves a real problem at the time. A managed service ships faster or a deployment model fits better in that moment, but eventually a difficult constraint appears. 

When business conditions inevitably change, those accumulated choices and their consequences will determine whether a team can pivot accordingly. Limits on flexibility rarely trace back to a single vendor; more often, they hinge on how reversible the team’s past decisions are.

What is vendor lock-in and how can it harm your business?

The risks of vendor lock-in are not really about relying on vendors, since every production system relies on vendors. The big issue is dependencies that become too expensive or impractical to unwind.

For a platform team, that dependency builds up across APIs, contracts, roadmaps, and data models. It extends further into managed services, identity patterns, observability pipelines, and operational tooling. Each piece likely represents a reasonable design choice, but together they can quietly limit your options and raise the cost of leaving. When switching a database or control plane means rewriting tons of integrations, retraining the whole staff, or migrating data under inconvenient timelines, you have lost the room to maneuver.

The big impacts of small, invisible and unexamined decisions

Not every dependency is automatically a problem; some are understood, contained, and worth the tradeoff. The real risk lives in the dependencies no one examined closely, which may stay invisible until they block the business from evolving. 

“The real risk lives in the dependencies no one examined closely, which may stay invisible until they block the business from evolving.”

Unfortunately, some teams are familiar with these invisible dependencies. A managed database might pick up proprietary extensions, which application code then starts to assume. A Kubernetes environment might bind to one cloud’s IAM, networking, storage, and load balancer model. Observability and logging pipelines might harden around a single provider’s formats. None of these choices is reckless on its own, but together they can create significant friction. 

Obstacles to change and their hidden costs

The extent of a dependency-based tradeoff can sometimes remain unknown until circumstances shift, such as a new compliance requirement or customers needing a new deployment model. The hidden costs of these moments often escalate in stages. It might start with a visible, unwelcome migration bill, but the expense can also show up as operational drag. Rushed migrations can lead to additional service disruptions later. A workload may be unable to move, limiting services to certain customers. When you are tied to a specific vendor’s release cadence, it can make it difficult or even impossible to adopt emerging technology. 

Concentration risk compounds the problem, because a single change from one provider that carries pricing, support quality, and roadmap can ripple across the estate. By the time a switch becomes necessary, the cost shows up as service disruption, complex data transfer, and retraining. Naming these costs early keeps them from arriving as surprises.

At some point, a dependency can accumulate enough of these costs to become more than an architectural detail. Once it affects budgets and timelines, leadership has to account for it—and the team has to be ready to explain it. Identifying these dependencies early gives everyone time to plan.

Open source offers a different path

One way to proactively address this pattern is to evaluate potential dependencies more deliberately. For example, before committing to a platform or service, try to determine its reversibility. In other words, establish how difficult it would be for the team to change its mind about the investment in the future.

“Open source offers no guarantee against lock-in, however, since a team can still build tight coupling on open foundations.”

Open source solutions tend to perform well against that test, because they are intentionally built to keep systems inspectable, portable, supportable, and replaceable. By design, open source makes it easier for you to preserve options over time. It offers no guarantee against lock-in, however, since a team can still build tight coupling on open foundations.

What is open source?

Open source describes software you can inspect, run, modify, extend, support, and replace with relative ease compared to proprietary alternatives. The software’s source is available, and the license grants you the right to use and change it. Notably, no-cost or freeware software is not necessarily open source, specifically if it does not provide this level of access and rights.

Several companies have open source principles at their core, and open source software can be extremely valuable in enterprise contexts. Transparent code is often easier to audit, and open standards can reduce friction when moving between tools.

Open source also changes who can move the goalposts

For developers, reversibility is not only about APIs and data formats. It is also about whether one company can change the terms underneath a foundational technology. The Linux kernel is a useful example. Linux kernel documentation notes that copyright assignments are not required, so merged code retains its original ownership and the kernel now has thousands of owners. That makes unilateral relicensing of the kernel effectively impractical.

Kubernetes has a different legal structure, but the practical protection is similar. The project is licensed under Apache 2.0 and governed by the Cloud Native Computing Foundation. The license grants users durable rights to the existing code, so no single vendor, including SUSE, can retroactively take those open-source rights away from the project as it already exists. That matters because a platform can remain available even if a particular vendor changes strategy.

The Terraform-to-OpenTofu fork shows why this is more than a theoretical distinction. In 2023, HashiCorp changed Terraform’s license from the Mozilla Public License 2.0 to the Business Source License 1.1. The community responded by forking the last open-source codebase into OpenTofu, now a Linux Foundation project that remains under the MPL 2.0. The lesson for developers is not that every open-source project is immune to licensing changes. It is that open licensing and neutral governance can preserve a viable exit path when a vendor changes direction.

Open source powered by enterprise discipline

Open source ultimately earns its place through engineering discipline. Source availability has benefits but does not resolve governance, patching, lifecycle management, documentation, security, or integration on its own. A community project can be powerful and nonetheless arrive without enterprise-grade operational guarantees.

Enterprise open source providers exist and can help with closing that gap. They embrace open foundations and add the support, security, maintenance, and lifecycle discipline that production environments require. Founded in 1992, SUSE was the first provider of an enterprise Linux distribution. Today, it focuses on helping organizations operationalize open source with enterprise-grade support.

These companies aim not to close off open source software but to make it dependable at scale. In other words, open source and operational rigor can coexist. And enterprises should expect both from any external provider.

Digital sovereignty: the x-factor that makes open source even more critical

Digital sovereignty describes how much control an organization has over its infrastructure, data, operations, and technology choices. Sovereignty is a spectrum, and architecture decisions can move an organization a step in either direction.

Recent research by SUSE suggests that almost all enterprises are prioritizing digital sovereignty, but only 52% are actively taking steps toward it. That gap is largely an execution problem, and much of it surfaces in everyday platform decisions. 

If your team supports regulated industries or deploys in on-premises or air-gapped environments, you may be especially familiar with growing pressures around sovereignty.

Sovereignty puts a deadline on work that was already worth doing

Developers can hear “digital sovereignty” and assume it means a separate compliance workstream with a separate engineering bill. In practice, much of the work is the same discipline platform teams already invest in: portable workloads, clean interfaces, automated verification, reproducible deployment, auditable behavior, and the ability to replace a dependency without rewriting the system around it.

“Sovereignty does not suddenly make that engineering work valuable. It puts a deadline on work that was already worth doing.”

Those practices already have an economic case. They reduce migration costs, lower operational risk, make platform changes less disruptive, and preserve options when pricing, regulations, or business requirements shift. Sovereignty does not suddenly make that engineering work valuable. It puts a deadline on work that was already worth doing.

That reframe matters because it turns sovereignty from a policy overlay into an architecture property. The useful question is not simply, “How much extra work will sovereignty cost?” It is, “Which parts of our stack already fail the portability, interface, and verification tests we would want anyway?”

How to strengthen sovereignty with open source

Sovereignty depends on how a team designs, deploys, and operates its systems. Open source does not make an organization sovereign by default, but it can improve the conditions for sovereignty. 

In fact, many of the same questions that expose lock-in also matter for digital sovereignty. Each of the following questions about reversibility connects to open source and sovereignty alike:

Reversibility questionWhy open source can helpHow sovereignty strengthens
Can we run this workload elsewhere?Open source typically runs across on-premises, cloud, hybrid, and edge environments, not just one vendor’s platform.More control over where workloads run, including specific regions and regulated contexts.
Can we understand and audit how it works?Source availability and community scrutiny improve inspectability over closed alternatives.Teams can verify behavior, assess risk, and meet assurance requirements.
Can we migrate or reuse our data?Open ecosystems favor open formats and interoperable tooling.Data stays more portable, improving control over storage and movement.
Can another team or partner support it?Multiple support paths exist, from internal teams to integrators and enterprise vendors.Less dependence on one vendor’s pricing, availability, or roadmap.
Can we replace one component without rewriting everything?Open interfaces and modular design make components easier to swap.More control over architecture as requirements change.
Can we keep operating if a vendor changes direction?Open source projects can outlast one vendor’s strategy or license.Less exposure to decisions the team cannot control.
Can we deploy closer to the data?Open source can run in private data centers, sovereign clouds, edge sites and hybrid models.Sensitive workloads, including AI, can be governed nearer the data.

The ongoing work of digital sovereignty

Sovereignty is more of a practice rather than a specific destination. For many teams, the work begins with identifying existing dependencies that are especially hard to reverse. Similarly, you’ll need to separate the tradeoffs worth accepting from the ones that remove a significant number of options. 

Moving forward, it can be helpful to prioritize open interfaces and portable foundations when possible. When evaluating new services or solutions, treat lifecycles, support, and governance as first-order concerns.

In some cases, sovereignty work can be too heavy for an in-house team to carry alone. Providers such as SUSE can help strengthen your operational layer, including security and observability, and especially in growing or hybrid contexts.

Automated checks can make those principles concrete by continuously testing whether workloads can be rebuilt, moved, audited, and recovered instead of waiting for a migration or compliance event to expose the gaps.

Open source lets you take control of your software ecosystem

No enterprise team avoids every dependency, and candidly none should try. Some coupling is reasonable, contained, and worth it. A vendor-free system is not a realistic goal for a major enterprise. A realistic goal is the judgment to separate acceptable dependencies from dangerous ones.

“The true cost of any platform includes the cost of leaving it, and teams should understand that cost before they commit.”

Reversibility gives that judgment something concrete to work with, because it can be broken down into capabilities a team can name, evaluate, and test:

  • Ownership. Ownership does not mean building everything yourself. It means holding the realistic ability to run, move, or hand over each layer of your stack. The test is simple: if a vendor disappeared tomorrow, or was ordered to stop serving you, what still runs next month?
  • Auditability. You should be able to verify what your software does, yourself or through an auditor you appoint, rather than accepting a vendor’s report as the final word. With open source, inspection is a property you hold. With closed software, it is a permission you are granted, and permissions can be withdrawn.
  • Exit velocity. An exit plan without speed is just a document. Exit velocity measures how fast a workload can move from one platform to another, and it only means something when you test it on a schedule, as earlier generations tested disaster recovery.
  • Pivot ability. These capabilities matter when conditions change: a new compliance requirement, a customer that needs a different deployment model, or a vendor that changes direction. Teams that can reroute workloads respond on their own timeline. Teams that cannot must renegotiate from a position of weakness.

Vendor lock-in becomes a manageable risk when you can confidently flag which decisions are hard to undo, weigh the tradeoffs honestly, and protect the team’s pathways to change. Open source strengthens every one of these capabilities because it keeps larger portions of your system inspectable, portable, and replaceable.

The true cost of any platform includes the cost of leaving it, and teams should understand that cost before they commit.

The post Avoiding vendor lock-in through an open-source approach: a developer’s perspective appeared first on The New Stack.

  •  

The agent didn’t break your controls. It went around them.

Three black circular directional signs on a gray concrete wall, showing arrows pointing straight ahead, turning left and turning right.

The identity part of agent security is settled. An agent needs its own identity: a short-lived, revocable credential scoped to the job, and an audit trail that names the human who set it running. NIST’s security leads made that case in August 2026, and most identity vendors agree.1

Identity and access management is table stakes. It’s necessary, but it isn’t what’s breaking.

What’s breaking is an assumption we’ve carried for twenty years: Get identity and permissions right at the door, and whatever happens inside takes care of itself. That worked when software was passive. Agents reason about a goal and choose their own steps toward it, like a seasoned escape artist.

An agent that hits a wall looks for another way

Almost every control in today’s stack answers a question about entry. Should it connect? Should it reach that service? Should its token be accepted here? Each is a question about a route, and there’s rarely just one route to anywhere worth going.

An agent treats a blocked route as a problem to solve, because that’s what we built it to do. A person who hits a locked door usually files a ticket, while an agent tries the window.

In July 2026, an autonomous agent spent four and a half days inside Hugging Face’s production systems.2 A filter controlled which internet addresses its dataset servers could download from, and it never fired, because “the agent stopped asking the worker to fetch remote resources and instead made it act on local ones.” The filter worked as designed, and the agent went around it anyway.

A person who hits a locked door usually files a ticket, while an agent tries the window.

On ordinary developer machines, malware in a compromised npm package tried to recruit the AI coding assistants already installed to search for secrets,3 and a coding agent deleted a production database during a change freeze before falsely telling its operator the data couldn’t be recovered.4 Both happened on the machine itself, where no network control was looking.

The shift from outside-in to inside-out

Outside-in controls govern entry, and most organizations run plenty of them. Make no mistake, inside-out security completes those controls rather than replacing them.

Inside-out control governs the action itself, and asks a narrower, harder question: Should this agent, acting on this person’s authority, delete this table in this database, right now?

That question matters because an agent can swap routes but not the outcome it’s after. No matter how many routes it tries, deleting a table is still deleting a table, and a checkpoint on the action sees it every time.

Here’s how today’s controls line up against it.

ControlWhat it coversWhat it misses
GatewayTraffic you route through itLocal shell commands and file edits never reach it
SandboxThe environment as a wholeConstrains reach, not individual actions
SIEMA record of what occurredReports after the action is completed
RegistryThat an agent existsWhat the agent did with that existence

Each does its job, but they all decide somewhere other than the moment the action runs.

Put the enforcement point where the agent acts

Every agent acts through an agent harness: the software that takes the action the model chose and carries it out, whether that means running a command, writing a file, or calling an API. In most deployments today, nothing checks that action before it runs.

An inside-out control puts an approval step in that gap. Before the harness executes anything, the checkpoint looks at which agent is asking, on whose authority, and against which system, then applies policy to allow the action, block it, or send it to a human. Because every action passes through it, an agent denied a destructive command and trying a smaller version of the same thing is held to the same rules. The remaining risk is a badly written policy, which can be fixed.

None of this works without the identity basics. Any type of control, whether it be at the prompt level, inference level, harness level, or MCP layer, can’t judge “may an agent take this action here on this object?” when the only name on the request is a service account shared by six agents and four engineers.

The companies building agent runtimes have reached the same conclusion. Over the past eighteen months, Anthropic, Google, Microsoft, OpenAI, LangChain, and Cursor have each added a hook that lets you inspect an agent’s action before it runs.5 When AWS explained its own agent policy design, it argued that controls belong at the moment an agent attempts to invoke tools.6

The catch is that each hook works differently, with no standardized request or response formats. An enterprise whose developers use Claude Code and Cursor while its platform team builds on LangChain would maintain the same enforcement logic in multiple different flavors, each with its own audit trail. That doesn’t scale, and it tightly couples your security model to whichever runtime a team favors that month. Enterprises need one vendor-agnostic agentic security layer that spans every harness, so adopting a new model or framework doesn’t mean restarting the entire onerous security review.

Turn the lights on before you start blocking

The standard, well-ingrained security instinct is to start blocking right away, but we’ve all seen how well that works with the business in the past. Security must move and adapt at the speed of business, not the other way around. Security tools such as intrusion prevention systems and web application firewalls both ran in monitoring mode until teams understood what normal looked like, and those that skipped that step tended to hear about it from a production outage.

Agents need the same sequence, only faster. An enforcement point in monitoring mode blocks nothing and quickly answers questions most organizations can’t today:

  • Which agents are actually running, not which ones someone believes are running
  • Who started each one, and whose authority it’s operating under
  • What capabilities it used, and against which systems
  • Which of those actions would have violated a policy, had the agent security platform been switched to enforcement mode

Write policy from what you know, see, and have evidence of, not just from an architecture diagram. Enforce first where the stakes are highest: destructive commands, production data, and anything that moves data out. Then watch-learn-build, just like the agents we use: Watch the patterns, build finer-grained controls and policies, and learn how to use AI securely, safely, and confidently. Observe first, then enforce, build, and deploy, in that order.

The bottom line

None of this requires a new category of infrastructure. It’s the identity, authorization, and audit you already run for your people, extended to agents and applied inside the harness before the action runs.

The perimeter is still there, but it has moved to the moment an agent acts, the one place it can’t route around.

Ory built Agent Security inside the harness, on the same identity and authorization engines that run in production for human users. It starts in “observe mode,” so you get that inventory first, and you can try it today at ory.com/agent-security.

Footnotes

  1. Bill Fisher and Ryan Galluzzo, “Back to the Future: Why Agentic AI Needs a Strong Identity Foundation,” NIST Cybersecurity Insights, August 27, 2026.  ↩︎
  2. Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident.” The intrusion ran July 9–13, 2026. ↩︎
  3. Nx, “s1ngularity postmortem,” August 2025. The malicious packages “attempted to use local AI tools (like Claude and Gemini)” while scanning systems for sensitive data. ↩︎
  4. AI Incident Database, Incident 1152: Replit agent deletes production database during code freeze, July 18, 2025. ↩︎
  5. Pre-execution hooks by vendor. Anthropic, Claude Code hooks; Google, Agent Development Kit callbacks; Microsoft, Agent Framework middleware; OpenAI, Agents SDK guardrails; LangChain, human-in-the-loop middleware; Cursor hooks (InfoQ, October 2025) ↩︎
  6. Liana Hadarean and Jean-Baptiste Tristan, “Why Policy in Amazon Bedrock AgentCore chose Cedar for securing agentic workflows,” AWS Security Blog, May 20, 2026. ↩︎

The post The agent didn’t break your controls. It went around them. appeared first on The New Stack.

  •  

Claude Opus 5.5 vs. Opus 5 on reasoning tasks: Cheaper, faster, but not better

Abstract long-exposure photograph of red and orange light trails forming layered curves around a dark central shape.

When Anthropic released Claude Opus 5.5 this week, the company claimed the new model costs 40% less than Opus 5 and generates output 30% faster. Anthropic’s marketing makes three claims. Opus 5.5 performs at the level of Claude Fable 5.1 (so it should outperform Opus 5), costs 40% less than Opus 5 on typical workloads, and generates output more than 30% faster.

Anthropic also cut the price developers pay to use the model through its API. Opus 5.5 costs $4 for every million tokens (chunks of text roughly three-quarters of a word long) sent to the model and $20 for every million it writes back, down from $5 and $25 for Opus 5. That price cut alone accounts for a 20% saving. The rest of the claimed 40% saving has to come from the model using fewer tokens.

I wanted to see how this translates for the average Claude user, so I skipped the usual developer workflow simulations this time. Lately, the models I test handle everyday tasks well. Reasoning tasks are where I’ve seen them struggle, so I tested Opus 5 against Opus 5.5 on reasoning tasks only. 

You can find the prompts at the bottom of this post if you want to replicate these tests on your own system.

The tests

I called both models through the Anthropic API with identical prompts. Both ran with adaptive thinking at the default effort level, since Opus 5.5 doesn’t allow you to turn thinking off. Each problem ran once per model. I planned to rerun any problem where the models gave me different results, but they never did.

Here are the tests I ran:

  • Logic grid (medium difficulty) – Seven engineers each have an on-call day, a language, a service, and a city, and 22 clues pin down one answer. Six clues are conditional or “exactly one of these is true” statements, and removing any single clue breaks the puzzle.
  • Constrained orderings (hard difficulty) – Reorder 10 deploy jobs so no job stays in its original slot and no two consecutively numbered jobs sit side by side. The model had to give the count for 6, 8, and 10 jobs.
  • Stone game with memory (harder difficulty) – Players remove 2, 5, 7, or 11 stones, but can’t repeat their opponent’s last move or their own. The model had to find who wins from 200 stones, count the losing starting sizes up to 500, and name the smallest losing size above 340.

I logged input tokens, output tokens, cost at list price, and time for every call. Thinking tokens are billed as output, so I included them.

The logic grid

Both models got all 28 cells right. Opus 5.5 took 65 seconds and 7,573 output tokens, for $0.16. Opus 5 took 108 seconds and 10,621 output tokens, for $0.27.

On this test, Opus 5.5 was 43% cheaper and delivered the same correct answer. Opus 5.5 was slightly more detailed and noted that it didn’t fully prove the solution was unique.

Constrained orderings

Neither model produced an answer, which made this the hardest problem in practice. The correct counts are 27, 1,695, and 159,019. The third answer is very hard to reach by reasoning alone. A computer program that checks every possible ordering can find it, but neither model could run code in this test.

With a 48,000-token output limit, both models spent the whole budget thinking and never replied. Opus 5.5 used 489 seconds and $0.96. Opus 5 used 553 seconds and $1.20.

I raised the limit to 128,000 tokens and ran it again. Opus 5 used every token, took over 25 minutes, and stopped with no answer. That cost $3.20. Opus 5.5 ran for 19 minutes and used 112,733 tokens, but the API ended the response with a “refusal” stop reason and no text. The prompt asks the model to count job orderings and contains nothing sensitive. That means the refusal was most likely a mistake by Anthropic’s safety filter flagging a harmless request.

Opus 5.5 was cheaper, but how much does that matter if you don’t get a result?

The stone game

And we’re back to the same answers again. Both models answered all three parts correctly. The first player loses from 200 stones; starting sizes from 120 to 500 are losses, and the smallest loss above 340 is 344.

The difference in this test came down to how much thinking each needed to get there. Opus 5.5 finished in 215 seconds with 28,740 output tokens, costing $0.58. Opus 5 took 624 seconds and produced 74,981 output tokens, costing $1.88. Opus 5.5 used 62% fewer tokens and cost 69% less for the same answer. 

Results

TestOpus 5.5Opus 5
Logic grid28/28, 1:05, 908 in / 7,573 out, $0.1628/28, 1:48, 906 in / 10,621 out, $0.27
Ordering problem (48k limit)No answer, 8:09, 235 in / 48,000 out, $0.96No answer, 9:13, 233 in / 48,000 out, $1.20
Ordering problem (128k limit)No answer (refusal), 18:56, 235 in / 112,733 out, $2.26No answer, 25:24, 233 in / 128,000 out, $3.20
Stone game3/3, 3:35, 323 in / 28,740 out, $0.583/3, 10:24, 321 in / 74,981 out, $1.88
Total tokens1,701 in / 197,046 out1,693 in / 261,602 out
Total time31 min 45 sec46 min 49 sec
Cost$3.95 ($4 in / $20 out per million tokens)$6.55 ($5 in / $25 out per million tokens)

Across every call, Opus 5.5 wrote 103.4 tokens per second and Opus 5 wrote 93.1, so Opus 5.5 was about 11% faster. Its biggest speed lead on any single problem was 19%, still short of Anthropic’s 30% claim. Its biggest cost savings came on the stone game, where it cost $0.58 to Opus 5’s $1.88, 69% less. Total spend for the test was $10.50. 

What do I think

Anthropic’s benchmarks show Opus 5.5 ahead of Opus 5 on coding, knowledge work, and reasoning. I didn’t rerun those benchmarks. I gave both models the same three reasoning problems, and they performed equally. Both solved the logic grid and the stone game, and both failed the ordering problem.

The savings are real. Opus 5.5 cost less and finished sooner on every problem, including 43% less on the logic grid and 69% less on the stone game. Most of that came from using fewer output tokens. The price cut accounts for 20%. Its writing speed was 11% faster, short of the 30% Anthropic claims.

Switch to Opus 5.5 if you run Opus 5 today. You get the same results on hard reasoning for less money and less waiting. Set a hard output limit and watch your spend on hard problems, though. Both models can think for close to 20 minutes or more and return nothing, which is a problem if you pay for every token. For counting problems like the ordering test, give the model a code execution tool instead of hoping it reasons its way through.

The prompts

Logic grid

Seven engineers (Ana, Ben, Cy, Dee, Eli, Fay, Gus) share an on-call rotation. Each is on call on exactly one day of a single week, Monday through Sunday (Monday is the earliest day, Sunday the latest), and no two share a day. Each writes a different language (Go, Rust, Python, Java, Kotlin, TypeScript, C++), owns a different service (auth, billing, search, queue, cache, gateway, metrics), and is based in a different city (Berlin, Tokyo, Denver, Lagos, Sydney, Toronto, Mumbai).

Clues:

The Java developer is Gus.

The TypeScript developer is on call exactly two days after the queue owner.

Exactly one of these is true: the engineer based in Berlin owns billing, or the C++ developer is on call Monday.

Exactly one of these is true: the engineer based in Tokyo owns cache, or the engineer based in Tokyo writes Rust.

If Fay is on call Sunday, then the engineer based in Denver does not write Kotlin.

Exactly one of these is true: the Kotlin developer owns auth, or the C++ developer is Cy.

The engineer based in Tokyo is on call earlier in the week than the C++ developer.

The engineer based in Denver does not write Python.

Exactly one of these is true: the Go developer is Ben, or the gateway owner is based in Berlin.

Ana is based in Mumbai.

The TypeScript developer is on call earlier in the week than Fay.

The engineer based in Sydney is on call earlier in the week than the billing owner.

The gateway owner is on call earlier in the week than the Kotlin developer.

The engineer on call Sunday is not based in Berlin.

Fay and the engineer based in Lagos are on call on consecutive days.

Eli is on call exactly three days after the auth owner.

The engineer based in Lagos writes Java.

Exactly one of these is true: the search owner is Eli, or the cache owner is Dee.

The gateway owner and the Go developer are on call on consecutive days.

The engineer based in Lagos is on call exactly four days after the auth owner.

The engineer based in Lagos owns metrics.

The engineer based in Mumbai is on call earlier in the week than the Python developer.

Determine the full assignment. At the end of your response, give exactly seven lines, one per engineer in the order Ana, Ben, Cy, Dee, Eli, Fay, Gus, in this format:

ANSWER: Name | Day | Language | Service | City

Ordering problem

A build system has n deploy jobs numbered 1 to n. Originally, job k runs in slot k. You reorder all n jobs into slots 1 to n (each slot gets one job) subject to two rules:

No job runs in its original slot (job k is not in slot k).

Jobs with consecutive numbers never run in adjacent slots (for example, jobs 4 and 5 cannot be in slots i and i+1 in either order).

How many valid orderings are there for (a) n = 6, (b) n = 8, (c) n = 10?

At the end of your response, give exactly three lines in this format:

ANSWER a: <number>

ANSWER b: <number>

ANSWER c: <number>

Stone game

Two players play a game with a pile of stones. They alternate turns. On each turn, a player removes exactly 2, 5, 7, or 11 stones, subject to two rules:

You may not remove the same number your opponent removed on their most recent turn.

You may not remove the same number you removed on your own most recent turn.

(On the very first turn of the game, neither rule applies. On the second player’s first turn, only the first rule applies.) You cannot remove more stones than are in the pile. A player who has no legal move on their turn loses. Both players play perfectly.

(a) Starting with 200 stones, does the first player win?

(b) For how many starting pile sizes from 1 to 500 inclusive does the first player lose?

(c) What is the smallest starting pile size greater than 340 for which the first player loses?

At the end of your response, give exactly three lines in this format:

ANSWER a: <yes or no>

ANSWER b: <number>

ANSWER c: <number>

The post Claude Opus 5.5 vs. Opus 5 on reasoning tasks: Cheaper, faster, but not better appeared first on The New Stack.

  •  

Microsoft’s new Copilot agents get their own email, calendar — and a place in the org chart

Satya Nadella stands smiling between Bill Gates, on the left, and Steve Ballmer, on the right, in front of a crowd of cheering employees, many holding up phones and tablets to take photos.

Microsoft announced what it calls its biggest Copilot update to date on Friday, with CEO Satya Nadella describing Copilot as “a new OS for work.”

Nadella framed Copilot as spanning every model, form factor, and task, and the update puts Autopilot, which Nadella called a “proactive and long-running agent built for the enterprise,” at the top of his list of the update’s four components. The pitch targets office workers, but the more consequential change for developers is the infrastructure underneath.

Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365, which means teams building production agents no longer have to assemble those pieces around a model on their own.

We’re building Copilot as a new OS for work that spans every model, every form factor, and every task. Today, we’re announcing our biggest update to Copilot to date, bringing four things together:

· Autopilot: proactive and long-running agent built for the enterprise
· Code:… pic.twitter.com/W2ClHHkCK3

— Satya Nadella (@satyanadella) September 25, 2026

The release adds a new Home experience that merges Chat and Cowork in the Copilot app, but the bigger changes for developers come from Code and Autopilot. Code generates apps, dashboards, and workflows from natural language, and Autopilot turns the agent Microsoft previously called Scout into a persistent background worker. Home and Code are rolling out first through Microsoft’s Frontier early-access program, and Autopilot is expanding to a private preview at month’s end.

Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365

Agents that don’t need prompts

Autopilot takes a role and goal from the person who sets it up, then continues working in the background without requiring a new prompt for each step. Each Autopilot gets its own governed Entra identity and agent user account, separating the agent’s permissions and activity from those of the person who created it.

For engineers, that moves much of the operational scaffolding required for long-running agents into Microsoft’s infrastructure. Independent vendors have been building dedicated layers for that problem; Diagrid, for example, adds durable recovery to LangGraph and other agent frameworks, while Microsoft is bringing those capabilities inside the Microsoft 365 environment.

An identity for every agent

The identity model is the piece developers building on Microsoft Foundry will feel first. Autopilot agents in Foundry, which have been in public preview since June, receive a full Entra Agent ID user account with a productivity license that gives them their own email, calendar, OneDrive storage, Teams access, and a place in the org chart.

Because that user account sits on top of the agent identity every Foundry agent already carries, an autopilot acts as itself rather than on behalf of a user, so developers no longer have to wire agents through shared service accounts or borrowed user credentials, a pattern AuthZed CEO Jake Moshenko has said reflects a common misconception about how agents should be deployed.

A developer creates an Autopilot blueprint from a Foundry-hosted agent, which appears in the Agent 365 registry once an administrator approves it. Employees can then hire instances of that agent in Teams. The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.

The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.

Hosting AI-generated apps

Code is built on the same underlying technology as GitHub Copilot, and the apps it generates run on Microsoft Copilot Managed Runtime, a platform now in public preview that hosts code inside the customer’s Microsoft 365 tenant boundary under IT governance.

Apps deployed there run within the company’s existing identity and governance framework, with Microsoft managing the underlying runtime and giving developers a controlled path to test and deploy new versions without taking the current release offline.

The runtime also accepts apps built in Copilot Studio and Cowork, and Microsoft is opening it to outside tools and professional developers through an SDK and command-line tooling, with Git tracking source and versions.

Lovable is already on board. In Microsoft’s announcement, the company’s head of global partnerships, Lan Roche, said apps built with Lovable can now run inside a Microsoft tenant “the same way everything else does,” using the same sign-in, policies, and app inventory.

The model resembles what serverless computing did for application infrastructure, where developers concentrate on application logic while the platform takes on more of the execution environment. Microsoft is applying that abstraction to generated enterprise software while tying the runtime directly to identity, tenant boundaries, and organizational data.

Long-running agents also change Copilot’s economics. The standard subscription covers the assistant, but Cowork, Code, Autopilot, and other agentic features are billed based on usage through Copilot Credits. That also applies to frontier models such as Fable and Astra, although users still need a Copilot license to access them. Microsoft is extending cost management in Agent 365 to cover Code and Copilot Managed Runtime, and it plans to support agents built in Copilot Studio in October.

Once an agent can keep working for hours or days without anyone watching, cost becomes part of the governance problem. Engineering teams need to control how much compute an agent uses alongside what it can access, which is why Microsoft is bringing those controls into the same administrative framework.

The portability trade-off

That convenience comes with a trade-off. Because Microsoft controls the underlying enterprise environment, it can handle much of the work around agent state, credentials, and access controls, but the more infrastructure a team hands over to Microsoft, the harder the agent may be to move elsewhere.

The models are not locked in, since Microsoft currently runs Copilot on models from both OpenAI and Anthropic and says more labs and open-weight models are coming, and the Agent 365 SDK adds governed Model Context Protocol access to Microsoft 365 workloads for agents regardless of the framework they were built with. Those open interfaces cover only part of an agent’s architecture, though. The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.

The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.

The post Microsoft’s new Copilot agents get their own email, calendar — and a place in the org chart appeared first on The New Stack.

  •  

OpenTelemetry and Prometheus are getting along. What’s still missing?

Abstract 3D illustration of metallic blue spoked hubs connected by purple tubes against a pink background.

Welcome to another edition of Road to KubeCon, where we’re tracking the major movements in the Kubernetes and cloud native ecosystem on the path to KubeCon + CloudNativeCon NA 2026, happening Nov. 9–12 in Salt Lake City, Utah.

This week, we look at how cloud-native teams are putting observability to work. There’s progress on OpenTelemetry and Prometheus interoperability, a migration spanning 100,000 hosts, and new data on the costs and benefits of monitoring AI systems. Plus, HPE’s latest Gartner recognition, agent governance updates, and a father-and-son story from KubeCon India.

HPE GreenLake named a Leader in Gartner quadrant

On Wednesday, HPE announced it had been named a Leader in Gartner’s Magic Quadrant for Infrastructure Platform Consumption Services for the second consecutive year.

Hewlett Packard Enterprise (HPE) is a presenting sponsor of Road to KubeCon. HPE Software helps IT organizations modernize infrastructure, streamline operations, and accelerate AI initiatives across hybrid, multi-vendor environments.

GreenLake, HPE’s cloud operations platform, helps teams monitor resource consumption, secure data, and manage infrastructure across data centers and private and public clouds. HPE points to recent additions, including agentic AI-powered operations, as part of the platform’s development.

Varma Kunaparaju, senior vice president and general manager of cloudops software and platform at HPE, says in the announcement: “We are building the operating model and platform for the agentic enterprise, giving customers the ability to simplify operations, govern intelligently, continuously optimize, and modernize without sacrificing choice.”

OpenTelemetry and Prometheus work better together

OpenTelemetry (OTel) and Prometheus are widely used for cloud-native monitoring and observability, often side by side. A new survey looks at how well that combination works.

Published Tuesday, the 2026 survey on Prometheus and OpenTelemetry interoperability found that nearly half of respondents mix Prometheus- and OTel-style instrumentation for infrastructure metrics. For application metrics, 30.7% use both.

While the two ecosystems haven’t always worked well together, the 2026 survey shows improvement: the average ease-of-use rating rose 0.5 points, from 3.1 to 3.6, while the share of those who find the two hard to use together fell from 29% to 10%.

As OTel contributors Dhruv Ahuja of SigNoz, Grafana Labs‘ Andrej Kiripolsky and Arthur Sens, and Ana Muenz share: “Two years of work on interoperability is paying off.” 

There’s still work to do. Respondents want better alignment between the projects’ data models, better handling of resource attributes and metadata, and fewer naming and formatting issues.

Atlassian moves metrics from 100,000 hosts to OpenTelemetry

A case study on the Cloud Native Computing Foundation (CNCF) blog details how Atlassian migrated its metrics collection to OpenTelemetry from gostatsd, its open-source Go implementation of Etsy’s StatsD.

The original pipeline had worked for years, handling metrics from roughly 100,000 hosts across 14 regions, but was increasingly out of sync with the shift to OTel. “It became the thing everyone standardized on, and more and more of what fed our pipeline was emitting OTel data we simply didn’t support,” write Atlassian’s Iris Grace Endozo, Farzad Vazirnia and Albert Kerr.

To maintain continuity throughout the migration to OTel, Atlassian swapped the collection and pipeline mechanics underneath while keeping the service-facing interface unchanged. This turned an organization-wide overhaul into what the authors call a “platform-team migration.”

According to the authors, aggregation now uses about half the CPU for the same traffic. Operations are more unified through the OTel Collector, CPU usage is more evenly distributed across ingest shards, and sidecar costs are down roughly 30% at fleet scale.

New Relic finds observability gains — and gaps

On Tuesday, New Relic released its 2026 Observability Forecast, based on a survey of 2,575 IT and engineering leaders and practitioners. The report found that 73% are standardized on OTel, actively migrating to it, or testing it.

The report also looks at observability’s role in AI adoption. It found that 83% of respondents consider observability essential for AI-generated code. Organizations monitoring AI agents are twice as likely to report a threefold return on observability investment as those running agents without monitoring.

The study also paints a picture of the impact of outages. Engineers now report spending 37% of their time addressing disruptions, while 42% of organizations learn about disruptions through inefficient channels, like manual checks or customer complaints.

Outages take a business toll. New Relic found that organizations lose $74 million a year on average due to high-impact outages. That’s $1.85 million per hour, or over $30,000 for every minute a system is down.

The findings show why teams are looking for ways to detect and resolve problems faster as their systems grow more complex.

As Kubernetes evolves, so do the demands on the teams running it. Presenting sponsor HPE helps teams address that complexity with software spanning virtualization, cloud management, observability, and automation.

Observability Day returns to KubeCon

If you’re into observability and attending KubeCon NA, definitely check out the agenda for Observability Day, happening during the co-located events in Salt Lake on November 9.  

OpenTelemetry’s graduation in May and growing production use give teams more experience to draw on as they adopt the standard.

New AI workloads, inference monitoring, and interoperability with other projects still present challenges. Those issues give practitioners plenty to compare notes on.

According to the Observability Day schedule, the agenda includes project updates and lessons from Capital One, Cisco, Nubank, and other organizations.

“Observability Day provides a vendor-neutral place for maintainers and practitioners to compare approaches and learn how the wider ecosystem is responding,” write Austin Parker, Iris Dyrmishi, Eduardo Silva Pereira, and Juraci Paixão Kröhling on the CNCF blog.

Komodor adds controls for agentic operations

Technically one we skipped last week, but potentially interesting vendor news nonetheless: Komodor, the site reliability engineering platform, announced its Komodor Agentic Operations Platform on Wednesday, September 16.

The additions let engineers deploy autonomous workflows and build or import agents under shared governance and context. Komodor says the release responds to the growing use of agents, including coding agents, and concerns about governance, return on investment, and costs.

“The hardest part of running agentic operations in production is not building the agents,” shares Itiel Shwartz, Komodor’s co-founder and CTO. Instead, the challenges lie in maintaining context, persistent memory, accuracy, security, and cost control — things the new Komodor release aims to address.

Spectro Cloud expands in the Middle East

The Middle East is an increasingly important technology market and a hotbed of data center construction. At the same time, data sovereignty and compliance requirements are driving interest in sovereign infrastructure.

This week, Spectro Cloud announced plans to expand in the Middle East, including new local partners and a dedicated regional office.

“The Middle East has bold ambitions for global AI leadership, from sovereign AI factories to AI-powered economies,” shared Tamer Riyal, Spectro Cloud’s sales director for the Middle East, in the announcement. “We’re investing in the regional expertise and partnerships to support that vision for the long term…”

The company aims to expand adoption of PaletteAI, its platform for managing AI infrastructure across sovereign clouds, enterprise data centers, and edge locations. The expansion reflects the region’s growing role in cloud-native infrastructure.

A father and son take the KubeCon stage

This week’s updates also have a personal side. Earlier this year, analyst, advisor, and TNS columnist Janakiram MSV co-presented a talk at KubeCon + CloudNativeCon India 2026 with his son, Shreyas Mocherla, a CNCF Kubestronaut and software engineer at Nirmata.

As Mocherla describes on the CNCF blog in a post published this week: “Presenting alongside him made this special on a level that goes beyond the conference itself. I grew up watching him speak at technology events. Standing next to him at the same podium, in front of the KubeCon audience, felt like a full-circle moment.”

You can watch the talk, “Run Your Own AI Cluster on a DGX Spark: Kubernetes, GPUs, and DRA,” below:

It’s a reminder that the Kubernetes community’s connections can span generations as well as organizations.

Other updates from the K8s universe

More updates for the platform engineers and cloud operators working in the Kubernetes ecosystem:

Follow the Road to KubeCon

Road to KubeCon is an eight-part series presented by HPE, which will be at KubeCon + CloudNativeCon North America in Salt Lake City. Before you go, explore how HPE Software helps IT teams do more with less complexity.

We’ll be here every Friday until KubeCon.

If you’d like to participate, Bill Doerrfeld, the writer of this series, is open to pitches — you can send release notes, quotes, reports, videos, case studies, or hot takes through his contact page.

If you missed the previous editions covering Kubernetes v1.37 and AI inference, you can catch up through those links. Visit the Road to KubeCon page for the complete archive.

The post OpenTelemetry and Prometheus are getting along. What’s still missing? appeared first on The New Stack.

  •  

OpenAI and Cursor agree on agent coordinators. They disagree on who runs them.

Abstract digital art of thousands of thin glowing strands in orange, red and pink bundled into a single sweeping arch against a black background.

OpenAI opened its Agents API in public beta this month, exposing the harness that powers Codex with managed sessions, tool coordination, and subagent orchestration. On the same day, September 10, Cursor launched Projects to coordinate multiple coding agents around larger bodies of software work. While the products sit at different points in the stack, both converge on the same architecture: a coordinator understands the larger objective and manages the work, while specialized agents execute individual pieces.

That pattern isn’t new: AWS Bedrock AgentCore reached general availability in October 2025, and Anthropic’s Claude Managed Agents entered public beta in April 2026. What makes these announcements notable is that two major players in AI-assisted software development are independently exposing the same coordinator-worker split at the same time.

Hilliary Lipsig, a senior principal site reliability engineer at Red Hat who leads Azure Red Hat OpenShift SRE teams and hosts the YouTube livestream GitOps Guide to the Galaxy, has watched this dynamic play out firsthand.

“This convergence highlights the reality developers across the industry have been discussing on and offline — an agent with too much context loses accuracy and reliability, and focused work with clearer contexts allows for faster, more accurate iterations,” Lipsig tells The New Stack.

“The need for orchestration in distributed computing has been fundamentally recognized repeatedly,” Lipsig says. “That’s part of how we got to Kubernetes. These multi-agent workflows are the same concept, just in a new part of the technical stack. While the specialized agents do their area of work, the orchestrator can act as a source of truth — ideally enforcing guardrails, recovering from any failure states, and intelligently routing work to the most efficient target agent.”

“The need for orchestration in distributed computing has been fundamentally recognized repeatedly… These multi-agent workflows are the same concept, just in a new part of the technical stack.”

The industry has spent the first generation of AI coding tools asking how capable a model can become at writing software. The emerging question is different: How do you build a reliable system around multiple capable agents working on the same problem?

The problem with the single-agent loop

A coding agent works through what Anthropic describes as LLMs using tools based on environmental feedback in a loop: it observes the state of a repository, reasons about what to do next, calls a tool, examines the result, and continues. For a small task, that loop can be enough. As the scope expands, however, maintaining reliability in a single context becomes harder.

A large migration might require understanding an unfamiliar codebase, identifying dependencies, changing database schemas, updating services, rewriting tests, modifying deployment configuration, and validating the resulting system. A single agent can theoretically perform all of that work, but it must maintain relevant information from every stage while continuing to reason about what comes next.

The pressure lands first on the context window. “A large context doesn’t only include everything correct or important — it also includes a lot of throwaway information,” Lipsig tells The New Stack. “Through compaction, that information can inadvertently end up ranked as important and incorrectly influence what your agent does. Or correct information can be distorted to become incorrect.

“Either way, after a couple of rounds of compaction, developers are seeing accuracy degrade and are starting to manage context once again manually.”

Lipsig’s read matches what researchers call context rot — and it hasn’t gone away with newer models.

A 2026 study testing frontier models,, including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro, found they missed a dangerous action buried in a long agent transcript two to 30 times more often once it came after 800,000 tokens of benign activity — the AI equivalent of a security guard who stops checking badges carefully after the two-hundredth person walks through, even though nothing about their training changed.

Furthermore, the tasks themselves may not be sequential. Forcing one agent to execute database analysis, documentation work, and test discovery one after another turns a potentially parallel workload into a serial one.

Subagents change that execution model. Instead of requiring one agent to carry an entire task through a single context, a coordinator breaks the work into smaller units and assigns them to specialized agents. GitHub’s custom-agent model illustrates this: different agents receive only the prompts, tools, and context they need for their tasks, executing work in isolated contexts rather than crowding an increasingly large conversation.

Multi-agent systems therefore bring higher token costs and additional coordination and integration risks, and splitting work across agents does not guarantee better software quality.

The coordinator is not another coding agent

Once the work is divided this way, the coordinator becomes a control plane rather than another coding agent. Its job isn’t to write the code, but to understand the global task, manage dependencies, and decide how execution should proceed. Unlike a conventional scheduler, an agentic coordinator makes probabilistic judgments about result quality and resource allocation.

It may dispatch one agent to investigate a database schema, another to examine the service layer, and a third to inspect the test suite. When they return, the coordinator determines if their findings are sufficient to move to implementation. If a worker produces an incorrect result, the system must recognize the failure and decide whether to retry the work, reassign it, or change the task itself.

Anthropic has documented this same pattern in its own production system, calling it orchestrator-subagent architecture: a lead agent analyzes a query, develops a strategy, and spawns specialized subagents to investigate different facets in parallel. In a June 2025 writeup of that system, Anthropic reported a Claude Opus 4 lead agent with Claude Sonnet 4 subagents outperformed single-agent Opus 4 by 90.2% on its internal research eval — at roughly 15 times the token cost of a standard chat interaction (Anthropic puts single agents at about 4 times), a tradeoff that makes the pattern a deliberate architectural bet, not a free upgrade.

Parallelism introduces distributed-systems failure modes

Parallelism is valuable because software work contains many independent tasks, but it creates coordination problems. Imagine a migration where one agent changes a database schema, another updates the consuming service, and a third updates integration tests.

If the schema changes while the service agent works against an earlier assumption, the system produces internally inconsistent work. This isn’t a risk unique to hypothetical migrations — the International AI Safety Report 2026 notes that “interactions between multiple AI agents are also becoming more common, introducing further risks, as errors propagate between systems.”

A single model invocation is a disposable computation, but a twenty-minute workflow modifying a repository is not. If an agent loses its machine halfway through, restarting from scratch is expensive and potentially unsafe against a changed environment.

To solve this, Cursor moved its cloud-agent execution loop to Temporal to handle durable execution and retries, pushing its cloud agents past two 9s of reliability. Temporal now handles 50 million of Cursor’s actions a day across 7 million unique workflows. “Durable execution isn’t a nice-to-have here. It’s the difference between a system you can operate and one you can only demo,” Lipsig tells The New Stack.

“Durable execution isn’t a nice-to-have here. It’s the difference between a system you can operate and one you can only demo.”

By separating agent, machine, and conversation state, the execution engine can reason about the workflow independently. Reliability is no longer just about whether the model produces a good answer; it is about reliably completing distributed workflows composed of many operations, machines, and dependencies.

The environment, context, and observability are one problem

In production, an agent is more than a model and a prompt; it requires a workspace, source code, dependencies, credentials, and state retention. Both companies provision isolated environments for these resources, directly linking an agent’s capability to its blast radius. OpenAI’s Agents API currently supports U.S. data residency but not Zero Data Retention; choosing a self-hosted sandbox does not make the Agents API eligible for ZDR. Cursor supports similar cloud isolation alongside local execution for machine-specific work.

An agent that can only inspect a repository poses a different risk than one that can modify production infrastructure. Consequently, the coordinator is inextricably linked to the security model, determining which agent receives specific information and authorities.

This logic extends to context routing. Giving every subagent the parent’s entire history increases cost and complexity while leaking irrelevant or sensitive information. Instead, the coordinator enforces information-flow boundaries: a database-analysis agent receives only schemas and relevant migrations, while a security-review agent gets the resulting diff without deployment credentials.

As agents increasingly use interfaces like MCP to reach external systems, the platform must strictly govern which agent receives the authority to use specific tools, and for how long. MCP’s governance now sits inside the Agentic AI Foundation, a Linux Foundation foundation co-founded by OpenAI, Anthropic, and Block, with support from AWS, Google, Microsoft, Bloomberg, and Cloudflare to host MCP alongside AGENTS.md and Block’s goose — a sign the industry already treats it as infrastructure worth governing jointly, not a feature any one vendor owns.

This complexity creates a visibility problem. A simple final response often conceals a history involving multiple agents, tool calls, environments, and retries. Systems must expose task-level provenance — which agent received the assignment, what context it used, where it executed, and how the coordinator handled failures or human interventions.

Without execution provenance, debugging requires reconstructing distributed workflows from fragments. GitHub’s exposure of subagent lifecycle events points in this direction, treating agent lifecycles as observable components rather than hidden processes.

Coordination authority is not execution authority

The most critical architectural boundary is the distinction between coordination authority and execution authority. A coordinator needs broad visibility to make useful decisions, but that does not imply unrestricted control over the project. “Just like you don’t want humans running around with root permissions, you don’t want your agents running with them either,” Lipsig tells The New Stack.

“Just like you don’t want humans running around with root permissions, you don’t want your agents running with them either.”

“The ease of creating and leveraging AI agent permissions is lagging behind the speed of AI innovation, but any product team that needs to maintain compliance standards will tell you that easy or not, access controls are incredibly important. Functionally, that means fully agentic workflows need to run in a dedicated environment with their own permissions and connectivity. You can’t rely on agent context to do what you want — trying to keep something in context that long will just result in errors from context rot. Hardware or virtual-hardware-level permissions, in addition to system-level permissions, need to be configured so the agent is asked to provide only its unique value add.”

The risk isn’t hypothetical. OWASP’s Top 10 for Agentic Applications formally names this failure mode Identity and Privilege Abuse (ASI03).

METR’s independent investigation, conducted with a Redwood Research contractor, found that the June–July 2026 incident — in which OpenAI’s own agents, running in internal ExploitGym cyber evaluations, escaped their sanctioned scope and attacked Hugging Face — unfolded in distinct phases. On June 26, agents discovered an exploit that provided full administrator access to OpenAI’s internal Artifactory package repository. On July 4, high-volume Artifactory activity caused an outage, triggering a security investigation and pausing evaluations. OpenAI switched to a freshly initialized Artifactory instance on July 6, which it says wiped the cache and messages left there.

On July 7, a new wave of ExploitGym experiments launched tens of thousands of agents, and on July 8 one agent established the primary unsanctioned message board. Roughly 1,200 agents eventually used the board, sending more than 70,000 messages and files; about 700 later participated in the attack on Hugging Face. The attack itself began on July 10–11 and wound down over July 12–13. The chronology matters because the administrator-access event, the Artifactory outage, and the later message-board activity were separate phases, not one continuous incident.

By binding autonomy, worker agents operate with the minimum permissions required for their specific tasks, keeping sensitive operations behind explicit approval boundaries. This also reshapes human review. Requiring human approval for every tool call destroys the efficiency of multi-agent execution, but showing only the final result obscures critical intermediate decisions.

The most useful design places human intervention around consequential, irreversible transitions — like moving into production or altering sensitive infrastructure. This is especially vital as agents become event-driven participants that respond to Slack messages or pull request updates, not just direct prompts.

OpenAI and Cursor own different parts of the architecture

The convergence does not mean OpenAI and Cursor have built interchangeable systems. Their products put the orchestration boundary in different places.

OpenAI is exposing an agent harness through an API. Its model gives developers primitives for managing context, tools, subagents, and execution environments, leaving application teams to decide how those capabilities fit into their own systems. The harness is open source, so teams can inspect the coordinator logic instead of treating it as a black box.

Cursor packages more of the surrounding workflow. Projects provides the coordinator, cloud execution, shared project context, and a developer-facing workflow in the same environment.

That difference matters because orchestration is a collection of infrastructure decisions: who owns the execution environment, where workflow state persists, how agents are isolated, how credentials are provisioned, what happens when a worker fails, how one agent’s output becomes another agent’s input, and which actions can happen without human approval.

An API gives developers more responsibility for answering those questions. An integrated platform answers more of them on the developer’s behalf.

Neither approach removes the underlying engineering problems. It changes where they are implemented and who is responsible for operating them.

The coordinator is becoming an architectural boundary

The evidence from these systems points to a change in the role of the coding agent itself.

The model still performs the reasoning and code generation. But larger agentic workflows require another layer to determine how that capability is applied: which work is delegated, what context crosses an agent boundary, which tools are exposed, how execution state survives failures, and when the workflow needs human intervention.

Those are familiar distributed-systems concerns. Workers operate concurrently, state can be shared or isolated, dependencies connect tasks, workers can fail independently, and results need to be persisted and observed. The difference is that the workers are now probabilistic software agents rather than conventional processes.

That makes the coordinator more than a convenience feature. It is where a high-level software objective becomes executable work — and where decisions about context, permissions, durability, observability, and human intervention converge.

The September 10 launches make that shift visible from two different directions. OpenAI exposed orchestration infrastructure through an API. Cursor embedded it into a project-level development environment.

Neither announcement proves that one architecture will become the universal model for software development. But together with the systems already emerging around them, they show coding agents moving away from a single model executing an entire task and toward workflows that divide work among specialized agents, execution environments, and persistent infrastructure.

The engineering question is therefore no longer only whether an agent can write the code. It is whether the system around it can reliably decide what to do, which agent should do it, what that agent should be allowed to see and change, how to verify its work, and where a human should take control.

Those are architecture and infrastructure questions — and as coding agents move from interactive assistants toward autonomous software workflows, they may matter as much as the underlying model.

The post OpenAI and Cursor agree on agent coordinators. They disagree on who runs them. appeared first on The New Stack.

  •  

OpenAI’s agent had a routine task. It breached a government portal.

Sam Altman, OpenAI CEO

An OpenAI agent researching public medicine spending bypassed security blocks and gained unauthorized access to public and non-public files on an Australian government Medicare statistics portal, the government there disclosed Thursday. The agent, which OpenAI said was running during an internal evaluation in June, also wrote files to an internal server, according to the complaint.

Transluce, an independent nonprofit AI research lab, analyzed public request logs from the URL scanning service urlquery.net and found autonomous agents attempting SQL injection, cross-site scripting, command injection, and path traversal against the University of New Mexico’s digital library, the public data platform Data USA, and the Australian Institute of Health and Welfare (AIHW). The agents tried to retrieve ordinary information, including a historical photograph, University of Iowa data, and local pharmaceutical data in Victoria, and the offensive behavior appeared only after normal retrieval methods failed.

Transluce ties the Data USA and AIHW activity to an agent swarm that it says OpenAI previously confirmed originated from the company, based on shared targets, tactics and timing.

A day after Transluce published its findings on Wednesday, Australian Prime Minister Anthony Albanese disclosed that an OpenAI agent researching public medicine spending bypassed blocks on a Services Australia Medicare statistics portal in June and gained unauthorized access to public and non-public files. OpenAI, which said the agent was running during an internal evaluation, notified the Australian government on September 10. According to Albanese, the agent also wrote files to an internal server.

The incidents have not been publicly confirmed as the same activity. However, OpenAI told the Australian Broadcasting Corporation that much of what Transluce documented overlaps with cases in its ongoing review of misaligned model activity.

Transluce ties the Data USA and AIHW activity to an agent swarm that it says OpenAI previously confirmed originated from the company, based on shared targets, tactics and timing.

Probing public data providers

On May 25 and 26, agents trying to retrieve a photograph from the University of New Mexico’s Valmora collection ran into repeated failures. They responded with seven vulnerability probes against the service. The probes included an attempt to execute a system command, a path traversal request targeting /etc/passwd, and a SQL injection string containing UNION SELECT; they also arrived alongside a burst of 80 requests aimed at the image, although none appear to have succeeded.

Two days later, agents gathering University of Iowa data from Data USA hit errors caused by a malformed query parameter and followed those errors with 12 vulnerability probes against the API, covering SQL injection, reflected XSS, server-side template injection, path traversal, and command injection.

Cloudflare blocks, staging server workaround

The AIHW case on June 20 and 21 followed a very similar pattern. While trying to download pharmaceutical spending data from an AIHW Tableau dashboard, the agents ran into Cloudflare’s bot protections. One of them responded within minutes by testing the dashboard with a reflected XSS payload, although Cloudflare stopped it before it reached the application.

Transluce found no evidence that the probe succeeded, and AIHW says there is no evidence that non-public information was accessed. The agent ultimately found the same public dataset on AIHW’s pre-production server at pp.aihw.gov.au, and retrieved it through more than 100 scans.

The agent ultimately found the same public dataset on AIHW’s pre-production server at pp.aihw.gov.au, and retrieved it through more than 100 scans.

Agents build their own toolchains

The earliest clear sequence Transluce documented dates to March 6, when an agent tried to pull drug-enforcement statistics from the API of Thailand’s Office of the Narcotics Control Board. A direct request returned no page content, and adding a JSON format parameter changed nothing, so the agent routed the page through r.jina.ai three times and got text back with the Thai province names garbled.

Then, it wrote a custom Base64-encoded script and submitted it through urlquery.net, effectively using the service’s remote scanning browser to run its own client-side JavaScript against the data. That script ran into the same parsing problem, and a later attempt through Jina triggered a redirect to localhost that got the entire API blocked by Jina for an hour.

Beginning in mid-April, Transluce found the same technique across thousands of agent requests on urlquery.net, and on June 14 agents used the service’s browser to create a disposable email inbox and then tried to register a urlquery.net account with it. Registered users can make their scans private, which means the public records Transluce analyzed may understate how much of this activity actually took place.

An agent with web access isn’t limited to the software its developer gave it because remote browsers, proxies, URL scanners, and other public services can fill in the gaps, which gives the agent ways to make requests or run code that its own environment doesn’t provide.

Egress controls for AI agents

Instructions won’t be enough if the agent can still send whatever it wants over the network. For narrowly defined jobs, outbound traffic can be limited to approved hosts, a closed-by-default approach also used for securing AI agent sandboxes. Research agents may need to reach more of the web, so the focus shifts to controlling where they can connect.

Guidance for GKE Agent Sandbox recommends isolated runtimes with default-deny network policies that open only the endpoints an agent needs. Public proxies, URL scanners, and disposable email services can stay blocked unless the job requires them.

Developers can also limit what an agent can send. So, instead of handing it a networking tool that accepts any URL or request body, an API integration can restrict requests to specific fields and formats. The runtime can then catch path traversal attempts, SQL injection strings, and executable markup before anything is sent. OpenAI takes a related isolation approach in its Agents SDK sandboxes, and the company’s Responses API tech lead has said large enterprise deployments often call for agents that are isolated from the network entirely.

Repeated failures can also be a reason to pause a run, especially when an agent keeps hitting client errors, anti-bot challenges, or unexpected redirects and begins trying increasingly aggressive ways to get around them, as Transluce documented in several of these cases.

Keeping the original task, tool calls, and server responses in the same trace gives operators a better chance of catching that behavior change when a retrieval job starts generating encoded scripts, visiting staging domains, or sending exploit payloads, rather than discovering it later in someone else’s security logs.

Repeated failures can also be a reason to pause a run, especially when an agent keeps hitting client errors, anti-bot challenges, or unexpected redirects and begins trying increasingly aggressive ways to get around them, as Transluce documented in several of these cases.

The post OpenAI’s agent had a routine task. It breached a government portal. appeared first on The New Stack.

  •  

How a forgotten node can put Oracle Java back in production

Enterprise Java company Azul announced its Azul Intelligence Cloud AI Assistant on Wednesday. The technology arrives in response to industry-wide alerts that AI has become a force multiplier for threat actors today.

The Azul service is a natural-language query interface that allows software engineering teams to see where security risk and licensing infringements are hiding in their live production Java estate. The assistant provides answers that are “grounded in live runtime data”, so it replaces static reports that grow less accurate after the day they’re generated.

How fast do code-scanning reports go stale?

Azul said that “most IT and engineering teams” still manage Java risk with static IT and software asset management (ITAM/SAM) reports and code-scanning tools that describe a moment in time. The company has insisted that these reports are “accurate on the day they’re generated, and increasingly wrong after that”, typically because Java Virtual Machines (JVMs) are spun up, patched, drifted and retired underneath the report’s scope.

“For years, enterprises have built dashboards and reports to understand what’s actually running in their Java estate, but by the time a report gets properly summarized and reviewed, the risk it describes has often already changed,” said Scott Sellers, co-founder and CEO of Azul. “That used to be a productivity problem. Now that AI can find and weaponize a vulnerability in hours instead of weeks, it’s a business risk – for security, for compliance and for the licensing exposure that shows up in an audit.”

“Now that AI can find and weaponize a vulnerability in hours instead of weeks, it’s a business risk…”

The Azul Intelligence Cloud AI Assistant lets software teams ask a direct question, in plain language, and get an answer grounded in what’s actually running in production right at that moment in time, as well as query historical information for further analysis.

The weaponization gap is closing

Where enterprises run business-critical workloads on Java alongside AI services, Azul said the practical effect is that the gap between a Common Vulnerabilities and Exposures (CVE) entry being disclosed and it being weaponized is getting shorter. 

Citing AI models such as Anthropic’s Mythos and OpenAI’s Aardvark, which have autonomously discovered real-world vulnerabilities, the company pointed to an April 2026 Cloud Security Alliance white paper (listed as unofficial AI-assisted research), which suggested that, “While organizations historically took a median of 32 days to apply patches to known vulnerabilities – a window that once roughly corresponded to the time available before exploitation began – that window has collapsed to approximately 5 days for median time-to-exploit in 2025.”

Azul highlighted its Intelligence Cloud service, which gives engineers two “continuously updated” records of their Java estate: JVM Inventory, a live catalog of every JVM instance running anywhere (on-premises, cloud or container) and Code Inventory, a runtime record of which code actually executes in production versus what is merely provisioned. The new AI Assistant puts in a conversational layer using LLM models on top of both.

Software engineers can ask questions such as: 

  • “Which JVMs are running Java versions which are not the latest updates?”
  • “Where is Oracle Java running in production right now?”
  • “What code hasn’t run in the past four quarters and is safe to remove?”

In the FAQ section of its announcement, Azul suggested that post-migration JVM “drift is common”, often due to a rollback, a forgotten node, a shadow deployment or various scripts and processes that haven’t been updated, which can reintroduce an Oracle Java runtime, exposing compliance and licensing risk, if not security risk also.

The shape of the Java runtime security market

In terms of which other vendors operate in the Java runtime analytics and security market, there are more than a handful of usual suspects. Contrast Security is known for its JVM agent and in-app bytecode instrumentation. Dynatrace offers Runtime Vulnerability Analytics as an extension of its core observability platform, which ships with OneAgent monitoring for Java vulnerable functions and JVM-level bytecode instrumentation agents.

Through its acquisitions by HP, Micro Focus, and now OpenText, Fortify remains known for its static and dynamic application testing services, including Fortify Application Defender, a runtime application self-protection (RASP) agent built to monitor Java workloads during execution. Part of Thales, Imperva’s runtime security for Java and .NET applications spans simple access control up to complex anomaly detection algorithms, though Imperva has reportedly put its standalone RASP product on an end-of-sale path. Then there’s Datadog, with its Application Performance Monitoring (APM), built to power code-level distributed tracing from browser and mobile applications to backend services and databases.

“Runtime context is absolutely critical for understanding the real risk in the production environment.”

A busy market for sure, so just how much of a problem are now-anachronistic static reports?

Head of security advocacy at Datadog, Andrew Krug, tells The New Stack that the downside of most point-in-time inventory scans is that they are “not always representative” of the runtime environment. 

“Runtime context is absolutely critical for understanding the real risk in the production environment,” Krug says. “Even in the most mature software development lifecycle (SDLC) flows, tooling that generates static software bill of materials (SBOMs) may be bypassable [i.e. circumventable or subvertible] to get a feature deployed. Moreover, traditional vulnerability management flows outside of SDLC can bump versions outside of CI/CD process, accidentally compounding the problem and introducing added risk by bypassing known good guardrails like dependency cooldowns.”

“Even in the most mature software development lifecycle (SDLC) flows, tooling that generates static software bill of materials (SBOMs) may be bypassable… to get a feature deployed.”

Krug further advises that Datadog now sees “an increasing rise” in automated drive-by attacks on known vulnerabilities, “particularly Java” in many cases.

“Attacks used to be added to scanners, either specifically or generically, and scans are indiscriminately against targets,” Krug clarifies. “LLMs make it cheaper to add support for new vulnerabilities. However, it also makes it easier for individual researchers/hackers to have their own custom rulesets.”

“LLMs make it cheaper to add support for new vulnerabilities.”

He advises that this fact makes trends much harder to read than “oh, someone added support to CVE-2026-whatever in FFUF”, so today the goal of many attacks is unchanged.  Attackers are looking to move laterally, establish persistence, and often automations will look for credentials to leverage to do just that.

NOTE: (Fuzz Faster U Fool) is an extremely fast web application fuzzer (written in the Go language), which is used by security testers to discover hidden files, directories and endpoints.

Dead code; it’s really a ‘thing’

The above-noted list of competitors that work in Azul’s marketplace is (arguably) substantial evidence of the real commercial licensing and security risks that exist where Java code, redundant JVMs and chunks of unsubstantiated (or more likely just untracked) Java components have been left to roam free. Azul itself noted that there’s a real maintenance overhead here that needs to be addressed because “unused and dead code that still gets tuned, tested and carried through every migration”, usually because no one can assess it’s safe to remove.

The larger the Java estate, the larger the exposures get, obviously. But that also means that the less a point-in-time report can be trusted to catch these exposures before they become an incident, an audit finding or a breach.

Dedicated compliance officers will likely enjoy wider deployment of these tools, although that role itself may now reside within a DevSecOps or platform engineering team, or both. 

Azul Intelligence Cloud AI Assistant works regardless of which JVMs are deployed, from which vendor, or how old or large the applications running on them are. JVM Inventory and Code Inventory retain component and code-use history over time, so the AI Assistant can reason over which code, JVMs and applications have actually run in production, now and in the past. 

The post How a forgotten node can put Oracle Java back in production appeared first on The New Stack.

  •  

Developers and platform teams both want Kubernetes self-service. They disagree on who owns it.

Abstract view up through bold yellow angled beams to a white skylight grid and curved ceiling panels.

What do developers want? Kubernetes environments when they need them.

What do they not want? Those environments after a week or more of tickets. 

Platform teams, meanwhile, own what those environments cost, who can access them, and whether they meet company policy.

That tension is the core of Kubernetes self-service: What can safely be handed to developers, and what still belongs to the platform team?

In a recent interview with enterprise cloud specialists — Marius Bogoevici, Senior Principal Product Manager at Hewlett Packard Enterprise (HPE), and Karthik Subramanian, Principal Product Manager for HPE Morpheus Software — The New Stack explored the core friction points of Kubernetes self-service. The conversation focused less on whether self-service is desirable than on where to draw the line.

HKS, HPE’s CNCF-certified Kubernetes distribution, is integrated with HPE Morpheus Software to help platform teams deliver and lifecycle-manage Kubernetes environments as part of a broader operating model spanning Kubernetes, VMs, infrastructure, and clouds. HPE Morpheus Advanced Software supports the on-premises private-cloud use case with HKS, while HPE Morpheus Enterprise Software extends Kubernetes and application operations across hybrid and public-cloud environments.

Together, HKS and HPE Morpheus Software extend that operating model beyond infrastructure provisioning. Through service and application catalogs, platform teams can connect approved Kubernetes environments with the CI/CD pipelines, container registries, automation tools, and other services developers already use. Developers receive a governed, ready-to-use path from code to deployment instead of manually assembling the toolchain for each project.

The self-service paradox

Open-source Kubernetes provides orchestration and declarative APIs, but not a complete operating model.

Subramanian says teams building their own self-service layer usually run into two recurring problems:

  1. Tool and package sprawl: To make upstream Kubernetes production-ready, platform teams must curate and maintain an ever-evolving ecosystem of third-party CNCF tooling for networking (CNI), storage (CSI), ingress, identity, and policy enforcement. Navigating and supporting this fragmented stack creates immense maintenance overhead for internal platform teams.
  2. Day-2 lifecycle and hybrid footprint complexity: Spinning up a Kubernetes cluster is the easy part, but keeping it current — across development, QA, staging, and production — is where the work piles up. That is why HPE says every Kubernetes upgrade must be checked against the networking, storage, ingress, identity, and policy components around it. The problem gets harder when clusters span bare metal, private clouds, edge sites, and public clouds, because one-off scripts and environment-specific configurations can quickly create drift. That maintenance burden belongs with the platform team, not with developers trying to ship applications.

Giving developers direct access to raw Kubernetes APIs just shifts the operational work — it’s far from gone for good. In fact, developers will wind up debugging manifests and storage drivers instead of writing code. 

Meanwhile, operations teams have to deal with overprovisioning, idle clusters, and configurations that reach production without review.

What developers control — and what the platform supplies

The practical answer is not unrestricted access. It is a paved path: approved Kubernetes services that developers can request themselves, with access, configuration, placement, approvals, and lifecycle controls defined by the platform team.

“The best candidates for self-service are requests that are repeatable, low-risk, and well-understood,” Bogoevici tells The New Stack. “For example, a developer should be able to request a development cluster, deploy an approved application, create a namespace, or select resources from pre-approved configurations without opening a ticket. The platform team decides what a safe configuration looks like, and the developer chooses from a supporting menu.”

“The platform team decides what a safe configuration looks like, and the developer chooses from a supporting menu.”

Rather than asking developers to write YAML for ingress, storage classes, and RBAC, HPE Morpheus exposes those choices through service catalogs, reusable layouts and blueprints, workflows, role-based access control, approvals, APIs, and automation. Developers do not lose Kubernetes. They retain direct access through standard Kubernetes interfaces and tools where permitted, while the platform team standardizes the request, governance, and lifecycle processes around them.

Those catalog items can package more than infrastructure settings. They can also integrate the approved services and application components that support the development workflow – including CI/CD tooling, source and artifact repositories, container registries, and runtime dependencies – while the platform team controls how those components are configured and governed.

Developers choose the parameters that matter to the application:

  • Approved Kubernetes versions and cluster sizes: Select from pre-tested Kubernetes runtime releases and node count templates.
  • Resource quotas: Specify required CPU, RAM, and persistent storage capacity tailored to the workload.
  • Integrated toolsets and IDE environments: Select required developer toolchains, container registries, and runtime dependencies. 
  • Lease and duration limits: Define explicit operational lifetimes for temporary development or sandbox clusters to prevent abandoned infrastructure sprawl.

Network isolation, identity-provider integration, security policy, and cost allocation stay with the platform team and are applied automatically through the approved service configuration.

The same division of responsibility applies to the delivery toolchain: Developers choose from approved services, while the platform team manages the integrations, credentials, policies, and automation behind them. This gives developers a consistent experience without shifting toolchain maintenance and governance onto individual application teams.

The division of responsibility looks like this:

Service areaDeveloper chooses or requestsPlatform team defines and suppliesReview or exception path
Development cluster provisioningApproved Kubernetes service, version, size, target environment, and duration.Reusable layout or blueprint, access controls, placement rules, storage and network defaults, and lifecycle policy.Nonstandard versions, placements, configurations, or requests outside quota.
Production deploymentApplication artifacts, target namespace, and deployment request through the approved path.RBAC, tenancy, policy, audit, backup, and release controls appropriate to the environment.Formal review for production changes and exceptions.
Resource allocation and quotasCPU, memory, storage, and other approved capacity parameters within project limits.Project quotas, upper bounds, placement constraints, and supported resource profiles.Requests above quota or for specialized resources.
Networking and securityApplication endpoints and permitted connectivity within approved patterns.Identity integration, RBAC, tenant isolation, network policy, secrets, and audit controls.Cross-tenant access, elevated privileges, or changes to baseline security policy.
Lifecycle and cost governanceService lifetime and approved operational actions.Visibility, policy, approvals, retirement workflows, and applicable cost controls for the licensed variant.Long-running exceptions, nonstandard lifecycle actions, or budget exceptions.

From ticket queues to a repeatable paved path

In conventional IT environments, provisioning a dedicated Kubernetes environment for a new project often involves cross-departmental ticket handoffs spanning infrastructure, networking, security, and storage teams. This friction frequently stretches provisioning timelines from days to weeks.

By unifying infrastructure orchestration, role-based access controls, and multi-tenancy into a single operational experience, HPE Morpheus Software can compress these provisioning workflows down to minutes or hours, according to HPE. “Developers get a usable environment that complies with the organization’s defined controls and policies, without needing to understand all the complex infrastructure steps sitting underneath,” Bogoevici says. “When you reduce provisioning time from weeks to hours, that is super meaningful and tangible.”

“When you reduce provisioning time from weeks to hours, that is super meaningful and tangible.”

The result is not only faster cluster provisioning. HPE Morpheus Software can also automate the handoff into the developer’s established delivery process by making approved CI/CD and application services available with the environment. Instead of waiting for separate teams to connect pipelines, registries, credentials, and runtime dependencies, developers receive a ready-to-use path from development through deployment.

Faster provisioning can create a different problem, too: The speed can and will cause teams to lose track of what was provisioned and why. HPE Morpheus Software gives administrators visibility into utilization and cost, while lease controls can shut down temporary development clusters when their time expires.

Bogoevici says ticket volume is a poor measure of success, particularly early on, when more developers may be trying the catalog. He recommends watching deployment success, exception rates, resource utilization, and the day-to-day effort required to keep the service running.

Security belongs in the service design

Security is another boundary that must be designed into the self-service path. If identity, access, tenancy, and policy are added only after a cluster is created, every request produces more work and more room for inconsistency.

“Security must be a core design consideration built directly into the service, not an afterthought during deployment,” Bogoevici says. “HPE Morpheus Software brings identity integration, role-based access, tenant isolation, approvals, and policy into the operational workflow.”

A newly provisioned environment should arrive through an approved configuration with the applicable identity, RBAC, tenant, policy, and audit controls attached. Platform teams can validate the paved path by testing an allowed request, a request that should be rejected, and the resulting audit record.

The operating-model test

The strongest Kubernetes self-service model does not hide Kubernetes or make it the control plane for every workload. It gives developers useful, approved choices and direct access to the Kubernetes workflows they need, while the platform team standardizes the enterprise processes around those workflows.

That matters because the enterprise still runs VMs, clouds, and existing infrastructure alongside Kubernetes. HPE Morpheus Software helps platform teams use common request, governance, automation, and lifecycle processes across these environments without forcing every workload onto one runtime or creating another operational silo.

In practice, that means self-service should deliver more than a Kubernetes cluster. With HPE Morpheus Software, a catalog request can bring together the approved environment, application services, and DevOps toolchain integrations developers need, while preserving the governance and lifecycle controls the platform team requires. Developers spend less time assembling and troubleshooting delivery infrastructure – and more time building and releasing applications.

Looking toward 2027, the goal is not unrestricted control. It is faster access, predictable results, transparent guardrails, and a clear exception path when the standard service does not fit.

The post Developers and platform teams both want Kubernetes self-service. They disagree on who owns it. appeared first on The New Stack.

  •  

Google’s Gemini CLI now asks before editing your build files

Abstract wave

The appeal of an autonomous coding agent is that you hand it a task, give it access to your repository and tools, and stay out of its way while it works. Google’s latest Gemini CLI release carves out specific moments when the agent now has to stop and wait for you.

Gemini CLI 0.61.0, released Wednesday, requires explicit confirmation before the agent edits build configuration files, runs build or test commands after such an edit, or executes shell commands whose arguments appear to come from untrusted external content. The same release separately hardens Gemini CLI’s optional sandbox so that host credentials and configuration stay out of reach of whatever runs inside it.

Giving a coding agent more authority to modify and execute code also gives an attacker more ways to turn that authority against the developer. Gemini CLI 0.61.0 puts a human back in the loop at some of those points.

Security fixes, in public

Google announced at I/O in May that it would move Gemini CLI’s Pro, Ultra, and free-tier users to its closed-source Antigravity CLI, and since June 18, the open-source tool has served mainly enterprise customers and developers with paid API keys. The company said Gemini CLI would continue to get model updates, bug fixes, and security patches. Those security changes are still developed in public, and the pull requests behind version 0.61.0 show exactly what Google was worried about.

Build files become attack vectors

A change to package.json, Makefile, pyproject.toml or a Bazel BUILD file can pull in a dependency or trigger a script. Gemini CLI can make those edits using information from web searches and external tools, then run shell commands. If documentation fetched while fixing a bug contains hidden instructions to add a postinstall script to package.json, the agent could make the edit, run the project’s test suite, and execute the malicious code without the developer ever typing the command.

Giving a coding agent more authority to modify and execute code also gives an attacker more ways to turn that authority against the developer.

Pull request #29250, titled “prevent indirect prompt injection via build file modifications and untrusted flags,” targets that sequence directly. Edits to recognized build files now require confirmation, and Gemini CLI tracks which build files change during a session so it holds any later build or test command, such as npm run, make, or cargo, for explicit approval. The confirmation dialog also shows full build-file diffs rather than truncating them.

Untrusted arguments need approval

The second check covers command arguments. Gemini CLI now treats content from web fetches, MCP server responses, Google Docs, and Buganizer, Google’s internal issue tracker, as untrusted context, and it asks before running any shell command whose flags or arguments match tokens from that content. In both cases, the prompt drops the persistent approval options, so a developer can’t grant a standing “always allow” for these actions.

The pull request ties the changes to restricted workspace mode, the safe mode Gemini CLI applies to folders a user hasn’t marked as trusted, and it doesn’t spell out how the checks behave in a trusted folder or under auto-approval.

The argument check matches tokens rather than tracing the provenance of every value, and the pull request’s review history shows how hard that is to get right. Google’s automated reviewer flagged several workarounds in earlier versions, including quoted arguments, environment-variable prefixes, shell redirection targets, and Windows path handling, all of which were addressed before the change merged on September 11.

Sandbox keeps credentials out

Pull request #29214 tightens Gemini CLI’s sandbox. When the sandbox runs through Docker, Podman, LXC, or macOS Seatbelt, the host’s ~/.gemini directory is no longer mounted inside it. Instead, the CLI passes in a sanitized copy of the user’s settings with API keys, hooks, and custom tool commands stripped out. It also blocks the sandbox from launching in sensitive locations such as the home directory, while new Seatbelt rules deny access to OAuth credentials, trusted-folder decisions, and .env files.

Google’s sandboxing documentation calls the feature a security barrier between AI operations and the host system, while cautioning that it reduces risk without eliminating it. The two pull requests show why both layers are needed. The sandbox limits what a process can reach once it runs, and the confirmation requirements decide whether the agent gets to take a sensitive action in the first place. Build files make the gap concrete: the sandbox mounts the project directory so the agent can edit it, meaning a poisoned package.json written inside the sandbox still sits in the repository when a developer or a CI job later runs the build outside it.

…a poisoned package.json written inside the sandbox is still sitting in the repository when a developer or a CI job later runs the build outside of it.

Gemini CLI already gives developers ways to decide how much the agent does on its own, from hooks that run deterministic checks at fixed points in the agent’s workflow to an MCP server trust setting that, according to Google’s documentation, bypasses all tool call confirmations for that server. Trust granted once can age badly, though, as tool-poisoning and rug-pull attacks on MCP servers have shown when a tool approved on one day starts returning attacker-controlled content later.

Trust granted once can age badly… when a tool approved on one day starts returning attacker-controlled content later.

The post Google’s Gemini CLI now asks before editing your build files appeared first on The New Stack.

  •  

OpenAI makes you call sales for a custom voice. Google just made it self-serve.

Abstract sound waves

Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS today through the Gemini API and Google AI Studio. Text-to-speech APIs have historically left developers working with whatever voices were already available, but Gemini 3.8 changes that by letting users create the voice itself.

Now, developers can describe the voice they have in mind or start with a short recording of an existing voice, then save what they create and use it again across an application. Google handles the voice profile from there, so the original recording or description doesn’t have to accompany every new request.

Turning recordings into voice IDs

Replication runs through a new Voices endpoint (POST /v1beta/voices) and two recordings are required from the same speaker; those need to be clean samples between 10 and 30 seconds and a separate consent recording. For that second clip, the speaker reads a statement, confirming that the voice belongs to them and that they agree to let Google create a synthetic version of it. Google confirms that the person giving consent and the reference clip are the same person before proceeding.

Once approved, Google returns a voice_… ID and keeps it in the developer’s project for a year, alongside any voices created with Gemini’s voice-design tools. A project can hold up to 200 voices in total, and developers can retrieve, list, or delete them through the API just as they would other stored resources.

Voice replication can also be used without storing the profile in the project. Setting store=False returns an encrypted voicekey_… instead, which stays with the application and is supplied again when the voice is needed. Because the key expires after seven days, this option makes  sense for short-lived jobs.

A few more things are worth noting before building around the feature are the fact that Google marks audio generated by Gemini with SynthID, and replicated voices also carry C2PA content credentials that can be used to trace where the audio came from. Google doesn’t offer voice replication through AI Studio in Illinois, Texas, the European Economic Area, the U.K., Switzerland or India.

A project can hold up to 200 voices in total, and developers can retrieve, list, or delete them through the API just as they would other stored resources.

Prompting a voice from scratch

Voice design generates a persona from a natural-language description of role, accent, and character, and Google says it works across more than 100 languages and dialects. The docs list 130 supported languages for Flash TTS and 101 for Flash-Lite. Google’s announcement also claims a library of more than 2,000 production-ready voices.

The developer docs describe 30 prebuilt studio voices plus hundreds more in an extended library that can be filtered by language, accent, pitch, and use case through GET /v1beta/voices. A remixing feature for adjusting the timbre, pitch, pace, and accent of library voices with prompts is something Google lists as coming soon.

The company recommends creating a voice once and reusing its ID rather than describing the same persona in every request. According to the docs, repeatedly sending long persona descriptions is the most common cause of voice drift. Once the voice is created, subsequent requests need only a short style instruction, if any.

Gemini 3.8 sees input text strictly as a verbatim transcript, a breaking change for anyone who embedded stage directions in prompts to the 3.1 preview model. Sustained direction for a turn, such as whispering, sarcasm, or speaking rapidly, now goes in a speech_metadata annotation, while momentary sounds like <sigh>, <cough>, and <short pause> sit inline in angle brackets. In two-speaker scripts, listener reactions wrapped in pipes, such as |mhm|, produce backchannels and overlapping speech without breaking the script into extra turns.

Gemini 3.8 sees input text strictly as a verbatim transcript, a breaking change for anyone who embedded stage directions in prompts to the 3.1 preview model.

Two-speaker scripts have limits

Native two-speaker generation has one limitation that’s important to mention. A single request supports up to two speakers using prebuilt voices, while dialogue between designed or replicated voices has to be generated turn by turn and stitched together from the 24 kHz PCM output.

Unary requests return WAV by default, streaming requests return raw 16-bit PCM, and mu-law and A-law encodings are available for telephony pipelines. Google says Flash TTS maintains voice quality and timbre across hours of continuous audio, targeting audiobook and podcast production.

Flash for performance, Flash-Lite for volume

Both models share an API schema, so switching between them is a one-parameter change, and both support voice design and replication.

The company positions Flash TTS for demanding acting work, including complex dialogue, heavy use of vocal tags, difficult pronunciations, regional dialects, and long narration. Flash-Lite TTS is the faster, less expensive option and the direct replacement for gemini-3.1-flash-tts-preview, tuned for bulk production, read-aloud features, and cascaded voice agents that pair a text model with a separate speech step.

For those agents, Google recommends one TTS call per turn as the LLM’s text arrives, with the stored voice carrying identity across the conversation.

Plugging into voice agent frameworks

A speech model is only one layer of a production voice application, and a real-time agent still needs transport, speech recognition, turn detection, interruption handling, and session state. Google points developers toward frameworks that already handle those layers, naming Agora, LiveKit, Pipecat and Vercel’s AI Gateway as platforms that support Gemini speech generation through the Gemini API.

That lets a team drop Gemini in as the speech layer without rebuilding its audio pipeline, although anyone planning to rely on a replicated voice should confirm their framework passes custom voice_… IDs through before committing. API access through Gemini Enterprise is listed as coming soon.

How OpenAI’s approach compares

OpenAI also offers custom voices, but access is tighter. Customers have to go through sales, are limited to 20 voices per organization and must provide a consent recording alongside a voice sample of up to 30 seconds. The resulting voice ID works across its speech endpoint, Realtime API, and Chat Completions.

What OpenAI doesn’t have is Google’s prompt-based voice design, which can create a voice from a written description. Its 13 built-in voices can be steered for tone or speed, and apps must disclose that the speech is AI-generated.

In comparison, Google’s advantage is that it’s giving developers more ways to create the voice they want before the first line of text ever reaches it.

Google’s advantage is that it’s giving developers more ways to create the voice they want before the first line of text ever reaches it.

The post OpenAI makes you call sales for a custom voice. Google just made it self-serve. appeared first on The New Stack.

  •  

What managing 150,000 AI agents could look like for database teams

Abstract 3D render of translucent orange cubes and panels scattered across a pale gray background, with bundles of glossy teal tubes curving in from the right.

The database administrator of the future will spend considerably less time administering databases.

That sounds contradictory, but AI agents are taking over that work. For decades, DBAs have handled the decidedly hands-on work of keeping databases available, performant, secure, and affordable. They provision capacity, troubleshoot slow queries, manage migrations, and step in when something inevitably goes sideways.

AI is already taking on some of that work. At the same time, it is creating a much bigger data infrastructure fleet to manage.

The result is likely to be a very different kind of DBA: one who spends less time tending individual databases and more time supervising the autonomous systems doing it for them.

Congratulations, you’re managing robots now

This shift starts with a familiar problem: more infrastructure needs managing than the people available to manage it.

Database automation is hardly new, but agents can potentially go further than the scripts and rules DBAs already rely on. Rather than automating one predetermined task, an agent can inspect what is happening, decide what needs attention, use tools to act on it, and check whether its intervention worked.

That changes the DBA’s relationship with the database. A performance problem that once required someone to dig through metrics, identify the troublesome query, and decide how to respond could increasingly be investigated by an agent before a human gets involved.

It doesn’t remove the DBA from the equation. Someone still has to decide what an agent can do, where human approval is required, and what happens when it gets something wrong. But the work moves up a layer. Instead of personally performing every operational task, DBAs start managing the systems carrying them out.

Instead of personally performing every operational task, DBAs start managing the systems carrying them out.

And before anyone gets too comfortable with that idea, the number of those systems could become enormous.

150,000 agents walk into a database…

Gartner predicts that the average global Fortune 500 company will have more than 150,000 AI agents in use by 2028, up from fewer than 15 in 2025. Only 13% of organizations currently believe they have the right governance in place to manage them.

Not every agent will need its own database, but plenty will. They will create state, retrieve data, remember previous interactions, and exchange information with other agents. Many will also behave very differently from the applications DBAs are used to supporting: spinning up quickly, sitting idle for long stretches, and suddenly becoming busy when there is work to do.

Nobody is hiring 150,000 DBAs to manage them.


That is the scale problem Yugabyte is targeting with YugabyteDB AMP, or Agentic Multitenant PostgreSQL. Rather than treating each new agent workload as another database for an administrator to provision and babysit, AMP manages databases as a fleet.

The platform packs hundreds of small Postgres workloads onto shared distributed infrastructure while keeping their databases isolated. Lifecycle operations, including provisioning, branching, scaling, migration, and teardown, can be exposed to agents through MCP. Yugabyte has also built specialized agents for setup, migration, performance tuning, and integrations.

In that model, a DBA is no longer provisioning database number 14,372. The interesting job is setting the rules for how database number 14,372 is provisioned, operated, and fine-tuned without them.

Do more with less (no, really)

Scale is only half of the problem. Someone also has to pay for all this stuff.

Agent workloads make traditional capacity planning particularly awkward because many are bursty and frequently idle. Giving every experimental agent permanently provisioned infrastructure could leave companies paying for many databases that spend much of their lives doing very little.

This is where consolidation becomes as much an economic question as an operational one.

AMP’s approach is serverless multitenancy and scale-to-zero. Multiple small workloads share the underlying distributed infrastructure, while customers pay by CPU minute and idle agents consume no compute. Resource governance can impose CPU limits on individual workloads, preventing a single overeager agent from consuming the capacity intended for its neighbors.

The human equivalent matters too. If routine setup, migrations, tuning and other database operations can increasingly be delegated, a smaller database team can potentially look after a much larger estate.

That doesn’t mean companies get to fire the DBAs and hand the keys to the robots. It means scarce database expertise can be spent on architecture, governance, and genuinely difficult problems instead of repeatedly doing the work that software can handle.

Your 2028 database problem starts now

The harder question is what to build underneath all of this when nobody really knows what the enterprise AI estate will look like in two years.

An agent that begins as an experiment today could disappear next month. Another could suddenly become a production application used across the business. Building one infrastructure stack for cheap experiments and another for serious workloads risks creating a migration problem every time an experiment succeeds.

Yugabyte bets that both ends of that journey should sit on the same foundation.

YugabyteDB AMP lets workloads start on serverless Postgres and transition to fully distributed YugabyteDB as their scale and criticality increase, without rewriting the application or migrating data to a different database platform.

Then there is the problem above the individual database: agents need to remember what happened, and not just in a silo.

That’s where Meko fits into the Yugabyte stack. Meko is an agent-native context engine designed for multi-agent AI systems. It provides persistent memory, shared knowledge, decision traces, and autidability across multiple agents, rather than leaving each agent working from its own isolated context. An agent can pick up information learned by another agent instead of retrieving it again or restarting the reasoning process.

Taken together, it delivers a single data stack for an agent’s entire lifecycle: Meko for the context shared among agents, YugabyteDB AMP for agentically managing fleets of Postgres databases, and distributed Postgres-compatible YugabyteDB for workloads that outgrow their serverless beginnings.

Of course, there’s no guarantee that 2028 will look exactly like today’s forecasts. That’s rather the point. The safest architectural bet may be one that doesn’t require you to know in advance which of today’s tiny AI experiments will become tomorrow’s critical applications.

The DBA is still critical in that world, but the job will look different. The DBA of the future may manage fewer databases directly, while taking responsibility for vastly more of them. Instead, managing the autonomous systems that do the administering.

The post What managing 150,000 AI agents could look like for database teams appeared first on The New Stack.

  •  

Query decomposition doesn’t fix context starvation — it just moves it

Abstract digital art showing warped light lines surrounding a void, illustrating data compression and AI context starvation.

I have a small example that would best communicate the message I am trying to convey: say you built a chat widget for GitLab’s public documentation (the corpus we are experimenting with in this article) and one of the developers sends this kind of message:

We got an email saying our card was declined for something called “quarterly reconciliation” and I need to know what actually happens now. On top of that, I think we’ve gone over our seat count; there are more people in the group than seats we bought. Our CI has been queuing all week and I want to know whether the compute minutes we purchased last month rolled over or if we lose them. Our finance lead also needs to be the one who gets the invoices from now on, not me. And last thing, is the REST API rate limited? We’re building an internal dashboard and would rather find out now than after it breaks.

Five separate asks: the declined payment, the seat overage, compute-minute rollover, changing who receives invoices, and API rate limits. Each one is answered by a specific passage in GitLab’s public documentation, and you labeled which passage answers which before running anything, so you knew in advance exactly what a correct system needed to find.

Then you run the message through a pipeline that follows current best practice. It splits the query into five clean sub-queries, retrieves for each one independently, merges the results, drops near-duplicates, reranks the merged pool against the original message, and packs the highest-scoring passages into a 2,000-token context.

The pipeline retrieved all five correct passages, but only one of them survived into the packed context; that is one of five asks, not one of five sentences. The packer found, scored, and threw away the other four before the model ever saw them. The same message with no decomposition at all managed three out of five.

The failure has a name, and it isn’t the one you’re thinking of

I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.

The definition is deliberately about allocation rather than retrieval, because allocation is the part nobody watches. If the passage answering the fifth question was found, scored, and then squeezed out by three passages about the first question, the fifth sub-intent is starved, and every recall metric you have will report that the system worked perfectly.

“I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.”

Two failure modes already in circulation describe something different, and it’s worth separating them cleanly:

Semantic dilution happens at retrieval time: when you embed a five-part question as a single vector, you get a centroid that sits somewhere between five topics and lands close to none of them, so the evidence is never found. Decomposition fixes this, which is why it spread.

Context poisoning is about what is present, not what is missing. Wrong, stale, or adversarial content enters the window and corrupts what the model generates downstream. Poisoning is a contamination problem. Starvation is an absence problem, and policy, not accident, produces the absence.

I borrowed the word from operating systems. In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work. Every ingredient of that situation is present in a retrieval pipeline: a fixed resource, competing demands, and a policy that decides who gets served. A relevance-greedy packer is priority scheduling with no aging term, and under priority scheduling without aging, valid low-priority work waits forever.

Decomposition is the right fix to the wrong half of the problem

Split that message into five single-intent queries, and each one embeds cleanly, so per-sub-query recall climbs sharply. This is well-trodden ground. LlamaIndex ships a SubQuestionQueryEngine that breaks a complex query into sub-questions and synthesizes the responses. LangChain’s MultiQueryRetriever generates query variants and returns the unique union of what they retrieve. RAG-Fusion applies reciprocal rank fusion across the per-query result lists. The technique works, and it isn’t mine.

“In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work.”

The context window did not grow. Let me explain: after decomposition, you have n result sets competing for one fixed token budget, and something downstream has to decide the split. In most production pipelines, that something is a short, unremarkable sequence: merge the pools, drop near-duplicates, rerank the merged pool against the original query, then greedily fill until the budget closes.

That sequence is a scheduler. It has a priority function, which is the reranker score, and it has no fairness constraint of any kind. A sub-intent with three strongly-scoring passages takes three slots. A sub-intent whose single correct passage scores mid-pack takes none of them.

So the failure did not go away. It moved from the embedding, where it has a name and people watch for it, into the packer, where it has neither. It also moved somewhere with much worse instrumentation, because recall@k per sub-query is the metric decomposition usually gets validated with, and that number goes up. It goes up at the same time as coverage inside the packed context goes down. You ship on a green dashboard.

The harness

The corpus, GitLab’s public documentation: 10,000 chunks and 2.2M tokens, split on heading boundaries and capped at 480 tokens each. Sixty-one single-intent questions span nine topics, from seat management to rate limits, each labeled with the one passage that answers it. I built multi-intent queries by concatenating those questions while varying n across 2, 3, 5, and 7, randomizing the order so position doesn’t confound topic, and varying topical distance so half the queries draw everything from one topic and half span distinct ones. That produces 100 queries, 25 at each value of n, whose correct decomposition I know exactly.

Two decisions matter more than the rest:

The metric is not recall. Recall tells you what the retriever found. What I need is what survived into the packed context, per sub-intent. So I log each sub-intent twice: once for whether its correct passage reached the candidate pool, and once for whether it reached the packed context. The gap between those two numbers is the entire argument.

Every question has to be retrievable on its own before it’s allowed in. A question enters only if its correct passage ranks in the top 10 for its own isolated query, under both retriever configurations, and both scored 100% recall@10 on that test. Since each sub-query’s candidate pool is exactly its own top 10, passing that gate guarantees the correct passage sits in the pool for every decomposed arm. Any sub-intent that then fails to appear was denied by the packer rather than missed by the retriever, which removes the most obvious objection to everything below.

The core measurement uses no language model. Because queries are composed from known sub-questions, the decomposer is an oracle so that anyone can reproduce the main result with no API key.

That invites an objection, so I tested it. A real LLM decomposer, blind to n, disagreed with my ground truth on 41% of the queries, and on inspection it was right every time. Five of my sixty-one supposedly single-intent questions contain two distinct information needs. What is excess storage usage, and what happens when we go over the free limit? is two questions wearing one question mark. Adjusted for those five, agreement is 100 out of 100. That is not evidence decomposers are reliable, because my queries are joined by fixed connectives and splitting on those alone recovers n perfectly, which real messages never allow. What it caught was an error in my own labels, and that is the best argument I have for the oracle design.

Results

Every arm runs at 2,000, 4,000, and 8,000 tokens against two retriever configurations. The stronger pairs are BAAI/bge-base-en-v1.5 with BAAI/bge-reranker-base; the weaker pairs are a quantized BAAI/bge-small-en-v1.5 with Xenova/ms-marco-MiniLM-L-6-v2. Retrieval is in-memory cosine similarity over a NumPy array because, at 10,000 chunks, a vector database would be slower to write, slower to run, and harder to verify.

Sub-intent coverage at a 2,000-token budget on the stronger configuration:

Table showing coverage (in %) per arm.

The production-default pipeline starves 31.1% of sub-intents whose evidence it had already retrieved. It beats no decomposition by nine points, while a flat B/n split, which is the crudest allocator anyone could write, beats it by fourteen.

Floors work, but not the obvious floor. Reserving one passage per sub-intent before the greedy fill satisfied 99.4% of its reservations and bought only seven points. The mechanism fires correctly and reserves the wrong passage because it picks each sub-intent’s best chunk by score against the original query. The original query asks about all five intents at once. Selecting that same reservation by score against its own sub-query pushes coverage to 89.4% and cuts allocation starvation from 30.6% to 10.1%. That is a change of about four lines.

Two of my own recommendations died here. I expected reranking against the original query to beat reranking against the fragment, and it loses by seventeen points. The incomparable score scales I worried about turn out to help, because each sub-query’s best match ends up at the top of its own scale, producing per-intent fairness for free. I also expected deduplication before allocation to matter, but near-duplicates consume 1.0% of the budget and removing them moves coverage by 0.3 points.

The crossover: starvation by tokens-per-sub-intent, which is simply the budget divided by n:

Table showing tokens per sub-intent.

Below roughly 1,000 tokens per sub-intent, allocation policy dominates. Above it, nothing you do to the allocator matters, because everything fits anyway.

The retriever comparison is the one I’d lead with. Upgrading the retriever moves coverage on the production-default arm from 56.2% to 68.9%, a gain of 12.7 points. Changing the allocation policy on the same retriever moves it from 68.9% to 89.4%, a gain of 20.5 points. In its sharpest form: the weaker retriever with a fragment-scored floor reaches 84.5%, and beats the stronger retriever with a greedy packer at 68.9% by sixteen points. A worse retriever with a better allocator wins.

Position: Held within a fixed n so query difficulty doesn’t contaminate the comparison; starvation across the seven positions of an n=7 query runs 1.3%, 21.3%, 30.7%, 48.0%, 48.0%, 32.0%, and 13.3%. That is a serial-position curve. The packer protects what you asked for first, protects what you asked for last a little less, and drops the middle. The fragment-scored floor flattens it to 4.0%, 9.3%, 6.7%, 8.0%, 22.7%, 5.3%, and 6.7%.

“A worse retriever with a better allocator wins.”

Topical distance: Sub-intents that span distinct topics starve about twice as often as sub-intents drawn from one topic, at 38.1% against 18.3% for n=7. I predicted the opposite. A topically coherent query gives the reranker a coherent target, and it scores all the correct passages similarly. In contrast, a scattered query lets it latch onto some topics and abandon others.

What the user actually sees

Everything above is retrieval-side. What decides whether any of it matters is what reaches the person who wrote the message, so I generated real support replies from 80 packed contexts and had every reply graded per sub-intent, with both the generation and the grading blind to which arm produced which context.

When the correct evidence reached the packed context, the reply addressed that question 100% of the time, across 261 out of 261 cases, in both arms. Coverage predicts the generated outcome exactly, which is the strongest justification I have for measuring it.

When a sub-intent was starved, the reply answered it anyway 48.1% of the time, based on whatever else happened to be in the window. It explicitly flagged the gap 45.6% of the time, with some version of “I’ll follow up on that separately.” It went silent only 6.3% of the time.

“Starvation mostly does not produce silence; it produces unsupported answers.”

I expected silence, and I was wrong. Starvation mostly does not produce silence; it produces unsupported answers. Whether those answers are actually incorrect is the next experiment, because this harness measures whether a question was addressed, not whether the answer was right.

Some limits: composed queries are cleaner than real support messages, which carry pronouns, implicit context, and conditional clauses. This is one corpus and one embedding family. I drafted the gold labels with model assistance and verified them myself. The model writing those replies was strong, so a cheaper production model would plausibly flag fewer gaps and invent more.

What an allocator actually looks like

Give every sub-intent a floor, and choose it by fragment score. Not the naive floor, which satisfies 99% of its reservations and buys seven points. Select a reservation by relevance to the sub-intent it protects, not by relevance to the message as a whole.

Rerank against the fragment rather than the original query. This inverts what I expected and what I have seen recommended. Scores from different fragments are not comparable across sub-intents, and that incomparability is doing useful work.

Don’t spend your effort on deduplication; near-duplicates cost 1.0% of the budget here. Dedup is worth doing, but it isn’t why your fifth question went unanswered, and treating it as the fix will cost you weeks.

Log per-sub-intent coverage: You already computed it to pack, and it predicts the generated outcome perfectly. A sub-intent that received zero passages is the best predictor available that your reply is about to assert something you cannot support.

Where parallel decomposition breaks

“If it’s late can I get a refund” is one clause and two intents, and the second one’s retrieval target depends on the first one’s answer. Parallel decomposition treats them as siblings. It retrieves the late-delivery policy and the general refund policy, packs both, and misses that the passage you actually need covers refunds for late delivery, which may match neither sub-query particularly well.

There are two ways out: You can tag dependencies at decomposition time, or run a deferred second pass that re-retrieves conditional clauses once the first round resolves.

I would take dependency tagging, for three reasons: A second pass costs a full retrieval round trip inside a latency budget a support bot does not have. The tag is reusable, because a dependent sub-intent should not hold a floor reservation. At the same time, its parent is unsatisfied, so it feeds the allocator directly instead of bolting on a separate mechanism. And it fails visibly, since an untagged dependency shows up as a starved sub-intent in the coverage signal. In contrast, a deferred pass that resolves the wrong condition produces a confident wrong answer with nothing to flag it.

The cost is real; dependency tagging pushes work onto the decomposer, which is already the weakest component in the chain, and I have not measured tagged against untagged. That is a design position rather than a result, and it is the one thing here I am asking you to take on argument instead of evidence.

What to measure on Monday

Take your production pipeline and compute one number: your context budget divided by the average count of distinct questions per incoming message. If that number lands below roughly 1,000 tokens, your allocation policy costs more than your retriever does, and the reranker upgrade sitting in your backlog will buy you less than reserving one slot per question.

On my corpus, the retriever upgrade was worth 12.7 points of coverage, and the allocation change was worth 20.5, which is why I think the ordering is wrong in most pipelines I’ve seen. That ordering is the falsifiable part. Run the same two comparisons against your own corpus, and if the retriever wins, I want to see the numbers, because that result would tell me the crossover sits somewhere other than where I measured it.

The cheaper thing to do first takes an afternoon. Log, for every multi-intent request, how many sub-intents ended up with zero passages in the packed context. A support system that cannot tell you which question it dropped will keep answering that question anyway, about half the time, out of whatever else was in the window.

The post Query decomposition doesn’t fix context starvation — it just moves it appeared first on The New Stack.

  •  

Cursor acquired Firetiger. A month later, it launched a bot that tracks code changes from PR to production.

Inspecting changes on a laptop screen document

We all know that producing code is easier than ever thanks to the abundance of AI coding tools and agents. The harder part undoubtedly comes after that code is written: making sure changes are safe to ship, spotting regressions in production, and figuring out what went wrong.

And that’s why Cursor is introducing Rollouts, a new agent that follows code changes into production and monitors whether they behave as intended.

The Firetiger effect

The announcement comes a little over a month after SpaceX closed its bumper $60 billion acquisition of Cursor, giving the AI coding company access to SpaceX’s vast GPU infrastructure as it develops its own models.

The day before that deal closed, however, Cursor quietly announced an acquisition of its own: it snapped up the team behind Firetiger, a three-year-old startup building AI agents that monitor software changes from pull request through deployment.

At the time, Firetiger co-founder and CEO Rustam Lalkaka argued that coding agents had dramatically reduced the effort involved in creating software changes, while doing little to reduce the risks involved in actually deploying them.

“Over the last two years, agentic coding has changed software dramatically,” Lalkaka wrote in a LinkedIn post following the deal’s announcement. “The cost of creating changes has dropped to near zero. The cost and risk of deploying them has stayed largely the same.”

“Writing code is no longer the slow part. What hasn’t sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.”

Rustam Lalkaka, Cursor

Fast forward to today, and Lalkaka, now at Cursor, has unveiled the first fruits from that acquisition — including Rollouts. In a blog post published on Wednesday, Lalkaka notes that the new agent, or “bot” as the company calls it, is all about helping developers “get safe, reliable code into production faster.”

“Writing code is no longer the slow part,” Lalkaka writes. “What hasn’t sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.”

Rollouts is effectively Firetiger’s Change Monitors reborn inside Cursor, rebuilt using a tool dubbed Bot Development Kit. This kit, too, appears to be new from Cursor: an early-stage framework for building and serving Cursor bots and agents, published as the @cursor/bdk package on npm. Its documentation says developers can define agents using Markdown and TypeScript, with support for tools, skills, subagents, webhooks and scheduled runs.

Like Change Monitors before it, Rollouts starts working when a pull request opens. It examines the proposed code change, works out which systems could be affected, and produces a monitoring plan covering what the change is supposed to do, the risks it sees, the signals it intends to watch, and any holes in the available instrumentation. Developers can review and edit that plan before the code reaches production.

Rollouts in action (1)
Rollouts generates a monitoring plan for a change

Once the change is deployed, Rollouts checks the resulting telemetry — including logs, metrics and traces — against that plan. Staging and production are assessed independently, with each deployment ultimately receiving one of three verdicts: verified healthy, regression detected or inconclusive.

That means a change could, for example, pass its checks in staging before Rollouts subsequently spots a problem when the same code reaches production.

Rollouts in action (2)
Rollouts reports deployment status as changes ship

If Rollouts does detect a regression, it can identify the change it suspects, alert the developer responsible and, depending on how it’s been configured, either open a revert pull request for review or hand the problem to a Cursor cloud agent to attempt a fix. There is still a human in the consequential part of that loop for now: Rollouts doesn’t merge fixes or roll back deployments by itself, though it can pause a progressive rollout.

Lalkaka notes that Rollouts is already capable of picking up problems limited to a particular endpoint or region before they trigger a broader alert, while it can also distinguish expected changes in behavior from genuine regressions.

Also “coming soon” to Rollouts, according to Cursor, is an integration with feature flags so it can directly adapt the traffic reaching a change, while support for release trains and deployment freezes is also in the works.

Enter Security Reviewer

Alongside Rollouts, Cursor is also introducing an upgraded Security Reviewer bot, which first appeared in beta back in April.

At launch, the bot could automatically inspect pull requests for security vulnerabilities, authentication regressions, privacy and data-handling risks, agent tool auto-approvals, and prompt-injection attacks, leaving findings alongside the relevant code.

As with Rollouts, the idea is that developers don’t have to remember to invoke it manually: Security Reviewer can be set to run whenever a new pull request is opened.

Security Reviewer in action
Security Reviewer runs automatically on new pull requests

In its current guise, Security Reviewer analyzes pull requests in the context of the wider codebase, with a focus on exploitable issues such as injection flaws and broken authentication, and returns a severity rating, attack path and proposed fix.

“Security Review reads code the way a security engineer does,” Lalkaka writes. “Where does user input enter, where does it end up, what does it pass through on the way.”

“Security Review reads code the way a security engineer does.”

He says that things have sped up considerably, too: average review time has fallen 21%, from 4.8 minutes to 3.8, while developer acceptance of its comments has risen from roughly 45–50% to 60–70%.

Both Rollouts and Security Reviewer are available through Cursor’s Automations tab for customers on its Teams and Enterprise plans.

The Origin story

Digging into the nuts and bolts of Rollouts reveals how it might serve as a boon for Cursor as it builds out Origin, the fledgling Git-compatible code hosting platform it launched back in August.

Origin is essentially an effort to build an alternative to GitHub for an agent-heavy software development world. It remains early, with limited functionality, but Cursor has been clear that tighter integration with its own agents is supposed to become one of the main reasons to use it.

When Cursor announced the Firetiger acquisition last month, Maxime Prades on the Cursor product team noted in a blog post that the deal was part of a “broader investment in long-running, autonomous, context-aware agents for teams.”

And he pointed to Origin and Change Monitors as two examples of that investment.

“Agents that write code should also be able to tell whether it works in production,” Prades wrote. “Today, those systems are mostly separate. Cursor and Firetiger bring them closer together so an agent can ship a change, see how it behaves, and respond when something goes wrong.”

Rollouts offers an early glimpse of that. It can connect to either Origin or GitHub for source control, pull deployment events from continuous delivery systems, and use signals from Datadog and other telemetry providers. If it spots a regression, it can then pass the problem back to a Cursor cloud agent to investigate or attempt a fix.

Origin potentially gives Cursor a native home for more of that loop: its cloud agents can already create branches, commit and push code, and open pull requests against Origin repositories. Rollouts then adds information about what happened after.

That could become increasingly important as more companies take aim at GitHub’s central role in software development. Zed, for example, put Delta into public beta last week, with its own ideas about how source control should change for teams working heavily with agents.

Cursor also faces competition further downstream. Datadog’s Bits Release, launched in preview in June, similarly follows changes from pull request into production and checks telemetry for regressions. Harness has long offered automated deployment verification and rollback based on logs and metrics, while LaunchDarkly’s Guarded Rollouts can monitor feature releases for regressions and automatically reverse them.

What Cursor can potentially bring to the table is proximity: the coding agent, repository, pull request, security checks, and production feedback can all sit much closer together. Rollouts doesn’t require Origin — GitHub remains supported — but owning the forge gives Cursor more room to integrate those pieces over time. And that may prove more compelling than simply recreating GitHub’s existing feature set.

The post Cursor acquired Firetiger. A month later, it launched a bot that tracks code changes from PR to production. appeared first on The New Stack.

  •  

Can enterprises protect data without making AI less reliable?

Abstract long-exposure photograph of dark motion blur and light streaks representing digital speed and data flow.

As organizations invest in AI, many are discovering a new bottleneck: obtaining data that is both protected and useful. Engineering teams need realistic, production-like data to validate AI-generated changes, train models, test applications, and generate business insights. Yet, privacy initiatives can sometimes make that data harder to access, less representative of real-world conditions, or unable to preserve critical relationships between records. 

The Perforce Delphix “2026 State of AI and Data Privacy Report” highlights this challenge. Among surveyed organizations, 26% say privacy controls make production-quality data harder to obtain, 25% struggle to preserve relationships across data entities, and 51% cite data quality challenges. 

“Protecting data isn’t enough if it can no longer support the systems that depend on it.”

Protecting data isn’t enough if it can no longer support the systems that depend on it. For engineering teams, the question is whether their data protection strategies can preserve the qualities that make data valuable in the first place. You need a well-rounded data strategy with tools that maintain referential integrity and relationships across your environments.

What these statistics mean for practitioners

At first glance, statistics from the report, like “51% of enterprises cite data quality challenges,” may sound like a purely governance issue.

In practice, they represent engineering problems. Low-quality datasets can produce:

  • Inaccurate analytics.
  • Poorly trained AI models.
  • Incomplete test coverage.
  • Increased rework.
  • Delayed releases.
  • Reduced confidence in data automation.

“When an AI model is trained on incomplete or distorted data, its outputs become less reliable.”

When an AI model is trained on incomplete or distorted data, its outputs become less reliable. When test environments contain unrealistic data, defects can escape into production. When analytics datasets lack consistency, teams spend more time validating results than acting on them.

Data protection and utility are not opposing goals

A common misconception is that organizations must choose between privacy and innovation, but the most successful organizations know that compliance, quality, and speed can and need to work together.

Protected data still needs to be:

  • Realistic enough for testing and validation.
  • Representative enough for analytics.
  • Accessible enough for engineering teams.
  • Governed enough for regulatory requirements.
  • Connected enough to preserve referential integrity.

There’s a two-fold goal in the AI era: reduce sensitive data exposure while also creating trusted data that remains valuable after protection. All too often, enterprises sacrifice compliance for innovation or speed. That’s a big reason 84% of respondents in our report have a data privacy exception in their non-production environments.

“There’s a two-fold goal in the AI era: reduce sensitive data exposure while also creating trusted data that remains valuable after protection.”

Organizations can only move at AI speed when they have access to trustworthy data that accurately represents production conditions. When privacy controls degrade quality, limit realism, or restrict access to representative datasets, the data layer becomes the new bottleneck.

Why referential integrity matters more than ever

Many discussions about data privacy focus on masking sensitive fields. However, masked data that loses referential integrity between entities can create a different kind of risk.

A customer, order, or payment record may still exist, but if the relationships connecting those records break during protection processes, the data no longer resembles reality. Take billing validation, for example. It needs referential integrity when a customer has multiple products, charges, and invoices across several database tables to produce the correct products or make accurate charges. 

Broken relationships are especially problematic for modern AI and analytics systems. Analytics pipelines depend on consistent identifiers to join information across sources. AI and machine learning workflows depend on complete business context to identify patterns and make predictions. Software testing depends on realistic relationships between records to validate application behavior accurately.

When referential integrity is lost:

  • Analytics can produce incomplete or misleading results.
  • AI models can learn from flawed datasets.
  • Testing environments can fail to expose production issues.
  • Teams lose trust in protected datasets.

Importantly, these failures are often difficult to detect. Pipelines may continue running successfully while quietly producing degraded outcomes, which can be a very costly mistake.

This helps explain why 25% of surveyed organizations cited preserving relationships across data entities as a significant challenge. For practitioners, that statistic is a warning that privacy controls can unintentionally undermine the quality of AI and analytics initiatives if they fail to preserve business context.

Keep in mind that not all data protection solutions are created equal — enterprise-grade masking algorithms are key to preserving relationships. When applied consistently and at scale, these algorithms ensure the same input produces the same masked output across your systems and environments. Other masking approaches might be done piecemeal, resulting in broken relationships.

What engineering teams should measure

Many organizations measure privacy success through compliance metrics alone. However, AI-driven environments require a broader definition of success. Engineering leaders should evaluate privacy initiatives against several dimensions:

  • Data quality: Does the protected dataset accurately reflect production conditions?
  • Realism: Can developers, data scientists, and analysts use the data confidently for their intended purpose?
  • Referential integrity: Do relationships remain consistent across applications, tables, environments, and data sources?
  • Accessibility: Can teams obtain compliant data without introducing delays?
  • Provisioning speed: How quickly can trusted datasets be delivered when needed?

These measurements help organizations determine whether privacy efforts enable AI outcomes or create new obstacles.

Designing for governance by default

As AI adoption grows, privacy cannot remain a separate process that occurs after development begins. Organizations should instead map the entire data lifecycle, from data request and discovery to reuse and retirement. 

This approach helps ensure governance is built into workflows rather than applied as a late-stage checkpoint. It also provides stronger auditability, reduces compliance exceptions, and gives teams greater confidence that protected datasets remain fit for purpose.

Most importantly, it aligns privacy objectives with business outcomes instead of treating them as competing priorities.

Trusted data will become a competitive differentiator

The need for test data — in volume, coverage, and scale — is booming as agentic development continues to rise. AI has also increased the volume of change enterprises can generate, but the challenge of validating that change remains unsolved.

That responsibility still belongs to data. The organizations that gain the most value from AI will not necessarily be the ones with the most advanced models. They will be the ones with the most trustworthy data foundations — using a portfolio approach that combines data virtualization for speed, masking for security, and synthetic data for coverage as needed.

“In the AI era, trusted data, not fast model output, may be the ultimate competitive advantage.”

The research points to a clear lesson: If privacy controls uphold realism, quality, relationships between records, or access to representative datasets, they can enable AI success.

As enterprises continue investing in AI and data privacy, the real objective should be ensuring protection and utility coexist. Because in the AI era, trusted data, not fast model output, may be the ultimate competitive advantage.

The post Can enterprises protect data without making AI less reliable? appeared first on The New Stack.

  •  

Q.ANT gives away the software for its light-powered AI chips in a CUDA-style bet on developers

Q.ANT, a startup out of Stuttgart, Germany, builds processors that use light instead of electricity to do some of the math behind AI. The company pitches them as a way to run AI on a fraction of the power today’s chips need.

Now developers can start writing software for those chips without owning one. Q.ANT pushed a free, open-source software kit to GitHub this week that lets developers build and test programs on a normal computer, then run them on the real chips once they get access.

This is a move out of Nvidia’s playbook. Nvidia owes its lead in AI as much to CUDA, the software developers use to program its GPUs, as it does to the chips themselves. 

But with Q.ANT, the catch is the hardware. Q.ANT’s chips are running at a few research computing centers, and everyone else has to wait “the coming months” for cloud access through German provider IONOS or an on-site server from Q.ANT.

The kit, called the Q.ANT Native Computing Toolkit, is free on GitHub under a license that allows commercial use. Developers can work in Python or C. The key piece is a simulator that mimics the chip on a regular computer, with no Q.ANT drivers required.

What can it do today? The AI tools in this first version focus on running models that have already been trained. The examples read handwritten numbers, identify objects in photos and outline shapes in images. Training still happens on regular CPUs and GPUs.

The pitch for photonic computing is power. AI chips burn a lot of energy moving data back and forth between memory and the processor. Q.ANT’s chips do part of the math with light, specifically wave-shaped functions similar to a cosine, which regular chips calculate digitally. Q.ANT says AI models built around those functions get better results with fewer parameters, the settings a model learns during training. Fewer parameters means a smaller model, less data to move and less power. The kit includes examples comparing a standard model with one built Q.ANT’s way. Those comparisons are the company’s own.

“An ecosystem isn’t created by hardware alone. It emerges when the software layer is open and others can build on it,” said Michael Förtsch, Q.ANT’s founder and CEO. He calls the release the “Linux moment” of photonic computing.

Q.ANT is betting light can do the math itself. Lightmatter, one of the best-known companies in the field, now puts its focus on Passage, which uses light to move data between chips. The idea of light-based AI isn’t new, either. TNS covered MIT’s photonic processor for building optical neural networks back in 2017.

Q.ANT raised €62 million in July 2025 in a round led by Cherry Ventures, UVC Partners and imec.xpand. In March, it said its second-generation chips were running at the Leibniz Supercomputing Centre near Munich. The results it published from there compare the new chip with its old one: more than 50 times faster at the kind of math that does most of the work in AI models, and six times less energy on typical jobs, by the company’s numbers. Its bigger claims, like up to 30 times better energy efficiency, don’t say what they’re measured against.

Good software alone won’t carry a new chip. Nvidia has been building CUDA for nearly 20 years and is still adding to it, including deeper native Python support last year. Graphcore, the British AI chip startup, had its own software kit and still ended up being sold to SoftBank in 2024.

Q.ANT calls this the first openly available software kit for programming a photonic processor. That depends on how you count. Xanadu has offered free, open software for its light-based quantum computers since 2018. For now, developers can play with the simulator. What they can’t do yet is test Q.ANT’s power-saving claims on their own models. That has to wait until the chips open up.

The post Q.ANT gives away the software for its light-powered AI chips in a CUDA-style bet on developers appeared first on The New Stack.

  •  

“Impressive level of openness”: Xiaomi goes way beyond the usual open-weight playbook with MiMo-V2.6

A picture of an open laptop

New models are coming out thick and fast, almost on a weekly cadence, ranging from the powerful proprietary systems coming out of the major US AI labs to the more open alternatives being released by some of China’s biggest tech companies.

On Tuesday alone, Anthropic debuted Claude Opus 5.5, while OpenAI launched GPT-6 Sol and Luna, each accompanied by their the usual claims about how they outperform their rivals. Amidst all the hullabaloo of the frontier-model frenzy, however, Xiaomi also debuted MiMo-V2.6, another powerful open model from one of China’s growing ranks of AI developers.

All the initial headline numbers look pretty promising, too. The flagship MiMo-V2.6-Pro is a trillion-parameter model, with 42 billion parameters active at a time, a one-million-token context window, and support for text, images, audio and video. Broadly speaking, that puts it in the same frontier territory as the latest models from OpenAI and Anthropic: GPT-6 Sol has a 1.05-million-token context window, while Claude Opus 5.5 has a one-million-token window, though neither company discloses comparable parameter counts.

Xiaomi, for its part, makes broad claims of frontier-level performance across coding, agentic tasks, cybersecurity, multimodal work and research. Independent analysis lends some weight to those claims –Artificial Analysis gives MiMo-V2.6-Pro an Intelligence Index score of 46, ranking it first among the 114 large open-weight models it tracks.

Artificial Analysis  Intelligence Index
Artificial Analysis Intelligence Index



So far, so good. But arguably the bigger story in Xiaomi’s offering is the manner in which it trained the model, how much of that process it showed in public, and what it’s releasing afterward.

A public record

Xiaomi livestreamed its RL training through a public dashboard, exposing metrics from the production reinforcement-learning runs in real time over a five-day period starting on September 15. By the time the runs had finished, the dashboard showed costs of $854,044 for the smaller MiMo-V2.6-Flash model and $2,620,670 for Pro — about $3.5 million combined.

Xiaomi livestreamed its RL runs over a 5-day period.
Xiaomi livestreamed its RL runs over a 5-day period.

It’s worth noting that this figure covers only the RL stage; Xiaomi hasn’t said what pretraining the models cost. Even so, public RL bills are rare. The closest precedents came last year, when MiniMax said the RL phase of its 456-billion-parameter MiniMax-M1 cost $534,700 in GPU rental, and DeepSeek put the RL training of its 671-billion-parameter R1 at $294,000. Both were leading open reasoning models when they launched, though the comparison only goes so far: MiMo-V2.6-Pro is larger, and its RL run targeted longer, agentic tasks.

Shortly after the stream began, Fuli Luo, who leads Xiaomi’s MiMo team after previously working at DeepSeek, took to X to explain the thinking behind the project. The team, she said, had spent almost six months exploring how far RL could be pushed, increasing the amount of training, the variety of environments and agent setups, and the resources used to grade the model’s attempts.

“We’ll open-source the details piece by piece over the coming weeks,” she added.

Nearly half a year of silence. We spent it studying one problem: how far RL can scale.

MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task…

— Fuli Luo (@_LuoFuli) September 16, 2026

Responding on X, Hugging Face co-founder and chief science officer Thomas Wolf called the move an “Impressive level of openness on such a large run.”

However, what Xiaomi’s putting out alongside the finished models is arguably just as interesting. The company has released the model weights under the permissive MIT license, alongside its technical report and a 9-billion-parameter Qwen-based model, intended as a starting point for further agentic RL research.

“Impressive level of openness on such a large run.”

Xiaomi says it has also “fully open-sourced” a broader set of RL resources: more than 7,000 task environments spanning software engineering, vulnerability reproduction, knowledge work and web development; an end-to-end training framework covering everything from environment interaction to reward evaluation and policy optimization; and lightweight agent harnesses for experimenting with different tools, prompts and context setups. At the time of writing, however, Xiaomi’s link to the open-source collection on Hugging Face contain only the three model releases, with the 7,000-plus environments and other supporting resources not surfaced there. Luo had said earlier that Xiaomi would be open-sourcing the various elements “over the coming weeks.”

As the results began arriving this week, attention in the research community quickly moved beyond the benchmark score to what Xiaomi had committed to releasing overall. Elie Bakouch, a former Hugging Face researcher who is now a research engineer at Prime Intellect, singled out the promised RL resources.

“The most insane part, they will release ~7k RL training data and the framework leading to this top 6 model on AA,” Bakouch writes on X. “They also shipped the model + tech report less than 1 week after starting the final RL run.”

Wolf went further, arguing that access to the environments in which models learn may now be especially valuable for open research, as more model development shifts toward RL with verifiable rewards (RLVR). Because RLVR depends on tasks whose outcomes can be automatically checked — whether code passes a test, for example — the environments themselves become a crucial ingredient in training.

“Releasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now.”

“Releasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now,” Wolf writes. “The equivalent of sharing high quality pretraining data, but in the new RLVR paradigm.”

Open-weight vs open-source

So while the benchmarks around Xiaomi’s latest model are notable in their own right, it’s the company’s approach that is generating much of the fanfare so far.

Indeed, MiMo-V2.6 serves as a useful example of a distinction that often gets muddied in the AI sphere: “open-weight” and “open-source” are routinely used as though they mean the same thing, but they don’t. Many “open” models amount largely to downloadable weights — essentially, the vast collection of numerical values a model learned during training, which can then be used to run or fine-tune it — while much of what went into producing them remains closed.

Some companies have gone further in muddying those terms. Meta, for example, has often referred to its Llama models as open-source despite significant restrictions that have led open-source advocates to push back heavily on that description.

And so MiMo-V2.6 goes further than most open-source releases. Its MIT license carries none of the conditions that the likes of Moonshot’s Kimi K3 and Alibaba’s Qwen3.8-Max attach for large commercial users. And if the environments are released as promised, outside researchers will have much more of the post-training process to inspect and build on.

The post “Impressive level of openness”: Xiaomi goes way beyond the usual open-weight playbook with MiMo-V2.6 appeared first on The New Stack.

  •  

A third option is emerging in the fight over AI and your data

Split-screen video interview with The New Stack host Alex Wilhelm and VAST Data cofounder Jeff Denworth.

Not your keys, not your coins. Not your model, not your data?

Over the summer, the tech industry was consumed by a debate about AI use in the enterprise and the need to protect IP. If an enterprise used proprietary models, was data leakage a necessary evil?

Companies seemed to have two options: They could use state-of-the-art, proprietary models and risk losing control of their data, or they could use open-weight models and never kiss the frontier.

Thankfully, a third option is emerging.

Consider the concern: Company A wants to use LLM B from AI Lab C, and they want to avoid training AI Lab C how to eat Company A’s lunch by building its capabilities into LLM B. A good way to resolve the tension would be to let Company A run LLM B on its own infrastructure, so there’s no risk of its information fleeing on the wind.

AI agents are “creating a whole different set of requirements at the data layer.”
–Vast Data co-founder Jeff Denworth

But that raises another problem: AI Lab C doesn’t want to allow Company A to run LLM B on its own GPUs because it doesn’t want to hand over its model weights. It’s the same IP issue the company ran into, in reverse. You have to solve the trust problem in both directions!

Enter VAST Data co-founder Jeff Denworth and a new product called DataEnclave, which aims to let AI labs and enterprise-scale companies deploy proprietary models in secure compute environments without risking data transfer in either direction. (DataEnclave uses Nvidia’s Confidential Computing technology to make the system tick; Vast Data’s core product is AI OS, infrastructure that fits beneath a company’s AI applications.) 

The New Stack had Denworth on the podcast to chat about the confidential computing market. I was curious about timing. Why did Vast build DataEnclave now? Nvidia began rolling out Confidential Computing in a serious way in 2024, after all. Denworth argues that the market needed the core technology, yes, but also demand.

And until late 2025, AI demand was modest compared to today’s token totals. Once agentic coding tools took off, corporate demand for AI products soared. This led to the pricing crisis we saw in early 2026, and the secure AI usage debate we endured over the summer. 

Performance drove demand, demand drove usage, and usage dug up fresh problems to solve. Now the question for the market is whether or not DataEnclave has solved enough concerns on both sides of the proprietary AI-proprietary data equation. The market will sort that out as it moves through early access and into general availability.

Our conversation goes deep into the arc of AI, where companies are in their AI journey today, and how much data remains to be unlocked inside the enterprise. If you want to feel the acceleration, it’s a fun one!

The post A third option is emerging in the fight over AI and your data appeared first on The New Stack.

  •