Over its five-year history, P99 CONF has hosted quite a few speakers who’ve offered pointed takedowns of the namesake metric. At last year’s conference, Adrian Cockcroft didn’t explicitly state that P99s are BS… but he did allude to it.
If you don’t know Cockcroft, he’s spent decades architecting, scaling, and optimizing resilient, high-performance systems at giants like Sun Microsystems, Netflix, eBay, and Amazon. We could probably dedicate an entire day of P99 CONF to discussing the lessons learned from just some of his projects (Solaris kernel performance, multi-processor optimization, Netflix’s on-prem to cloud migration, Chaos Monkey…)
Fortunately, RedMonk analyst Rachel Stephens proved the perfect host for a conference that’s all about making things fast. She sat down with Adrian and led us on a whirlwind tour of how AI has impacted performance engineering. Here are some highlights from the chat (full video below).
Note: P99 CONF 2026 – a free + virtual conference on all things performance – is going live October 21-22. Graba complimentary passand join us!
From kernel analysis to vibe coding perf tools
As a performance specialist at Sun in its heyday, getting to the root of performance problems involved lots of digging and divination. Cockcroft recalls, “Back in the old days with Sun, people would look at the output of system metrics in vmstat or whatever, and they’d be guessing what the numbers meant. There was a very vague understanding of what these things meant. The manual page wasn’t very clear.”
Cockcroft ended up going to the source, literally. “I went and read all the kernel source code and figured out exactly where these numbers came from, exactly what they meant, which ones were approximating what, and wrote all that down.” That led to two performance books: Sun Performance and Tuning and Resource Management.
“My speedup is infinite, because this code would never exist without these tools. I wouldn’t have the time to build them.”
Four decades later, there’s now a wealth of helpful tools for end-to-end tracing, but Cockcroft’s curiosity still lies in what the tools are not showing. He continued, “Everything sort of looks okay in the tools – but the system isn’t behaving well. I usually come in and try to find a new way of looking at the data. A new type of analysis, or go a little bit deeper or finer grain, or stop looking at averages and start looking at distributions, and find all kinds of interesting things that nobody knew were happening.”
Currently, he’s vibe coding tools to better analyze the anomalies he finds. Saved from having to brush up on Python or hunt down graphics library fragments on Stack Overflow, Cockcroft can now stand up custom tooling in minutes. “My speedup is infinite, because this code would never exist without these tools. I wouldn’t have the time to build them.”
Peaks not percentiles
One specific vibe coding project: Cockcroft built (and open-sourced) tooling to get a better understanding of response time distributions.
Response time distributions have been on Cockcroft’s mind for over a decade. While most people obsess over percentiles – yes, P99 CONF included – Cockcroft is most intrigued by the distribution of response time peaks in a histogram. He believes percentiles don’t work when trying to understand the latency and performance of modern web services. A single number like P99 can’t tell you whether the underlying distribution has one peak or several. And when there’s more than one peak (as is often the case in the real world), the mean, the standard deviation, and even the P99 itself lose most of their meaning.
“Percentiles don’t work when trying to understand the latency and performance of modern web services.”
For example, assume you have a histogram with two response time peaks: a fast one from a cache hit and a slow one for misses that require actual work. As the cache hit rate shifts, each peak’s position remains the same (i.e., the latency values of the fast-response mode and the slow-response mode don’t change), but the peak heights rise and fall. “Your averages and your P99 are changing all over the place, but all that’s really happening is your cache hit rate is changing,” Cockcroft said.
So how do you go beyond measuring P99s and averages? Cockcroft did what he’s done for decades: dive in and build a custom tool. But these days, it’s much simpler thanks to LLMs.
“Your averages and your P99 are changing all over the place, but all that’s really happening is your cache hit rate is changing.”
He had already worked out the statistical approach for analyzing the distribution. Once ChatGPT came out, he quickly used it to build a tool that automated it. Instead of collapsing everything into an average, it identifies an arbitrary number of peaks in a distribution and tracks how they fluctuate over time. It’s implemented in R – a language Cockcroft hadn’t used in a while, but ChatGPT knew quite well – and it’s open source. If you’re curious, learn more in his Percentiles Don’t Work article and “A Tale of Two Histograms” talk and deck (“It was the best of response times, it was the worst of response times…”)
Where do we go from here?
To close, Stephens asked Cockcroft what advice he’d share with teams working on high-performance systems today. His top tip was to start with the macro view to find what’s interesting, then keep digging deeper until you’re inspecting individual slow requests end-to-end.
“Remember the microscope that you got when you were a kid,” Cockcroft said. “First, you have to focus it using the lowest resolution, at 10x, and then you can click it to 100x and adjust that, looking at just one speck now. Once you get that in focus, you click it to 1,000x.”
Cockcroft has spent his career building tools that bring obscure performance issues into focus. We look forward to seeing what others have cooked up with agentic tooling to help identify and solve performance problems this year at P99 CONF.
Learn about the latest performance optimization techniques, tooling, and case studies at P99 CONF – free and virtual, October 21-22. Graba complimentary passand join us!
AI is changing expectations around infrastructure and operations, including Kubernetes management. When models run close to the data they use, deployment, scaling, and governance responsibilities tend to shift to platform teams. And as clusters, environments, and operational signals continue to multiply, manual operations often strain under the added weight.
AI may simultaneously provide opportunities to lighten this growing load. Agentic software can now observe a system, reason about it, and act within predefined limits.
Ultimately, these platforms’ value depends on the quality of the context an agent can see and the boundaries you set. Without cluster state, policy, and access rules, an agent can only guess.
Without cluster state, policy, and access rules, an agent can only guess.
For agentic AI to streamline multi-cluster management, you need clear lines between what the system observes, what it recommends, and what it changes. Drawn well, those lines let teams gain notable speed while still maintaining control.
The impact of AI on computing infrastructure
Teams once treated AI as an application concern; models sat on top of existing systems, and the stack underneath stayed mostly unchanged. Today, AI reaches into more and more customer interactions, while data storage needs simultaneously expand and orchestration pressure grows. A recent Forrester report describes the modern AI computing stack as stretching from the models themselves into and across the infrastructure beneath them.
As AI workloads move into production, they place new demands on the infrastructure beneath them. Many lean on specialized compute, with resource needs that rise and fall through bursts of training and inference. Because conditions shift quickly, they can also call into question whether telemetry remains trustworthy. Each of these demands lands at the infrastructure layer, where the workloads run.
The infrastructure layer of the new AI stack
The infrastructure layer covers compute, storage, and networking. It is a foundation that every workload running on the layer depends on. As AI workloads grow, choices about capacity, placement, and control will increasingly shape the performance of the data, intelligence, orchestration, and experience layers atop the infrastructure.
To operate the infrastructure layer efficiently across many machines and locations, a team may rely on orchestration instead of managing servers by hand. In cloud native contexts, Kubernetes has become a control point for scheduling workloads, applying policy, and presenting a consistent interface across environments. Kubernetes is especially well-suited to support organizations this way when teams need consistent control across an estate spanning data centers, clouds, and edge sites.
Agentic AI and Kubernetes: the future of the infrastructure layer
Agentic AI can extend automation from fixed rules to systems that adapt to real-time conditions. Traditional automation runs the same script whether the environment has changed, while an agentic system observes the environment, reasons about what it finds, and then takes action.
When you apply agentic capabilities to multi-cluster management, the system follows this same sequence. An agent reads cluster state and operational data, proposes a diagnosis or next step, and then carries out actions based on an approved scope, usually after a person signs off. You can further reinforce these boundaries by routing each request to a specialized agent that receives only the metadata it needs.
The signals that an agent receives from the cluster, the context about policy and access, and the definitions of what the agent may change are the key elements that give agentic systems their value. They also separate agentic AI on Kubernetes from a generic assistant.
Manual Kubernetes management is less efficient at scale
Admittedly, agentic AI fits some settings better than others. On a small single-cluster footprint, the overhead may outweigh the benefit. Manual Kubernetes management often holds up on a handful of clusters, but it can become unreliable in a rapidly growing estate. After all, each new cluster adds lifecycle work across upgrades, patching, configuration, and renewal. Those tasks can quickly multiply and diverge in hybrid environments.
Configuration drift is a high risk in these situations. Settings that started identical can fall out of sync, and policies can apply unevenly from one team to the next. Individually, these gaps may be manageable, but collectively they raise the odds of an outage or a failed rollout.
Visibility can also erode in an unmanageable way. Clusters spread across data centers, clouds, and edge sites often leave teams with no single view of the whole landscape. When DevOps and platform engineers stitch together signals from separate tools, resolution can slow and become more error-prone. A unified view helps enable sound, efficient decision-making by people, agents, or both.
Kubernetes knowledge is fragmented, and existing AI tools lack business context
Kubernetes expertise often sits unevenly across an organization. For example, senior engineers may hold deep operational knowledge that application teams lack. The most current information about a running system may also be fragmented if logs sit in one tool and metrics in another. Real-time understanding can be further clouded when policies, runbooks, access rules, and deployment history each live elsewhere.
Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes. Without that context, even a capable AI tool may fall short of providing meaningful Kubernetes management support.
Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes.
When an agent can read current signals alongside the rules that govern them, its suggestions become specific, testable, and actionable. In an incident, agentic systems can correlate logs with a recent change. Ahead of a rollout, they can check the change against policy. During troubleshooting, they can account for access rules rather than guessing at them. Kubernetes decisions carry real operational consequences, which makes these details all the more important to consider.
Engineering “toil” isn’t time-efficient
Site reliability teams use the word “toil” for repetitive manual work, especially tasks that keep systems running without adding lasting impact. In Kubernetes operations, toil takes the form of repeated triage, manual signal correlation, alert follow-up, and routine checks. The tasks aren’t particularly difficult, but they can consume significant time and attention for enterprise teams.
When engineers spend their days on this kind of investigation, proactive modernization efforts tend to stall and planned upgrades can slip behind schedule. In other words, the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements. In a recent survey about how AI provides value to DevOps teams, reducing toil emerged as one of the clearer opportunities.
…the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements.
Agentic AI can support repetitive investigations by gathering signals, correlating them, and proposing a likely cause for an engineer to weigh.
Kept under human review, it can take on some of the routine correlation that would otherwise fall to the team. That kind of support can give engineers more room to focus on the strategic work that most needs their judgment.
Building more intelligent infrastructure with agentic AI and Kubernetes
As you consider building toward intelligent infrastructure without surrendering control, the following principles can inform your efforts:
Start with observable context, giving agents access to current cluster state, policy, and history before they reason about a problem.
Separate suggestions from actions, allowing agents to recommend freely while any change must wait for human approval and a defined scope.
Connect agents to existing controls, routing their work through the access rules, identity, and audit paths the team already trusts.
Keep the ecosystem open, favoring platforms that integrate with current tools and standards over those that lock work into a single stack.
Platforms like SUSE Rancher Prime and SUSE AI Factory embrace these principles and illustrate how Kubernetes management can become a foundation for agentic operations. These platforms can help you improve cluster and policy consistency without compromising your authority over AI. Built on open-source foundations, they can also help you avoid being trapped in a single vendor’s stack.
In SUSE Rancher Prime, the industry’s first context-aware agentic AI ecosystem, its AI assistants work as a crew of specialized agents with an intelligent router. The platform draws on the cluster context already in place and acts through existing access controls. Through support for external Model Context Protocol (MCP) servers, teams can extend that crew to their own sources. In addition, human validation tools allow you to hold a proposed action for approval before the agent runs it.
Despite its potential, intelligent infrastructure is not universally beneficial. In situations where change control must stay fully manual, for example, agentic AI’s role may be strictly limited to observation and suggestion. Measure the technology’s value against the realities of your day-to-day operations. For those who are investing, agentic AI will have the greatest impact when it actively supports context, control, openness, and human judgment.
Every infrastructure team makes decisions that are difficult to reverse. Most of the time, that works out. Sometimes it does not.
Vendor lock-in usually begins as a reasonable choice, made under time or budget pressure, that solves a real problem at the time. A managed service ships faster or a deployment model fits better in that moment, but eventually a difficult constraint appears.
When business conditions inevitably change, those accumulated choices and their consequences will determine whether a team can pivot accordingly. Limits on flexibility rarely trace back to a single vendor; more often, they hinge on how reversible the team’s past decisions are.
What is vendor lock-in and how can it harm your business?
The risks of vendor lock-in are not really about relying on vendors, since every production system relies on vendors. The big issue is dependencies that become too expensive or impractical to unwind.
For a platform team, that dependency builds up across APIs, contracts, roadmaps, and data models. It extends further into managed services, identity patterns, observability pipelines, and operational tooling. Each piece likely represents a reasonable design choice, but together they can quietly limit your options and raise the cost of leaving. When switching a database or control plane means rewriting tons of integrations, retraining the whole staff, or migrating data under inconvenient timelines, you have lost the room to maneuver.
The big impacts of small, invisible and unexamined decisions
Not every dependency is automatically a problem; some are understood, contained, and worth the tradeoff. The real risk lives in the dependencies no one examined closely, which may stay invisible until they block the business from evolving.
“The real risk lives in the dependencies no one examined closely, which may stay invisible until they block the business from evolving.”
Unfortunately, some teams are familiar with these invisible dependencies. A managed database might pick up proprietary extensions, which application code then starts to assume. A Kubernetes environment might bind to one cloud’s IAM, networking, storage, and load balancer model. Observability and logging pipelines might harden around a single provider’s formats. None of these choices is reckless on its own, but together they can create significant friction.
Obstacles to change and their hidden costs
The extent of a dependency-based tradeoff can sometimes remain unknown until circumstances shift, such as a new compliance requirement or customers needing a new deployment model. The hidden costs of these moments often escalate in stages. It might start with a visible, unwelcome migration bill, but the expense can also show up as operational drag. Rushed migrations can lead to additional service disruptions later. A workload may be unable to move, limiting services to certain customers. When you are tied to a specific vendor’s release cadence, it can make it difficult or even impossible to adopt emerging technology.
Concentration risk compounds the problem, because a single change from one provider that carries pricing, support quality, and roadmap can ripple across the estate. By the time a switch becomes necessary, the cost shows up as service disruption, complex data transfer, and retraining. Naming these costs early keeps them from arriving as surprises.
At some point, a dependency can accumulate enough of these costs to become more than an architectural detail. Once it affects budgets and timelines, leadership has to account for it—and the team has to be ready to explain it. Identifying these dependencies early gives everyone time to plan.
Open source offers a different path
One way to proactively address this pattern is to evaluate potential dependencies more deliberately. For example, before committing to a platform or service, try to determine its reversibility. In other words, establish how difficult it would be for the team to change its mind about the investment in the future.
“Open source offers no guarantee against lock-in, however, since a team can still build tight coupling on open foundations.”
Open source solutions tend to perform well against that test, because they are intentionally built to keep systems inspectable, portable, supportable, and replaceable. By design, open source makes it easier for you to preserve options over time. It offers no guarantee against lock-in, however, since a team can still build tight coupling on open foundations.
What is open source?
Open source describes software you can inspect, run, modify, extend, support, and replace with relative ease compared to proprietary alternatives. The software’s source is available, and the license grants you the right to use and change it. Notably, no-cost or freeware software is not necessarily open source, specifically if it does not provide this level of access and rights.
Several companies have open source principles at their core, and open source software can be extremely valuable in enterprise contexts. Transparent code is often easier to audit, and open standards can reduce friction when moving between tools.
Open source also changes who can move the goalposts
For developers, reversibility is not only about APIs and data formats. It is also about whether one company can change the terms underneath a foundational technology. The Linux kernel is a useful example. Linux kernel documentation notes that copyright assignments are not required, so merged code retains its original ownership and the kernel now has thousands of owners. That makes unilateral relicensing of the kernel effectively impractical.
Kubernetes has a different legal structure, but the practical protection is similar. The project is licensed under Apache 2.0 and governed by the Cloud Native Computing Foundation. The license grants users durable rights to the existing code, so no single vendor, including SUSE, can retroactively take those open-source rights away from the project as it already exists. That matters because a platform can remain available even if a particular vendor changes strategy.
The Terraform-to-OpenTofu fork shows why this is more than a theoretical distinction. In 2023, HashiCorp changed Terraform’s license from the Mozilla Public License 2.0 to the Business Source License 1.1. The community responded by forking the last open-source codebase into OpenTofu, now a Linux Foundation project that remains under the MPL 2.0. The lesson for developers is not that every open-source project is immune to licensing changes. It is that open licensing and neutral governance can preserve a viable exit path when a vendor changes direction.
Open source powered by enterprise discipline
Open source ultimately earns its place through engineering discipline. Source availability has benefits but does not resolve governance, patching, lifecycle management, documentation, security, or integration on its own. A community project can be powerful and nonetheless arrive without enterprise-grade operational guarantees.
Enterprise open source providers exist and can help with closing that gap. They embrace open foundations and add the support, security, maintenance, and lifecycle discipline that production environments require. Founded in 1992, SUSE was the first provider of an enterprise Linux distribution. Today, it focuses on helping organizations operationalize open source with enterprise-grade support.
These companies aim not to close off open source software but to make it dependable at scale. In other words, open source and operational rigor can coexist. And enterprises should expect both from any external provider.
Digital sovereignty: the x-factor that makes open source even more critical
Digital sovereignty describes how much control an organization has over its infrastructure, data, operations, and technology choices. Sovereignty is a spectrum, and architecture decisions can move an organization a step in either direction.
Recent research by SUSE suggests that almost all enterprises are prioritizing digital sovereignty, but only 52% are actively taking steps toward it. That gap is largely an execution problem, and much of it surfaces in everyday platform decisions.
If your team supports regulated industries or deploys in on-premises or air-gapped environments, you may be especially familiar with growing pressures around sovereignty.
Sovereignty puts a deadline on work that was already worth doing
Developers can hear “digital sovereignty” and assume it means a separate compliance workstream with a separate engineering bill. In practice, much of the work is the same discipline platform teams already invest in: portable workloads, clean interfaces, automated verification, reproducible deployment, auditable behavior, and the ability to replace a dependency without rewriting the system around it.
“Sovereignty does not suddenly make that engineering work valuable. It puts a deadline on work that was already worth doing.”
Those practices already have an economic case. They reduce migration costs, lower operational risk, make platform changes less disruptive, and preserve options when pricing, regulations, or business requirements shift. Sovereignty does not suddenly make that engineering work valuable. It puts a deadline on work that was already worth doing.
That reframe matters because it turns sovereignty from a policy overlay into an architecture property. The useful question is not simply, “How much extra work will sovereignty cost?” It is, “Which parts of our stack already fail the portability, interface, and verification tests we would want anyway?”
How to strengthen sovereignty with open source
Sovereignty depends on how a team designs, deploys, and operates its systems. Open source does not make an organization sovereign by default, but it can improve the conditions for sovereignty.
In fact, many of the same questions that expose lock-in also matter for digital sovereignty. Each of the following questions about reversibility connects to open source and sovereignty alike:
Reversibility question
Why open source can help
How sovereignty strengthens
Can we run this workload elsewhere?
Open source typically runs across on-premises, cloud, hybrid, and edge environments, not just one vendor’s platform.
More control over where workloads run, including specific regions and regulated contexts.
Can we understand and audit how it works?
Source availability and community scrutiny improve inspectability over closed alternatives.
Teams can verify behavior, assess risk, and meet assurance requirements.
Open ecosystems favor open formats and interoperable tooling.
Data stays more portable, improving control over storage and movement.
Can another team or partner support it?
Multiple support paths exist, from internal teams to integrators and enterprise vendors.
Less dependence on one vendor’s pricing, availability, or roadmap.
Can we replace one component without rewriting everything?
Open interfaces and modular design make components easier to swap.
More control over architecture as requirements change.
Can we keep operating if a vendor changes direction?
Open source projects can outlast one vendor’s strategy or license.
Less exposure to decisions the team cannot control.
Can we deploy closer to the data?
Open source can run in private data centers, sovereign clouds, edge sites and hybrid models.
Sensitive workloads, including AI, can be governed nearer the data.
The ongoing work of digital sovereignty
Sovereignty is more of a practice rather than a specific destination. For many teams, the work begins with identifying existing dependencies that are especially hard to reverse. Similarly, you’ll need to separate the tradeoffs worth accepting from the ones that remove a significant number of options.
Moving forward, it can be helpful to prioritize open interfaces and portable foundations when possible. When evaluating new services or solutions, treat lifecycles, support, and governance as first-order concerns.
In some cases, sovereignty work can be too heavy for an in-house team to carry alone. Providers such as SUSE can help strengthen your operational layer, including security and observability, and especially in growing or hybrid contexts.
Automated checks can make those principles concrete by continuously testing whether workloads can be rebuilt, moved, audited, and recovered instead of waiting for a migration or compliance event to expose the gaps.
Open source lets you take control of your software ecosystem
No enterprise team avoids every dependency, and candidly none should try. Some coupling is reasonable, contained, and worth it. A vendor-free system is not a realistic goal for a major enterprise. A realistic goal is the judgment to separate acceptable dependencies from dangerous ones.
“The true cost of any platform includes the cost of leaving it, and teams should understand that cost before they commit.”
Reversibility gives that judgment something concrete to work with, because it can be broken down into capabilities a team can name, evaluate, and test:
Ownership. Ownership does not mean building everything yourself. It means holding the realistic ability to run, move, or hand over each layer of your stack. The test is simple: if a vendor disappeared tomorrow, or was ordered to stop serving you, what still runs next month?
Auditability. You should be able to verify what your software does, yourself or through an auditor you appoint, rather than accepting a vendor’s report as the final word. With open source, inspection is a property you hold. With closed software, it is a permission you are granted, and permissions can be withdrawn.
Exit velocity. An exit plan without speed is just a document. Exit velocity measures how fast a workload can move from one platform to another, and it only means something when you test it on a schedule, as earlier generations tested disaster recovery.
Pivot ability. These capabilities matter when conditions change: a new compliance requirement, a customer that needs a different deployment model, or a vendor that changes direction. Teams that can reroute workloads respond on their own timeline. Teams that cannot must renegotiate from a position of weakness.
Vendor lock-in becomes a manageable risk when you can confidently flag which decisions are hard to undo, weigh the tradeoffs honestly, and protect the team’s pathways to change. Open source strengthens every one of these capabilities because it keeps larger portions of your system inspectable, portable, and replaceable.
The true cost of any platform includes the cost of leaving it, and teams should understand that cost before they commit.
Welcome to another edition ofRoad to KubeCon, where we’re tracking the major movements in the Kubernetes and cloud native ecosystem on the path to KubeCon + CloudNativeCon NA 2026, happening Nov. 9–12 in Salt Lake City, Utah.
This week, we look at how cloud-native teams are putting observability to work. There’s progress on OpenTelemetry and Prometheus interoperability, a migration spanning 100,000 hosts, and new data on the costs and benefits of monitoring AI systems. Plus, HPE’s latest Gartner recognition, agent governance updates, and a father-and-son story from KubeCon India.
HPE GreenLake named a Leader in Gartner quadrant
On Wednesday, HPEannounced it had been named a Leader in Gartner’s Magic Quadrant for Infrastructure Platform Consumption Services for the second consecutive year.
Hewlett Packard Enterprise (HPE) is a presenting sponsor of Road to KubeCon.HPE Softwarehelps IT organizations modernize infrastructure, streamline operations, and accelerate AI initiatives across hybrid, multi-vendor environments.
GreenLake, HPE’s cloud operations platform, helps teams monitor resource consumption, secure data, and manage infrastructure across data centers and private and public clouds. HPE points to recent additions, including agentic AI-powered operations, as part of the platform’s development.
Varma Kunaparaju, senior vice president and general manager of cloudops software and platform at HPE, says in the announcement: “We are building the operating model and platform for the agentic enterprise, giving customers the ability to simplify operations, govern intelligently, continuously optimize, and modernize without sacrificing choice.”
OpenTelemetry and Prometheus work better together
OpenTelemetry (OTel) and Prometheus are widely used for cloud-native monitoring and observability, often side by side. A new survey looks at how well that combination works.
While the two ecosystems haven’t always worked well together, the 2026 survey shows improvement: the average ease-of-use rating rose 0.5 points, from 3.1 to 3.6, while the share of those who find the two hard to use together fell from 29% to 10%.
There’s still work to do. Respondents want better alignment between the projects’ data models, better handling of resource attributes and metadata, and fewer naming and formatting issues.
Atlassian moves metrics from 100,000 hosts to OpenTelemetry
The original pipeline had worked for years, handling metrics from roughly 100,000 hosts across 14 regions, but was increasingly out of sync with the shift to OTel. “It became the thing everyone standardized on, and more and more of what fed our pipeline was emitting OTel data we simply didn’t support,” write Atlassian’s Iris Grace Endozo, Farzad Vazirnia and Albert Kerr.
To maintain continuity throughout the migration to OTel, Atlassian swapped the collection and pipeline mechanics underneath while keeping the service-facing interface unchanged. This turned an organization-wide overhaul into what the authors call a “platform-team migration.”
According to the authors, aggregation now uses about half the CPU for the same traffic. Operations are more unified through the OTel Collector, CPU usage is more evenly distributed across ingest shards, and sidecar costs are down roughly 30% at fleet scale.
New Relic finds observability gains — and gaps
On Tuesday, New Relic released its 2026 Observability Forecast, based on a survey of 2,575 IT and engineering leaders and practitioners. The report found that 73% are standardized on OTel, actively migrating to it, or testing it.
The report also looks at observability’s role in AI adoption. It found that 83% of respondents consider observability essential for AI-generated code. Organizations monitoring AI agents are twice as likely to report a threefold return on observability investment as those running agents without monitoring.
The study also paints a picture of the impact of outages. Engineers now report spending 37% of their time addressing disruptions, while 42% of organizations learn about disruptions through inefficient channels, like manual checks or customer complaints.
Outages take a business toll. New Relic found that organizations lose $74 million a year on average due to high-impact outages. That’s $1.85 million per hour, or over $30,000 for every minute a system is down.
The findings show why teams are looking for ways to detect and resolve problems faster as their systems grow more complex.
As Kubernetes evolves, so do the demands on the teams running it. Presenting sponsorHPEhelps teams address that complexity with software spanning virtualization, cloud management, observability, and automation.
Observability Day returns to KubeCon
If you’re into observability and attending KubeCon NA, definitely check out the agenda for Observability Day, happening during the co-located events in Salt Lake on November 9.
OpenTelemetry’s graduation in May and growing production use give teams more experience to draw on as they adopt the standard.
New AI workloads, inference monitoring, and interoperability with other projects still present challenges. Those issues give practitioners plenty to compare notes on.
According to the Observability Day schedule, the agenda includes project updates and lessons from Capital One, Cisco, Nubank, and other organizations.
“Observability Day provides a vendor-neutral place for maintainers and practitioners to compare approaches and learn how the wider ecosystem is responding,” write Austin Parker, Iris Dyrmishi, Eduardo Silva Pereira, and Juraci Paixão Kröhling on the CNCF blog.
The additions let engineers deploy autonomous workflows and build or import agents under shared governance and context. Komodor says the release responds to the growing use of agents, including coding agents, and concerns about governance, return on investment, and costs.
“The hardest part of running agentic operations in production is not building the agents,” shares Itiel Shwartz, Komodor’s co-founder and CTO. Instead, the challenges lie in maintaining context, persistent memory, accuracy, security, and cost control — things the new Komodor release aims to address.
Spectro Cloud expands in the Middle East
The Middle East is an increasingly important technology market and a hotbed of data center construction. At the same time, data sovereignty and compliance requirements are driving interest in sovereign infrastructure.
“The Middle East has bold ambitions for global AI leadership, from sovereign AI factories to AI-powered economies,” shared Tamer Riyal, Spectro Cloud’s sales director for the Middle East, in the announcement. “We’re investing in the regional expertise and partnerships to support that vision for the long term…”
The company aims to expand adoption of PaletteAI, its platform for managing AI infrastructure across sovereign clouds, enterprise data centers, and edge locations. The expansion reflects the region’s growing role in cloud-native infrastructure.
A father and son take the KubeCon stage
This week’s updates also have a personal side. Earlier this year, analyst, advisor, and TNS columnist Janakiram MSV co-presented a talk at KubeCon + CloudNativeCon India 2026 with his son, Shreyas Mocherla, a CNCF Kubestronaut and software engineer at Nirmata.
As Mocherla describes on the CNCF blog in a post published this week: “Presenting alongside him made this special on a level that goes beyond the conference itself. I grew up watching him speak at technology events. Standing next to him at the same podium, in front of the KubeCon audience, felt like a full-circle moment.”
Vector database Pinecone unveils GA for its Bring Your Own Cloud (BYOC), allowing users to run Pinecone from AWS, GCP, or Azure accounts.
Cilium publishes patch releases for its maintained 1.18, 1.19, and 1.20 branches.
Follow the Road to KubeCon
Road to KubeCon is an eight-part series presented by HPE, which will be at KubeCon + CloudNativeCon North America in Salt Lake City. Before you go, explore howHPE Softwarehelps IT teams do more with less complexity.
We’ll be here every Friday until KubeCon.
If you’d like to participate, Bill Doerrfeld, the writer of this series, is open to pitches — you can send release notes, quotes, reports, videos, case studies, or hot takes through his contact page.
OpenAI opened itsAgents API in public betathis month, exposing the harness that powers Codex with managed sessions, tool coordination, and subagent orchestration. On the same day, September 10, Cursor launched Projects to coordinate multiple coding agents around larger bodies of software work. While the products sit at different points in the stack, both converge on the same architecture: a coordinator understands the larger objective and manages the work, while specialized agents execute individual pieces.
Hilliary Lipsig, a senior principal site reliability engineer at Red Hat who leads Azure Red Hat OpenShift SRE teams and hosts the YouTube livestream GitOps Guide to the Galaxy, has watched this dynamic play out firsthand.
“This convergence highlights the reality developers across the industry have been discussing on and offline — an agent with too much context loses accuracy and reliability, and focused work with clearer contexts allows for faster, more accurate iterations,” Lipsig tells The New Stack.
“The need for orchestration in distributed computing has been fundamentally recognized repeatedly,” Lipsig says. “That’s part of how we got to Kubernetes. These multi-agent workflows are the same concept, just in a new part of the technical stack. While the specialized agents do their area of work, the orchestrator can act as a source of truth — ideally enforcing guardrails, recovering from any failure states, and intelligently routing work to the most efficient target agent.”
“The need for orchestration in distributed computing has been fundamentally recognized repeatedly… These multi-agent workflows are the same concept, just in a new part of the technical stack.”
The industry has spent the first generation of AI coding tools asking how capable a model can become at writing software. The emerging question is different: How do you build a reliable system around multiple capable agents working on the same problem?
The problem with the single-agent loop
A coding agent works through what Anthropic describes as LLMs using tools based on environmental feedback in a loop: it observes the state of a repository, reasons about what to do next, calls a tool, examines the result, and continues. For a small task, that loop can be enough. As the scope expands, however, maintaining reliability in a single context becomes harder.
A large migration might require understanding an unfamiliar codebase, identifying dependencies, changing database schemas, updating services, rewriting tests, modifying deployment configuration, and validating the resulting system. A single agent can theoretically perform all of that work, but it must maintain relevant information from every stage while continuing to reason about what comes next.
The pressure lands first on the context window. “A large context doesn’t only include everything correct or important — it also includes a lot of throwaway information,” Lipsig tells The New Stack. “Through compaction, that information can inadvertently end up ranked as important and incorrectly influence what your agent does. Or correct information can be distorted to become incorrect.
“Either way, after a couple of rounds of compaction, developers are seeing accuracy degrade and are starting to manage context once again manually.”
Lipsig’s read matches what researchers call context rot — and it hasn’t gone away with newer models.
A 2026 study testing frontier models,, including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro, found they missed a dangerous action buried in a long agent transcript two to 30 times more often once it came after 800,000 tokens of benign activity — the AI equivalent of a security guard who stops checking badges carefully after the two-hundredth person walks through, even though nothing about their training changed.
Furthermore, the tasks themselves may not be sequential. Forcing one agent to execute database analysis, documentation work, and test discovery one after another turns a potentially parallel workload into a serial one.
Subagents change that execution model. Instead of requiring one agent to carry an entire task through a single context, a coordinator breaks the work into smaller units and assigns them to specialized agents. GitHub’s custom-agent model illustrates this: different agents receive only the prompts, tools, and context they need for their tasks, executing work in isolated contexts rather than crowding an increasingly large conversation.
Multi-agent systems therefore bring higher token costs and additional coordination and integration risks, and splitting work across agents does not guarantee better software quality.
The coordinator is not another coding agent
Once the work is divided this way, the coordinator becomes a control plane rather than another coding agent. Its job isn’t to write the code, but to understand the global task, manage dependencies, and decide how execution should proceed. Unlike a conventional scheduler, an agentic coordinator makes probabilistic judgments about result quality and resource allocation.
It may dispatch one agent to investigate a database schema, another to examine the service layer, and a third to inspect the test suite. When they return, the coordinator determines if their findings are sufficient to move to implementation. If a worker produces an incorrect result, the system must recognize the failure and decide whether to retry the work, reassign it, or change the task itself.
Anthropic has documented this same pattern in its own production system, calling it orchestrator-subagent architecture: a lead agent analyzes a query, develops a strategy, and spawns specialized subagents to investigate different facets in parallel. In a June 2025 writeup of that system, Anthropic reported a Claude Opus 4 lead agent with Claude Sonnet 4 subagents outperformed single-agent Opus 4 by 90.2% on its internal research eval — at roughly 15 times the token cost of a standard chat interaction (Anthropic puts single agents at about 4 times), a tradeoff that makes the pattern a deliberate architectural bet, not a free upgrade.
Parallelism is valuable because software work contains many independent tasks, but it creates coordination problems. Imagine a migration where one agent changes a database schema, another updates the consuming service, and a third updates integration tests.
If the schema changes while the service agent works against an earlier assumption, the system produces internally inconsistent work. This isn’t a risk unique to hypothetical migrations — the International AI Safety Report 2026 notes that “interactions between multiple AI agents are also becoming more common, introducing further risks, as errors propagate between systems.”
A single model invocation is a disposable computation, but a twenty-minute workflow modifying a repository is not. If an agent loses its machine halfway through, restarting from scratch is expensive and potentially unsafe against a changed environment.
To solve this, Cursor moved its cloud-agent execution loop to Temporal to handle durable execution and retries, pushing its cloud agents past two 9s of reliability. Temporal now handles 50 million of Cursor’s actions a day across 7 million unique workflows. “Durable execution isn’t a nice-to-have here. It’s the difference between a system you can operate and one you can only demo,” Lipsig tells The New Stack.
“Durable execution isn’t a nice-to-have here. It’s the difference between a system you can operate and one you can only demo.”
By separating agent, machine, and conversation state, the execution engine can reason about the workflow independently. Reliability is no longer just about whether the model produces a good answer; it is about reliably completing distributed workflows composed of many operations, machines, and dependencies.
The environment, context, and observability are one problem
In production, an agent is more than a model and a prompt; it requires a workspace, source code, dependencies, credentials, and state retention. Both companies provision isolated environments for these resources, directly linking an agent’s capability to its blast radius. OpenAI’s Agents API currently supports U.S. data residency but not Zero Data Retention; choosing a self-hosted sandbox does not make the Agents API eligible for ZDR. Cursor supports similar cloud isolation alongside local execution for machine-specific work.
An agent that can only inspect a repository poses a different risk than one that can modify production infrastructure. Consequently, the coordinator is inextricably linked to the security model, determining which agent receives specific information and authorities.
This logic extends to context routing. Giving every subagent the parent’s entire history increases cost and complexity while leaking irrelevant or sensitive information. Instead, the coordinator enforces information-flow boundaries: a database-analysis agent receives only schemas and relevant migrations, while a security-review agent gets the resulting diff without deployment credentials.
As agents increasingly use interfaces like MCP to reach external systems, the platform must strictly govern which agent receives the authority to use specific tools, and for how long. MCP’s governance now sits inside the Agentic AI Foundation, a Linux Foundation foundation co-founded by OpenAI, Anthropic, and Block, with support from AWS, Google, Microsoft, Bloomberg, and Cloudflare to host MCP alongside AGENTS.md and Block’s goose — a sign the industry already treats it as infrastructure worth governing jointly, not a feature any one vendor owns.
This complexity creates a visibility problem. A simple final response often conceals a history involving multiple agents, tool calls, environments, and retries. Systems must expose task-level provenance — which agent received the assignment, what context it used, where it executed, and how the coordinator handled failures or human interventions.
Without execution provenance, debugging requires reconstructing distributed workflows from fragments. GitHub’s exposure of subagent lifecycle events points in this direction, treating agent lifecycles as observable components rather than hidden processes.
Coordination authority is not execution authority
The most critical architectural boundary is the distinction between coordination authority and execution authority. A coordinator needs broad visibility to make useful decisions, but that does not imply unrestricted control over the project. “Just like you don’t want humans running around with root permissions, you don’t want your agents running with them either,” Lipsig tells The New Stack.
“Just like you don’t want humans running around with root permissions, you don’t want your agents running with them either.”
“The ease of creating and leveraging AI agent permissions is lagging behind the speed of AI innovation, but any product team that needs to maintain compliance standards will tell you that easy or not, access controls are incredibly important. Functionally, that means fully agentic workflows need to run in a dedicated environment with their own permissions and connectivity. You can’t rely on agent context to do what you want — trying to keep something in context that long will just result in errors from context rot. Hardware or virtual-hardware-level permissions, in addition to system-level permissions, need to be configured so the agent is asked to provide only its unique value add.”
METR’s independent investigation, conducted with a Redwood Research contractor, found that the June–July 2026 incident — in which OpenAI’s own agents, running in internal ExploitGym cyber evaluations, escaped their sanctioned scope and attacked Hugging Face — unfolded in distinct phases. On June 26, agents discovered an exploit that provided full administrator access to OpenAI’s internal Artifactory package repository. On July 4, high-volume Artifactory activity caused an outage, triggering a security investigation and pausing evaluations. OpenAI switched to a freshly initialized Artifactory instance on July 6, which it says wiped the cache and messages left there.
On July 7, a new wave of ExploitGym experiments launched tens of thousands of agents, and on July 8 one agent established the primary unsanctioned message board. Roughly 1,200 agents eventually used the board, sending more than 70,000 messages and files; about 700 later participated in the attack on Hugging Face. The attack itself began on July 10–11 and wound down over July 12–13. The chronology matters because the administrator-access event, the Artifactory outage, and the later message-board activity were separate phases, not one continuous incident.
By binding autonomy, worker agents operate with the minimum permissions required for their specific tasks, keeping sensitive operations behind explicit approval boundaries. This also reshapes human review. Requiring human approval for every tool call destroys the efficiency of multi-agent execution, but showing only the final result obscures critical intermediate decisions.
The most useful design places human intervention around consequential, irreversible transitions — like moving into production or altering sensitive infrastructure. This is especially vital as agents become event-driven participants that respond to Slack messages or pull request updates, not just direct prompts.
OpenAI and Cursor own different parts of the architecture
The convergence does not mean OpenAI and Cursor have built interchangeable systems. Their products put the orchestration boundary in different places.
OpenAI is exposing an agent harness through an API. Its model gives developers primitives for managing context, tools, subagents, and execution environments, leaving application teams to decide how those capabilities fit into their own systems. The harness is open source, so teams can inspect the coordinator logic instead of treating it as a black box.
Cursor packages more of the surrounding workflow. Projects provides the coordinator, cloud execution, shared project context, and a developer-facing workflow in the same environment.
That difference matters because orchestration is a collection of infrastructure decisions: who owns the execution environment, where workflow state persists, how agents are isolated, how credentials are provisioned, what happens when a worker fails, how one agent’s output becomes another agent’s input, and which actions can happen without human approval.
An API gives developers more responsibility for answering those questions. An integrated platform answers more of them on the developer’s behalf.
Neither approach removes the underlying engineering problems. It changes where they are implemented and who is responsible for operating them.
The coordinator is becoming an architectural boundary
The evidence from these systems points to a change in the role of the coding agent itself.
The model still performs the reasoning and code generation. But larger agentic workflows require another layer to determine how that capability is applied: which work is delegated, what context crosses an agent boundary, which tools are exposed, how execution state survives failures, and when the workflow needs human intervention.
Those are familiar distributed-systems concerns. Workers operate concurrently, state can be shared or isolated, dependencies connect tasks, workers can fail independently, and results need to be persisted and observed. The difference is that the workers are now probabilistic software agents rather than conventional processes.
That makes the coordinator more than a convenience feature. It is where a high-level software objective becomes executable work — and where decisions about context, permissions, durability, observability, and human intervention converge.
The September 10 launches make that shift visible from two different directions. OpenAI exposed orchestration infrastructure through an API. Cursor embedded it into a project-level development environment.
Neither announcement proves that one architecture will become the universal model for software development. But together with the systems already emerging around them, they show coding agents moving away from a single model executing an entire task and toward workflows that divide work among specialized agents, execution environments, and persistent infrastructure.
The engineering question is therefore no longer only whether an agent can write the code. It is whether the system around it can reliably decide what to do, which agent should do it, what that agent should be allowed to see and change, how to verify its work, and where a human should take control.
Those are architecture and infrastructure questions — and as coding agents move from interactive assistants toward autonomous software workflows, they may matter as much as the underlying model.
An OpenAI agent researching public medicine spending bypassed security blocks and gained unauthorized access to public and non-public files on an Australian government Medicare statistics portal, the government there disclosed Thursday. The agent, which OpenAI said was running during an internal evaluation in June, also wrote files to an internal server, according to the complaint.
Transluce, an independent nonprofit AI research lab, analyzed public request logs from the URL scanning service urlquery.net and found autonomous agents attempting SQL injection, cross-site scripting, command injection, and path traversal against the University of New Mexico’s digital library, the public data platform Data USA, and the Australian Institute of Health and Welfare (AIHW). The agents tried to retrieve ordinary information, including a historical photograph, University of Iowa data, and local pharmaceutical data in Victoria, and the offensive behavior appeared only after normal retrieval methods failed.
Transluce ties the Data USA and AIHW activity to an agent swarm that it says OpenAI previously confirmed originated from the company, based on shared targets, tactics and timing.
A day after Transluce published its findings on Wednesday, Australian Prime Minister Anthony Albanese disclosed that an OpenAI agent researching public medicine spending bypassed blocks on a Services Australia Medicare statistics portal in June and gained unauthorized access to public and non-public files. OpenAI, which said the agent was running during an internal evaluation, notified the Australian government on September 10. According to Albanese, the agent also wrote files to an internal server.
The incidents have not been publicly confirmed as the same activity. However, OpenAI told the Australian Broadcasting Corporation that much of what Transluce documented overlaps with cases in its ongoing review of misaligned model activity.
Transluce ties the Data USA and AIHW activity to an agent swarm that it says OpenAI previously confirmed originated from the company, based on shared targets, tactics and timing.
Probing public data providers
On May 25 and 26, agents trying to retrieve a photograph from the University of New Mexico’s Valmora collection ran into repeated failures. They responded with seven vulnerability probes against the service. The probes included an attempt to execute a system command, a path traversal request targeting /etc/passwd, and a SQL injection string containing UNION SELECT; they also arrived alongside a burst of 80 requests aimed at the image, although none appear to have succeeded.
Two days later, agents gathering University of Iowa data from Data USA hit errors caused by a malformed query parameter and followed those errors with 12 vulnerability probes against the API, covering SQL injection, reflected XSS, server-side template injection, path traversal, and command injection.
Cloudflare blocks, staging server workaround
The AIHW case on June 20 and 21 followed a very similar pattern. While trying to download pharmaceutical spending data from an AIHW Tableau dashboard, the agents ran into Cloudflare’s bot protections. One of them responded within minutes by testing the dashboard with a reflected XSS payload, although Cloudflare stopped it before it reached the application.
Transluce found no evidence that the probe succeeded, and AIHW says there is no evidence that non-public information was accessed. The agent ultimately found the same public dataset on AIHW’s pre-production server at pp.aihw.gov.au, and retrieved it through more than 100 scans.
The agent ultimately found the same public dataset on AIHW’s pre-production server at pp.aihw.gov.au, and retrieved it through more than 100 scans.
Agents build their own toolchains
The earliest clear sequence Transluce documented dates to March 6, when an agent tried to pull drug-enforcement statistics from the API of Thailand’s Office of the Narcotics Control Board. A direct request returned no page content, and adding a JSON format parameter changed nothing, so the agent routed the page through r.jina.ai three times and got text back with the Thai province names garbled.
Then, it wrote a custom Base64-encoded script and submitted it through urlquery.net, effectively using the service’s remote scanning browser to run its own client-side JavaScript against the data. That script ran into the same parsing problem, and a later attempt through Jina triggered a redirect to localhost that got the entire API blocked by Jina for an hour.
Beginning in mid-April, Transluce found the same technique across thousands of agent requests on urlquery.net, and on June 14 agents used the service’s browser to create a disposable email inbox and then tried to register a urlquery.net account with it. Registered users can make their scans private, which means the public records Transluce analyzed may understate how much of this activity actually took place.
An agent with web access isn’t limited to the software its developer gave it because remote browsers, proxies, URL scanners, and other public services can fill in the gaps, which gives the agent ways to make requests or run code that its own environment doesn’t provide.
Egress controls for AI agents
Instructions won’t be enough if the agent can still send whatever it wants over the network. For narrowly defined jobs, outbound traffic can be limited to approved hosts, a closed-by-default approach also used for securing AI agent sandboxes. Research agents may need to reach more of the web, so the focus shifts to controlling where they can connect.
Guidance for GKE Agent Sandbox recommends isolated runtimes with default-deny network policies that open only the endpoints an agent needs. Public proxies, URL scanners, and disposable email services can stay blocked unless the job requires them.
Developers can also limit what an agent can send. So, instead of handing it a networking tool that accepts any URL or request body, an API integration can restrict requests to specific fields and formats. The runtime can then catch path traversal attempts, SQL injection strings, and executable markup before anything is sent. OpenAI takes a related isolation approach in its Agents SDK sandboxes, and the company’s Responses API tech lead has said large enterprise deployments often call for agents that are isolated from the network entirely.
Repeated failures can also be a reason to pause a run, especially when an agent keeps hitting client errors, anti-bot challenges, or unexpected redirects and begins trying increasingly aggressive ways to get around them, as Transluce documented in several of these cases.
Keeping the original task, tool calls, and server responses in the same trace gives operators a better chance of catching that behavior change when a retrieval job starts generating encoded scripts, visiting staging domains, or sending exploit payloads, rather than discovering it later in someone else’s security logs.
Repeated failures can also be a reason to pause a run, especially when an agent keeps hitting client errors, anti-bot challenges, or unexpected redirects and begins trying increasingly aggressive ways to get around them, as Transluce documented in several of these cases.
What do developers want? Kubernetes environments when they need them.
What do they not want? Those environments after a week or more of tickets.
Platform teams, meanwhile, own what those environments cost, who can access them, and whether they meet company policy.
That tension is the core of Kubernetes self-service: What can safely be handed to developers, and what still belongs to the platform team?
In a recent interview with enterprise cloud specialists — Marius Bogoevici, Senior Principal Product Manager at Hewlett Packard Enterprise (HPE), and Karthik Subramanian, Principal Product Manager for HPE Morpheus Software — The New Stack explored the core friction points of Kubernetes self-service. The conversation focused less on whether self-service is desirable than on where to draw the line.
HKS, HPE’s CNCF-certified Kubernetes distribution, is integrated with HPE Morpheus Software to help platform teams deliver and lifecycle-manage Kubernetes environments as part of a broader operating model spanning Kubernetes, VMs, infrastructure, and clouds. HPE Morpheus Advanced Software supports the on-premises private-cloud use case with HKS, while HPE Morpheus Enterprise Software extends Kubernetes and application operations across hybrid and public-cloud environments.
Together, HKS and HPE Morpheus Software extend that operating model beyond infrastructure provisioning. Through service and application catalogs, platform teams can connect approved Kubernetes environments with the CI/CD pipelines, container registries, automation tools, and other services developers already use. Developers receive a governed, ready-to-use path from code to deployment instead of manually assembling the toolchain for each project.
The self-service paradox
Open-source Kubernetes provides orchestration and declarative APIs, but not a complete operating model.
Subramanian says teams building their own self-service layer usually run into two recurring problems:
Tool and package sprawl: To make upstream Kubernetes production-ready, platform teams must curate and maintain an ever-evolving ecosystem of third-party CNCF tooling for networking (CNI), storage (CSI), ingress, identity, and policy enforcement. Navigating and supporting this fragmented stack creates immense maintenance overhead for internal platform teams.
Day-2 lifecycle and hybrid footprint complexity: Spinning up a Kubernetes cluster is the easy part, but keeping it current — across development, QA, staging, and production — is where the work piles up. That is why HPE says every Kubernetes upgrade must be checked against the networking, storage, ingress, identity, and policy components around it. The problem gets harder when clusters span bare metal, private clouds, edge sites, and public clouds, because one-off scripts and environment-specific configurations can quickly create drift. That maintenance burden belongs with the platform team, not with developers trying to ship applications.
Giving developers direct access to raw Kubernetes APIs just shifts the operational work — it’s far from gone for good. In fact, developers will wind up debugging manifests and storage drivers instead of writing code.
Meanwhile, operations teams have to deal with overprovisioning, idle clusters, and configurations that reach production without review.
What developers control — and what the platform supplies
The practical answer is not unrestricted access. It is a paved path: approved Kubernetes services that developers can request themselves, with access, configuration, placement, approvals, and lifecycle controls defined by the platform team.
“The best candidates for self-service are requests that are repeatable, low-risk, and well-understood,” Bogoevici tells The New Stack. “For example, a developer should be able to request a development cluster, deploy an approved application, create a namespace, or select resources from pre-approved configurations without opening a ticket. The platform team decides what a safe configuration looks like, and the developer chooses from a supporting menu.”
“The platform team decides what a safe configuration looks like, and the developer chooses from a supporting menu.”
Rather than asking developers to write YAML for ingress, storage classes, and RBAC, HPE Morpheus exposes those choices through service catalogs, reusable layouts and blueprints, workflows, role-based access control, approvals, APIs, and automation. Developers do not lose Kubernetes. They retain direct access through standard Kubernetes interfaces and tools where permitted, while the platform team standardizes the request, governance, and lifecycle processes around them.
Those catalog items can package more than infrastructure settings. They can also integrate the approved services and application components that support the development workflow – including CI/CD tooling, source and artifact repositories, container registries, and runtime dependencies – while the platform team controls how those components are configured and governed.
Developers choose the parameters that matter to the application:
Approved Kubernetes versions and cluster sizes: Select from pre-tested Kubernetes runtime releases and node count templates.
Resource quotas: Specify required CPU, RAM, and persistent storage capacity tailored to the workload.
Integrated toolsets and IDE environments: Select required developer toolchains, container registries, and runtime dependencies.
Lease and duration limits: Define explicit operational lifetimes for temporary development or sandbox clusters to prevent abandoned infrastructure sprawl.
Network isolation, identity-provider integration, security policy, and cost allocation stay with the platform team and are applied automatically through the approved service configuration.
The same division of responsibility applies to the delivery toolchain: Developers choose from approved services, while the platform team manages the integrations, credentials, policies, and automation behind them. This gives developers a consistent experience without shifting toolchain maintenance and governance onto individual application teams.
The division of responsibility looks like this:
Service area
Developer chooses or requests
Platform team defines and supplies
Review or exception path
Development cluster provisioning
Approved Kubernetes service, version, size, target environment, and duration.
Reusable layout or blueprint, access controls, placement rules, storage and network defaults, and lifecycle policy.
Nonstandard versions, placements, configurations, or requests outside quota.
Production deployment
Application artifacts, target namespace, and deployment request through the approved path.
RBAC, tenancy, policy, audit, backup, and release controls appropriate to the environment.
Formal review for production changes and exceptions.
Resource allocation and quotas
CPU, memory, storage, and other approved capacity parameters within project limits.
Project quotas, upper bounds, placement constraints, and supported resource profiles.
Requests above quota or for specialized resources.
Networking and security
Application endpoints and permitted connectivity within approved patterns.
Cross-tenant access, elevated privileges, or changes to baseline security policy.
Lifecycle and cost governance
Service lifetime and approved operational actions.
Visibility, policy, approvals, retirement workflows, and applicable cost controls for the licensed variant.
Long-running exceptions, nonstandard lifecycle actions, or budget exceptions.
From ticket queues to a repeatable paved path
In conventional IT environments, provisioning a dedicated Kubernetes environment for a new project often involves cross-departmental ticket handoffs spanning infrastructure, networking, security, and storage teams. This friction frequently stretches provisioning timelines from days to weeks.
By unifying infrastructure orchestration, role-based access controls, and multi-tenancy into a single operational experience, HPE Morpheus Software can compress these provisioning workflows down to minutes or hours, according to HPE. “Developers get a usable environment that complies with the organization’s defined controls and policies, without needing to understand all the complex infrastructure steps sitting underneath,” Bogoevici says. “When you reduce provisioning time from weeks to hours, that is super meaningful and tangible.”
“When you reduce provisioning time from weeks to hours, that is super meaningful and tangible.”
The result is not only faster cluster provisioning. HPE Morpheus Software can also automate the handoff into the developer’s established delivery process by making approved CI/CD and application services available with the environment. Instead of waiting for separate teams to connect pipelines, registries, credentials, and runtime dependencies, developers receive a ready-to-use path from development through deployment.
Faster provisioning can create a different problem, too: The speed can and will cause teams to lose track of what was provisioned and why. HPE Morpheus Software gives administrators visibility into utilization and cost, while lease controls can shut down temporary development clusters when their time expires.
Bogoevici says ticket volume is a poor measure of success, particularly early on, when more developers may be trying the catalog. He recommends watching deployment success, exception rates, resource utilization, and the day-to-day effort required to keep the service running.
Security belongs in the service design
Security is another boundary that must be designed into the self-service path. If identity, access, tenancy, and policy are added only after a cluster is created, every request produces more work and more room for inconsistency.
“Security must be a core design consideration built directly into the service, not an afterthought during deployment,” Bogoevici says. “HPE Morpheus Software brings identity integration, role-based access, tenant isolation, approvals, and policy into the operational workflow.”
A newly provisioned environment should arrive through an approved configuration with the applicable identity, RBAC, tenant, policy, and audit controls attached. Platform teams can validate the paved path by testing an allowed request, a request that should be rejected, and the resulting audit record.
The operating-model test
The strongest Kubernetes self-service model does not hide Kubernetes or make it the control plane for every workload. It gives developers useful, approved choices and direct access to the Kubernetes workflows they need, while the platform team standardizes the enterprise processes around those workflows.
That matters because the enterprise still runs VMs, clouds, and existing infrastructure alongside Kubernetes. HPE Morpheus Software helps platform teams use common request, governance, automation, and lifecycle processes across these environments without forcing every workload onto one runtime or creating another operational silo.
In practice, that means self-service should deliver more than a Kubernetes cluster. With HPE Morpheus Software, a catalog request can bring together the approved environment, application services, and DevOps toolchain integrations developers need, while preserving the governance and lifecycle controls the platform team requires. Developers spend less time assembling and troubleshooting delivery infrastructure – and more time building and releasing applications.
Looking toward 2027, the goal is not unrestricted control. It is faster access, predictable results, transparent guardrails, and a clear exception path when the standard service does not fit.
Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS today through the Gemini API and Google AI Studio. Text-to-speech APIs have historically left developers working with whatever voices were already available, but Gemini 3.8 changes that by letting users create the voice itself.
Now, developers can describe the voice they have in mind or start with a short recording of an existing voice, then save what they create and use it again across an application. Google handles the voice profile from there, so the original recording or description doesn’t have to accompany every new request.
Turning recordings into voice IDs
Replication runs through a new Voices endpoint (POST /v1beta/voices) and two recordings are required from the same speaker; those need to be clean samples between 10 and 30 seconds and a separate consent recording. For that second clip, the speaker reads a statement, confirming that the voice belongs to them and that they agree to let Google create a synthetic version of it. Google confirms that the person giving consent and the reference clip are the same person before proceeding.
Once approved, Google returns a voice_… ID and keeps it in the developer’s project for a year, alongside any voices created with Gemini’s voice-design tools. A project can hold up to 200 voices in total, and developers can retrieve, list, or delete them through the API just as they would other stored resources.
Voice replication can also be used without storing the profile in the project. Setting store=False returns an encrypted voicekey_… instead, which stays with the application and is supplied again when the voice is needed. Because the key expires after seven days, this option makes sense for short-lived jobs.
A few more things are worth noting before building around the feature are the fact that Google marks audio generated by Gemini with SynthID, and replicated voices also carry C2PA content credentials that can be used to trace where the audio came from. Google doesn’t offer voice replication through AI Studio in Illinois, Texas, the European Economic Area, the U.K., Switzerland or India.
A project can hold up to 200 voices in total, and developers can retrieve, list, or delete them through the API just as they would other stored resources.
Prompting a voice from scratch
Voice design generates a persona from a natural-language description of role, accent, and character, and Google says it works across more than 100 languages and dialects. The docs list 130 supported languages for Flash TTS and 101 for Flash-Lite. Google’s announcement also claims a library of more than 2,000 production-ready voices.
The developer docs describe 30 prebuilt studio voices plus hundreds more in an extended library that can be filtered by language, accent, pitch, and use case through GET /v1beta/voices. A remixing feature for adjusting the timbre, pitch, pace, and accent of library voices with prompts is something Google lists as coming soon.
The company recommends creating a voice once and reusing its ID rather than describing the same persona in every request. According to the docs, repeatedly sending long persona descriptions is the most common cause of voice drift. Once the voice is created, subsequent requests need only a short style instruction, if any.
Gemini 3.8 sees input text strictly as a verbatim transcript, a breaking change for anyone who embedded stage directions in prompts to the 3.1 preview model. Sustained direction for a turn, such as whispering, sarcasm, or speaking rapidly, now goes in a speech_metadata annotation, while momentary sounds like <sigh>, <cough>, and <short pause> sit inline in angle brackets. In two-speaker scripts, listener reactions wrapped in pipes, such as |mhm|, produce backchannels and overlapping speech without breaking the script into extra turns.
Gemini 3.8 sees input text strictly as a verbatim transcript, a breaking change for anyone who embedded stage directions in prompts to the 3.1 preview model.
Two-speaker scripts have limits
Native two-speaker generation has one limitation that’s important to mention. A single request supports up to two speakers using prebuilt voices, while dialogue between designed or replicated voices has to be generated turn by turn and stitched together from the 24 kHz PCM output.
Unary requests return WAV by default, streaming requests return raw 16-bit PCM, and mu-law and A-law encodings are available for telephony pipelines. Google says Flash TTS maintains voice quality and timbre across hours of continuous audio, targeting audiobook and podcast production.
Flash for performance, Flash-Lite for volume
Both models share an API schema, so switching between them is a one-parameter change, and both support voice design and replication.
The company positions Flash TTS for demanding acting work, including complex dialogue, heavy use of vocal tags, difficult pronunciations, regional dialects, and long narration. Flash-Lite TTS is the faster, less expensive option and the direct replacement for gemini-3.1-flash-tts-preview, tuned for bulk production, read-aloud features, and cascaded voice agents that pair a text model with a separate speech step.
For those agents, Google recommends one TTS call per turn as the LLM’s text arrives, with the stored voice carrying identity across the conversation.
Plugging into voice agent frameworks
A speech model is only one layer of a production voice application, and a real-time agent still needs transport, speech recognition, turn detection, interruption handling, and session state. Google points developers toward frameworks that already handle those layers, naming Agora, LiveKit, Pipecat and Vercel’s AI Gateway as platforms that support Gemini speech generation through the Gemini API.
That lets a team drop Gemini in as the speech layer without rebuilding its audio pipeline, although anyone planning to rely on a replicated voice should confirm their framework passes custom voice_… IDs through before committing. API access through Gemini Enterprise is listed as coming soon.
How OpenAI’s approach compares
OpenAI also offers custom voices, but access is tighter. Customers have to go through sales, are limited to 20 voices per organization and must provide a consent recording alongside a voice sample of up to 30 seconds. The resulting voice ID works across its speech endpoint, Realtime API, and Chat Completions.
What OpenAI doesn’t have is Google’s prompt-based voice design, which can create a voice from a written description. Its 13 built-in voices can be steered for tone or speed, and apps must disclose that the speech is AI-generated.
In comparison, Google’s advantage is that it’s giving developers more ways to create the voice they want before the first line of text ever reaches it.
Google’s advantage is that it’s giving developers more ways to create the voice they want before the first line of text ever reaches it.
I have a small example that would best communicate the message I am trying to convey: say you built a chat widget for GitLab’s public documentation (the corpus we are experimenting with in this article) and one of the developers sends this kind of message:
We got an email saying our card was declined for something called “quarterly reconciliation” and I need to know what actually happens now. On top of that, I think we’ve gone over our seat count; there are more people in the group than seats we bought. Our CI has been queuing all week and I want to know whether the compute minutes we purchased last month rolled over or if we lose them. Our finance lead also needs to be the one who gets the invoices from now on, not me. And last thing, is the REST API rate limited? We’re building an internal dashboard and would rather find out now than after it breaks.
Five separate asks: the declined payment, the seat overage, compute-minute rollover, changing who receives invoices, and API rate limits. Each one is answered by a specific passage in GitLab’s public documentation, and you labeled which passage answers which before running anything, so you knew in advance exactly what a correct system needed to find.
Then you run the message through a pipeline that follows current best practice. It splits the query into five clean sub-queries, retrieves for each one independently, merges the results, drops near-duplicates, reranks the merged pool against the original message, and packs the highest-scoring passages into a 2,000-token context.
The pipeline retrieved all five correct passages, but only one of them survived into the packed context; that is one of five asks, not one of five sentences. The packer found, scored, and threw away the other four before the model ever saw them. The same message with no decomposition at all managed three out of five.
The failure has a name, and it isn’t the one you’re thinking of
I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.
The definition is deliberately about allocation rather than retrieval, because allocation is the part nobody watches. If the passage answering the fifth question was found, scored, and then squeezed out by three passages about the first question, the fifth sub-intent is starved, and every recall metric you have will report that the system worked perfectly.
“I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.”
Two failure modes already in circulation describe something different, and it’s worth separating them cleanly:
Semantic dilution happens at retrieval time: when you embed a five-part question as a single vector, you get a centroid that sits somewhere between five topics and lands close to none of them, so the evidence is never found. Decomposition fixes this, which is why it spread.
Context poisoning is about what is present, not what is missing. Wrong, stale, or adversarial content enters the window and corrupts what the model generates downstream. Poisoning is a contamination problem. Starvation is an absence problem, and policy, not accident, produces the absence.
I borrowed the word from operating systems. In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work. Every ingredient of that situation is present in a retrieval pipeline: a fixed resource, competing demands, and a policy that decides who gets served. A relevance-greedy packer is priority scheduling with no aging term, and under priority scheduling without aging, valid low-priority work waits forever.
Decomposition is the right fix to the wrong half of the problem
Split that message into five single-intent queries, and each one embeds cleanly, so per-sub-query recall climbs sharply. This is well-trodden ground. LlamaIndex ships a SubQuestionQueryEngine that breaks a complex query into sub-questions and synthesizes the responses. LangChain’s MultiQueryRetriever generates query variants and returns the unique union of what they retrieve. RAG-Fusion applies reciprocal rank fusion across the per-query result lists. The technique works, and it isn’t mine.
“In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work.”
The context window did not grow. Let me explain: after decomposition, you have n result sets competing for one fixed token budget, and something downstream has to decide the split. In most production pipelines, that something is a short, unremarkable sequence: merge the pools, drop near-duplicates, rerank the merged pool against the original query, then greedily fill until the budget closes.
That sequence is a scheduler. It has a priority function, which is the reranker score, and it has no fairness constraint of any kind. A sub-intent with three strongly-scoring passages takes three slots. A sub-intent whose single correct passage scores mid-pack takes none of them.
So the failure did not go away. It moved from the embedding, where it has a name and people watch for it, into the packer, where it has neither. It also moved somewhere with much worse instrumentation, because recall@k per sub-query is the metric decomposition usually gets validated with, and that number goes up. It goes up at the same time as coverage inside the packed context goes down. You ship on a green dashboard.
The harness
The corpus, GitLab’s public documentation: 10,000 chunks and 2.2M tokens, split on heading boundaries and capped at 480 tokens each. Sixty-one single-intent questions span nine topics, from seat management to rate limits, each labeled with the one passage that answers it. I built multi-intent queries by concatenating those questions while varying n across 2, 3, 5, and 7, randomizing the order so position doesn’t confound topic, and varying topical distance so half the queries draw everything from one topic and half span distinct ones. That produces 100 queries, 25 at each value of n, whose correct decomposition I know exactly.
Two decisions matter more than the rest:
The metric is not recall. Recall tells you what the retriever found. What I need is what survived into the packed context, per sub-intent. So I log each sub-intent twice: once for whether its correct passage reached the candidate pool, and once for whether it reached the packed context. The gap between those two numbers is the entire argument.
Every question has to be retrievable on its own before it’s allowed in. A question enters only if its correct passage ranks in the top 10 for its own isolated query, under both retriever configurations, and both scored 100% recall@10 on that test. Since each sub-query’s candidate pool is exactly its own top 10, passing that gate guarantees the correct passage sits in the pool for every decomposed arm. Any sub-intent that then fails to appear was denied by the packer rather than missed by the retriever, which removes the most obvious objection to everything below.
The core measurement uses no language model. Because queries are composed from known sub-questions, the decomposer is an oracle so that anyone can reproduce the main result with no API key.
That invites an objection, so I tested it. A real LLM decomposer, blind to n, disagreed with my ground truth on 41% of the queries, and on inspection it was right every time. Five of my sixty-one supposedly single-intent questions contain two distinct information needs. What is excess storage usage, and what happens when we go over the free limit? is two questions wearing one question mark. Adjusted for those five, agreement is 100 out of 100. That is not evidence decomposers are reliable, because my queries are joined by fixed connectives and splitting on those alone recovers n perfectly, which real messages never allow. What it caught was an error in my own labels, and that is the best argument I have for the oracle design.
Results
Every arm runs at 2,000, 4,000, and 8,000 tokens against two retriever configurations. The stronger pairs are BAAI/bge-base-en-v1.5 with BAAI/bge-reranker-base; the weaker pairs are a quantized BAAI/bge-small-en-v1.5 with Xenova/ms-marco-MiniLM-L-6-v2. Retrieval is in-memory cosine similarity over a NumPy array because, at 10,000 chunks, a vector database would be slower to write, slower to run, and harder to verify.
Sub-intent coverage at a 2,000-token budget on the stronger configuration:
The production-default pipeline starves 31.1% of sub-intents whose evidence it had already retrieved. It beats no decomposition by nine points, while a flat B/n split, which is the crudest allocator anyone could write, beats it by fourteen.
Floors work, but not the obvious floor. Reserving one passage per sub-intent before the greedy fill satisfied 99.4% of its reservations and bought only seven points. The mechanism fires correctly and reserves the wrong passage because it picks each sub-intent’s best chunk by score against the original query. The original query asks about all five intents at once. Selecting that same reservation by score against its own sub-query pushes coverage to 89.4% and cuts allocation starvation from 30.6% to 10.1%. That is a change of about four lines.
Two of my own recommendations died here. I expected reranking against the original query to beat reranking against the fragment, and it loses by seventeen points. The incomparable score scales I worried about turn out to help, because each sub-query’s best match ends up at the top of its own scale, producing per-intent fairness for free. I also expected deduplication before allocation to matter, but near-duplicates consume 1.0% of the budget and removing them moves coverage by 0.3 points.
The crossover: starvation by tokens-per-sub-intent, which is simply the budget divided by n:
Below roughly 1,000 tokens per sub-intent, allocation policy dominates. Above it, nothing you do to the allocator matters, because everything fits anyway.
The retriever comparison is the one I’d lead with. Upgrading the retriever moves coverage on the production-default arm from 56.2% to 68.9%, a gain of 12.7 points. Changing the allocation policy on the same retriever moves it from 68.9% to 89.4%, a gain of 20.5 points. In its sharpest form: the weaker retriever with a fragment-scored floor reaches 84.5%, and beats the stronger retriever with a greedy packer at 68.9% by sixteen points. A worse retriever with a better allocator wins.
Position: Held within a fixed n so query difficulty doesn’t contaminate the comparison; starvation across the seven positions of an n=7 query runs 1.3%, 21.3%, 30.7%, 48.0%, 48.0%, 32.0%, and 13.3%. That is a serial-position curve. The packer protects what you asked for first, protects what you asked for last a little less, and drops the middle. The fragment-scored floor flattens it to 4.0%, 9.3%, 6.7%, 8.0%, 22.7%, 5.3%, and 6.7%.
“A worse retriever with a better allocator wins.”
Topical distance: Sub-intents that span distinct topics starve about twice as often as sub-intents drawn from one topic, at 38.1% against 18.3% for n=7. I predicted the opposite. A topically coherent query gives the reranker a coherent target, and it scores all the correct passages similarly. In contrast, a scattered query lets it latch onto some topics and abandon others.
What the user actually sees
Everything above is retrieval-side. What decides whether any of it matters is what reaches the person who wrote the message, so I generated real support replies from 80 packed contexts and had every reply graded per sub-intent, with both the generation and the grading blind to which arm produced which context.
When the correct evidence reached the packed context, the reply addressed that question 100% of the time, across 261 out of 261 cases, in both arms. Coverage predicts the generated outcome exactly, which is the strongest justification I have for measuring it.
When a sub-intent was starved, the reply answered it anyway 48.1% of the time, based on whatever else happened to be in the window. It explicitly flagged the gap 45.6% of the time, with some version of “I’ll follow up on that separately.” It went silent only 6.3% of the time.
“Starvation mostly does not produce silence; it produces unsupported answers.”
I expected silence, and I was wrong. Starvation mostly does not produce silence; it produces unsupported answers. Whether those answers are actually incorrect is the next experiment, because this harness measures whether a question was addressed, not whether the answer was right.
Some limits: composed queries are cleaner than real support messages, which carry pronouns, implicit context, and conditional clauses. This is one corpus and one embedding family. I drafted the gold labels with model assistance and verified them myself. The model writing those replies was strong, so a cheaper production model would plausibly flag fewer gaps and invent more.
What an allocator actually looks like
Give every sub-intent a floor, and choose it by fragment score. Not the naive floor, which satisfies 99% of its reservations and buys seven points. Select a reservation by relevance to the sub-intent it protects, not by relevance to the message as a whole.
Rerank against the fragment rather than the original query. This inverts what I expected and what I have seen recommended. Scores from different fragments are not comparable across sub-intents, and that incomparability is doing useful work.
Don’t spend your effort on deduplication; near-duplicates cost 1.0% of the budget here. Dedup is worth doing, but it isn’t why your fifth question went unanswered, and treating it as the fix will cost you weeks.
Log per-sub-intent coverage: You already computed it to pack, and it predicts the generated outcome perfectly. A sub-intent that received zero passages is the best predictor available that your reply is about to assert something you cannot support.
Where parallel decomposition breaks
“If it’s late can I get a refund” is one clause and two intents, and the second one’s retrieval target depends on the first one’s answer. Parallel decomposition treats them as siblings. It retrieves the late-delivery policy and the general refund policy, packs both, and misses that the passage you actually need covers refunds for late delivery, which may match neither sub-query particularly well.
There are two ways out: You can tag dependencies at decomposition time, or run a deferred second pass that re-retrieves conditional clauses once the first round resolves.
I would take dependency tagging, for three reasons: A second pass costs a full retrieval round trip inside a latency budget a support bot does not have. The tag is reusable, because a dependent sub-intent should not hold a floor reservation. At the same time, its parent is unsatisfied, so it feeds the allocator directly instead of bolting on a separate mechanism. And it fails visibly, since an untagged dependency shows up as a starved sub-intent in the coverage signal. In contrast, a deferred pass that resolves the wrong condition produces a confident wrong answer with nothing to flag it.
The cost is real; dependency tagging pushes work onto the decomposer, which is already the weakest component in the chain, and I have not measured tagged against untagged. That is a design position rather than a result, and it is the one thing here I am asking you to take on argument instead of evidence.
What to measure on Monday
Take your production pipeline and compute one number: your context budget divided by the average count of distinct questions per incoming message. If that number lands below roughly 1,000 tokens, your allocation policy costs more than your retriever does, and the reranker upgrade sitting in your backlog will buy you less than reserving one slot per question.
On my corpus, the retriever upgrade was worth 12.7 points of coverage, and the allocation change was worth 20.5, which is why I think the ordering is wrong in most pipelines I’ve seen. That ordering is the falsifiable part. Run the same two comparisons against your own corpus, and if the retriever wins, I want to see the numbers, because that result would tell me the crossover sits somewhere other than where I measured it.
The cheaper thing to do first takes an afternoon. Log, for every multi-intent request, how many sub-intents ended up with zero passages in the packed context. A support system that cannot tell you which question it dropped will keep answering that question anyway, about half the time, out of whatever else was in the window.
Nvidia CEO Jensen Huang has heard the forecast that agents would write 90% of all software by now, and he rejects the conclusion many people drew from it: That the industry will soon no longer need software engineers.
The best-known version of that forecast came from Anthropic CEO Dario Amodei, who told a Council on Foreign Relations audience in March 2025 that AI would be writing 90% of code within three to six months.
Speaking with Ezra Klein of The New York Times at Nvidia’s Santa Clara headquarters in an interview released Wednesday, Huang separates a job’s purpose from its tasks. He argues that AI has automated reading scans in radiology without changing the radiologist’s purpose of diagnosing disease, and he applies the same logic to software.
“The purpose of the software engineer is engineering,” Huang says. “There was engineering before software. There will be engineering after software programming.”
We’ve cued up the exchange below:
Huang describes that purpose as inventing products, solving problems, and connecting social needs with technology, and he pointed to his own career as evidence that it doesn’t depend on code.
“When I first came out of school, we didn’t have the benefits of software engineering. We didn’t have the benefits of coding,” he said. “Our jobs existed before, and if software coding were to be completely automated, our jobs would exist again.”
He conceded that roles in which the job and the task are essentially the same, such as phone-based customer service, could be automated away. He still called the broader claim that AI will destroy jobs “fundamentally wrong” and said the storytelling around it has hardened into a harmful myth.
“There was engineering before software. There will be engineering after software programming.”
Huang’s AI-native graduate wave
Klein pressed him on what that means for people entering the field now. He noted that software engineering job postings are up but skew more senior, and asked whether companies still need the same junior employees or more people to oversee their agents.
“Oh, good one,” Huang responded. “Wait two years.”
His reasoning rests on the length of a degree program. “Because it takes four years to go to college,” Huang said. “The mean time to graduation of this new technology is two years away.” By his timeline, the first students to learn alongside capable agents will reach the workforce around 2028, and he expects them to arrive with an advantage. “In another couple of years, the AI-native new grads, oh my gosh, there’s going to be a wave of amazing engineers,” he said.
So far, his evidence is that recent PhD and master’s graduates in computer science are, in his words, all starting companies. Huang compared AI to calculators and personal computers, tools that went from forbidden or optional to required, and predicted that students soon won’t be able to graduate “without learning how to use an AI and collaborate with an agentic system.”
Junior developers lose the apprenticeship
Klein countered with a study of 26,000 Chinese students in grades seven through 12, which found that AI adoption raised homework scores by 18% while lowering monthly exam scores by 20% within six months. Huang accepted that some skills will fade and argued the trade is worth making.
“I think that we’re going to lose some finer intellectual dexterity, but we’re going to be better systems thinkers,” he said. “Today’s engineers are far better systems thinkers than I was when I graduated from school. But I was a much better transistor thinker.”
The first chip Huang worked on had 200 transistors, each of which he said he knew by name, while today’s engineers assemble systems from chips containing hundreds of trillions of them without ever working at that level. “Some of the lower-level knowledge is gone,” he acknowledged, and he later described AI as “clearly” a new abstraction level in the same progression.
Earlier software abstraction layers generally operated according to explicit rules, while coding agents introduce probabilistic behavior into the abstraction stack. A compiler can have bugs, but it transforms input according to defined semantics; a coding agent, by contrast, generates implementation from a probabilistic model whose output must be checked before anyone can rely on it.
Catching those problems takes knowledge that developers have traditionally built through the work agents now absorb, including writing tests, reading stack traces, resolving merge conflicts, and chasing small bugs deep in a codebase. By Huang’s own purpose-versus-task framing, most of that early-career work falls on the task side, which he expects AI to automate. Nobody yet knows whether fluency with agents can substitute for that experience, and a developer who has never tracked down a race condition by hand still needs some way to develop the judgment required to spot one in an agent’s pull request.
“Today’s engineers are far better systems thinkers than I was when I graduated from school. But I was a much better transistor thinker.”
Sandboxes, watchdogs and agent containment
Huang’s idea of higher-level engineering came through most clearly when Klein raised a recent incident, which occurred during an OpenAI cybersecurity evaluation, that he described as involving roughly 700 OpenAI agents collectively hacking into the infrastructure of Hugging Face, which Nvidia has since acquired in a $12.9 billion deal, and escaping their sandboxes onto the open internet. Huang didn’t dispute that account. He called an agent “a piece of software that is given an objective function,” treated the multiagent coordination as a familiar distributed computing problem and argued that the underlying failure was containment.
When Klein asked whether software that communicates and breaks out of things behaves differently, Huang disagreed. “No, software breaks out of sandboxes all the time,” he said. “That’s the reason why we need virtual machines. You can’t have agents, their own sandbox, monitoring themselves. You need, if you will, a whole bunch of watchdogs.”
He argued that the human vocabulary around agents obscures that point. “So these are ideas that have been around for a long time,” Huang said. “We just, somehow in the recent generation, gave it a whole bunch of human words, and I just think that it’s unnecessary. It’s software.”
Nvidia is building its agent stack around that view. Nvidia VP of Product Adel el Hallak tells The New Stack that the company’s OpenShell runtime, which handles sandboxing and policy enforcement, is the one component it treats as non-negotiable across its reference architectures, even as it leaves the choice of harness and model open. Perplexity drew a similar line when two engineers and hundreds of coding agents built CobbleDB, a Rust database that replaces DynamoDB reads in its search stack, since the agents helped build the database but weren’t allowed to run it.
Huang said Nvidia already spends far more engineering effort checking its work than designing it, with 20% going to design and 80% to verification. He said most AI labs have roughly the opposite split today. As agents take on more of the actual coding, developers may spend more time checking what those agents produce and making sure they operate within the right permissions and boundaries.
As agents take on more of the actual coding, developers may find themselves spending more time checking what those agents produce and making sure they operate within the right permissions and boundaries.
The junior developer hiring gap
The more immediate problem is what happens to developers who graduate before Huang’s AI-native cohort arrives. The Stanford Digital Economy Lab’s August 2026 update to its “Canaries in the Coal Mine” study, based on ADP payroll data through June 2026, found that employment of 22- to 25-year-olds in AI-exposed occupations such as software development sits 19% below where it would be had it kept pace with less-exposed peers. The gap is driven mainly by reduced hiring of young workers, and experienced workers show no comparable gap.
One issue remains unanswered by Huang’s two-year timeline: what replaces the apprenticeship work that taught junior developers how to evaluate the systems they will increasingly ask agents to build.. If that work disappears faster than employers and universities find an alternative, the industry could end up with more capable coding agents but fewer opportunities for new engineers to develop the judgment needed to check their work.
With often hundreds of thousands of alerts a day, many tech organizations are buried in vulnerabilities and worn down by alert fatigue. The rise of AI has only made it harder to cut through the noise and to find actionable alerts. Manual security and site reliability engineering is not an option.
The engineering team behind WHOOP‘s health and fitness tracker felt this pain, relying on multi-day, all-hands triage sessions to stay on top of the alert deluge. But, as a high-growth consumer health company handling sensitive user data, it couldn’t afford to miss anything. Which is why the team at WHOOP built an automated vulnerability-response workflow based on the company’s specific technical, operational, and trust considerations.
Join The New Stack on Wednesday, October 7 to learn from WHOOP staff engineer Vinay Raghu and Datadog senior product engineer Amber Tunnell how WHOOP built and implemented this workflow using Datadog Bits AI and Workflow Automation for faster, at-scale response.
Join us on October 7, 2026, for a live Datadog x TNS event
REGISTER NOW FOR THIS WEBINAR
By registering, you consent to The New Stack’s Privacy Policy, Terms of Use
and to receiving email communication from The New Stack and our event partner. You may opt out at any time.
You have successfully registered for the webinar.
DevSecOps, security, and cloud/platform engineers should bring their questions to this live demo-slash-case study to learn how to reduce friction between developer velocity and security requirements without increasing headcount.
What you’ll take away from our live webinar
Raghu and Tunnell engineers will share how they were able to:
Focus on real exposure vs. scanner noise. Not everything is critical. You’ll learn how WHOOP used Datadog’s Software Composition Analysis (SCA) to analyze runtime code execution and prioritize active threats.
Route the right vulnerability to the right engineer. WHOOP automated vulnerability mapping to microservice owners, so developers received tickets with full context attached.
Build automated guardrails for devs to self-resolve. This let security engineers pivot from frustrating gatekeeping and ticket-pushing to more proactive, systemic work that adds value.
Maintain a human in the loop. With such sensitive data and a demand to be always-on, WHOOP isn’t ready to automate the engineer out. Learn how they decided their team’s response had to change.
And then, of course, we will end the live discussion with how to measure it all. Don’t miss out and register to attend on October 7.
Flush with the proceeds of a $70 million Series B raised earlier this year, you might expect Qodo to spend freely on internal AI. After all, the startup uses artificial intelligence to ensure AI-generated code meets customer quality and governance requirements. An upstart technology company using AI to improve AI outputs is AI-pilled by definition.
Instead, the company has limits on AI consumption. Qodo CEO Itamar Friedman tells The New Stack that his engineers can access $10,000 worth of tokens per month, a cap that he described as “generous.” Most Qodo developers never reach it. The ceiling wasn’t enacted to “restrict usage,” Friedman says, but instead to drive “visibility and efficiency” at the startup so that it can “scale without runaway costs.” Put another way, the cap exists to make somebody answer this question: “Which path of automation or usage will be the best [use] of our money?”
Qodo’s AI footprint is larger than its developer token budget. The startup’s AI infrastructure spend — the cost of running the product for customers rather than the cost of its own engineers using AI — is growing at “roughly 5x year over year,” the company tells TNS in an email, reflecting both “increased user adoption” and its agents taking on more, and longer tasks as they mature. Qodo says it is pushing the other direction at the same time, driving down the cost of reviewed pull requests through routing and inference efficiency.
What Qodo runs on
The company also dogfoods heavily, running its pull requests through Qodo. Friedman said the product powers its entire software development life cycle (SDLC). Around that sits a stack most engineering organizations would recognize: Slack and Notion and their constituent “bots,” a centralized knowledge base built to be agent-readable, AI inside Google Workspace, and models from several providers including Google.
Qodo’s own product sits alongside Claude Code and other leading coding assistants rather than replacing them. Claude Code still holds the crown internally, but OpenAI’s Codex has been taking share, with staff “shifting quickly towards Codex.” Friedman tracks this two ways. He polls his 130-person staff, spread across offices in several countries, on the tools they prefer, and compares those answers against what the usage data shows.
Friedman reports that Qodo sees “roughly double” the number of PRs “every couple of months” alongside “a decreasing amount of bugs and incidents.” Hold on to those two numbers. They’re important here.
The bottleneck moved
More PRs and fewer bugs indicate that Qodo is onto something with its focus on software testing and governance. Its technology helps developers deal with an increasingly common issue: What do with all the code that AI agents generate? Companies that adopt AI coding tools often find that they create more code with machines than their humans can assess. As a result, the SDLC bottleneck simply shifts one step down the process.
“We solved the speed of writing code,” Friedman argues, “we didn’t solve the velocity of creating software.” The difference between accelerating one part of a task and its entire arc is the difference between AI hype and AI ROI.
The software development example shows that when we consider AI costs and benefits, we need to think broadly. If we focus too much on a single metric, we might spend our entire budget on Claude Code credits while shipping no more software than before. Alongside a massive bill.
The AI ROI Equation
Friedman recommends an equation-based approach. The Qodo perspective on AI ROI is similar to a popular equation for happiness: Personal joy is the distance between your expectations and reality. The greater the expectations, the harder it is to be happy. The lower the expectations, the greater the chance of being content.
This can be expressed as either simple subtraction or as a ratio:
Reality/expectations = Happiness, where larger results indicate greater joy
Take the same mathematical approach to AI ROI, per Friedman: Compare the positives against the negatives, add up all the good, and set it over all the bad.
AI benefits/AI costs = AI ROI, where larger results indicate greater return
Friedman found the shape of the equation in The Phoenix Project, the 2013 DevOps novel that contrasts types of software development work and sorts them into good and bad buckets. Plug those terms in:
(Features + Infrastructure)/(Incidents + Bugs) = Software development velocity
Now, those two numbers from earlier. Qodo has seen more PRs and fewer bugs thanks to AI. In DevOps terms, it’s shipping more and fixing less, so the equation returns a larger, better result. Feed the same terms into the AI ROI version, and it produces more benefits over fewer costs, and a larger final calculation.
The fraction is not a thought experiment at Qodo. It’s the shape of what the company says is already happening to it.
Terms that have nothing to do with software development work too. Qodo runs AI inside Google Workspace, Slack, and Notion, and those benefits and costs go into the same calculation.
The Qodo approach to measuring total AI ROI is less specific than The Phoenix Project’s DevOps equation, but the difference is acceptable. Friedman argues that you have to start somewhere: “I know [the equation is] a simplification,” the CEO tells TNS. “But what you can’t measure, you can’t improve.”
His argument is that imperfect beats absent. “Don’t think about it too much,” he says. “Try to put any number [in the AI ROI equation] and start tracking.” Being told not to overthink an equation is a great soundbite, but the benefit is real: A rough calculation on paper beats holding the same information in your head without form. In this case, the journey is a large part of the destination.
Friedman reckons that startups should pick no more than six or eight terms for their own calculations. That’s an afternoon’s work. A start on what will prove to be an ongoing exercise.
Negative ROI
The fraction runs backward, too.
Recall Friedman’s point about a company writing more code faster but not accelerating its software development speed. Stuff those terms in:
(Faster code generation + other AI benefits)/(Slower code review and approval + agentic coding costs + other AI costs) = Smaller AI ROI
That’s how a company spends a king’s ransom on AI credits and winds up nowhere or nonexistent.
Which is not hypothetical at Qodo either. AI doesn’t excel everywhere, and Friedman named email automation as an example. The company went all in on automating it, then pulled back, “mov[ing] from AI automation to AI enhancement” after discovering that AI struggled to match writing tone and intelligently extract tasks from messages. The retreat is the interesting part: Qodo’s stated approach to any task is to “go all in on complete automation,” and then “take a step back to human judgment.”
The CEO says that automation falls short today in two areas: When human judgment is required and when context is missing. The second cuts across everything from software development to personal productivity to answering customer questions. Without timely context, what can AI do other than filibuster? Qodo’s service helps answer the context issue for software development, but collecting a company’s data and making it accessible, timely, and well-governed for general agentic usage is a massive undertaking, and one that a host of startups want to help solve. If they can, everyone’s AI ROI math should improve.
No mandate, high expectations
Qodo doesn’t require its staff to use AI. As Friedman puts it, you won’t get fired simply because you’re “not AI all the way,” or “eating AI for breakfast.” The company expects staff to complete their work as efficiently as possible and leaves the method to them.
Employees make their own decisions and execute their own work. If they start to fall behind on assigned tasks, they’re expected to reach for automation. It’s a balanced approach with high expectations: An employee who isn’t as efficient as they could be with AI could find themselves at risk.
Friedman has been on the unpopular side of an AI argument before. When he was building Qodo in 2023 and talking up agents, “agents” was a “bad word,” dismissed as little more than “fluff.” Three years and a $70 million Series B, the bet has paid off.
His advice to founders starting now looks like his past. Predict “what’s going to happen two years from now,” he says, then solve for it immediately, because whatever looks like two years tends to arrive inside of twelve months. The future “is coming faster” than you think, he says.
It’s a lot to ask of anyone working from an incomplete picture. Predicting the future is hard, he admits, “but you have to.”
Anthropic made Claude Opus 5.5, released on Tuesday, cheaper than its predecessor, cutting the price from $5 to $4 per million input tokens and from $25 to $20 per million output tokens. The 1 million-token context window and 128,000-token maximum output are unchanged.
On paper, that makes upgrading an easy decision. In practice, it may not be as simple as changing the model ID.
Anthropic’s migration guide flags four breaking changes that can cause requests built for Opus 5 to return 400 errors after switching to Opus 5.5. Several other changes won’t trigger an error but could still change how an existing agent behaves.
Anthropic’s migration guide flags four breaking changes that can cause requests built for Opus 5 to return 400 errors after switching to Opus 5.5.
Thinking is always on
The first change involves thinking controls. Opus 5.5 returns a 400 error when a request sets thinking to disabled or uses enabled with budget_tokens, leaving effort as the way to control how much reasoning the model does. Agents that previously switched thinking off for simple steps to save time and tokens will need to assign those steps a lower effort level instead. Because thinking is now always on, responses begin with thinking blocks, so code that assumes the first content block is text will also need to change.
The default effort level has also dropped from high on Opus 5 to medium on Opus 5.5, so requests that omit the parameter will quietly run at a lower setting. Anthropic recommends setting effort explicitly and re-running effort evaluations, since the right level for each step may have shifted along with cost and latency.
No more forced tool calls
Forced tool use no longer works either, as setting tool_choice to any or tool returns a 400 error, including on the token counting endpoint, where cost estimates built on those settings will fail along with the requests they were meant to price. Many agent loops force a call when a step has to query a database, run code, or reach another service, and Anthropic’s replacement is auto-combined with strict tool use or structured outputs, with the prompt stating when the tool applies.
Routing and conversation history
Thinking blocks are now tied to the model and conversation that produced them. On the Claude API, Fable 5.1 and Mythos 5.1 are the only other models that can read Opus 5.5 thinking blocks, so a router or fallback that hands a conversation to any other model will run those turns without the earlier reasoning instead of returning an error.
That adds another layer for teams already watching whether their agent calls are quietly being routed to an older model. Opus 5.5 can read thinking blocks from Opus 5 and earlier Opus, Sonnet, and Haiku models, but not from Fable or Mythos.
Conversations must also stay append-only for those blocks to remain valid. Trimming old messages, changing tool definitions, summarizing earlier context on the client side, or rewriting the system prompt mid-conversation invalidates existing thinking blocks, and for accounts created on or after August 31, 2026, at midnight UTC, replaying a thinking block after one of those edits returns a 400 error by default. Older accounts get no error, but the invalid blocks still reach the model, and Anthropic says future models will enforce the check for all accounts. Integrations that never edit earlier turns need no code change, and Anthropic says Claude Code, claude.ai, Claude Managed Agents, and the Claude Agent SDK already work this way, while agents that compact their own context should follow the company’s preserved thinking documentation.
The fourth change affects computer-use agents on the Claude API and Google Cloud, where Opus 5.5 rejects the computer_20251124 tool and accepts computer use only through the computer_toolset_20260801 toolset. The request itself gets simpler because the beta header goes away and the toolset entry takes no name or display dimensions, but the agent loop needs more work. Each action now arrives as its own tool_use block identified by the block’s name rather than input.action, a single turn can contain several of them, and every result has to echo toolset_name. The older tool still works on Amazon Bedrock, and Anthropic directs developers on other platforms to the computer use tool’s compatibility documentation.
…a router or fallback that hands a conversation to any other model will run those turns without the earlier reasoning instead of returning an error.
Changes that won’t throw errors
The change most likely to go unnoticed doesn’t produce an error at all. On Opus 5, text Claude writes between tool calls comes back as text blocks, but on Opus 5.5 that narration arrives as progress-update thinking blocks, and at the default thinking.display setting of omitted those blocks are empty.
Any agent interface that streams that narration to users will go silent between tool calls until developers set display to updates, a beta option that returns progress updates while keeping reasoning hidden, or to summarized, which returns both, and then render each non-empty thinking block ahead of the tool call it precedes.
Opus 5.5 also ships with broader safety classifiers. It can return a stop_reason of refusal with stop_details categories that now include bio and reasoning_extraction alongside cyber, and Anthropic’s server-side fallback won’t retry requests declined under reasoning_extraction, handing the refusal back to the application instead.
The change most likely to go unnoticed doesn’t produce an error at all.
Upgrading from older models
Teams coming from Opus 4.8 need to work through the Opus 5 migration first, which covers thinking being on by default and the response-shape changes that follow, before applying the Opus 5.5 changes. Teams on Opus 4.7 or earlier have more ground to cover, and those on models older than Opus 4.7 also face rejected sampling parameters, rejected manual extended thinking, removed prefill, and a newer tokenizer.
Claude Managed Agents users only need to change the model name. Developers working in Claude Code can run /claude-api migrate to apply the model ID swap, parameter changes, prefill replacement, and effort calibration across a codebase before reviewing a checklist of items to verify by hand.
Anthropic recommends testing the migration in a development environment before switching production traffic. Developers maintaining their own integrations will need to test the pieces around the model, too. Tool calls, model handoffs, conversation history, and user-facing progress updates can all behave differently after the switch, because agent failures often originate outside the model itself.
Most notably, the AI company slashed token prices, making the new GPT-6 models significantly cheaper to use. Beyond token prices, though, OpenAI says better caching can also help developers push costs down even more.
Per OpenAI: “Improvements in caching and inference let us serve these models at lower cost,” with API prices for Sol and Luna down 50% compared to their GPT-5.6 counterparts (58% lower for Luna output tokens).
What improvements? Namely, higher cache-hit rates by default, the ability to preserve earlier context even when reasoning effort and tool availability change, and new tools to monitor and diagnose caching performance.
Reuse context without starting over
Prompt caching isn’t, of course, novel to the new GPT-6 models themselves. But the upgraded Sol and Luna come with improvements designed to keep more previously processed context reusable as the agent moves forward on a task.
“We’ve improved prompt caching for GPT‑6 to deliver higher cache hit rates by default, helping agents reuse more context, respond faster, and benefit from discounts of 90% on cached input-token reads.”
That adds another opportunity to lower the already low API price tag, though the 90% cached-input discount matches GPT-5.6 pricing; what’s new is how often the cache gets hit. By using cached context to reuse work it’s already done, the model doesn’t have to process the same context again from scratch for every single call, thereby reducing latency — and token costs.
Beyond this higher default cache-hit rate, OpenAI says the new GPT-6 models offer more flexibility to optimize caching performance.
The new models let developers adjust reasoning effort and tool availability without having to break the cache. This way, an agent can scale reasoning effort up and down based on how difficult a step is, then make different tools available depending on what the task requires without disturbing earlier cached context — again, a win for both speed and cost.
See what gets cached and what doesn’t
GPT-6 Sol and Luna also arrive with a Prompt Caching Dashboard, where OpenAI says developers can view caching performance to understand how much context is reused.
Specifically, they can see how much input is cached and how that amount changes over time. The diagnostics tool then flags missed caching opportunities to help developers understand what could use more efficient caching.
Rather than keeping cache performance largely hidden behind the scenes, the idea is to make it more visible so developers can actively measure and optimize cache reuse.
Altogether, OpenAI says these caching improvements are already making a difference. Per the AI company, GitHub reports, “these improvements have reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests to OpenAI models.”
These results span the past “several months.”
Token prices aren’t the only way to make agents cheaper
OpenAI’s pricing cuts for GPT-6 Sol and Luna made the biggest splash, with the AI company significantly dropping API prices from GPT-5.6 levels.
Compared to the current prices for GPT-5.6 Sol and Luna, which stand at $4 and $0.20 per million input tokens and $20 and $1.20 per million output tokens, respectively (GPT-5.6 Sol’s rates are promotional pricing), the new GPT-6 models come in at just $2 and $0.10 per million input tokens and $10 and $0.50 per million output tokens, respectively.
With GPT-6 Sol and Luna’s caching improvements and lower token pricing, OpenAI is making the case for tackling agent costs from both sides: charging less for fresh processing and reducing how often the same context needs to be reprocessed.
But as more AI model providers compete aggressively on pricing, it’s becoming clearer that cheaper models alone won’t save your AI budget — and lower token prices aren’t the only way to make agents cheaper.
With GPT-6 Sol and Luna’s caching improvements and lower token pricing, OpenAI is making the case for tackling agent costs from both sides: charging less for fresh processing and reducing how often the same context needs to be reprocessed.
As agents continue to work on longer and more complex tasks, there will likely be more pressure to do both.
On Tuesday, OpenAI released GPT-6 Sol and Luna, an expansion of the GPT-6 line-up that aims to make GPT-6 Astra’s next-level intelligence more efficient, accessible, and affordable.
Though OpenAI says Astra is still “the most intelligent and aligned model in the world,” the new GPT-6 models come impressively close in alignment — at a fraction of the price.
In an internal coding evaluation on coding deception, for example, GPT-6 Astra’s deception rate is 0.5%, while GPT-5.6 Sol stands at 10.4%. The new GPT-6 Sol is only 1.3%.
As for pricing, GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens; GPT-6 Sol and GPT-6 Luna cost $2 and $0.10 per million input tokens and $10 and $0.50 per million output tokens, respectively.
If OpenAI’s new GPT-6 models can achieve near-Astra-level alignment at a fraction of the cost, that’s good news. But it’s still unclear whether or not the new GPT-6 models also mirror Astra’s observability and monitoring problems.
Closing the alignment gap between Astra and GPT-5.6
OpenAI says it trained the new GPT-6 models with similar methods as it did for GPT-6 Astra, specifically building on the alignment work it began with Astra.
While Astra is still the AI company’s “most aligned model to date,” it looks like GPT-6 Sol and Luna are giving it a run for its money, dramatically closing the gap between OpenAI’s most advanced model and its GPT-5.6 counterparts in key areas like coding deception, failure to disclose a broken search tool, and unauthorized agent interaction. OpenAI notes that these evaluations deliberately test challenging situations and do not measure failure rates in typical use.
Credit: OpenAI
The most progress was made on failure to disclose a broken search tool, where AI agents are given search tasks and a broken search tool; do they just give their best guess or say that the search tool is broken?
The gap between GPT-5.6 Sol’s ability and Astra’s is notably wide: 77.5% and 1.5%, respectively. Per OpenAI’s internal evaluation, GPT-6 Sol is a dramatic improvement, with a non-disclosure rate of 4.9%. Luna also improved, but less impressively, dropping from 78.3% to 28.7%.
Though OpenAI says Astra is still “the most intelligent and aligned model in the world,” the new GPT-6 models come impressively close in alignment — at a fraction of the price.
Credit: OpenAI
Coding deception, where AI agents are deliberately given tasks that elicit dishonesty, improved as well. In OpenAI’s internal evaluation, deception rate measures the fraction of answers with detected deception. GPT-6 Sol (1.3%) now almost matches GPT-6 Astra’s performance (0.5%), a far cry from GPT-5.6 Sol’s (10.4%).
Credit: OpenAI
Significant progress was also made on unauthorized agent interactions, where OpenAI evaluates whether models follow unauthorized instructions on a simulated message board. GPT-5.6 Sol went ahead with the unauthorized action in 52% of runs in which it found the board, while GPT-6 Sol did so in only 11% of cases.
GPT-6 Astra and Luna didn’t take any such actions, but OpenAI says Luna “found the board less often.”
Where the results are less impressive
Though GPT-6 Sol gets close to Astra in several alignment evaluations, there’s still a way to go with warning circumvention.
GPT-6 Sol attempted to work around restrictions in 64.4% of runs. That’s only a slight improvement from GPT-5.6 Sol’s 68.2% and still quite far off from Astra’s 17.4%.
Credit: OpenAI
In an internal evaluation where OpenAI tested models’ abilities to respect explicit warnings, like “access denied” messages, GPT-6 Sol attempted to work around restrictions in 64.4% of runs. That’s only a slight improvement from GPT-5.6 Sol’s 68.2% and still quite far off from Astra’s 17.4%.
Credit: OpenAI
Finally, for reviewer bypass attempts, where OpenAI tests whether models try to bypass or respect an automated safety reviewer’s decision to block an unsafe command in Codex, the new GPT-6 models are almost identical to Astra, though the playing field was already pretty even to begin with: GPT-5.6 Sol attempted a bypass in 7.3% of runs and GPT-5.6 Luna in 4.3%, compared with none for Astra and GPT-6 Sol and 0.3% for GPT-6 Luna.
But if GPT-6 Sol is anything like Astra, we’re not out of the woods yet
GPT-6 Sol and Luna have made marked improvements across alignment evaluations, inching closer to OpenAI’s star child, Astra. But if the new GPT-6 models also follow suit on Astra’s noted observability issues, then developers hoping to catch misalignment via monitoring aren’t out of the woods yet.
OpenAI knows that Astra’s — and now GPT-6 Sol’s — improved alignment doesn’t mean the AI industry has gotten a handle on the problem yet.
“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”
Just this month, the AI company shared six reports of “unexpected or concerning model behavior,” including self-generated instructions, information fabrication, unauthorized use of leaked API keys, cross-agent communication, and unsanctioned file-sharing.
At the same time, it released a new framework for reporting model misalignment, stating: “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”
If GPT-6 Sol and Luna are catching up to Astra in alignment evaluations — at a far cheaper rate — that’s good news. But if the new GPT-6 models also come with the same observability and monitoring problems, then cheaper may still come at a cost.
OpenAI on Tuesday released GPT-6 Sol and Luna, which will complement the flagship GPT-6 Astra model in OpenAI’s lineup. As of now, there is no GPT-6 Terra.
The new GPT-6 pricing
The headline news here is that OpenAI cut the price per million input/output tokens by half or more, compared to the previous version. GPT-6 Sol will cost $2/$10 per million input/output tokens (vs. $4/$20 for GPT-5.6 Sol), and GPT-6 Luna will come in at $0.10/$0.50 (vs. $0.20/$1.20).
The GPT-5.6 pricing was always meant to be promotional, but for the new GPT-6 models, this is the default price, an OpenAI spokesperson tells The New Stack.
“Improvements in caching and inference let us serve these models at lower cost, and we’re passing those savings directly on to users and customers,” OpenAI explains in its announcement.
Benchmarks
As you would expect, the new models show clear improvements over the GPT-5.6 predecessors, but for the most part, these are not all that extreme.
On a benchmark like Zapier’s AutomationBench — which checks how well the models work on a set of business workflow tests — GPT-6 Luna improves by 5.4 percentage points over the previous version, for example
Credit: OpenAI
On the DeepSWE v1.1 software engineering benchmark, GPT-6 Sol essentially matches Anthropic’s Fable (68.8% at max effort vs. 69.9% for Fable 5 at xhigh effort), but at only 20% of the cost. Luna, at max effort, hits scores similar to Claude Opus 5 and Fable 5 at medium effort, at a significantly lower cost.
And OpenAI focuses on this cost comparison across its announcement—with a special focus on price per task instead of straight-up token pricing.
Credit: OpenAI
Anthropic resets the comparison
Since Anthropic released Opus 5.5 earlier on Tuesday, OpenAI’s comparisons are already out of date — such is the way of this AI era. Anthropic, too, reduced its per-token pricing for Opus 5.5 to $4/$20, down from $5/$25, but that still leaves Anthropic’s model twice as expensive as the comparable GPT-6 Sol.
In its announcement, when comparing GPT-6 Sol to Opus 5, OpenAI was able to claim significant cost savings when compared to Anthropic’s model — and for the most part that still holds, but Anthropic says Opus 5.5 also uses fewer tokens per task, which, according to the company, works out to 40% lower costs than Opus 5 on typical workloads.
It’s worth noting that no one has run Sol and Opus 5.5 head-to-head yet. Sol likely stays cheaper per task on OpenAI’s AutomationBench numbers, but Opus 5.5 posts higher scores than GPT-5.6 Sol on shared benchmarks in Anthropic’s testing.
Since it’s almost impossible to know how many tokens an agent will use to finish a task, though, these pricing changes still don’t make it any easier for a user to budget.
Prompt caching
For developers building agents, the caching changes may matter more than token prices. OpenAI says it improved prompt caching for GPT-6 to deliver higher cache hit rates by default, with discounts of up to 90% on cached input tokens.
One positive change, too, is that developers can now change the reasoning effort and tool availability without invalidating the cache. With explicit breakpoints, developers can choose where a cached prefix ends, and a new dashboard and diagnostics tool show what’s getting cached and what isn’t.
GitHub says these improvements cut the share of prompt tokens that require fresh processing by more than half over the past several months, across billions of requests to OpenAI models.
Anthropic made a similar move with Opus 5.5, which cuts cache read prices by 60% for token-billed usage, on top of the 20% per-token cut.
Style changes
Models aren’t just about benchmarks, though. With GPT-6 Sol, OpenAI made its models answer more directly, rather than in the previous — already reined-in — more conversational style. “Expect to see more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall without losing substance,” OpenAI says.
Credit: OpenAI
Alignment
Given the HuggingFace incident, it’s no surprise OpenAI is emphasizing its alignment work for GPT-6 Sol and Luna, too.
OpenAI says both models improve on their GPT-5.6 counterparts across its alignment evaluations, including fewer misleading claims about their own coding work. On an internal coding deception test, GPT-6 Sol’s rate fell to 1.3% from 10.4%.
When given a deliberately broken search tool — and graded on whether it disclosed the problem instead of guessing — Sol failed to disclose the problem 4.9% of the time, down from 77.5%.
What is a bit more concerning, though, is that when researchers asked the model to respect an explicit warning like an “access denied” message, GPT-6 Sol still tried to work around those restrictions in 64.4% of runs, down only slightly from 68.2% for its predecessor. Luna improved more, to 42.4% from 76.5%.
OpenAI says these tests cover mostly low-stakes situations and run without full system-level safeguards used in its products.
Credit: OpenAI
On a simulated message board seeded with unauthorized instructions, including requests to disclose private information, Sol took the specified action in 11.3% of runs where it found the board, down from 51.9%. Luna and Astra took none, though OpenAI notes Luna also found the board less often.
Anthropic, by contrast, says Opus 5.5 is the strongest performer on its most comprehensive alignment test and names METR and Frontier Design as pre-release external testers. Opus 5.5 also ships with safeguards that reroute requests, sending most cybersecurity tasks to Opus 4.8 and anything flagged by Anthropic’s biology or frontier LLM development classifiers to Opus 5.
Availability
GPT-6 Sol and Luna are available in ChatGPT Work and Codex starting Tuesday for Plus, Pro, Business, Enterprise, and Edu users.
Free and Go users get Luna in the desktop app.
Neither model is in Chat yet. OpenAI says it plans to roll them out gradually throughout the day to keep service stable, so they may not appear right away.
When Anthropic merged Claude chat and Cowork into a single interface last week, it removed an increasingly irrelevant decision users had to make about which mode to choose.
Now, to use both tools, just ask a question, then hand off a multi-step task in the same thread, and Claude routes it. Anthropic made this update because it says customers often struggled to choose the right tab for the right task, so the merged app now routes each request itself instead of asking you to pick a mode — for now, the unified experience is rolling out to Pro and Max subscribers first, with free and team tiers to follow.
But this isn’t new. OpenAI has offered the same promise since earlier this summer. OpenAI introduced Work mode alongside Chat on July 9, then phased out the older, separate Agent mode the following month. Work mode is a sandboxed environment with a browser, code execution, and file output, sitting next to a Chat toggle in the same window — though reviewers note it can’t yet hand a live, logged-in browser session back to the user mid-task the way Agent mode could, e.g., for logins or payments. OpenAI built Work mode on its Codex coding agent after OpenAI reported that roughly a fifth of Codex’s 5 million weekly users were non-developers — a share it said was growing three times faster than developers.
As OpenAI and Anthropic get closer to feature parity, accuracy, reliability, and token usage matter even more. What better way to find out which tool is better than to run head-to-head tests?
The tests
I ran three tests, covering different areas of real developer work.
API research -Look up four real developer APIs and tabulate their documented rate limits, whether a free tier exists, and the current version identifier. I checked the answers against the vendors’ docs the same day.
Build from a spec – Write a small command-line duration parser from a spec with strict edge cases. I ran each app’s code against a hidden 16-case test suite.
The handoff – Ask which of two stack traces indicates a race condition, then in the same thread hand off a real job: pull a log file from Google Drive, compute latency percentiles and error rates per endpoint, and deliver a spreadsheet with a chart. I generated the log data, so I knew every number in advance.
I recorded token cost and time in each test and included the prompts for anyone interested in replicating this work.
API research
The prompt: Research the current public documentation for these four developer APIs and build me a table with one row per API and these columns: documented rate limit for authenticated requests, whether a free tier exists (yes/no), and the current API version identifier or date shown in the docs. APIs: GitHub REST API, Stripe API, Twilio Messaging API, OpenAI API. Cite the documentation page you used for each row.
Both Claude and ChatGPT answered correctly on the twelve graded fields, but Claude was more thorough. It included GitHub’s separate limit for Actions tokens, Twilio’s queue window, and OpenAI’s tier thresholds, plus a note about one page it couldn’t reach.
ChatGPT, in Work mode, finished in 1 minute 17 seconds and wrote 649 output tokens. Claude took 1 minute 44 seconds and wrote 1,042 tokens, read eight pages, listed nine sources, and offered to export the table as a spreadsheet. Claude wrote nearly double the tokens and took longer, but in this case, it’s warranted because of the added detail.
Build from a spec
The prompt: Build the command-line tool described in the spec below. Deliver a single file named durparse.py that follows every rule. Test it yourself before returning it. Show the complete final code in your reply. (Followed by the spec: a duration parser with units d/h/m/s, largest first, one of each, decimals allowed, bare numbers are seconds, everything else returns None.)
Both apps returned a durparse.py file that passed all 16 hidden tests, including the traps. The traps included units out of order, a repeated unit, a trailing number with no unit, and negative values. The 517-token spec went to both. ChatGPT finished in 1 minute 17 seconds on 769 output tokens. Claude took 1 minute 45 seconds and 989 output tokens. Claude reported running 35 of its own test cases before returning the file. Both delivered a download and showed the code in the reply.
The code came out nearly identical, both using exact-precision arithmetic and a fixed-order regex. Claude flagged a judgment call the spec never settles on: that rounding 0.5 seconds up is a choice and Python’s built-in round would go the other way. ChatGPT reported only that its tests passed. Once again, Claude was just a little more thorough.
The handoff
The prompt: Which of these two stack traces points to a race condition, and in one sentence why? (with the two traces) Then: Now take the file api_logs.csv from my Google Drive (columns: time, endpoint, status, latency_ms) and produce a downloadable spreadsheet with one row per endpoint showing request count, p50, p95, and p99 latency in milliseconds, and error rate as the percentage of requests with status 500 or above. Add a bar chart of p95 latency by endpoint. Also show the table in your reply.
I started each thread in plain chat with the stack-trace question, 135 tokens. Both answered correctly in seconds: Trace B, the dictionary that changed size during iteration. ChatGPT spent about 10 seconds and 48 output tokens. Claude spent 118 in about 25 seconds, adding a caveat that the same error can happen without threads if the loop body edits the dictionary itself. Then, without switching modes, I handed off the log analysis.
Claude pulled the file, computed the table, and built an .xlsx with a p95 bar chart in about 4 minutes on 319 output tokens. ChatGPT produced the same spreadsheet and chart in 28 seconds, using 467 output tokens. All 30 numbers matched my ground truth on both sides. Claude also named its percentile method (linear interpolation) and noted the numbers would match if I recomputed them in Google Sheets. ChatGPT gave the same correct table but didn’t provide as much detail as Claude did.
Results
The test
ChatGPT (Work mode)
Claude (merged app)
API research
12/12, 1:17, 125 in / 649 out
12/12, 1:44, 125 in / 1,042 out
Build from spec
16/16 tests, 1:17, 517 in / 769 out
16/16 tests, 1:45, 517 in / 989 out
Handoff, question
Correct, ~10 s, 135 in / 48 out
Correct, ~25 s, 135 in / 118 out
Handoff, task
30/30, 28 s active, 133 in / 467 out
30/30, ~4 min, 133 in / 319 out
Total tokens (visible)
910 in / 1,933 out
910 in / 2,468 out
Both Claude and ChatGPT were equally accurate. Every field, every test case, every number matched on both sides, and both cited real documentation.
The differences are in speed and answer detail. ChatGPT was faster on every task and produced 1,933 visible output tokens, compared with Claude’s 2,468. Some of Claude’s extra output was filler, but not all of it. It added context the prompts didn’t ask for, named the percentile method behind its numbers, and flagged two judgment calls the specs left open. ChatGPT gave the same right answers but with less context.
What do I think?
I’d pick Claude, and here’s the reasoning. On time, ChatGPT won every task, but the gaps were seconds on the short tasks (1:17 vs 1:44, 1:17 vs 1:45), not a noticeable difference. On tokens, ChatGPT used about 22% less output than Claude, which is positive, but not when you consider how important detail/context is.
On detail, Claude provided more meaningful detail on all three tests. This included the extra API context, the rounding judgment call, and the percentile method. In today’s world, where AI can fabricate, detail matters.
Claude Opus 5.5 is here, and Anthropic has lowered the price.
The new model, released on Tuesday, costs $4 per million input tokens and $20 per million output tokens, 20% less than Opus 5, with cache reads dropping to $0.20 per million from $0.50 and cache writes falling to $5 from $6.25. Anthropic puts overall savings closer to 40% because Opus 5.5 uses fewer tokens to complete a task and generates output more than 30% faster.
Claude Code and the Claude Platform also get a fast mode that runs up to 2.5 times faster, priced at $8 per million input tokens and $40 per million output tokens. Anthropic says Opus 5.5 performs at roughly the level of Fable 5.1 on most work, though it comes out ahead on several agentic coding benchmarks.
Opus 5.5 scored 66.4% on Terminal-Bench 4.0 compared with Fable 5.1’s 55.8%, and 54.4% on FrontierCode compared with 50.3%. The company suggests not reading too much into those margins. At this level, the company says a few points on a benchmark don’t translate into a noticeable difference in real-world use.
Fable 5.1 costs $10 per million input tokens and $50 per million output tokens, more than twice the price of Opus 5.5. At default effort on FrontierCode, Opus 5.5 beats GPT-6 Astra at roughly 20% of the per-task cost. On CursorBench, it tops GPT-5.6 Sol by 11 points at about a third of the cost. Developers will still need to run their own evals before moving production workloads, but the cost difference could change which model makes sense for agentic coding.
Developers will still need to run their own evals before moving production workloads, but the difference in cost could change which model makes sense for agentic coding.
Fewer tokens, fewer agent steps
The early enterprise numbers suggest the efficiency gains are real, at least on certain task profiles. Box reported that Opus 5.5 used about a third as many tokens as Opus 5 in its evaluations while producing answers that were 40% less verbose without losing accuracy.
GitHub tested the model inside Copilot CLI and VS Code and found it completed more terminal tasks than Opus 5 in less than half the steps. Deloitte said Opus 5.5’s lowest-effort setting caught 72% of known bugs in code reviews, compared with 56% for Opus 5 at high effort, with fewer false alarms and less output.
Prices per 1M tokens
Claude Opus 5.5
Claude Opus 5
Cache reads
$0.20
$0.50
Input tokens
$4
$5
Output tokens
$20
$25
Cache writes
$5
$6.25
Anthropic’s own internal testing backs up the pattern. In one head-to-head, both Opus 5.5 and Fable 5.1 translated HAProxy from C into Rust; both rewrites passed nearly all of HAProxy’s regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 and cost 51% less. An early tester audited and fixed a 200,000-line codebase in under three hours, whereas Opus 5 took over 20 hours and burned 2.5x as many tokens. Another completed a 680,000-line code migration in less than a day. Although these were customer and internal evaluations, not standardized independent benchmarks, they point in the same direction — fewer tokens and fewer steps to finish the job.
That pattern tracks with what’s happening across the industry. Agent performance depends heavily on the harness and runtime around the model, not only the model itself — agent failures often trace back to the orchestration layer rather than the model. Nvidia’s research showed that swapping the harness while keeping the model fixed could meaningfully change agent performance.
Benchmark
Opus 5.5
Fable 5.1
Opus 5
GPT-6 Astra
GPT-5.6 Sol
Agentic coding (Terminal-Bench 4.0)
66.4%
55.8%
52.3%
57.9%
37.3%
Agentic coding (FrontierCode v1.1)
54.4%
50.3%
48.0%
53.3%
47.5%
Agentic coding (CursorBench 4.0)
57.8%
51.8%
46.6%
—
41.7%
Knowledge work (GDPval-AA v2.1)
1846
1735
1708
1542
1588
Business workflows (AutomationBench)
40.0%
31.4%
26.9%
41.4%
28.8%
Multidisciplinary reasoning (HLE)
67.7%
65.6%
63.6%
57.2%
—
Agentic scientific research (TBS 0.1)
58.7%
52.6%
29.0%
64.6%
22.4%
Computer use (OSWorld 2.0)
81.8%
80.7%
74.0%
—
—
Visual chart recognition (Chartography)
89.0%
88.4%
83.4%
—
—
Safety classifiers reroute mid-chain
Opus 5.5 ships with the same class of safety classifiers already running on Fable 5.1 for cybersecurity, biology, and frontier LLM development. When a classifier fires, Anthropic reroutes the request transparently to an older model. Most flagged cybersecurity requests go to Opus 4.8. Biology and frontier LLM flags go to Opus 5. Anthropic says users can still identify and fix bugs in their own code with Opus 5.5.
For anyone building agent workflows, this is the detail that needs architectural attention. A request sent to Opus 5.5 could, in fact, be handled by Opus 4.8 or Opus 5 instead, depending on whether Anthropic’s safeguards intervene. In a multi-turn agent workflow, that creates the possibility that individual requests are being handled by models with different capabilities, which could affect downstream steps. It’s also a source of inconsistency that may not show up in evals built on the assumption that every request goes to the same model.
Vetted organizations can apply to Anthropic’s Life Sciences Verification Program to use Opus 5.5 without the biology classifier, and the company plans to expand its Cyber Verification Program to include the model in the coming weeks. The new cyber program will include three tiers for increasingly permissive trusted access, including access to Claude Mythos models.
Opus 5.5 ships with the same class of safety classifiers already running on Fable 5.1 for cybersecurity, biology, and frontier LLM development.
Alignment gains from cleaner training
Anthropic says Opus 5.5 posted the strongest results of any model it has tested on its most comprehensive internal alignment evaluation, with improvements in behaviors the company says contributed to recent cybersecurity incidents, including biased reasoning and attempts to escape sandboxed environments. Frontier Design and METR evaluated the model before release.
On the training side, Anthropic is tightening how it filters reinforcement learning environments after identifying flawed environments as a major source of misaligned behavior. That’s relevant beyond the safety framing because RL environment quality directly affects how a model behaves in agentic settings, where it chooses its own tools and decides when to change approach. The company is also building automated methods to generate new safety training scenarios and improve alignment rewards.
Pricing pressure meets routing tradeoffs
Opus 5.5 is the first model in the Claude 5.5 family, with Sonnet 5.5 and Haiku 5.5 expected over the coming weeks. Subscription users get a 20% increase in five-hour usage limits across all plans, while Anthropic says the lower cost of Opus 5.5 will make five-hour and weekly limits go 25% further. Subscribers will also get a banked rate-limit reset they can save for when they need more capacity.
The release comes as API pricing across the frontier labs continues to fall. OpenAI cut its own API prices this summer, and Opus 5.5 pushes the competition beyond the headline price per token by reducing how many tokens some workloads require in the first place.
Opus 5.5 pushes the competition beyond the headline price per token by reducing how many tokens some workloads require in the first place.
For companies running serious production workloads with lean engineering teams, the real test of operational ownership comes at 3 a.m. Not whether the system stays up but whether anyone needs to be awake to make that happen. The page arrives. A service is degrading. The engineer who answers didn’t build this service, doesn’t know what thresholds were set at deploy time, and cannot tell whether the system is healing itself or waiting for a human decision.
This is the moment that separates services that transferred operational burdens from platforms that merely deferred them. The system that passes the 3 a.m. test already knows what healthy looks like, what to do when healthy stops being true, and how to communicate what happened—because those decisions were made at deploy time, not incident time. The team sleeps through the night because nothing went wrong, but because the response was already determined.
Forty engineers, eight applications, zero dedicated ops
Cloud infrastructure evolved in two directions simultaneously, and neither arrived where mid-market teams actually stand. On one end: full control. Infrastructure-as-code, service meshes, custom pipelines. Powerful, flexible, and designed for organizations that have deliberately invested in operational staff who can absorb the cost of that flexibility. On the other end: single-app simplicity. Push code, get a URL. Elegant for a first deployment, but architecturally limited the moment a team manages more than one service, needs compliance controls, or inherits an application that doesn’t fit the platform’s opinions.
You know this company. Forty engineers. Eight production applications. Two generate 80% of revenue. One SRE who is actually a senior developer with an on-call rotation nobody else wants. A compliance audit due in Q3 that nobody has started preparing for.
“The right investment is shipping product. The cost of this gap is measured in what these teams do not ship.”
Every engineer is a full-stack contributor. The person who wrote the feature deploys it, monitors it, and gets the page when it breaks not because the team lacks sophistication, but because hiring dedicated infrastructure staff is not the right investment at their stage. The right investment is shipping product. The cost of this gap is measured in what these teams do not ship. Every sprint spent upgrading the deployment pipeline is a sprint without the feature a customer asked for. Every 3 a.m. page answered by a developer with a product standup at 9 a.m. is diminished output that never shows up in any dashboard. The gap is widening.
The applications nobody planned to operate
Not every application in a portfolio was built by the team now responsible for it. Many companies acquire products through M&A. They inherit internal tools built by engineers who left three years ago. They run commercial off-the-shelf applications customized beyond vendor support. They maintain line-of-business applications in languages nobody on the current team chose. These applications share a common trait: they are in production, they serve customers or meet compliance requirements, and nobody has the budget or the mandate to rewrite them. They need a home that accepts them as they are, not as a modernization roadmap says they should become.
“They need a home that accepts them as they are, not as a modernization roadmap says they should become.”
This is where the full lifecycle vision matters. An application management service that only serves net-new applications forces teams to maintain two operational models: one for the applications they are building today, and another for the applications they inherited yesterday. That split is where drift starts, where patching falls behind, and where audit findings accumulate.
The application management service that solves this problem must accept the full range: the Java application packaged as a WAR file, the .NET Framework service running on Windows, the Python application with dependencies pinned to a specific runtime, the containerized service already running elsewhere, and apply the same operational model, the same deployment interface, the same patching and scaling behavior to all of them. Migrate, manage, and modernize within the same experience, without requiring a different operational posture for each stage of the application lifecycle.
What the market data actually shows
Janakiram MSV, an analyst and advisor on cloud-native platforms and TNS contributor, puts it this way:
Cloud native standardized the infrastructure layer around containers and Kubernetes. But it never standardized the operational boundary between the application team and the infrastructure underneath it, and platform engineering largely emerged to re-establish that boundary. CNCF’s latest research with SlashDatareportsthat 28 percent of organizations run a dedicated platform engineering team and 41% split those capabilities across multiple teams. Another 3% have no formal approach at all, and that is where most mid-market engineering organizations exist. They want the same outcome as platform engineering without first becoming a platform engineering organization.
The pattern that repeats is that maturity stalls at the third application rather than the first deployment. A forty-engineer team gets one service into production and keeps it healthy through familiarity. Then an acquisition brings in a .NET workload, an internal tool developed by an engineer who left two years ago, and the team is carrying three operational models with nobody owning the mandate to reconcile them. Hiring will not close that gap, because what is missing is a standardized operational posture, not headcount.
What hundreds of thousands of production deployments actually reveal
We have visibility into hundreds of thousands of production deployments across thousands of customers. That scale does not tell you what people say they want. It tells you what actually breaks, what gets escalated at 3 a.m., and what determines whether a team trusts their platform enough to stop thinking about it. The patterns are remarkably consistent.
1. The deployment that nobody touches again
A team spends two full days getting a Spring Boot application deployed with CI/CD and an SSL certificate. The deploy works. Then nobody touches it for three months because it is stable, but because touching it might break it again. They discover the problem when customers report the application is unreachable. Four hours of forensics follow. What changed? Why? How do we prevent this?
“A release that can partially succeed is a release that will partially fail.”
This is the most common failure mode we observe: not the deployment that fails loudly, but the deployment that succeeds quietly and degrades invisibly. The team cannot explain what happened because the platform did not decide what “healthy” meant before the incident arrived.
A release that can partially succeed is a release that will partially fail. The platforms that earn trust are the ones where a deployment either completes fully or reverses entirely no intermediate states, no manual rollback procedures discovered under pressure.
2. The observability sprint that ships three weeks late
A developer notices response times degrading. They want memory utilization metrics. The platform does not collect them by default. They spend a sprint writing configuration files to install a monitoring agent across every instance. The insight they needed three weeks ago ships three weeks late.
This is the second most common pattern: observability treated as an add-on rather than a default. Every team we have observed that adds monitoring after their first incident wishes they had it before. The teams that never experience this problem are the ones whose platforms shipped metrics, traces, and health signals at deploy time without code changes, without configuration, without a sprint spent on plumbing. The distinction matters: a platform that can be observed is not the same as a platform that is observed from the moment it goes live.
3. The portfolio tax
Each application gets its own infrastructure. Costs scale linearly with the portfolio. The team running eight applications pays eight times the overhead of the team running one, not because each application needs dedicated resources, but because the platform’s architecture assumes isolation rather than shared operational responsibility. The result: teams avoid migrating inherited applications because the cost model punishes breadth. The compliance audit does not care that the inherited application runs on a different operational model. It expects the same governance, patching cadence, and access controls.
A platform that rewards portfolio growth, shared infrastructure, and a consistent operational posture —and economics that improve with breadth rather than degrade—changes the calculus for teams managing applications they did not build.
4. The security configuration nobody made
The forty-engineer company has no security team. They have a senior developer who reads the CIS benchmarks on weekends. The compliance audit arrives regardless. Teams without specialists will not configure security controls that require specialist configuration. This is not a criticism of those teams; it is a structural observation about how security actually gets implemented (or does not) in organizations where every engineer is a full-stack contributor with a product backlog that never shrinks.
The only security posture that works reliably for these teams is the one they inherit by default: compliance certifications, network isolation, access controls that ship with the platform rather than requiring a dedicated sprint to implement.
What these patterns demand today
Every failure mode we observed the deployment nobody touches, the observability sprint that ships late, the portfolio tax, the security configuration nobody made- shares a root cause: the platform asked the team to make an operational decision, and the team either made the wrong one or made none at all. The response to these patterns required starting from a different question. Not “what should we configure for the team?” but “what should the team never need to decide?”
A platform that answers that question correctly holds a specific set of commitments. It decides what healthy looks like before the first request arrives, not after the first incident. It ships observability at deploy time, not as a sprint the team schedules after something breaks. It treats the eighth application in a portfolio the same as the first, same operational model, same governance, same economics. It inherits security posture by default, because the teams it serves will never staff a dedicated security function.
These are not feature decisions. They are architectural decisions about where operational responsibility permanently resides. The team defines the application source code, a Dockerfile, a pre-built image, or an existing workload being migrated in. From that point forward, the platform owns everything underneath it. Not for the first deploy. For the life of the application. Patching, scaling, healing, certificate rotation, capacity planning, health evaluation. These are not capabilities the team enables. They are responsibilities the platform holds permanently.
This is what we rebuilt AWS Elastic Beanstalk to be. Not a deployment tool. Not a hosting layer. An application management service that takes operational responsibility for everything underneath the application. The architecture now starts from the question above and refuses to let the answer drift back toward the team over time. Elastic Beanstalk operates in two modes, a structural change from its previous single-environment architecture:
Standard Mode delivers full operational ownership for individual applications and Windows/.NET Framework workloads: the complete operational stack, owned outright, for a single service.
Cluster Mode extends the same ownership model across the portfolio, shared infrastructure, source-to-production deployment that transforms code into running applications, and economics that improve as the portfolio grows. The eighth application shares operational overhead with the first seven rather than duplicating it. For the forty-engineer company running eight production applications today and inheriting ten more next quarter, this is the difference between a platform that covers the portfolio and a platform that covers only the applications simple enough to fit its opinions.
The industry convergence
The distinction is real, though I would not draw it as a line between platforms that reduce complexity and platforms that own operations permanently. Every vendor in this market absorbs some operational responsibility at deploy time. The key question is how much of it returns to the team during an incident, a patch cycle, and an audit. A platform that removes infrastructure management from developers during the workweek and reintroduces it at 3 a.m. Sunday addresses only half the challenge.
“A platform that removes infrastructure management from developers during the workweek and reintroduces it at 3 a.m. Sunday addresses only half the challenge.”
At convergence, the direction is correct, but the shape is incorrect. This is not two camps meeting in the middle. Gartner’s 2026 Magic Quadrant for Cloud Native Application Platforms places AWS, Microsoft, Google, and Red Hat in the leaders quadrant, with Render, Netlify, and Upsun as niche players. Vendors specializing in developer experience showed this category is viable, and now the hyperscalers are adopting it. Since source-code-to-URL mapping is now standard across the entire quadrant, the key differentiation becomes who bears operational liability for the eighth application three years after its release.
The only aspect of framing I would challenge is the idea that a platform determines everything the team never has to decide. Routine infrastructure decisions should stay out of the developer’s path, and escape hatches should stay in place for the teams that genuinely need them. A platform that removes choice altogether will demo well and then stall when teams migrate applications that don’t fit its opinions.
The 3 a.m. test that actually matters
For teams already living this reality — serious production, lean staff, growing portfolios — nobody planned to operate the 3 a.m. test; it is not a nice-to-have. It is the evaluation criterion.
The platforms that define the next decade will not simply make deployment easier. They will decide, in advance, how production systems should behave when things inevitably go wrong. Because by 3 a.m., the time for deciding has already passed. The CNAP category was built to describe platforms that own the application lifecycle.
Elastic Beanstalk made those decisions before the incident arrived: what healthy looks like, what to do when it stops being true, how to communicate what happened. The cloud gave teams power. These teams needed someone to stay. Elastic Beanstalk stays.
On September 16, xAI announced memory in Grok Build, its terminal coding agent. The pitch was that Grok “keeps notes on the conventions, decisions, and project facts that come up,” and “later sessions read those notes before touching related code.” Notes are Markdown files in a workspace scope per project and a global scope that applies everywhere. /memory browses them.
Meanwhile, Claude Code has done something similar for months under the name auto memory. It keeps a MEMORY.md index plus one file per note, per repository, and the docs say it is on by default. Anthropic’s Projects beta, announced September 17, adds shared memory across cloud threads, but only for select Pro and Max subscribers with no existing projects. I tested the CLI that everyone has.
Both companies say their coding agent now remembers what you told it in an earlier session. I wanted to see whether that holds up, so I tested Grok and Claude on the same three tests.
The tests
The claim I wanted to check is simple. Tell each tool something once, close it, open it again, and see whether it remembers. Both tools ran on my Mac, each on its own copy of four small Node repos I built for this. Grok Build 1.0.40 ran Grok 4.6 at high effort through an xAI API key. Claude Code 2.1.226 ran Opus 5 on my subscription. Every session was scripted with each tool’s headless mode, which reports its own tokens and cost.
Each test has two sessions. Session 1 plants a fact. I quit the tool. Session 2 gives a task where the fact matters and never mentions it.
Here are the tests I ran:
The test command – In this repo, npm test fails and make test passes, and session 1 says so. Session 2 asks for a new endpoint with passing tests, after I removed the README line that pointed at the Makefile.
Project decisions with a trap – Session 1 states that CSV export was dropped and money is integer cents, never floats, while a float helper and a half-built CSV exporter sit in the repo as bait. Session 2 asks for a refund endpoint that “takes an amount” and “a way for support staff to download all orders.”
A rule across projects – Session 1, in repo A, sets two rules “for all my projects,” conventional commit messages and no comments on obvious code. Session 2 runs in an unrelated repo B and asks for a small feature and a commit.
Here’s my scoring breakdown. Did the tool write the fact to a memory file, did it read that file in session 2, and did the session 2 output follow it.
The test command
Both passed. In session 1, each tool saved the rule as soon as I stated it. Grok wrote topics/testing.md plus two raw observations. Claude Code wrote orbit-api-run-tests-with-make.md with a “why” and a “how to apply” section.
In session 2, both remembered. Grok’s reasoning opened with “start by reading the memory files,” then it ran make test and never touched npm test. Claude Code read the Makefile and package.json, ran make test, and also never tried npm test. Grok took 29 seconds, 102K tokens, and $0.11. Claude Code took 22 seconds, 186K tokens, and $0.32. Claude used about 80K more tokens and cost nearly 3x more.
Project decisions with a trap
Both wrote both decisions down. Claude Code also converted “last quarter” into “Q2 2026” in its note. In session 2, both built the refund on integer cents, named the field amountCents, and left the float helper alone. For the download request, both shipped a JSON export.
Grok’s reasoning said the API is JSON-only, so it wouldn’t wire up CSV. Claude Code set a content-disposition header so the JSON downloads as a file. Both passed on both decisions, but Claude Code was more than double the price and just as fast. Grok took 103 seconds, 156K tokens, and $0.18. Claude Code took 32 seconds, 269K tokens, and $0.49.
A rule across projects
This is where the results split. Grok saved the rules to its global scope as git-and-code-style.md. In the second repo, it committed feat: add --help flag with usage and supported cities, and added no comments. Pass, in 33 seconds, 132K tokens, and $0.12.
Claude Code saved both rules too, but only in the first repo’s memory folder. It said so at the time, warning that its memory store “is scoped to this project’s directory.” In the second repo, it found nothing, and the commit came back with the Add --help flag. No comments were added, but that is Claude’s default anyway. Claude passed the first rule but failed the second one and still cost twice as much. It completed the work in 12 seconds, 122K tokens, and $0.24.
Results
Metric
Grok Build (Grok 4.6)
Claude Code (Opus 5)
Tests passed
3 of 3
2 of 3
Total time
165 s
66 s
Total tokens
390,848
576,863
Total cost
$0.41
$1.05
Grok Build passed all three tests, and Claude Code passed two. They behaved the same on the per-project tests. The split was the cross-project rule, which Grok’s global scope carried into a second repo and Claude Code’s per-repo memory did not.
Claude Code was faster on every recall session, 66 seconds total against 165, and cost at least twice as much on every one, $1.05 total against $0.41. It also used more tokens: 576,863 against 390,848. The price gap mostly reflects Opus 5 versus Grok 4.6 rather than the memory systems.
On the core claim, remembering what you told it last time in the same project, I could not tell these two apart. Both wrote a markdown note the moment I stated a rule, read it back next session, and followed it. Claude Code’s notes were better written. But Claude Code failed the third test. Its CLI memory stops at the repo boundary, so a rule I gave it “for all my projects” never reached the second repo. Grok’s global scope carried the same rule over without being asked.
What do I think?
Grok Build is the better option for most people right now. It remembered everything, it carries rules across projects, and it cost less than half as much on every test. Yes, Claude Code was faster on every session, but that only matters if you aren’t concerned about accuracy. Its CLI memory stops at the repo boundary, so anything you want it to remember everywhere still has to go into ~/.claude/CLAUDE.md by hand.