❌

Vue lecture

The rise of agentic AI on Kubernetes: unleashing the new infrastructure layer

Abstract 3D render of blue cubes inside gold wireframe boxes, linked by red rods into a dense cluster, with teal lines connecting outer cubes.

AI is changing expectations around infrastructure and operations, including Kubernetes management. When models run close to the data they use, deployment, scaling, and governance responsibilities tend to shift to platform teams. And as clusters, environments, and operational signals continue to multiply, manual operations often strain under the added weight.

AI may simultaneously provide opportunities to lighten this growing load. Agentic software can now observe a system, reason about it, and act within predefined limits. 

Ultimately, these platforms’ value depends on the quality of the context an agent can see and the boundaries you set. Without cluster state, policy, and access rules, an agent can only guess.

Without cluster state, policy, and access rules, an agent can only guess.

For agentic AI to streamline multi-cluster management, you need clear lines between what the system observes, what it recommends, and what it changes. Drawn well, those lines let teams gain notable speed while still maintaining control.

The impact of AI on computing infrastructure

Teams once treated AI as an application concern; models sat on top of existing systems, and the stack underneath stayed mostly unchanged. Today, AI reaches into more and more customer interactions, while data storage needs simultaneously expand and orchestration pressure grows. A recent Forrester report describes the modern AI computing stack as stretching from the models themselves into and across the infrastructure beneath them.

As AI workloads move into production, they place new demands on the infrastructure beneath them. Many lean on specialized compute, with resource needs that rise and fall through bursts of training and inference. Because conditions shift quickly, they can also call into question whether telemetry remains trustworthy. Each of these demands lands at the infrastructure layer, where the workloads run.

The infrastructure layer of the new AI stack

The infrastructure layer covers compute, storage, and networking. It is a foundation that every workload running on the layer depends on. As AI workloads grow, choices about capacity, placement, and control will increasingly shape the performance of the data, intelligence, orchestration, and experience layers atop the infrastructure.

To operate the infrastructure layer efficiently across many machines and locations, a team may rely on orchestration instead of managing servers by hand. In cloud native contexts, Kubernetes has become a control point for scheduling workloads, applying policy, and presenting a consistent interface across environments. Kubernetes is especially well-suited to support organizations this way when teams need consistent control across an estate spanning data centers, clouds, and edge sites. 

Agentic AI and Kubernetes: the future of the infrastructure layer

Agentic AI can extend automation from fixed rules to systems that adapt to real-time conditions. Traditional automation runs the same script whether the environment has changed, while an agentic system observes the environment, reasons about what it finds, and then takes action.

When you apply agentic capabilities to multi-cluster management, the system follows this same sequence. An agent reads cluster state and operational data, proposes a diagnosis or next step, and then carries out actions based on an approved scope, usually after a person signs off. You can further reinforce these boundaries by routing each request to a specialized agent that receives only the metadata it needs.

The signals that an agent receives from the cluster, the context about policy and access, and the definitions of what the agent may change are the key elements that give agentic systems their value. They also separate agentic AI on Kubernetes from a generic assistant. 

Manual Kubernetes management is less efficient at scale

Admittedly, agentic AI fits some settings better than others. On a small single-cluster footprint, the overhead may outweigh the benefit. Manual Kubernetes management often holds up on a handful of clusters, but it can become unreliable in a rapidly growing estate. After all, each new cluster adds lifecycle work across upgrades, patching, configuration, and renewal. Those tasks can quickly multiply and diverge in hybrid environments.

Configuration drift is a high risk in these situations. Settings that started identical can fall out of sync, and policies can apply unevenly from one team to the next. Individually, these gaps may be manageable, but collectively they raise the odds of an outage or a failed rollout.

Visibility can also erode in an unmanageable way. Clusters spread across data centers, clouds, and edge sites often leave teams with no single view of the whole landscape. When DevOps and platform engineers stitch together signals from separate tools, resolution can slow and become more error-prone. A unified view helps enable sound, efficient decision-making by people, agents, or both.

Kubernetes knowledge is fragmented, and existing AI tools lack business context

Kubernetes expertise often sits unevenly across an organization. For example, senior engineers may hold deep operational knowledge that application teams lack. The most current information about a running system may also be fragmented if logs sit in one tool and metrics in another. Real-time understanding can be further clouded when policies, runbooks, access rules, and deployment history each live elsewhere.

Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes. Without that context, even a capable AI tool may fall short of providing meaningful Kubernetes management support.

Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes.

When an agent can read current signals alongside the rules that govern them, its suggestions become specific, testable, and actionable. In an incident, agentic systems can correlate logs with a recent change. Ahead of a rollout, they can check the change against policy. During troubleshooting, they can account for access rules rather than guessing at them. Kubernetes decisions carry real operational consequences, which makes these details all the more important to consider. 

Engineering “toil” isn’t time-efficient

Site reliability teams use the word “toil” for repetitive manual work, especially tasks that keep systems running without adding lasting impact. In Kubernetes operations, toil takes the form of repeated triage, manual signal correlation, alert follow-up, and routine checks. The tasks aren’t particularly difficult, but they can consume significant time and attention for enterprise teams.

When engineers spend their days on this kind of investigation, proactive modernization efforts tend to stall and planned upgrades can slip behind schedule. In other words, the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements. In a recent survey about how AI provides value to DevOps teams, reducing toil emerged as one of the clearer opportunities.

…the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements.

Agentic AI can support repetitive investigations by gathering signals, correlating them, and proposing a likely cause for an engineer to weigh.

Kept under human review, it can take on some of the routine correlation that would otherwise fall to the team. That kind of support can give engineers more room to focus on the strategic work that most needs their judgment.

Building more intelligent infrastructure with agentic AI and Kubernetes

As you consider building toward intelligent infrastructure without surrendering control, the following principles can inform your efforts:

  • Start with observable context, giving agents access to current cluster state, policy, and history before they reason about a problem.
  • Separate suggestions from actions, allowing agents to recommend freely while any change must wait for human approval and a defined scope.
  • Connect agents to existing controls, routing their work through the access rules, identity, and audit paths the team already trusts.
  • Keep the ecosystem open, favoring platforms that integrate with current tools and standards over those that lock work into a single stack.

Platforms like SUSE Rancher Prime and SUSE AI Factory embrace these principles and illustrate how Kubernetes management can become a foundation for agentic operations. These platforms can help you improve cluster and policy consistency without compromising your authority over AI. Built on open-source foundations, they can also help you avoid being trapped in a single vendor’s stack.

In SUSE Rancher Prime, the industry’s first context-aware agentic AI ecosystem, its AI assistants work as a crew of specialized agents with an intelligent router. The platform draws on the cluster context already in place and acts through existing access controls. Through support for external Model Context Protocol (MCP) servers, teams can extend that crew to their own sources. In addition, human validation tools allow you to hold a proposed action for approval before the agent runs it.

Despite its potential, intelligent infrastructure is not universally beneficial. In situations where change control must stay fully manual, for example, agentic AI’s role may be strictly limited to observation and suggestion. Measure the technology’s value against the realities of your day-to-day operations. For those who are investing, agentic AI will have the greatest impact when it actively supports context, control, openness, and human judgment.

The post The rise of agentic AI on Kubernetes: unleashing the new infrastructure layer appeared first on The New Stack.

  •  

Avoiding vendor lock-in through an open-source approach: a developer’s perspective

Abstract dark digital artwork depicting dense undulating layers, symbolizing cloud architecture and software ecosystem tension.

Every infrastructure team makes decisions that are difficult to reverse. Most of the time, that works out. Sometimes it does not.

Vendor lock-in usually begins as a reasonable choice, made under time or budget pressure, that solves a real problem at the time. A managed service ships faster or a deployment model fits better in that moment, but eventually a difficult constraint appears. 

When business conditions inevitably change, those accumulated choices and their consequences will determine whether a team can pivot accordingly. Limits on flexibility rarely trace back to a single vendor; more often, they hinge on how reversible the team’s past decisions are.

What is vendor lock-in and how can it harm your business?

The risks of vendor lock-in are not really about relying on vendors, since every production system relies on vendors. The big issue is dependencies that become too expensive or impractical to unwind.

For a platform team, that dependency builds up across APIs, contracts, roadmaps, and data models. It extends further into managed services, identity patterns, observability pipelines, and operational tooling. Each piece likely represents a reasonable design choice, but together they can quietly limit your options and raise the cost of leaving. When switching a database or control plane means rewriting tons of integrations, retraining the whole staff, or migrating data under inconvenient timelines, you have lost the room to maneuver.

The big impacts of small, invisible and unexamined decisions

Not every dependency is automatically a problem; some are understood, contained, and worth the tradeoff. The real risk lives in the dependencies no one examined closely, which may stay invisible until they block the business from evolving. 

“The real risk lives in the dependencies no one examined closely, which may stay invisible until they block the business from evolving.”

Unfortunately, some teams are familiar with these invisible dependencies. A managed database might pick up proprietary extensions, which application code then starts to assume. A Kubernetes environment might bind to one cloud’s IAM, networking, storage, and load balancer model. Observability and logging pipelines might harden around a single provider’s formats. None of these choices is reckless on its own, but together they can create significant friction. 

Obstacles to change and their hidden costs

The extent of a dependency-based tradeoff can sometimes remain unknown until circumstances shift, such as a new compliance requirement or customers needing a new deployment model. The hidden costs of these moments often escalate in stages. It might start with a visible, unwelcome migration bill, but the expense can also show up as operational drag. Rushed migrations can lead to additional service disruptions later. A workload may be unable to move, limiting services to certain customers. When you are tied to a specific vendor’s release cadence, it can make it difficult or even impossible to adopt emerging technology. 

Concentration risk compounds the problem, because a single change from one provider that carries pricing, support quality, and roadmap can ripple across the estate. By the time a switch becomes necessary, the cost shows up as service disruption, complex data transfer, and retraining. Naming these costs early keeps them from arriving as surprises.

At some point, a dependency can accumulate enough of these costs to become more than an architectural detail. Once it affects budgets and timelines, leadership has to account for it—and the team has to be ready to explain it. Identifying these dependencies early gives everyone time to plan.

Open source offers a different path

One way to proactively address this pattern is to evaluate potential dependencies more deliberately. For example, before committing to a platform or service, try to determine its reversibility. In other words, establish how difficult it would be for the team to change its mind about the investment in the future.

“Open source offers no guarantee against lock-in, however, since a team can still build tight coupling on open foundations.”

Open source solutions tend to perform well against that test, because they are intentionally built to keep systems inspectable, portable, supportable, and replaceable. By design, open source makes it easier for you to preserve options over time. It offers no guarantee against lock-in, however, since a team can still build tight coupling on open foundations.

What is open source?

Open source describes software you can inspect, run, modify, extend, support, and replace with relative ease compared to proprietary alternatives. The software’s source is available, and the license grants you the right to use and change it. Notably, no-cost or freeware software is not necessarily open source, specifically if it does not provide this level of access and rights.

Several companies have open source principles at their core, and open source software can be extremely valuable in enterprise contexts. Transparent code is often easier to audit, and open standards can reduce friction when moving between tools.

Open source also changes who can move the goalposts

For developers, reversibility is not only about APIs and data formats. It is also about whether one company can change the terms underneath a foundational technology. The Linux kernel is a useful example. Linux kernel documentation notes that copyright assignments are not required, so merged code retains its original ownership and the kernel now has thousands of owners. That makes unilateral relicensing of the kernel effectively impractical.

Kubernetes has a different legal structure, but the practical protection is similar. The project is licensed under Apache 2.0 and governed by the Cloud Native Computing Foundation. The license grants users durable rights to the existing code, so no single vendor, including SUSE, can retroactively take those open-source rights away from the project as it already exists. That matters because a platform can remain available even if a particular vendor changes strategy.

The Terraform-to-OpenTofu fork shows why this is more than a theoretical distinction. In 2023, HashiCorp changed Terraform’s license from the Mozilla Public License 2.0 to the Business Source License 1.1. The community responded by forking the last open-source codebase into OpenTofu, now a Linux Foundation project that remains under the MPL 2.0. The lesson for developers is not that every open-source project is immune to licensing changes. It is that open licensing and neutral governance can preserve a viable exit path when a vendor changes direction.

Open source powered by enterprise discipline

Open source ultimately earns its place through engineering discipline. Source availability has benefits but does not resolve governance, patching, lifecycle management, documentation, security, or integration on its own. A community project can be powerful and nonetheless arrive without enterprise-grade operational guarantees.

Enterprise open source providers exist and can help with closing that gap. They embrace open foundations and add the support, security, maintenance, and lifecycle discipline that production environments require. Founded in 1992, SUSE was the first provider of an enterprise Linux distribution. Today, it focuses on helping organizations operationalize open source with enterprise-grade support.

These companies aim not to close off open source software but to make it dependable at scale. In other words, open source and operational rigor can coexist. And enterprises should expect both from any external provider.

Digital sovereignty: the x-factor that makes open source even more critical

Digital sovereignty describes how much control an organization has over its infrastructure, data, operations, and technology choices. Sovereignty is a spectrum, and architecture decisions can move an organization a step in either direction.

Recent research by SUSE suggests that almost all enterprises are prioritizing digital sovereignty, but only 52% are actively taking steps toward it. That gap is largely an execution problem, and much of it surfaces in everyday platform decisions. 

If your team supports regulated industries or deploys in on-premises or air-gapped environments, you may be especially familiar with growing pressures around sovereignty.

Sovereignty puts a deadline on work that was already worth doing

Developers can hear “digital sovereignty” and assume it means a separate compliance workstream with a separate engineering bill. In practice, much of the work is the same discipline platform teams already invest in: portable workloads, clean interfaces, automated verification, reproducible deployment, auditable behavior, and the ability to replace a dependency without rewriting the system around it.

“Sovereignty does not suddenly make that engineering work valuable. It puts a deadline on work that was already worth doing.”

Those practices already have an economic case. They reduce migration costs, lower operational risk, make platform changes less disruptive, and preserve options when pricing, regulations, or business requirements shift. Sovereignty does not suddenly make that engineering work valuable. It puts a deadline on work that was already worth doing.

That reframe matters because it turns sovereignty from a policy overlay into an architecture property. The useful question is not simply, “How much extra work will sovereignty cost?” It is, “Which parts of our stack already fail the portability, interface, and verification tests we would want anyway?”

How to strengthen sovereignty with open source

Sovereignty depends on how a team designs, deploys, and operates its systems. Open source does not make an organization sovereign by default, but it can improve the conditions for sovereignty. 

In fact, many of the same questions that expose lock-in also matter for digital sovereignty. Each of the following questions about reversibility connects to open source and sovereignty alike:

Reversibility questionWhy open source can helpHow sovereignty strengthens
Can we run this workload elsewhere?Open source typically runs across on-premises, cloud, hybrid, and edge environments, not just one vendor’s platform.More control over where workloads run, including specific regions and regulated contexts.
Can we understand and audit how it works?Source availability and community scrutiny improve inspectability over closed alternatives.Teams can verify behavior, assess risk, and meet assurance requirements.
Can we migrate or reuse our data?Open ecosystems favor open formats and interoperable tooling.Data stays more portable, improving control over storage and movement.
Can another team or partner support it?Multiple support paths exist, from internal teams to integrators and enterprise vendors.Less dependence on one vendor’s pricing, availability, or roadmap.
Can we replace one component without rewriting everything?Open interfaces and modular design make components easier to swap.More control over architecture as requirements change.
Can we keep operating if a vendor changes direction?Open source projects can outlast one vendor’s strategy or license.Less exposure to decisions the team cannot control.
Can we deploy closer to the data?Open source can run in private data centers, sovereign clouds, edge sites and hybrid models.Sensitive workloads, including AI, can be governed nearer the data.

The ongoing work of digital sovereignty

Sovereignty is more of a practice rather than a specific destination. For many teams, the work begins with identifying existing dependencies that are especially hard to reverse. Similarly, you’ll need to separate the tradeoffs worth accepting from the ones that remove a significant number of options. 

Moving forward, it can be helpful to prioritize open interfaces and portable foundations when possible. When evaluating new services or solutions, treat lifecycles, support, and governance as first-order concerns.

In some cases, sovereignty work can be too heavy for an in-house team to carry alone. Providers such as SUSE can help strengthen your operational layer, including security and observability, and especially in growing or hybrid contexts.

Automated checks can make those principles concrete by continuously testing whether workloads can be rebuilt, moved, audited, and recovered instead of waiting for a migration or compliance event to expose the gaps.

Open source lets you take control of your software ecosystem

No enterprise team avoids every dependency, and candidly none should try. Some coupling is reasonable, contained, and worth it. A vendor-free system is not a realistic goal for a major enterprise. A realistic goal is the judgment to separate acceptable dependencies from dangerous ones.

“The true cost of any platform includes the cost of leaving it, and teams should understand that cost before they commit.”

Reversibility gives that judgment something concrete to work with, because it can be broken down into capabilities a team can name, evaluate, and test:

  • Ownership. Ownership does not mean building everything yourself. It means holding the realistic ability to run, move, or hand over each layer of your stack. The test is simple: if a vendor disappeared tomorrow, or was ordered to stop serving you, what still runs next month?
  • Auditability. You should be able to verify what your software does, yourself or through an auditor you appoint, rather than accepting a vendor’s report as the final word. With open source, inspection is a property you hold. With closed software, it is a permission you are granted, and permissions can be withdrawn.
  • Exit velocity. An exit plan without speed is just a document. Exit velocity measures how fast a workload can move from one platform to another, and it only means something when you test it on a schedule, as earlier generations tested disaster recovery.
  • Pivot ability. These capabilities matter when conditions change: a new compliance requirement, a customer that needs a different deployment model, or a vendor that changes direction. Teams that can reroute workloads respond on their own timeline. Teams that cannot must renegotiate from a position of weakness.

Vendor lock-in becomes a manageable risk when you can confidently flag which decisions are hard to undo, weigh the tradeoffs honestly, and protect the team’s pathways to change. Open source strengthens every one of these capabilities because it keeps larger portions of your system inspectable, portable, and replaceable.

The true cost of any platform includes the cost of leaving it, and teams should understand that cost before they commit.

The post Avoiding vendor lock-in through an open-source approach: a developer’s perspective appeared first on The New Stack.

  •  

OpenTelemetry and Prometheus are getting along. What’s still missing?

Abstract 3D illustration of metallic blue spoked hubs connected by purple tubes against a pink background.

Welcome to another edition of Road to KubeCon, where we’re tracking the major movements in the Kubernetes and cloud native ecosystem on the path to KubeCon + CloudNativeCon NA 2026, happening Nov. 9–12 in Salt Lake City, Utah.

This week, we look at how cloud-native teams are putting observability to work. There’s progress on OpenTelemetry and Prometheus interoperability, a migration spanning 100,000 hosts, and new data on the costs and benefits of monitoring AI systems. Plus, HPE’s latest Gartner recognition, agent governance updates, and a father-and-son story from KubeCon India.

HPE GreenLake named a Leader in Gartner quadrant

On Wednesday, HPE announced it had been named a Leader in Gartner’s Magic Quadrant for Infrastructure Platform Consumption Services for the second consecutive year.

Hewlett Packard Enterprise (HPE) is a presenting sponsor of Road to KubeCon. HPE Software helps IT organizations modernize infrastructure, streamline operations, and accelerate AI initiatives across hybrid, multi-vendor environments.

GreenLake, HPE’s cloud operations platform, helps teams monitor resource consumption, secure data, and manage infrastructure across data centers and private and public clouds. HPE points to recent additions, including agentic AI-powered operations, as part of the platform’s development.

Varma Kunaparaju, senior vice president and general manager of cloudops software and platform at HPE, says in the announcement: “We are building the operating model and platform for the agentic enterprise, giving customers the ability to simplify operations, govern intelligently, continuously optimize, and modernize without sacrificing choice.”

OpenTelemetry and Prometheus work better together

OpenTelemetry (OTel) and Prometheus are widely used for cloud-native monitoring and observability, often side by side. A new survey looks at how well that combination works.

Published Tuesday, the 2026 survey on Prometheus and OpenTelemetry interoperability found that nearly half of respondents mix Prometheus- and OTel-style instrumentation for infrastructure metrics. For application metrics, 30.7% use both.

While the two ecosystems haven’t always worked well together, the 2026 survey shows improvement: the average ease-of-use rating rose 0.5 points, from 3.1 to 3.6, while the share of those who find the two hard to use together fell from 29% to 10%.

As OTel contributors Dhruv Ahuja of SigNoz, Grafana Labs‘ Andrej Kiripolsky and Arthur Sens, and Ana Muenz share: “Two years of work on interoperability is paying off.” 

There’s still work to do. Respondents want better alignment between the projects’ data models, better handling of resource attributes and metadata, and fewer naming and formatting issues.

Atlassian moves metrics from 100,000 hosts to OpenTelemetry

A case study on the Cloud Native Computing Foundation (CNCF) blog details how Atlassian migrated its metrics collection to OpenTelemetry from gostatsd, its open-source Go implementation of Etsy’s StatsD.

The original pipeline had worked for years, handling metrics from roughly 100,000 hosts across 14 regions, but was increasingly out of sync with the shift to OTel. “It became the thing everyone standardized on, and more and more of what fed our pipeline was emitting OTel data we simply didn’t support,” write Atlassian’s Iris Grace Endozo, Farzad Vazirnia and Albert Kerr.

To maintain continuity throughout the migration to OTel, Atlassian swapped the collection and pipeline mechanics underneath while keeping the service-facing interface unchanged. This turned an organization-wide overhaul into what the authors call a “platform-team migration.”

According to the authors, aggregation now uses about half the CPU for the same traffic. Operations are more unified through the OTel Collector, CPU usage is more evenly distributed across ingest shards, and sidecar costs are down roughly 30% at fleet scale.

New Relic finds observability gains — and gaps

On Tuesday, New Relic released its 2026 Observability Forecast, based on a survey of 2,575 IT and engineering leaders and practitioners. The report found that 73% are standardized on OTel, actively migrating to it, or testing it.

The report also looks at observability’s role in AI adoption. It found that 83% of respondents consider observability essential for AI-generated code. Organizations monitoring AI agents are twice as likely to report a threefold return on observability investment as those running agents without monitoring.

The study also paints a picture of the impact of outages. Engineers now report spending 37% of their time addressing disruptions, while 42% of organizations learn about disruptions through inefficient channels, like manual checks or customer complaints.

Outages take a business toll. New Relic found that organizations lose $74 million a year on average due to high-impact outages. That’s $1.85 million per hour, or over $30,000 for every minute a system is down.

The findings show why teams are looking for ways to detect and resolve problems faster as their systems grow more complex.

As Kubernetes evolves, so do the demands on the teams running it. Presenting sponsor HPE helps teams address that complexity with software spanning virtualization, cloud management, observability, and automation.

Observability Day returns to KubeCon

If you’re into observability and attending KubeCon NA, definitely check out the agenda for Observability Day, happening during the co-located events in Salt Lake on November 9.  

OpenTelemetry’s graduation in May and growing production use give teams more experience to draw on as they adopt the standard.

New AI workloads, inference monitoring, and interoperability with other projects still present challenges. Those issues give practitioners plenty to compare notes on.

According to the Observability Day schedule, the agenda includes project updates and lessons from Capital One, Cisco, Nubank, and other organizations.

“Observability Day provides a vendor-neutral place for maintainers and practitioners to compare approaches and learn how the wider ecosystem is responding,” write Austin Parker, Iris Dyrmishi, Eduardo Silva Pereira, and Juraci Paixão Kröhling on the CNCF blog.

Komodor adds controls for agentic operations

Technically one we skipped last week, but potentially interesting vendor news nonetheless: Komodor, the site reliability engineering platform, announced its Komodor Agentic Operations Platform on Wednesday, September 16.

The additions let engineers deploy autonomous workflows and build or import agents under shared governance and context. Komodor says the release responds to the growing use of agents, including coding agents, and concerns about governance, return on investment, and costs.

“The hardest part of running agentic operations in production is not building the agents,” shares Itiel Shwartz, Komodor’s co-founder and CTO. Instead, the challenges lie in maintaining context, persistent memory, accuracy, security, and cost control — things the new Komodor release aims to address.

Spectro Cloud expands in the Middle East

The Middle East is an increasingly important technology market and a hotbed of data center construction. At the same time, data sovereignty and compliance requirements are driving interest in sovereign infrastructure.

This week, Spectro Cloud announced plans to expand in the Middle East, including new local partners and a dedicated regional office.

“The Middle East has bold ambitions for global AI leadership, from sovereign AI factories to AI-powered economies,” shared Tamer Riyal, Spectro Cloud’s sales director for the Middle East, in the announcement. “We’re investing in the regional expertise and partnerships to support that vision for the long term…”

The company aims to expand adoption of PaletteAI, its platform for managing AI infrastructure across sovereign clouds, enterprise data centers, and edge locations. The expansion reflects the region’s growing role in cloud-native infrastructure.

A father and son take the KubeCon stage

This week’s updates also have a personal side. Earlier this year, analyst, advisor, and TNS columnist Janakiram MSV co-presented a talk at KubeCon + CloudNativeCon India 2026 with his son, Shreyas Mocherla, a CNCF Kubestronaut and software engineer at Nirmata.

As Mocherla describes on the CNCF blog in a post published this week: “Presenting alongside him made this special on a level that goes beyond the conference itself. I grew up watching him speak at technology events. Standing next to him at the same podium, in front of the KubeCon audience, felt like a full-circle moment.”

You can watch the talk, “Run Your Own AI Cluster on a DGX Spark: Kubernetes, GPUs, and DRA,” below:

It’s a reminder that the Kubernetes community’s connections can span generations as well as organizations.

Other updates from the K8s universe

More updates for the platform engineers and cloud operators working in the Kubernetes ecosystem:

Follow the Road to KubeCon

Road to KubeCon is an eight-part series presented by HPE, which will be at KubeCon + CloudNativeCon North America in Salt Lake City. Before you go, explore how HPE Software helps IT teams do more with less complexity.

We’ll be here every Friday until KubeCon.

If you’d like to participate, Bill Doerrfeld, the writer of this series, is open to pitches — you can send release notes, quotes, reports, videos, case studies, or hot takes through his contact page.

If you missed the previous editions covering Kubernetes v1.37 and AI inference, you can catch up through those links. Visit the Road to KubeCon page for the complete archive.

The post OpenTelemetry and Prometheus are getting along. What’s still missing? appeared first on The New Stack.

  •  

Developers and platform teams both want Kubernetes self-service. They disagree on who owns it.

Abstract view up through bold yellow angled beams to a white skylight grid and curved ceiling panels.

What do developers want? Kubernetes environments when they need them.

What do they not want? Those environments after a week or more of tickets. 

Platform teams, meanwhile, own what those environments cost, who can access them, and whether they meet company policy.

That tension is the core of Kubernetes self-service: What can safely be handed to developers, and what still belongs to the platform team?

In a recent interview with enterprise cloud specialists — Marius Bogoevici, Senior Principal Product Manager at Hewlett Packard Enterprise (HPE), and Karthik Subramanian, Principal Product Manager for HPE Morpheus Software — The New Stack explored the core friction points of Kubernetes self-service. The conversation focused less on whether self-service is desirable than on where to draw the line.

HKS, HPE’s CNCF-certified Kubernetes distribution, is integrated with HPE Morpheus Software to help platform teams deliver and lifecycle-manage Kubernetes environments as part of a broader operating model spanning Kubernetes, VMs, infrastructure, and clouds. HPE Morpheus Advanced Software supports the on-premises private-cloud use case with HKS, while HPE Morpheus Enterprise Software extends Kubernetes and application operations across hybrid and public-cloud environments.

Together, HKS and HPE Morpheus Software extend that operating model beyond infrastructure provisioning. Through service and application catalogs, platform teams can connect approved Kubernetes environments with the CI/CD pipelines, container registries, automation tools, and other services developers already use. Developers receive a governed, ready-to-use path from code to deployment instead of manually assembling the toolchain for each project.

The self-service paradox

Open-source Kubernetes provides orchestration and declarative APIs, but not a complete operating model.

Subramanian says teams building their own self-service layer usually run into two recurring problems:

  1. Tool and package sprawl: To make upstream Kubernetes production-ready, platform teams must curate and maintain an ever-evolving ecosystem of third-party CNCF tooling for networking (CNI), storage (CSI), ingress, identity, and policy enforcement. Navigating and supporting this fragmented stack creates immense maintenance overhead for internal platform teams.
  2. Day-2 lifecycle and hybrid footprint complexity: Spinning up a Kubernetes cluster is the easy part, but keeping it current — across development, QA, staging, and production — is where the work piles up. That is why HPE says every Kubernetes upgrade must be checked against the networking, storage, ingress, identity, and policy components around it. The problem gets harder when clusters span bare metal, private clouds, edge sites, and public clouds, because one-off scripts and environment-specific configurations can quickly create drift. That maintenance burden belongs with the platform team, not with developers trying to ship applications.

Giving developers direct access to raw Kubernetes APIs just shifts the operational work — it’s far from gone for good. In fact, developers will wind up debugging manifests and storage drivers instead of writing code. 

Meanwhile, operations teams have to deal with overprovisioning, idle clusters, and configurations that reach production without review.

What developers control — and what the platform supplies

The practical answer is not unrestricted access. It is a paved path: approved Kubernetes services that developers can request themselves, with access, configuration, placement, approvals, and lifecycle controls defined by the platform team.

“The best candidates for self-service are requests that are repeatable, low-risk, and well-understood,” Bogoevici tells The New Stack. “For example, a developer should be able to request a development cluster, deploy an approved application, create a namespace, or select resources from pre-approved configurations without opening a ticket. The platform team decides what a safe configuration looks like, and the developer chooses from a supporting menu.”

“The platform team decides what a safe configuration looks like, and the developer chooses from a supporting menu.”

Rather than asking developers to write YAML for ingress, storage classes, and RBAC, HPE Morpheus exposes those choices through service catalogs, reusable layouts and blueprints, workflows, role-based access control, approvals, APIs, and automation. Developers do not lose Kubernetes. They retain direct access through standard Kubernetes interfaces and tools where permitted, while the platform team standardizes the request, governance, and lifecycle processes around them.

Those catalog items can package more than infrastructure settings. They can also integrate the approved services and application components that support the development workflow – including CI/CD tooling, source and artifact repositories, container registries, and runtime dependencies – while the platform team controls how those components are configured and governed.

Developers choose the parameters that matter to the application:

  • Approved Kubernetes versions and cluster sizes: Select from pre-tested Kubernetes runtime releases and node count templates.
  • Resource quotas: Specify required CPU, RAM, and persistent storage capacity tailored to the workload.
  • Integrated toolsets and IDE environments: Select required developer toolchains, container registries, and runtime dependencies. 
  • Lease and duration limits: Define explicit operational lifetimes for temporary development or sandbox clusters to prevent abandoned infrastructure sprawl.

Network isolation, identity-provider integration, security policy, and cost allocation stay with the platform team and are applied automatically through the approved service configuration.

The same division of responsibility applies to the delivery toolchain: Developers choose from approved services, while the platform team manages the integrations, credentials, policies, and automation behind them. This gives developers a consistent experience without shifting toolchain maintenance and governance onto individual application teams.

The division of responsibility looks like this:

Service areaDeveloper chooses or requestsPlatform team defines and suppliesReview or exception path
Development cluster provisioningApproved Kubernetes service, version, size, target environment, and duration.Reusable layout or blueprint, access controls, placement rules, storage and network defaults, and lifecycle policy.Nonstandard versions, placements, configurations, or requests outside quota.
Production deploymentApplication artifacts, target namespace, and deployment request through the approved path.RBAC, tenancy, policy, audit, backup, and release controls appropriate to the environment.Formal review for production changes and exceptions.
Resource allocation and quotasCPU, memory, storage, and other approved capacity parameters within project limits.Project quotas, upper bounds, placement constraints, and supported resource profiles.Requests above quota or for specialized resources.
Networking and securityApplication endpoints and permitted connectivity within approved patterns.Identity integration, RBAC, tenant isolation, network policy, secrets, and audit controls.Cross-tenant access, elevated privileges, or changes to baseline security policy.
Lifecycle and cost governanceService lifetime and approved operational actions.Visibility, policy, approvals, retirement workflows, and applicable cost controls for the licensed variant.Long-running exceptions, nonstandard lifecycle actions, or budget exceptions.

From ticket queues to a repeatable paved path

In conventional IT environments, provisioning a dedicated Kubernetes environment for a new project often involves cross-departmental ticket handoffs spanning infrastructure, networking, security, and storage teams. This friction frequently stretches provisioning timelines from days to weeks.

By unifying infrastructure orchestration, role-based access controls, and multi-tenancy into a single operational experience, HPE Morpheus Software can compress these provisioning workflows down to minutes or hours, according to HPE. “Developers get a usable environment that complies with the organization’s defined controls and policies, without needing to understand all the complex infrastructure steps sitting underneath,” Bogoevici says. “When you reduce provisioning time from weeks to hours, that is super meaningful and tangible.”

“When you reduce provisioning time from weeks to hours, that is super meaningful and tangible.”

The result is not only faster cluster provisioning. HPE Morpheus Software can also automate the handoff into the developer’s established delivery process by making approved CI/CD and application services available with the environment. Instead of waiting for separate teams to connect pipelines, registries, credentials, and runtime dependencies, developers receive a ready-to-use path from development through deployment.

Faster provisioning can create a different problem, too: The speed can and will cause teams to lose track of what was provisioned and why. HPE Morpheus Software gives administrators visibility into utilization and cost, while lease controls can shut down temporary development clusters when their time expires.

Bogoevici says ticket volume is a poor measure of success, particularly early on, when more developers may be trying the catalog. He recommends watching deployment success, exception rates, resource utilization, and the day-to-day effort required to keep the service running.

Security belongs in the service design

Security is another boundary that must be designed into the self-service path. If identity, access, tenancy, and policy are added only after a cluster is created, every request produces more work and more room for inconsistency.

“Security must be a core design consideration built directly into the service, not an afterthought during deployment,” Bogoevici says. “HPE Morpheus Software brings identity integration, role-based access, tenant isolation, approvals, and policy into the operational workflow.”

A newly provisioned environment should arrive through an approved configuration with the applicable identity, RBAC, tenant, policy, and audit controls attached. Platform teams can validate the paved path by testing an allowed request, a request that should be rejected, and the resulting audit record.

The operating-model test

The strongest Kubernetes self-service model does not hide Kubernetes or make it the control plane for every workload. It gives developers useful, approved choices and direct access to the Kubernetes workflows they need, while the platform team standardizes the enterprise processes around those workflows.

That matters because the enterprise still runs VMs, clouds, and existing infrastructure alongside Kubernetes. HPE Morpheus Software helps platform teams use common request, governance, automation, and lifecycle processes across these environments without forcing every workload onto one runtime or creating another operational silo.

In practice, that means self-service should deliver more than a Kubernetes cluster. With HPE Morpheus Software, a catalog request can bring together the approved environment, application services, and DevOps toolchain integrations developers need, while preserving the governance and lifecycle controls the platform team requires. Developers spend less time assembling and troubleshooting delivery infrastructure – and more time building and releasing applications.

Looking toward 2027, the goal is not unrestricted control. It is faster access, predictable results, transparent guardrails, and a clear exception path when the standard service does not fit.

The post Developers and platform teams both want Kubernetes self-service. They disagree on who owns it. appeared first on The New Stack.

  •  

Vulnerability alert fatigue nearly swamped WHOOP. But its fix still keeps a human in charge.

Abstract image of thin vertical ribs against a bright orange background. In the center, a blurred rectangular glow shifts from green on the left, through dark red, to pink on the right, as if seen through ribbed glass.

With often hundreds of thousands of alerts a day, many tech organizations are buried in vulnerabilities and worn down by alert fatigue. The rise of AI has only made it harder to cut through the noise and to find actionable alerts. Manual security and site reliability engineering is not an option. 

The engineering team behind WHOOP‘s health and fitness tracker felt this pain, relying on multi-day, all-hands triage sessions to stay on top of the alert deluge. But, as a high-growth consumer health company handling sensitive user data, it couldn’t afford to miss anything. Which is why the team at WHOOP built an automated vulnerability-response workflow based on the company’s specific technical, operational, and trust considerations.

Join The New Stack on Wednesday, October 7 to learn from WHOOP staff engineer Vinay Raghu and Datadog senior product engineer Amber Tunnell how WHOOP built and implemented this workflow using Datadog Bits AI and Workflow Automation for faster, at-scale response. 

Join us on October 7, 2026, for a live Datadog x TNS event

REGISTER NOW FOR THIS WEBINAR
By registering, you consent to The New Stack’s Privacy Policy, Terms of Use and to receiving email communication from The New Stack and our event partner. You may opt out at any time.

DevSecOps, security, and cloud/platform engineers should bring their questions to this live demo-slash-case study to learn how to reduce friction between developer velocity and security requirements without increasing headcount. 

What you’ll take away from our live webinar

Raghu and Tunnell engineers will share how they were able to:

  1. Focus on real exposure vs. scanner noise. Not everything is critical. You’ll learn how WHOOP used Datadog’s Software Composition Analysis (SCA) to analyze runtime code execution and prioritize active threats.
  2. Route the right vulnerability to the right engineer. WHOOP automated vulnerability mapping to microservice owners, so developers received tickets with full context attached.
  3. Build automated guardrails for devs to self-resolve. This let security engineers pivot from frustrating gatekeeping and ticket-pushing to more proactive, systemic work that adds value. 
  4. Maintain a human in the loop. With such sensitive data and a demand to be always-on, WHOOP isn’t ready to automate the engineer out. Learn how they decided their team’s response had to change.

And then, of course, we will end the live discussion with how to measure it all. Don’t miss out and register to attend on October 7.

The post Vulnerability alert fatigue nearly swamped WHOOP. But its fix still keeps a human in charge. appeared first on The New Stack.

  •  

The cloud reduced operational complexity. But many teams need someone to own it completely.

Dark abstract digital wave with subtle color distortion, representing application lifecycle and cloud native infrastructure management.

For companies running serious production workloads with lean engineering teams, the real test of operational ownership comes at 3 a.m. Not whether the system stays up but whether anyone needs to be awake to make that happen. The page arrives. A service is degrading. The engineer who answers didn’t build this service, doesn’t know what thresholds were set at deploy time, and cannot tell whether the system is healing itself or waiting for a human decision. 

This is the moment that separates services that transferred operational burdens from platforms that merely deferred them. The system that passes the 3 a.m. test already knows what healthy looks like, what to do when healthy stops being true, and how to communicate what happened—because those decisions were made at deploy time, not incident time. The team sleeps through the night because nothing went wrong, but because the response was already determined.

Forty engineers, eight applications, zero dedicated ops

Cloud infrastructure evolved in two directions simultaneously, and neither arrived where mid-market teams actually stand. On one end: full control. Infrastructure-as-code, service meshes, custom pipelines. Powerful, flexible, and designed for organizations that have deliberately invested in operational staff who can absorb the cost of that flexibility. On the other end: single-app simplicity. Push code, get a URL. Elegant for a first deployment, but architecturally limited the moment a team manages more than one service, needs compliance controls, or inherits an application that doesn’t fit the platform’s opinions.

You know this company. Forty engineers. Eight production applications. Two generate 80% of revenue. One SRE who is actually a senior developer with an on-call rotation nobody else wants. A compliance audit due in Q3 that nobody has started preparing for.

“The right investment is shipping product. The cost of this gap is measured in what these teams do not ship.”

Every engineer is a full-stack contributor. The person who wrote the feature deploys it, monitors it, and gets the page when it breaks not because the team lacks sophistication, but because hiring dedicated infrastructure staff is not the right investment at their stage. The right investment is shipping product. The cost of this gap is measured in what these teams do not ship. Every sprint spent upgrading the deployment pipeline is a sprint without the feature a customer asked for. Every 3 a.m. page answered by a developer with a product standup at 9 a.m. is diminished output that never shows up in any dashboard. The gap is widening.

The applications nobody planned to operate

Not every application in a portfolio was built by the team now responsible for it. Many companies acquire products through M&A. They inherit internal tools built by engineers who left three years ago. They run commercial off-the-shelf applications customized beyond vendor support. They maintain line-of-business applications in languages nobody on the current team chose. These applications share a common trait: they are in production, they serve customers or meet compliance requirements, and nobody has the budget or the mandate to rewrite them. They need a home that accepts them as they are, not as a modernization roadmap says they should become.

“They need a home that accepts them as they are, not as a modernization roadmap says they should become.”

This is where the full lifecycle vision matters. An application management service that only serves net-new applications forces teams to maintain two operational models: one for the applications they are building today, and another for the applications they inherited yesterday. That split is where drift starts, where patching falls behind, and where audit findings accumulate.

The application management service that solves this problem must accept the full range: the Java application packaged as a WAR file, the .NET Framework service running on Windows, the Python application with dependencies pinned to a specific runtime, the containerized service already running elsewhere, and apply the same operational model, the same deployment interface, the same patching and scaling behavior to all of them. Migrate, manage, and modernize within the same experience, without requiring a different operational posture for each stage of the application lifecycle.

What the market data actually shows

Janakiram MSV, an analyst and advisor on cloud-native platforms and TNS contributor, puts it this way:

Cloud native standardized the infrastructure layer around containers and Kubernetes. But it never standardized the operational boundary between the application team and the infrastructure underneath it, and platform engineering largely emerged to re-establish that boundary. CNCF’s latest research with SlashData reports that 28 percent of organizations run a dedicated platform engineering team and 41% split those capabilities across multiple teams. Another 3% have no formal approach at all, and that is where most mid-market engineering organizations exist. They want the same outcome as platform engineering without first becoming a platform engineering organization.

The pattern that repeats is that maturity stalls at the third application rather than the first deployment. A forty-engineer team gets one service into production and keeps it healthy through familiarity. Then an acquisition brings in a .NET workload, an internal tool developed by an engineer who left two years ago, and the team is carrying three operational models with nobody owning the mandate to reconcile them. Hiring will not close that gap, because what is missing is a standardized operational posture, not headcount.

What hundreds of thousands of production deployments actually reveal

We have visibility into hundreds of thousands of production deployments across thousands of customers. That scale does not tell you what people say they want. It tells you what actually breaks, what gets escalated at 3 a.m., and what determines whether a team trusts their platform enough to stop thinking about it. The patterns are remarkably consistent.

1. The deployment that nobody touches again

A team spends two full days getting a Spring Boot application deployed with CI/CD and an SSL certificate. The deploy works. Then nobody touches it for three months because it is stable, but because touching it might break it again. They discover the problem when customers report the application is unreachable. Four hours of forensics follow. What changed? Why? How do we prevent this?

“A release that can partially succeed is a release that will partially fail.”

This is the most common failure mode we observe: not the deployment that fails loudly, but the deployment that succeeds quietly and degrades invisibly. The team cannot explain what happened because the platform did not decide what “healthy” meant before the incident arrived.

A release that can partially succeed is a release that will partially fail. The platforms that earn trust are the ones where a deployment either completes fully or reverses entirely no intermediate states, no manual rollback procedures discovered under pressure.

2. The observability sprint that ships three weeks late

A developer notices response times degrading. They want memory utilization metrics. The platform does not collect them by default. They spend a sprint writing configuration files to install a monitoring agent across every instance. The insight they needed three weeks ago ships three weeks late.

This is the second most common pattern: observability treated as an add-on rather than a default. Every team we have observed that adds monitoring after their first incident wishes they had it before. The teams that never experience this problem are the ones whose platforms shipped metrics, traces, and health signals at deploy time without code changes, without configuration, without a sprint spent on plumbing. The distinction matters: a platform that can be observed is not the same as a platform that is observed from the moment it goes live.

3. The portfolio tax

Each application gets its own infrastructure. Costs scale linearly with the portfolio. The team running eight applications pays eight times the overhead of the team running one, not because each application needs dedicated resources, but because the platform’s architecture assumes isolation rather than shared operational responsibility. The result: teams avoid migrating inherited applications because the cost model punishes breadth. The compliance audit does not care that the inherited application runs on a different operational model. It expects the same governance, patching cadence, and access controls.

A platform that rewards portfolio growth, shared infrastructure, and a consistent operational posture —and economics that improve with breadth rather than degrade—changes the calculus for teams managing applications they did not build.

4. The security configuration nobody made 

The forty-engineer company has no security team. They have a senior developer who reads the CIS benchmarks on weekends. The compliance audit arrives regardless. Teams without specialists will not configure security controls that require specialist configuration. This is not a criticism of those teams; it is a structural observation about how security actually gets implemented (or does not) in organizations where every engineer is a full-stack contributor with a product backlog that never shrinks.

The only security posture that works reliably for these teams is the one they inherit by default: compliance certifications, network isolation, access controls that ship with the platform rather than requiring a dedicated sprint to implement.

What these patterns demand today

Every failure mode we observed the deployment nobody touches, the observability sprint that ships late, the portfolio tax, the security configuration nobody made- shares a root cause: the platform asked the team to make an operational decision, and the team either made the wrong one or made none at all. The response to these patterns required starting from a different question. Not “what should we configure for the team?” but “what should the team never need to decide?”

A platform that answers that question correctly holds a specific set of commitments. It decides what healthy looks like before the first request arrives, not after the first incident. It ships observability at deploy time, not as a sprint the team schedules after something breaks. It treats the eighth application in a portfolio the same as the first, same operational model, same governance, same economics. It inherits security posture by default, because the teams it serves will never staff a dedicated security function.

These are not feature decisions. They are architectural decisions about where operational responsibility permanently resides. The team defines the application source code, a Dockerfile, a pre-built image, or an existing workload being migrated in. From that point forward, the platform owns everything underneath it. Not for the first deploy. For the life of the application. Patching, scaling, healing, certificate rotation, capacity planning, health evaluation. These are not capabilities the team enables. They are responsibilities the platform holds permanently.

This is what we rebuilt AWS Elastic Beanstalk to be. Not a deployment tool. Not a hosting layer. An application management service that takes operational responsibility for everything underneath the application. The architecture now starts from the question above and refuses to let the answer drift back toward the team over time. Elastic Beanstalk operates in two modes, a structural change from its previous single-environment architecture:

Standard Mode delivers full operational ownership for individual applications and Windows/.NET Framework workloads: the complete operational stack, owned outright, for a single service.

Cluster Mode extends the same ownership model across the portfolio, shared infrastructure, source-to-production deployment that transforms code into running applications, and economics that improve as the portfolio grows. The eighth application shares operational overhead with the first seven rather than duplicating it. For the forty-engineer company running eight production applications today and inheriting ten more next quarter, this is the difference between a platform that covers the portfolio and a platform that covers only the applications simple enough to fit its opinions.

The industry convergence

The distinction is real, though I would not draw it as a line between platforms that reduce complexity and platforms that own operations permanently. Every vendor in this market absorbs some operational responsibility at deploy time. The key question is how much of it returns to the team during an incident, a patch cycle, and an audit. A platform that removes infrastructure management from developers during the workweek and reintroduces it at 3 a.m. Sunday addresses only half the challenge.

“A platform that removes infrastructure management from developers during the workweek and reintroduces it at 3 a.m. Sunday addresses only half the challenge.”

At convergence, the direction is correct, but the shape is incorrect. This is not two camps meeting in the middle. Gartner’s 2026 Magic Quadrant for Cloud Native Application Platforms places AWS, Microsoft, Google, and Red Hat in the leaders quadrant, with Render, Netlify, and Upsun as niche players. Vendors specializing in developer experience showed this category is viable, and now the hyperscalers are adopting it. Since source-code-to-URL mapping is now standard across the entire quadrant, the key differentiation becomes who bears operational liability for the eighth application three years after its release.

The only aspect of framing I would challenge is the idea that a platform determines everything the team never has to decide. Routine infrastructure decisions should stay out of the developer’s path, and escape hatches should stay in place for the teams that genuinely need them. A platform that removes choice altogether will demo well and then stall when teams migrate applications that don’t fit its opinions.

The 3 a.m. test that actually matters

For teams already living this reality — serious production, lean staff, growing portfolios — nobody planned to operate the 3 a.m. test; it is not a nice-to-have. It is the evaluation criterion.

The platforms that define the next decade will not simply make deployment easier. They will decide, in advance, how production systems should behave when things inevitably go wrong. Because by 3 a.m., the time for deciding has already passed. The CNAP category was built to describe platforms that own the application lifecycle.

Elastic Beanstalk made those decisions before the incident arrived: what healthy looks like, what to do when it stops being true, how to communicate what happened. The cloud gave teams power. These teams needed someone to stay. Elastic Beanstalk stays.

Check the latest AWS release notes for Elastic Beanstalk here.

The post The cloud reduced operational complexity. But many teams need someone to own it completely. appeared first on The New Stack.

  •  

One engineer shipped 2,000 PRs a month to production. Verification is the key.

Abstract dark digital wireframe mesh with chromatic glitch effects, representing AI agent verification and virtualized software environments.

Lauren Tan, an engineer on the Grok team at SpaceXAI, who previously worked at Cursor and Meta, recently published a guide to her personal agent workflow: pstack. The attention-grabbing number is that pstack has let her ship 2,000 pull requests (PRs) a month to production with high confidence. That’s one engineer shipping nearly 100 PRs per working day.

The number is incredible and undeniably an outlier, but the direction is not a surprise. I have argued previously that coding agents would enable teams to generate ten times the code with similar headcounts. What is surprising is that these are not just code output numbers. These are actual changes landing in production.

According to Tan, the most critical piece of that workflow is verification. A verification skill lets an agent check its own work and keep going until the task is done, and she treats it as “critical infrastructure” rather than one skill among many. 

The verification skill rests on something underneath it: a rich runtime the agent can drive, inspect, and get structured answers from. For a single application, that runtime is the application itself, started on demand. For a system made of dozens or hundreds of services, no such runtime exists by default, and providing one that keeps up with hundreds of parallel agents is the hard part.

Verification is the whole game, and the math says so

Her argument for agentic verification is a throughput argument. An agent that can check its own output keeps working until the task is done. An agent that can’t hand you a diff and wait makes you the slowest component in the loop. That is why she claims strong verification skills can multiply a team’s output by 100 to 1,000 times.

“An agent that can check its own output keeps working until the task is done. An agent that can’t hand you a diff and wait makes you the slowest component in the loop.”

At 2,000 pull requests a month, reviewing every change by hand would allow about five minutes per PR across a full working month. Human review cannot be the verification layer at that volume. Whatever does the checking has to run without a person in the loop, and it has to run in parallel with the agents generating the work.

The model assumes the agent can run the whole application

The verification skill she describes generates a command line interface (CLI) and a feature map for the application. The CLI lets an agent start the app, navigate it, inspect state, and read structured JSON results back. Each agent gets a complete copy of the application and can test a change end to end.

She is direct about how much rests on that runtime: “I personally feel that agentic verification is so important that I would unironically suggest building your own rich debugging tools, or even choosing a different tech stack, in order to have unfair advantages and extreme productivity in building software.”

Her approach to providing a runtime for her agent works because the application fits in one process. A frontend, a compiler, or a single service with a database can start from a CLI in seconds and be thrown away afterward.

“I would unironically suggest building your own rich debugging tools, or even choosing a different tech stack, in order to have unfair advantages and extreme productivity in building software.”

For teams building complex distributed applications, their system does not have that property. The application is the interaction between an order service, a payments service, an inventory service, a queue, several databases, and a handful of third-party APIs. At larger shops, the count runs into the thousands. A pull request to one service is only verified by exercising the calls it makes and receives. The CLI can start the changed service. It cannot start the system.

None of the existing runtimes survive hundreds of parallel agents

Local runtimes with mocks are cheap and can run fully parallel using worktrees or CDEs. Their problem is fidelity. Mocks encode what a dependency did the last time someone looked, and they drift the moment the real service changes. An agent that verifies against mocks closes its loop against fiction, and the failure shows up after merge.

A full copy of the stack per change is faithful and isolated. But its cost scales with the number of services times the number of concurrent changes, and at hundreds of agents that cost is untenable. Time is the bigger problem. A full stack takes minutes to provision, and the loop she describes has the agent testing every iteration of a change while it is still working on it. An environment that is ready after the agent has moved on to its next attempt is no use to it.

Shared staging is faithful and cheap because there is one of it, and that is the whole problem. A single mutable environment cannot host hundreds of concurrent changes. Agents overwrite each other’s deployments, a broken change from one agent becomes failed tests for every other agent, and the loop-closing property that makes the workflow valuable disappears.

Each existing verification runtime plotted on a chart showing its realism of dependencies against concurrency.

What verification needs when the callers are agents

Read Lauren’s workflow as a requirements document, and five properties emerge:

  • The change has to run against real dependencies or the verification means nothing.
  • Hundreds of concurrent changes have to be unable to see each other.
  • The cost of an environment has to scale with the size of the change, not the size of the system.
  • Environments have to come up in seconds, because an agent waiting on provisioning is parallelism you paid for but didn’t use.
  • All of it has to be reachable through the CLI or MCP server the agent already uses, because the caller is an agent.

The first and third requirements pull in opposite directions. Realism pushes toward complete copies of the system. Cost pushes toward sharing as much as possible. Shared staging resolves that tension by giving up isolation, and a per-change full stack resolves it by giving up cost efficiency. A design that satisfies all five has to share and isolate at the same time.

Virtualized full-stack environments share the system and isolate the change

The architecture that does this treats an environment as a view of a running system rather than a copy. One shared set of stable services runs continuously, deployed from the main branch and kept healthy the way production is. When an agent needs to verify a change, it runs only the service it modified, on its own machine or as a lightweight deployment in the cluster, and joins it to the shared stack as a new isolated environment.

“The architecture that does this treats an environment as a view of a running system rather than a copy.”

From inside that environment, the changed service is the version of record, and every other call falls through to the shared stable versions. The agent sees a complete, realistic system, and so do the other hundred agents, each seeing a system that differs from the baseline by only the delta of its own change. Requests carry their environment identity as they cross service boundaries, which keeps one agent’s traffic from reaching another agent’s version under test. Stateful side effects that cannot be shared safely, like queue topics or writable databases, get a per-environment copy where needed.

Diagram showing "agent 1 env" and "agent 2 env" interacting with the shared cluster.

The cost model follows directly. An environment costs one or two running services instead of sixty; it is ready in the time a single service takes to start, and you can create and destroy it from inside the agent’s own loop. This is the pattern Signadot packages for Kubernetes, with the shared stable stack running in the team’s existing cluster.

The loop, end to end, with an agent as the actor

Put the two halves together, and the workflow that enables her to ship 2,000 PRs a month to production carries over to a distributed system almost unchanged. An agent picks up a task and changes one service. It asks for an environment for that change and gets one in the time it takes its service to start. It then drives real requests through the system’s entry point and watches them traverse the real dependency graph, with only its own service running new code. It reads structured results, fixes what failed, and runs again. When the checks pass, it opens the PR, and the environment goes away at merge.

“Environments stop being something the platform team hands out and become something agents create, use, and discard as needed.”

For the platform team, the unit of work changes. Today it provisions environments, whether that means keeping a shared staging alive or stamping out full copies of it. In this model, it runs one shared stable stack and the layer that virtualizes it: context propagation across every service, isolation for the stateful dependencies that cannot be shared, and the tooling that creates and tears down environments. Environments stop being something the platform team hands out and become something agents create, use, and discard as needed.

Verification capacity is the new ceiling on throughput

Lauren Tan’s post is not a story about one unusually productive engineer. It shows what happens when agents run the full loop, writing a change, verifying it, and iterating without a person in between. The verification infrastructure is the foundation that the entire loop stands on.

In distributed applications, that infrastructure has to be a runtime environment that gives every agent real dependencies, keeps hundreds of concurrent changes from seeing each other, costs a change rather than a copy of the system, and is ready in the seconds an agent is willing to wait. That is what turns agent parallelism into shipped code rather than a longer review queue. That model of runtime environments is exactly what we built Signadot to enable.

The post One engineer shipped 2,000 PRs a month to production. Verification is the key. appeared first on The New Stack.

  •  

Kubernetes can run AI inference. But can it count the real cost?

Abstract 3D illustration of interconnected purple geometric nodes and gold lines representing a distributed network or cloud infrastructure.

Welcome to another edition of Road to KubeCon, where we’re tracking the Kubernetes and cloud-native ecosystem on the way into KubeCon + Cloud Native Con NA 2026, to be held in Salt Lake City, Utah, November 9-12.

This week, we look back at the past week of significant movements in the Kubernetes space. Most notably, we see interesting advances in cloud-native architectures for AI inference. We take a look at that, plus a new Gartner quadrant, new Kubernetes hardening updates, and important CNCF project updates.

HPE challenges server virtualization platforms

On Monday, Gartner published its Magic Quadrant for Server Virtualization Platforms, a guide comparing solution providers in the server virtualization market. The quadrant names HPE as a Challenger based on Ability to Execute and Completeness of Vision.

Hewlett Packard Enterprise (HPE) is a presenting sponsor of Road to KubeCon. HPE Software helps IT organizations modernize infrastructure, streamline operations, and accelerate AI initiatives across hybrid, multi-vendor environments.

According to the HPE newsroom, the recognition reflects ongoing momentum behind HPE Morpheus Software, its virtualization and cloud operations portfolio. HPE was positioned in the Challengers quadrant alongside Canonical and Oracle.

The announcement comes as enterprises rethink their virtualization strategies. Rather than simply swapping in another hypervisor, HPE argues that organizations increasingly need unified governance and ways to provision, orchestrate, observe and secure workloads — including VMs, containers and AI workloads — across clouds.

Kubernetes hardens container storage

On Wednesday, Red Hat’s Nispriha Jagan and Neeraj Krishna wrote on the Kubernetes project blog about two new storage security features shipped as Alpha in Kubernetes v1.37, which included 67 enhancements.

The notable security features are new bind mount options and emptyDir permissions. The enhancements come as multiple security findings have surfaced regarding emptyDir volumes, one of the most common writable volume types. The additions are made possible by low-level Linux security mechanisms.

According to the authors, these enhancements give users native controls to harden Kubernetes workload storage better. “Supporting noexec, nodev, and nosuid gives users a native way to harden volume mounts to match security benchmarks and policy,” the authors write.

KubeCon adds AI Inference + Agentic track

Last month, CNCF announced it will feature an AI Inference + Agentic track at KubeCon + CloudNativeCon North America 2026, exploring the intersection of generative AI and cloud native infrastructure. Attendees can explore the sessions here.

The added track underscores the growing use of Kubernetes for production AI workloads, particularly as the focus shifts from training models to serving them in production. It also reflects emerging practices for building agentic systems around protocols like MCP and A2A, as well as infrastructure such as AI gateways.

China Merchants Bank unifies AI inference on Kubernetes

China Merchants Bank, a leading Chinese commercial bank, recently showcased its cloud-native AI infrastructure at a CNCF event in China. Its infrastructure team won the CNCF End User Case Study Contest with an architecture combining Kubernetes with several cloud native projects:

  • Kueue, for job queueing and quotas,
  • KEDA, for event-based auto-scaling,
  • Prometheus, for systems monitoring and metrics,
  • HAMi, for sharing accelerator capacity across Kubernetes workloads,
  • and Fluid, for accelerating access to datasets.

The bank has a large pool of nearly 10,000 accelerator cards used for AI computation. These are heterogeneous, meaning they are not all the same type or configuration.

According to the CNCF announcement, the architecture unified management of 99% of its AI compute resources, while increasing average utilization from 35% to more than 60%. It also cut the cost of processing 1 million tokens by 60% under comparable conditions.

The case study shows how cloud-native infrastructure can improve utilization and efficiency for AI training and inference, even in regulated areas like financial services.

Industry take: Can AI inference on Kubernetes handle token cost issues?

Interest in AI inference on cloud native infrastructure is palpable. However, this week Val Bercovici, chief AI officer at WEKA, an AI-native data platform, questions whether Kubernetes’ existing resource model fits the changing economics of large-scale AI inference.

Bercovici tells The New Stack: “With AI inference, it’s cost per token, and that cost depends on state Kubernetes was never designed to manage: request mix, KV cache occupancy, the balance of prefill and decode, and how memory and bandwidth are consumed inside the accelerator after a pod is already running.”

“My view is that Kubernetes doesn’t go away,” Bercovici says. “But unless its resource model evolves, it becomes a tax on inference economics.” 

He foresees a new scheduling and memory layer to emerge around Kubernetes that can compute what a token actually costs to serve. Then platforms could make more informed, cost-based decisions about how inference workloads are scheduled and served.

As Kubernetes evolves, so do the demands on the teams running it. Presenting sponsor HPE helps teams address that complexity with software spanning virtualization, cloud management, observability, and automation.

Move over, platform engineering. Hey, agentic engineering.

A new Weave Intelligence report, State of AI in Platform Engineering Volume 2, authored by Sam Barlien, Luca Galante, and Florian Lipp, surveyed 242 platform engineering leaders on the before-and-after effects of introducing agentic AI into platform engineering.

38% of teams are shipping at least twice as much as before AI. When assessing ROI across the software delivery life cycle, 20% report efficiency gains and 11% report operational savings. Yet only 8% report a transformative, structural shift. Meanwhile, 29% are still prototyping without realized gains, with some outliers reporting negative results.

The biggest roadblock to scaling AI usage? A lack of platform readiness, including APIs, deterministic pathways, and standardization. Weave’s takeaway is that platform engineering must increasingly account for AI readiness and agentic experience as agents become another key platform consumer.

OpenTelemetry Kubernetes attributes processor reaches v1.0.0

On Wednesday, OpenTelemetry, the graduated CNCF project and open standard for telemetry, announced the v1.0.0 release and distribution of its Kubernetes attributes processor. It’s a helpful feature that uses the Kubernetes API to add Kubernetes metadata, such as stability, distributions, warnings, issues, and other metrics, to resource attributes.

According to the release notes, written by Elastic’s Christos Markou and Datadog’s Pablo Baeyens, the feature has been in progress in the OpenTelemetry Collector SIG since late 2025, based on a roadmap of users’ most-requested features. Existing attribute processors should review the breaking changes and migration guide.

DigitalOcean opens Spot GPU node pools

Technically, this occurred the week before last, but didn’t make the digest. As of September 9, DigitalOcean Kubernetes’ (DOKS) Spot GPU Node Pools entered public preview. According to the release notes, the feature runs worker nodes on interruptible GPU capacity at a “lower, variable rate than on-demand GPU nodes.” This could offer a cost-effective option for fault-tolerant workloads.

Other updates from the K8s universe

More updates from the infrastructure-heads, platform engineers, and multi-cloud operators working in the Kubernetes ecosystem: 

Follow the Road to KubeCon

Road to KubeCon is an eight-part series presented by HPE, which will be at KubeCon + CloudNativeCon North America in Salt Lake City. Before you go, explore how HPE Software helps IT teams do more with less complexity.

We’ll be here every Friday until KubeCon.

If you’d like to participate, Bill Doerrfeld, the writer of this series, is open to pitches — you can send release notes, quotes, reports, videos, case studies, or hot takes through his contact page.

If you didn’t catch the inaugural edition covering the Kubernetes v1.37 release, check it out here. You can also visit the Road to KubeCon page for the complete archive.

The post Kubernetes can run AI inference. But can it count the real cost? appeared first on The New Stack.

  •  

How buildpacks help enterprises finally operate container security controls at scale

Dark abstract geometric grid texture symbolizing standardized container security controls.

Container security controls often fail less because organizations lack standards, scanners, or best practices, and more. After all, every service tends to use its own Dockerfile, building images in slightly different ways.

Buildpacks are often pitched as a way to make containerization more developer-friendly, standardized, and maintainable by eliminating the Dockerfile pain points. They can also help improve security posture in practice, especially in larger, multi-team environments.

In this article, we will explore how Buildpacks provide a standard application-image path where platform teams can centrally govern build inputs, runtime images, Software Bill of Materials (SBOM) production, artifact metadata, and update workflows–making container controls easier to enforce and operate at scale.

The patch that never reached production

What does a patch integration process look like in the real world?

In more than one enterprise, I’ve seen the story look roughly like this: a critical vulnerability is fixed in the organization’s approved runtime image. In a perfect world, a patched base image is detected centrally, rebuilt, tested, published, and automatically propagated to every affected application. But we do not live in a perfect world. A week later, some production workloads still run on the vulnerable base.

“A container security control isn’t useful on its own. It must be applied consistently, observable in the running estate, and maintainable when images, dependencies, and vulnerabilities change.”

Typical reasons include:

  • Inconsistent Dockerfiles using different base image tags, Linux versions, and runtime versions, making it harder to determine which service is affected;
  • Application images may start from a patched base but reinstall vulnerable dependencies because developers manually install packages;
  • Uneven rebuild cadences: some teams rebuild more frequently, others only rebuild when the application code changes, so a service with no recent feature work may run an old vulnerable image for months;
  • No centralized inventory, SBOMs, image metadata, or deployment tracking, so organizations cannot reliably determine which applications remain vulnerable or measure patch propagation. 

This story has not yet found its happy ending because the organization is only halfway into establishing a vulnerability management process. A container security control in place, such as “all services must use approved patched base images,” isn’t useful on its own. It must be applied consistently, observable in the running estate, and maintainable when images, dependencies, and vulnerabilities change.

That’s the stage where many teams fall short.

Why container security controls drift

Container security controls are hard to implement largely because of established containerization practices. In many organizations, no single governed build path exists. Each repository produces its own image, usually through a Dockerfile maintained by the application team.

Dockerfiles offer flexibility. They also push many security decisions onto developers: which base images to use, which packages to install, how to configure the runtime user, how minimal the image should be, how to track patches, which CI policies to apply. Over time, as services multiply and teams change, these choices drift. It’s not unusual to see different repositories using different base images, update cycles, and even different interpretations of what “secure” means in practice.

“The problem is expecting every developer to have enough container expertise to do this consistently across hundreds of repositories.”

Dockerfiles are not the problem by themselves. A well-written Dockerfile can produce a minimal, hardened image. The problem is expecting every developer to have enough container expertise to do this consistently across hundreds of repositories.

This approach also makes patch propagation unreliable. Platform teams may publish a patched base image, but each application team must still notice the update, modify its Dockerfile, rebuild, test, and redeploy. Some do it quickly. Others do it late or not at all. The control exists, but adoption remains uneven.

Buildpacks help close this gap by moving common containerization decisions out of individual repositories and into a shared, governed build path.

The shift from repository-specific builds to a governed build platform

Buildpacks turn application source code into a production-ready OCI container image without a Dockerfile. They detect the application type, select the required buildpacks, provide the necessary runtime and dependencies, and produce a runnable image.

For container security, the main benefit is not simply removing Dockerfiles. Buildpacks tend to provide a standard way to construct images, at least for typical application stacks.

“Buildpacks turn application source code into a production-ready OCI container image without a Dockerfile.”

This gives platform and security teams a central point for enforcing controls. Instead of asking every team to choose an approved base image, configure the runtime, manage layers, generate metadata, and track updates, the organization can encode much of this work in shared builders and buildpacks. Developers own their code, dependencies, and service behavior. Producing a compliant image becomes a platform responsibility.

Because Cloud Native Buildpacks is a CNCF graduated project, its specifications and reference implementations undergo rigorous community review and long-term maintenance, making them suitable as the foundation for enterprise security controls.

With that foundation in place, applying and maintaining specific container security controls becomes easier at scale.

Four container security controls buildpacks make it easier to operate

Standardization of build inputs

The first control is standardizing what goes into the build.

In Cloud Native Buildpacks, the key unit is the builder. A builder packages the buildpacks, lifecycle, build-time base image, and runtime base image used to create the final application image. This establishes the builder as a controlled definition of how application images are produced. Developers do not choose a random base image on Docker Hub to build their applications; the Buildpacks ecosystem defines a set of build and run images.

This builder-based approach introduces an important concept. A developer cannot change the base OS layer in a builder with a single line of code because compatibility isn’t guaranteed. With one line of code, however, they can swap builders and still get a compatible, working image.  

This is where standardization takes place. Developers keep using their preferred languages and frameworks, but the platform team builds images from a controlled set of approved builders. Extension paths for cases where buildpacks require customization can also be standardized.

“By controlling the builder, the organization also controls the buildpacks, runtime image family, lifecycle version, and build paths allowed in CI/CD.”

The security benefit is straightforward. By controlling the builder, the organization also controls the buildpacks, runtime image family, lifecycle version, and build paths allowed in CI/CD. The policy is defined once at the platform level and then applied across many services.

Best container security practices by default

After the builder standardization, the next question is: what does the resulting image look like? This is where buildpacks help. They improve the output by applying several container security practices as part of the normal image creation flow:

  • Non-root build and execution. Cloud Native Buildpacks require buildpack code to run as a non-root user. Platforms, such as pack, also produce images configured to run applications as the non-root user defined by the run image. Non-root defaults reduce the blast radius of a compromise and make container escape, filesystem tampering, and exploitation of the image build process more difficult.
  • Separation of build and runtime environments. Buildpacks mark layers as build-only or launch-time. Only launch layers are included in the final image, so compilers, npm tooling, build caches, and similar tools remain outside production, reducing the attack surface.
  • Restricted modifications of the base image. Because buildpacks run without root privileges, they cannot simply install OS packages or modify the base filesystem. OS changes must be provided through the controlled build/run images or image extensions.
  • Isolation of sensitive build privileges. When using an untrusted builder, sensitive lifecycle phases can run separately from the phases that execute buildpack code. This prevents an untrusted buildpack from receiving capabilities such as registry credentials and access to the container daemon.

Some builder providers offer additional security features. For example, Paketo buildpacks for Spring Boot provide base images without a shell. Another example is BellSoft’s hardened builder for Paketo buildpacks based on BellSoft Hardened Images.  

These defaults reduce the manual effort and make secure container builds part of a standard process.

SBOM Generation

A Software Bill of Materials (SBOM) lists all software components in an application or image. An SBOM is required for vulnerability tracking, license reviews, and audits, and many regulations require it.

Teams often add SBOM generation as a separate CI step, with separate tools, formats, storage rules, and ownership. That makes coverage uneven, especially across many repositories.

Cloud Native Buildpacks reduce this friction by emitting an SBOM for all dependencies they provide, in formats such as CycloneDX, SPDX, or Syft JSON. As a result, platform and security teams get consistent image inventory data without requiring every application team to build its own SBOM process.

For a typical Java service, that might mean SBOM entries for the JRE version, Spring Boot version, and key libraries pulled in at build time. For Node.js services, it might include the Node runtime version and major npm dependencies. The exact content will vary, but the pattern is the same: SBOM data comes from the build itself, not a separate, manually maintained process.

Patching at scale

As discussed earlier, the hard part is often not writing or receiving a patch, but getting it into the running workloads. Buildpacks shorten that path by using shared builders, buildpacks, and run images. Platform teams can then update these common inputs once, instead of waiting for every application team to repeat the same change.

Another powerful feature of Buildpacks that promotes rapid patching is rebasing. When OS-level fixes become available, the runtime base layers can be replaced with layers from a newer run image without rebuilding the application from source.

Kubernetes tools such as kpack can automate this process. They track image resources and trigger rebuilds when the source, builder, buildpacks, or stack changes.

Rebasing has limits: it updates only the run-image layers. Dependencies added by buildpacks, such as a JRE or Node.js runtime, usually require a rebuild with updated buildpacks or dependency versions. For a Node.js service, a rebuild might mean picking up a new Node.js runtime version and updated npm packages, not just a newer OS base.

So, we are not talking about magical automatic patching here. The key benefit is a more centralized, repeatable patch-propagation path across the image estate—though teams still need to handle testing, rollouts, and exceptions.

The new patch path after buildpacks

Let’s return to the original story. We left the organization in a state where vulnerability management lacked consistent patch integration. Suppose the enterprise has migrated their workflows to Buildpacks — how has the process changed?

Once again, a critical vulnerability is discovered. But after adopting buildpacks, the response looks different: the remediation starts with shared build inputs.

The platform team updates the approved run image, builder, or buildpacks. Images are then rebased or rebuilt:

  • Rebase when the fix affects the run-image OS layers;
  • Rebuild when the affected component is runtime, like a JRE or Node.js runtime, or application dependencies.

Of course, Buildpacks do not remove the need for testing or deployment controls. Application teams still own their code, dependencies, and compatibility testing. SRE teams still promote, deploy, monitor, and, when necessary, roll back the patched images.

“Buildpacks give organizations one controlled way to build container images instead of leaving every repository to define its own process.”

The most important change is the patch path. Platform teams maintain the shared build inputs and automation. Security defines scan policies and exception rules. Compliance defines the evidence that must be retained.

This path also promotes shifting security left, as it becomes easier for developers to meet the in-house security requirements and container security best practices.

FunctionResponsibility
PlatformBuilders, run images, buildpacks, and rebuild automation
SecurityVulnerability policy and exceptions
Application teamsCode, dependencies, and compatibility testing
SREPromotion, rollout, monitoring, and rollback
ComplianceAudit evidence and retention

Some workloads still need custom images. But you don’t need to choose between Buildpacks and Dockerfiles; you can use both. When required, you can create a custom buildpack and extend the build-time base image with a Dockerfile. The result is not a rigid “Buildpacks only” model, but a governed image strategy with controlled customizations. 

Buildpacks give organizations one controlled way to build container images instead of leaving every repository to define its own process. This makes security controls–such as approved images, safe defaults, SBOMs, and patching–easier to scale. The result is reduced drift, clearer ownership, and faster updates.

If you’d like to get started with Buildpacks, try them on one service and compare the workflow with your current Dockerfile process.

The post How buildpacks help enterprises finally operate container security controls at scale appeared first on The New Stack.

  •  

How to attach an owner to every cloud resource you find

Dark abstract 3D rendering of intertwined rings symbolizing complex cloud resource governance and technical debt.

The engineer who knew why that cloud instance existed has left the company. The instance is still running, the bill keeps growing, and the team must now decide whether it’s safe to shut down. This is a bad time to discover that its ownership history was someone’s memory.

“Good resource governance has three pillars: continuously synced inventory, policy that blocks resources without tagged owners, and an audit trail that survives every reorg.”

A cost review flags an EC2 instance nobody remembers provisioning. Someone scours Slack for the resource ID, finds nothing, burns half a day chasing dead ends, and eventually stumbles across an exhausted engineer who mumbles the mantra that ends most of these investigations: “I think that’s from the project Priya was running before she left.” 

Nobody follows up, because nobody knows how to reach Priya anymore. The instance stays up, because tearing down a fence when you don’t know what it’s protecting you from is a good way to find out the hard way. It’s the platform team equivalent of emotional baggage; they all seem to accumulate some.

Processing this baggage doesn’t require knowing Priya’s replacement, writing better documentation, or hoping the next reorg is more rigorous. 

“Processing this baggage doesn’t require knowing Priya’s replacement, writing better documentation, or hoping the next reorg is more rigorous.”

Solving the problem requires three things that already exist: queries, policies, and logs; and maybe just a little bit of therapy.

1. The query that tells you what’s missing an owner

CloudQuery’s asset inventory syncs continuously across every provider a team runs on, into tables you can query directly: aws_ec2_instances, gcp_compute_instances, azure_compute_virtual_machines, and so on. Finding every resource without an assigned owner is as easy as:

SQL
SELECT resource_id, 'aws' AS provider, 'ec2_instance' AS resource_type
FROM aws_ec2_instances
WHERE tags ->> 'owner' IS NULL
UNION ALL
SELECT resource_id, 'gcp', 'compute_instance'
FROM gcp_compute_instances
WHERE labels ->> 'owner' IS NULL
UNION ALL
SELECT resource_id, 'azure', 'virtual_machine'
FROM azure_compute_virtual_machines
WHERE tags ->> 'owner' IS NULL
ORDER BY provider;

Run it regularly, and you have a (hopefully short) boring list to refer to when the CFO asks who deployed an expensive instance; instead of being thrown into a frenzied goose chase at 4:59 p.m. on a Friday.

2. The policy that prevents it recurring

A query tells you what’s already missing an owner; but how do you prevent the next ownerless resource from being deployed? The answer is a policy, and env zero evaluates Open Policy Agent rules against every plan before it applies. A rule that requires an owner tag on every new resource looks something like this:

package env0

# METADATA
# title: require owner tag
# description: A resource can't be created without a declared owner.
deny[format(rego.metadata.rule())] {
	resource := input.resource_changes[_]
	resource.change.actions[_] == "create"
	not resource.change.after.tags.owner
}

format(meta) := meta.description

Add that to the project’s policy set, and a plan that creates a resource without an owner tag doesn’t just get a warning; it doesn’t get created.

3. The record that outlives its creator

A tag tells you who owns something today. It doesn’t tell you anything about who asked for it, why, or who signed off. By the time it matters, the person who could’ve answered from memory may no longer be reachable. An audit entry records that at the moment of creation, instead of reconstructing it afterward from the scraps of recollection scattered around the rest of the team. Here’s an example:

{
  "event": "resource.created",
  "resource_id": "i-0a1b2c3d4e5f",
  "requested_by": "j.chen@company.com",
  "approved_by": "platform-lead@company.com",
  "approval_ref": "ENV-4471",
  "stated_purpose": "load test environment, Q3 capacity planning",
  "timestamp": "2026-08-14T09:12:03Z"
}

This entry answers the question this whole piece opened with, without needing Priya or Slack. env zero keeps this record attached to the resource for as long as the resource exists, specifically so it outlasts any individual’s tenure.

Why this keeps happening

Employee attrition is an age-old challenge that’s only accelerating in the modern era. US private-sector voluntary turnover runs 22 to 25% a year, so a hundred-person org loses twenty-odd people every year, each one taking a small, specific piece of “why this exists” with them. Replacing a mid-level employee costs six to nine months of salary, more than double that for senior specialists. 

“Tags were supposed to survive this. In practice they rot the way everything else does.”

That figure doesn’t touch what the departure does to everyone else’s mental model of what’s actually running. Tags were supposed to survive this. In practice they rot the way everything else does: two teams merge and bring incompatible schemas, provisioning that runs on tribal knowledge accumulates configuration drift for the same reason it accumulates ambiguous ownership, and a resource tagged under a policy that’s been replaced twice since its inception isn’t much better documented than one that isn’t tagged at all.

We wrote in April about the hour it once took our own team to answer, “what are we actually running across both clouds?” That solved a point-in-time problem. The query, the policy, and the audit entry above stop the same story from playing out again next June.

The org chart will keep changing. The record doesn’t have to.

None of this stops people from leaving or teams from reorganizing; pretending otherwise is how platform teams end up rebuilding the same spreadsheet every eighteen months. What changes is whether the next “what are we actually running, and who owns it” conversation takes hours of archaeology across three teams, or is a query that already has the answer attached. Institutional knowledge decays at a fairly predictable rate. A system of record shouldn’t.

The post How to attach an owner to every cloud resource you find appeared first on The New Stack.

  •  

The AI-native SDLC won’t be one process 

Five fuzzy pom-poms — white, pink, magenta, purple, and teal — arranged in a horizontal row across the upper portion of a dark navy background, each casting a small shadow, with colored light washing the backdrop in magenta and teal.

Anthropic recently published its AI-Native SDLC Playbook. Its central claim is that “code is no longer the bottleneck.” When agents can produce an implementation in minutes, the constraint moves to everything around the build phase: planning, review, verification, deployment, and governance.

The risk, if organizations get this wrong, is producing ten times the changes at the same quality per change or worse, with no way to identify which changes are the bad ones. The traditional answer is that a person looks at each one, and that is exactly what stops working at this volume.

The risk… is producing ten times the changes at the same quality per change or worse, with no way to identify which changes are the bad ones.

The playbook gets the foundations right. What it doesn’t capture is that an organization’s process has nuance: it is really a family of processes that vary with the change at hand, not a single flow every change travels through.

The spec-driven wave

The playbook is part of a broader wave of spec-driven development tooling, including Amazon’s Kiro and GitHub’s Spec Kit. The tools share a common shape. Written artifacts drive the work: an intent document becomes a spec, a plan, a diff, and review findings, all committed to version control. Policy is enforced by deterministic mechanisms such as hooks, rather than by instructions in a prompt. Agents check their own work before a human sees it. Humans own the approvals.

But each of these tools also prescribes a particular process: a fixed sequence of stages that produce fixed artifacts, and that every change travels through. Adopting the tool means adopting its process.

One organization runs many processes

No real organization runs a single process. The right process for a change depends on the risk it carries and the accountability it requires. A documentation fix, a dependency upgrade, and a schema migration in a payments service should not travel the same path. They need different levels of verification, different approvers, and different records. In regulated domains, the process itself is part of the compliance obligation: auditors expect a record of who approved each change and based on what evidence. What must be recorded differs by the type of change.

When a tool prescribes one process, teams route around it for changes that don’t fit, which is the worst outcome because the real process becomes invisible.

When a tool prescribes one process, teams route around it for changes that don’t fit, which is the worst outcome because the real process becomes invisible. Or the vendor keeps adding configuration until the tool becomes a workflow engine that nobody fully understands.

The tool should not prescribe a process. It should give the organization a way to define its own.

Processes as state machines

A better model is to define each process as a state machine. The states are facts about a change: reviewed, validated against its dependencies, approved for production. Those facts live in systems no single tool owns: the repository, CI, the cluster, the tracker. So a process cannot be a program that executes steps. It is a set of rules that react to observations about those systems. Each rule specifies:

  1. The facts it requires before it can fire.
  2. Its gate: fire automatically, or wait for a person’s approval.
  3. The permission it grants when it fires, such as merging or deploying.

The definition is this set of rules, stored as data and reviewed like code. An organization runs many small machines, one per risk class.

At runtime, this behaves nothing like a workflow engine. No component tracks “we are on step four”: the process advances when a fact appears in the system that owns it, and rules react. Events that arrive late, twice, or after a restart are handled like any other, because rules only react to current state. The gate is one of a rule’s conditions, so you can hold firing during an incident or a release freeze without editing any definition.

Gates need enforcement. A gate implemented as a prompt instruction depends on the model following it. The agent harness can provide the determinism required to run the state machine and enforce its gates, stopping the agent between actions until a gate is answered, while the infrastructure enforces the rest.

The process adapts to the change

One fixed process definition per repository is not enough: every change in that repository would still travel the same path, regardless of its risk. The path a change takes should depend on what the change is, and this routing comes from classifying the change, not from the author choosing a path. The organization defines classification using signals it already has: the paths a change touches, the repository it lives in, a label on its tracking issue, etc.

The definitions themselves also need to change over time, and that has to be safe. Because a process definition is data, editing it is itself a change, and it goes through its own gated process. Loosening an approval gate on the release process gets reviewed the way a schema migration does, not edited the way a config file does.

Click to enlarge graphic.

Concretely, consider three changes to the same service:

  • A documentation fix is classified by the paths it touches. Its process has two states: the build passes, and it merges. No person is involved.
  • A dependency upgrade skips design review, but its process requires compatibility evidence: the upgraded service runs its integration tests against real dependencies. A major version bump adds an approval that a patch bump does not.
  • A schema migration in the payments service is classified by the component it touches, no matter what kind of change it claims to be. Its process adds states the others never see: review by a payments owner, validation against production-shaped data, and a release approval from someone accountable for that domain.

Each transition, in each path, is logged with who approved it and on what evidence.

Tenets

The tenets these processes should follow:

  • Autonomy is granted per action, and grows over time. Each transition is set to fire automatically, require approval, or hold. As agents prove reliable on a class of change, that setting is relaxed, so the process absorbs agent improvements without redesign.
  • Human attention is spent only where judgment is needed. Agent effort keeps getting cheaper; supervision hours do not. A person is brought into the loop only when the decision requires human judgment, and is given the context to decide quickly.
  • Evidence comes from outside the agent. An agent’s own report never moves a change forward. Transitions fire on facts from systems the agent cannot write to, such as test results and validation in a realistic environment.
  • The process record is the audit trail. The definition is the written policy, and the transition log shows who approved each step, on what evidence, under which version of the policy.

Quality at scale

The playbook and its peers get the foundations right. What is missing is the ability for an organization to define its own processes, vary them by the risk of each change, and evolve them safely. The goal is not fewer humans in the loop. It is spending human judgment only where it is needed, backed by evidence agents cannot produce about themselves, so that quality holds while throughput multiplies.

We are building these ideas at Signadot and acting as our own guinea pigs, running our own development through this process. If you’re experimenting with these ideas too, we’d love to talk!

The post The AI-native SDLC won’t be one process  appeared first on The New Stack.

  •  

Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots

Abstract dark red digital grid texture representing AI code verification loops and developer workflow checks.

Red Hat released Red Hat AI 3.5 this week, a move designed to let software engineering teams run AI with the same operational rigor as enterprise apps on mission-critical infrastructure.

Echoing both the “pilots-to-production” and “single control plane for infrastructure, models, and agents” narratives playing out across much of the tech industry, Red Hat’s key move here seems to be its expansion of platform capabilities to run enhanced multi-tenancy for AI service providers. 

Crucially, the new AI platform release is built to run AI use cases that require complete hardware-to-software isolation (needed when AI workloads have to wrangle sensitive data, proprietary models, and regulated information), as well as handle priority-aware service requests (where mission-critical workloads execute in favor of lower-grade tasks) with native multi-tenancy on a shared GPU infrastructure. 

Red Hat’s Senior Director of Product for Red Hat AI is Tushar Katarki. He tells The New Stack that running enterprise AI without safety controls is “like driving a supercar blindfolded” in real terms.

“With Red Hat AI 3.5, we are delivering the operational guardrails, verifiable trust, and multi-tenant controls needed to run AI as a mission-critical service rather than an unpredictable experiment,” Katarki says. “You can’t scale what you can’t measure, and you certainly shouldn’t deploy what you can’t verify. By unifying pre-deployment safety benchmarking, real-time observability, and GPU resource management, we are giving platform teams the power to turn isolated AI pilots into a fully governed enterprise architecture.”

Every GPU request now becomes a priority decision

Applied mathematician, data scientist, and fractional CMO Joshua Estrin, Ph.D., tells The New Stack that Red Hat’s work is of the time and of the moment; primarily because “every GPU request now becomes a priority decision”, so one developer’s internal experiment cannot be treated with the same urgency as a financial close.

“Looking at the state of AI infrastructure players out there now, Red Hat has clearly seen that priority-aware multi-tenancy lets companies use expensive compute more efficiently, but efficiency without isolation is just a faster way to create a security and reliability crisis,” Estrin says.

“Every GPU request now becomes a priority decision.”

He thinks that the winners in this market (he tags Nvidia, Nutanix, Suse with Rancher, HPE Ezmeral and VMware Cloud Foundation under Broadcom as usual suspects) will be the organizations that can “share capacity while still proving what happened where” in live production.

“That means proving whose workload actually executed and ran, who had access, what it cost, and what happens when demand spikes. Those answers rarely come from the infrastructure diagram; they get settled in the boardroom, usually after someone’s critical workflow has slowed down, but regardless, this sums up where AI infrastructure is now,” adds Estrin.

What are the elevated multi-tenancy pain points?

To unpack what’s happening here, let’s remind ourselves that GPUs are expensive, obviously. As organizations move to live production use cases of agentic AI, they will want to maximize their GPU state’s ability to serve multiple workloads across multiple teams, multiple customers, multiple apps, and so on. 

This all means that the breadth of AI infrastructure efficiency becomes the new agentic bottleneck.

The priority-aware services above for native multi-tenancy on shared GPU infrastructure are important right now; this function dynamically allocates GPU capacity based on workload priority. When lower-priority workloads can be run on spare (or cheaper) capacity (rather than separately provisioned GPU resources having to be spun up for lesser jobs), then everyone gets to go home earlier on Friday.

“The breadth of AI infrastructure efficiency becomes the new agentic bottleneck.”

But that’s not all the balls being juggled here; Red Hat mentioned isolation too, and that’s a concurrently complex AI infrastructure discipline challenge. GPU compute resources managed through isolation techniques enable AI services to run without accessing or interfering with another service’s data, models, or compute environment. 

In other words, this combines hardware consolidation with strong tenant isolation.

What new Red Hat technologies are on offer?

Red Hat says this release lets developers verify models before deployment through EvalHub, enabling risk-focused safety benchmarking and regulatory compliance certifications. 

New observability dashboards give platform teams metrics for a real-time view of inference health, GPU utilization, and AI model performance. Non-admin users can access dashboards for per-user token consumption showback (another term for token tracking) and distributed inference workloads.

Also new is shared GPU control for multi-tenant inference. So-called “fair-share GPU scheduling” manages resource allocation across tenants, while priority-aware serving provides admission control and priority-based request routing to protect real-time inference. As suggested above. it also allows background workloads to use available capacity.

VP of product management at Nutanix, Anindo Sengupta, tells The New Stack that running multi-tenant AI at scale does indeed require secure tenant isolation.

“The essential isolation is best achieved through virtualization,” Sengupta says. “For specialized at-scale AI workloads, the choice could be to run Kubernetes on bare metal. On top of that, to create real value, agents need access to both LLMs that run on containers and enterprise systems (databases, business systems, etc.) that run on traditional infrastructure. For hybrid AI to run efficiently, the platform must manage both these environments in a performant way, with a common operating model.”

Observability & model-as-a-service showback

Built-in observability and MaaS showback in Red Hat’s latest release are present to provide per-user token metering, performance dashboards for models and agents, MLflow visual agentic tracing, and GPU utilization dashboards for operational and usage transparency.

For efficient GPU memory management, the general availability of CPU offloading and the developer preview of storage offloading allow models to handle longer conversations and larger documents without additional GPU hardware. 

Red Hat AI Hub also introduces agent templates and starter kits with pre-configured reference implementations for common enterprise patterns, including code review, document processing, and research workflows.

Red Hat is hoping its Red Hat AI 3.5 version release gets us to a state where GPU-based AI resources are viewed as a policy-controlled infrastructure pool.

Power without interconnect bandwidth is botched.

As enterprise AI pilots succeed and initial results show some returns, IT teams must then address the need to deliver at scale. But scaling AI across the business demands the same operational rigor as any mission-critical infrastructure: verified safety before deployment, precise resource controls across shared GPU environments, governed agent behavior and transparent usage metrics.

Yoram Novick, CEO of sovereign AI edge cloud provider Zadara, has previously been on the record on this exact topic. He has said that when teams need to scale AI, “Simply adding more GPUs without ensuring adequate interconnect bandwidth can lead to diminishing returns” in the modern AI era. 

Overall, with its ability to direct priority-aware inference, tenant isolation, capacity sharing, and observability, Red Hat hopes its Red Hat AI 3.5 release gets us to a state where GPU-based AI resources are viewed as a policy-controlled infrastructure pool.

The post Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots appeared first on The New Stack.

  •  

After nine years as HashiCorp CEO, Dave McJannet now wants to “unblock” enterprise AI agents

Retro 3D-rendered computer with a two-icon logo on screen, keyboard, and mouse on a purple background

Ask a traditional enterprise application for a customer address or today’s revenue figures and, broadly speaking, it follows a predictable route its developers have already mapped out: authenticate the user, query the right system, return the result. Given the same underlying data, you’ll get the same answer each time.

Ask an AI agent the same question, and the journey is much harder to forecast. It might consult one system, decide it needs more context from another, make a dozen tool calls, pass information through a language model and only then produce an answer. Run the same request again, and it may take a different route altogether.

And in an enterprise, what happens along that route can matter just as much as the answer: which systems the agent accesses, what data it sees, what actions it takes and how much it spends.

That distinction — between predetermined software, and applications that make probabilistic decisions on the fly — sits at the heart of a new company from a founder who knows a thing or two about bringing order to a new generation of infrastructure.

AI agents are hard to govern

Dome Systems co-founder David McJannet left HashiCop in August 2025
Dome Systems co-founder David McJannet left HashiCop in August 2025

Dome Systems was co-founded at the turn of the year by David McJannet, who spent close to a decade leading Terraform-creator HashiCorp through the cloud era, culminating in its blockbuster 2021 IPO and subsequent $6.4 billion sale to IBM in 2025. McJannet is joined at the helm by Marc Holmes, who spent more than six years at HashiCorp as chief marketing officer.

In an interview with The New Stack, McJannet lays out his company’s thesis on AI agent governance, arguing that enterprises are now running into the same kind of problem that they did with cloud infrastructure: adoption comes first, then the real spadework begins of putting the right controls in place across security, operations and finance.

“It’s actually a very different architecture, and that is what unlocks the power of these new [agentic] applications.”

Part of the challenge, he says, is that agents are built very differently from the enterprise applications of yore, which companies spent years learning how to control.

“It’s actually a very different architecture, and that is what unlocks the power of these new [agentic] applications,” McJannet explains.

He points to self-driving cars as an example: a model takes in live inputs and interacts with the vehicle’s systems as conditions change, because no developer can reasonably pre-program every possible situation a car might encounter on the road.

“It’s making judgments along the way, as opposed to trying to look up the historical maps of the world and make a real-time decision,” McJannet continues.

An enterprise agent can behave in much the same way: call one tool, assess the result, decide it needs another, and keep going until the task is complete. That flexibility lets agents tackle work that would be difficult to script exhaustively in advance — but it also makes their behaviour harder for enterprises to govern.

And this gets to the heart of what McJannet is striving for with Dome.

Table stakes for the agent era

The company launched out of stealth back in April with $14 million in seed funding, with McJannet having departed HashiCorp the previous August after the IBM transition concluded.

Dome’s starting point is that an agent combines three things: code, a model, and the backend systems or tools it interacts with. Bringing those pieces together under one platform, McJannet says, is “table stakes” for applying meaningful constraints to what the agent can do.

“If you don’t have an integrated platform, you can’t enforce controls across everything that the agent is doing,” McJannet says.

“If you don’t have an integrated platform, you can’t enforce controls across everything that the agent is doing.”

And so Dome’s platform is built around those three elements. An agent registry keeps track of the agents themselves; an MCP gateway controls the tools they can call; and a model broker/router governs which models they can use and how requests are routed.

The setup starts by registering the agent and giving it an identity, establishing who is allowed to call it, and connecting the backend tools it can reach — Zendesk, in this example.

Dome registers an agent, verifies its caller and connects the tools it can use.
Dome registers an agent, verifies its caller and connects the tools it can use.

Next, Dome connects a model provider, groups available models into a pool with routing and failover rules, then combines the agent, its tools and its models behind a single gateway. That gateway becomes the point through which Dome can apply the policies governing what the agent is allowed to do.

Dome connects a model provider, creates a model pool and brings the agent behind a gateway.
Dome connects a model provider, creates a model pool and brings the agent behind a gateway.

Once those pieces are connected, teams can set permissions on each call, use guards to inspect responses, apply quotas to cap spending, and keep a common audit trail across the agent’s activity.

Today, McJannet says, enterprises are often piecing all of this together themselves. A standalone model broker might be brought in to control spending, while a separate tool gateway handles security and operational concerns. Some are then building their own agent registry to tie those systems together.

Moreover, buying those capabilities separately leaves enterprises with another integration problem to solve. A model router might govern one part of an agent’s activity and a tool gateway another, while the agent itself continues moving between them.

“If you just provide the tool gateway or just the model router, it doesn’t allow you to have this kind of system of control,” he says.

That is also where Dome’s latest move enters the fray. After spending its first months in early access, the company is now opening the platform to self-service users for the first time, allowing teams to sign up with little more than a credit card, bypassing the typically arduous enterprise sales process.

Dome goes self-serve

Self-serve is relatively unusual route for this kind of enterprise infrastructure product. Dome is publishing its prices, offering a free tier and letting practitioners get started without first going through a sales process, while keeping the traditional enterprise route open for larger customers.

The thinking is partly about who McJannet expects to use the product. Rather than limiting access to buyers who are already deep into a procurement process, for example, self-serve enables individual practitioners to be able to discover, try and use the platform themselves.

“”We want to make the barrier as low as possible to have people come on board,” McJannet says, adding that Dome had already seen a number of self-service sign-ups ahead of the launch.

Separately, its pricing reflects a belief about where value will ultimately sit in this market. McJannet regards model routing and tool connectivity as baseline capabilities, with the more valuable piece being the controls that sit across the agent as a whole — think permissions, data redaction and spending quotas.

It’s also worth noting that while Dome’s main target user will be platform engineering teams inside large enterprises, typically working alongside operations and security, self-serve also creates an opening for another kind of user: the small company, perhaps even only one or two people, building an agent and trying to sell into an enterprise. The sort of scenario that aligns with the fabled one-person unicorn promised by many in the AI realm.

Indeed, McJannet says developers can get far building the application itself, only to hit a wall when a prospective enterprise customer begins its security and operations review. How is identity enforced? Who can see the data the agent reaches? What happens when it calls other agents? Can its activity be reconstructed afterwards?

Some builders, he says, have asked whether they can “certify” their agents on Dome because “my agent won’t get deployed until I can satisfy these infrastructure elements.” McJannet is careful to add that Dome doesn’t currently run such a certification program, but it’s clearly one route the company could venture down.

“If you register that agent on Dome, all the infrastructure elements are taken care of,” McJannet says.

‘Unblocking AI agents’: Lessons from the cloud era

That division between developers eager to ship, and enterprise teams worried about what happens after, is also where McJannet sees the strongest parallel with his years at HashiCorp.

During McJannet’s tenure, HashiCorp increasingly positioned itself around helping large organizations standardize how cloud infrastructure was provisioned, secured and connected. That included the 2020 launch of HashiCorp Cloud Platform (HCP), which offered its infrastructure tools as managed cloud services.

More broadly, McJannet’s account of early cloud adoption begins with developers swiping a credit card and deploying directly to Amazon because cloud infrastructure allowed them to build applications that had previously been impractical. The applications were compelling enough that enterprises adopted cloud despite resistance from operations and security teams, and what followed was a second phase: companies needed common services for provisioning, credentials, networking and other controls before cloud could become routine across the organization.

Platform engineering teams became the people responsible for reconciling those two demands: allowing developers to build while giving security, operations and finance enough control to permit those applications into production. McJannet believes agents are now creating the same tension.

“You’ve got this queue of cool apps that developers build that the ops and security teams are just not comfortable letting flourish in their environments.”

“You’ve got this queue of cool apps that developers build that the ops and security teams are just not comfortable letting flourish in their environments,” he says. “And so, inevitably, it has to go that same direction where the platform engineering team has to figure out [a way] to get to say ‘yes’.”

Dome’s bet is that enterprises will eventually prefer one system spanning the entire agent to a patchwork of gateways, routers and security products. In McJannet’s telling, that common control layer is what gives enterprises a way to limit how far an agent can roam while still letting it act autonomously.

“You have to have this control layer that provides this corridor where we can constrain the behavior of that new type of application architecture,” he says. “Because without that, you cannot unblock the deployment of AI applications.”

“That’s the part that we’re trying to answer — how do we unblock agents at scale?”

There is still plenty for Dome to prove. The company isn’t naming customers at this stage; McJannet says none of the enterprises it has worked with are yet willing to be identified publicly, though he says Dome has spent the past eight months talking to dozens of them.

Ultimately, McJannet believes the cloud era showed that new applications only become commonplace once enterprises have the controls to let them through. Dome is his attempt to solve that problem for agents.

“I think that’s the part that we’re trying to answer — how do we unblock agents at scale?”

The post After nine years as HashiCorp CEO, Dave McJannet now wants to “unblock” enterprise AI agents appeared first on The New Stack.

  •  

Your agent context needs a development lifecycle

Abstract 3D digital visualization of dark cyan and black geometric blocks, representing AI agent context and software architecture.

Skills, agent configurations, prompt instructions, and rules files. These artifacts now determine what your coding agents produce. They shape every line of generated code, every architectural decision, every convention the agent follows or ignores. They are, functionally, software.

Nobody treats them that way. Teams write a skill, commit it to a repo, and never test whether it still works after a model update. Agent configurations are copied and pasted across teams without versioning. Rules files drift out of sync with the codebase they describe. When something breaks, the signal is a developer noticing weird output and complaining on Slack.

“If context is the new code, what is its software development lifecycle?”

Patrick coined a framework for what’s missing: the Context Development Lifecycle. The CDLC is not about context window management or fitting more tokens into a prompt. It’s about managing the quality of the pieces that go into the context window. Is a skill up to date? Does the model actually react to it correctly? Are you providing context the model already knows? If context is the new code, what is its software development lifecycle?

The four phases of the context development lifecycle

The CDLC has four phases that map directly onto what we already do with code.

Generate is where everyone starts. Writing skills, building prompt configurations, setting up agent rules. It’s the equivalent of writing code, and it’s where most of the time goes today.

Evaluate is testing. Checking whether the linting on the front matter is correct or whether the syntax is too long. At the sophisticated end, you run scenarios: load a skill, ask a specific question, check whether the agent produces the expected result. You test across models and versions. You check whether you’re writing in context the model already knows, which wastes tokens. You verify that the skill activates on the right trigger words.

This is a TDD loop for context. Write the skill, write the scenario, check the output, iterate.

Distribute is shipping. At the simplest level, it’s committing a skill to a repo. At the mature end, teams publish skills to an installable registry with versioning, discoverability, and access controls. Pasting a skill into a Slack channel is not distribution, the same way emailing a .jar file is not dependency management.

Observe is production monitoring. Is the skill being used? Is it producing the right results? How many turns does the agent take before a developer intervenes? Where are developers overriding the agent or correcting its output? It’s observability for your context.

Don’t skip the testing

The maturity curve here is identical to what happened with software development practices over the past two decades. Organizations generate and distribute first. They skip evaluation entirely. They ship skills to production, meaning to the developers using them, and wait to see what happens.

It’s directly comparable to teams skipping test-driven development, despite being told to do it.
They don’t know the pain, so they go immediately to production.

The pain arrives when a skill works on one model version but breaks on the next, or when it triggers on the wrong question and gives a developer confidently wrong instructions. When a convention that the skill enforced was correct six months ago, but the codebase has since moved on, these are the same failure modes we see in untested code. Regressions, false positives, stale assumptions.

“You cannot scale code quality by asking humans to review more carefully. You scale it by investing in the guardrails that codify your standards at both ends.”

Every codebase has patterns that AI consistently gets wrong. Convention blindness, hallucinated APIs, cargo-cult code, over-engineering. At Aviator, these are called Invariants, or the AI slop register. Both the skill register and the AI slop register exemplify the same underlying principle at both ends of the development lifecycle: catalog your engineering standards and feed them to agents.

Before code generation, that means skills. At code review, it’s a catalog of patterns AI consistently gets wrong in your codebase. The slop register informs automated checks that catch what slipped through. You cannot scale code quality by asking humans to review more carefully. You scale it by investing in the guardrails that codify your standards at both ends.

From 1x to 50x

A developer who optimizes their own agent loop gets better individual results, but the improvement stays with them. When they fix a skill, nobody else benefits. When they discover a failure mode, nobody else learns from it. The ROI is 1x.

Patrick frames the scaling question using two metrics that sit atop traditional DORA measures.

The first is human touch: how often does a developer need to intervene in a given agent workflow? Every intervention is a signal that context is missing or wrong. Reducing human touchpoints is a direct measure of how autonomous your agentic coding loop actually is, and it’s often correlated with cost, since more turns mean more agent spend.

“Reducing human touchpoints is a direct measure of how autonomous your agentic coding loop actually is.”

The second is the reuse multiplier: how many developers benefit when you improve a single skill? If one developer fixes a skill and only they benefit from it, that’s 1x. If that fix goes into a shared registry and 50 developers get it, that’s 50x. 

These two metrics together force an organization toward shared infrastructure. You can’t reduce human touches at scale without shared, well-tested context. You can’t get a reuse multiplier without distribution and versioning.

Your platform team already knows how

The organizational structure for this already exists. Platform teams have spent a decade building the infrastructure that enables development teams to ship code reliably: version control, CI/CD pipelines, artifact registries, security scanning, dependency management, and access controls. The playbook transfers almost directly.

What a platform team does for code repositories, it does for skills. Provide a registry. Configure access control and group permissions. Set up the evaluation infrastructure. Run security scanning and report findings. Build the dashboards that show which skills are performing well and which are degrading. Track ownership so that when a skill breaks after a model update, there’s someone responsible for fixing it.

What the platform team does not do is write the skills or fix them when they break. The team that owns the domain owns the skill. The platform team provides the governance layer and the tooling, the same division of responsibility that works for code.

“Don’t build the tool. Build the tool that builds the tool.”

The orphaned-skills problem is already emerging. A developer writes a skill, shares it, moves to another team, and now nobody maintains it. A model update breaks it, and the platform team inherits the problem by default. This is orphaned GitHub repos all over again. The solution is the same: ownership policies, maintenance requirements, deprecation paths.

Patrick draws the layers concisely: “Don’t build the tool. Build the tool that builds the tool. The platform team builds the tool for people building the tool that builds the tool.

Closing the loop with observability

The least developed phase in most organizations is observation, though it’s the most important. Without it, there’s no learning system. You’re generating and distributing context manually, hoping it works, and fixing things when somebody files a complaint.

Agent observability is still early. Standards are forming. Agent MD is broadly adopted. Skill and plugin standards are newer. The tooling isn’t mature, but the pattern is clear: instrument your agents, centralize the signals, and analyze them across teams.

We’ve argued before that production feedback loops are the missing piece in AI-assisted development. When something breaks in production, trace it back to the change, identify the error category, and feed it back into both the prompting and verification layers. The CDLC observability phase extends that idea from code quality into context quality. The signals collected from agent logs, developer corrections, and turn counts feed directly back into generating better skills, writing more targeted evals, and distributing improved versions.

Self-improving agentic development

The fully closed loop looks like a system where agent logs feed into analysis that identifies gaps. Those gaps generate new skills or updates to existing ones. The updated skills run through evaluations before they ship. They distribute through a registry with version control. And the cycle repeats.

Patrick is realistic about the end state. The “dark factory” vision, where agents produce code with zero human involvement, is what he calls “a noble direction, but a risky game.” The teams that get closest — the ones that can confidently say that they don’t read the code anymore — are the ones that invested heavily in the context, testing, and observability infrastructure that makes their agents reliable enough to need fewer human touches per cycle.

The post Your agent context needs a development lifecycle appeared first on The New Stack.

  •  

The 3 roles AI agents play in your developer platform

Dark abstract 3D metallic waves illustrating complex AI agent roles in developer platform architecture.

Engineering organizations are trying to deliver as fast as technology allows, bringing agentic AI into their developer platforms and working out how to use AI agents to maximize engineering productivity.

But when it comes to AI agents in an agentic developer platform, every team we talk with sees the role of those agents a little differently. Following hundreds of calls with our customers, we captured three types of roles.

Role 1: AI agents as platform consumers

In this role, an agent is basically a user of the platform. It uses the platform as part of its task to read context and to run actions. When agents first showed up, a lot of companies said they treat their AI agents like employees, and in a development platform, that makes the agent just another engineering resource consuming it.

Diagram showing an AI agent reading context and running actions from an agentic SDLC platform.

A typical case: an engineer asks Claude Code to add an endpoint to the payments service. Before it writes any code, the agent pulls the service owner, dependencies, and the standards it must meet from the platform, then spins up a preview environment via a self-service action and runs the tests.

“A lot of companies said they treat their AI agents like employees, and in a development platform that makes the agent just another engineering resource consuming it.”

For this to work, the agent has to reason over real, current information about your systems, starting with the service catalog and extending to ownership, dependencies, standards, and current state. If you get the context wrong, there’s a good chance an agent will get overconfident and do the wrong thing. Most teams solve this one agent at a time by providing local context. But then the same information ends up connected to agents in fragile ways, not to mention that none of them are governed. Compare that to a context lake that provides every agent with a single governed source of truth. The platform also has to be reachable the way an agent works, which means being API- and MCP-first.

What does it require from the platform? An API and MCP-first interface, a governed context layer the agent reads from, and a set of self-service actions it can call.

Role 2: AI agents as internal platform components

Platforms that can register agents and run them within workflows are using AI agents as part of a full business process. The agent runs within the platform, is triggered by an event rather than requested by a person, and sits in the orchestration engine alongside the deterministic steps.

Diagram showing an AI agent as an internal platform component

Run a nightly scan that flags vulnerable dependencies across 40 services. The platform pulls the remediation agent from the registry and runs it once per service, so every owning team wakes up to an open PR awaiting review.

What does it require from the platform? An orchestration layer to run the agents, a registry to pull the right agent from, an identity per agent so the action is logged against the agent rather than a borrowed human credential, and a human-in-the-loop step where the risk is significant.

Role 3: AI as a resource with its own lifecycle (a.k.a AgenticOps)

In this role, the agent is a resource like any other, as are the LLMs, MCP servers, and the skills that come with it. The platform provisions them, governs them, and hands them back, just as it does with a service, a database, or an environment.

Diagram showing AI agents as a resource with a lifecycle, being requested by an engineer.

For example, say an engineer needs an on-call triage agent. They pick the model, the tools, and the environment it runs in, either through a form or by describing what they need, and the platform provisions everything needed, a bit like a vending machine, with the addition of a well-governed agent, in the right standards. 

“That makes it a golden path problem. A golden path is the route that, by default, gets a team a resource the right way.”

That makes it a golden path problem. A golden path is the route that, by default, gets a team the right resource, and the agent lifecycle needs one: request it, get it provisioned and registered, and publish it for the next team. It is also what customers ask us for most, with an agent and skill registry raised by 47% of the organizations we spoke with through early 2026. I went into this in more depth in our golden paths post.

What does it require from the platform? A self-service path that provisions the runtime, issues the identity and scoped credentials, wires in the approved context, and registers the agent on the way out, plus a route to publish it for the next team.

The three roles at a glance

RoleWhat the agent isExampleWhat it requires from the platform
Role 1: AI agents as platform consumersA user of the platform, reading context and running actions as part of its taskClaude Code pulls the service owner, dependencies, and standards, then spins up a preview environment and runs the tests.A governed context layer, such as a context lake, and self-service actions it can call.l
Role 2: AI agents as internal platform componentsA step inside a workflow, triggered by an event rather than requested by a personA nightly scan flags a vulnerable dependency across 40 services, and the remediation agent opens a PR for each one.An orchestration layer to run agents in, a registry to pull the right agent from, an identity per agent, and a human in the loop where the risk is real
Role 3: AI as reusable building blocks (AgenticOps)A resource the platform provisions, governs, and hands backAn engineer requests an on-call triage agent, picks the model and tools, and gets one back already registeredA self-service path that provisions the runtime, issues the identity and scoped credentials, wires in the approved context, and registers the agent

Sometimes the three roles connect

The chained case is an interesting one. An engineer requests a triage agent through role 3; it’s added to the agent registry as part of the creation workflow, and a week later, an incident workflow calls it as a component (role 2). When it runs, it reads service ownership and recent deploys out of the same context lake, which is role 1. It is the same agent throughout, and which role it is in depends on when you look at it.

“It is the same agent throughout, and which role it is in depends on when you look at it.”

What does a platform that covers all three look like?

That is what we built Port for. An agent can use Port as a user through a service account, reading the context lake and running self-service actions. Agents run within Port workflows as part of a business process. And AgenticOps runs as self-service workflows, so a team can request an agent and receive a registered one. All three sit on the same catalog, the same context, and the same audit trail.

If you want the full picture, it is in our playbook, From Agentic Chaos to an AI-Native SDLC.

The post The 3 roles AI agents play in your developer platform appeared first on The New Stack.

  •  

Microsoft just released Agent Lightning v1.0. Here’s why it matters for platform engineers.

Glowing purple and blue waveforms flow across a dark gradient background.

Agentic reinforcement learning has been suffering from a disconnect, an uncoupling, and a misarticulation. The polarity arises from how a training engine handles resource management, compared with how a post-training live production harness does.

Microsoft wants to ensure the production harness is engaged from the start to oversee infrastructure services and agent interactions, during both initial training and subsequent reinforcement learning actions.

Redmond’s Microsoft Research division first introduced the Agent Lightning framework in August 2025 as an infrastructure concept for agent optimization to address structural challenges in post-training LLM-based agents as they enter reinforcement learning processes. Microsoft subsequently launched the Agent Lightning v1.0 release with a commit tagged on GitHub on August 16. 

Using Agent Lightning v1.0 on what Microsoft has called “modest compute” 6K training examples, reinforcement learning improves Qwen3.5-9B on OpenAI’s SWE-bench Verified benchmark from 41.8% to 56.4%, an absolute 14.6-point gain.

“Using Agent Lightning v1.0 on ‘modest compute’ 6K training examples, reinforcement learning improves Qwen3.5-9B on OpenAI’s SWE-bench Verified benchmark from 41.8% to 56.4%, an absolute 14.6-point gain.”

Who owns the interaction loop?

In traditional agentic reinforcement learning, the training engine owns the interaction loop, i.e., steps including observing the environment, selecting an action based on the policy, executing the action, receiving a numerical reward, storing and updating the policy… and so on. 

In harnessed agentic reinforcement learning, the harness owns context construction, tool execution, and the agent–environment loop. The training engine observes only a sequence of LLM request–response pairs across a service boundary. Therefore, developers do not need to reimplement the agent loop within the training environment.

According to a collected group of software engineers at Microsoft, because an AI model harness owns the loop that governs infrastructure access and operations procedures during harnessed agentic reinforcement learning, it “introduces challenges” including retokenization (re-segmenting text into new tokens during active training), sample merging, advantage calculation, loss normalization, and training backend scheduling.

All of which, if not addressed, can result in ineffective or unstable training.

How Agent Lightning v1.0 turns the tables

“[In Agent Lightning v1.0] the harness, rather than the trainer, owns context construction, tool execution, and the agent–environment interaction loop, while the training system observes and optimizes the resulting model calls across a service boundary. This formulation preserves the harness’s deployment-time context policy, tool protocols, and execution semantics without requiring its agent loop to be reimplemented inside the RL framework,” explained the Redmond team.

For coding agents, the team has said that it finds existing reinforcement learning frameworks provide limited support, including a lack of data and complete training scripts, and a reliance on large-scale computational resources. To address this gap, with Agent Lightning v1.0, Microsoft provides a “complete data-cleaning pipeline” and reproducible training scripts built on open-source datasets and models.

Is this the end of the training time liability?

For users inside Microsoft environments, this might feel like good news. Machine learning specialists and platform engineers may have a good harness or a good set of reinforcement learning tools; they would rarely have both.

Rather than having to hard-code retokenization, advantage calculation, and reward shaping every time they wanted to train an agent to call external services in reinforcement learning procedures, they can keep their existing agent architecture as an asset, rather than treating it as a training-time liability.

Nebraska-based software engineering researcher Md Rashedul “Rashed” Hasan tells The New Stack that training on the exact production harness matters, but not only for its efficiency and benchmark gains. 

Training through the real harness keeps semantics intact

“It reduces train–serve mismatch,” Hasan says. “If you train inside a simplified trainer loop and deploy inside a different harness, tool protocols, context policy, and recovery behavior can all drift. Training through the real harness keeps those semantics intact, so gains are more likely to transfer to production behavior, not only to a lab environment.”

Hasan thinks that Microsoft has clearly named the paradigm, kept the core framework small, and shipped what he defines as a “concrete coding-agent pipeline” with open data and scripts. 

“In terms of who this will appeal to, it’s application and platform engineers who already have a production agent harness (such as a coding assistant or support-triage agent), and want to improve the underlying model with reinforcement learning without rewriting deployment logic to fit a training framework. Reinforcement learning and machine learning platform teams would also use it when they need a thin, reproducible testbed for harnessed agentic reinforcement learning,” Hasan clarifies.

He agrees that the SWE-bench Verified lift with only 6K examples is “useful proof” that harnessed reinforcement learning can move a hard-coding benchmark without forcing teams to reimplement their agent stack within the trainer.

“Environment setup for coding agents, reward design, evaluation fidelity, and the retokenization, sample-merging, advantage, and loss-normalization details are still easy to get wrong. Adoption will also depend on whether teams can integrate this proxy pattern into existing orchestration, observability, and safety controls,” cautions Hasan.

“Environment setup for coding agents, reward design, evaluation fidelity, and the retokenization, sample-merging, advantage, loss-normalization details are still easy to get wrong. Adoption will also depend on whether teams can integrate this proxy pattern into existing orchestration, observability, and safety controls.”

Killing train-serve skew, the oldest & most expensive bug in machine learning

Colorado-based data science professional Priyank Jain tells The New Stack that training through the production harness to strip the framing away is really all about “killing train-serve skew”, which is the oldest and most expensive bug in applied machine learning.

“Models rarely blow up in production because the math was wrong,” Jain says. “They blow up because the training setup quietly lied to them about what production actually looks like. Training through the same harness that serves the agent is the right instinct, and honestly it’s overdue. The accuracy bump is nice, but the real win is that the thing you optimized is finally the thing you shipped.”

In terms of what Microsoft is getting right here, Jain says it shows Redmond wants to “meet developers where they already are,” i.e., allowing users to improve an existing agent with reinforcement learning without rewriting their deployment just to please a training framework.

“That’s a genuinely good call, and it opens this up to folks who aren’t reinforcement learning specialists. What I’d worry about is that it also lowers the bar for people to run reinforcement learning they don’t fully understand. The hard part was never wiring up the loop; it’s designing a reward that survives contact with a model actively looking for the path of least resistance. Make that part easy, and you’ll get more people optimizing the wrong thing faster,” Jain advises.

“3500 lines of code is small enough that an infrastructure engineer can actually read it before trusting it… and that’s a good thing.”

Just 3,500 lines of core Python code

Built with “simplicity as its first principle”, the entire framework consists of around 3,500 lines of core Python code.

Infrastructure engineering developer and founder-developer of Tooldex, a platform that autodiscovers MCP servers across LLM agents, Ria Banerjee, tells The New Stack that the simplicity element here is a positive, i.e. 3500 lines of code is small enough that an “infrastructure engineer can actually read it before trusting it,” and that’s a good thing.

“But in terms of who would actually use this tool, the honest answer is fewer teams than the framing suggests,” Banerjee says. “Realistically, you would need a GPU cluster and a Kubernetes cluster under your control. An app developer who built a support-triage agent on LangChain isn’t running that. So the real audience is platform teams that already employ infrastructure engineers.”

“Because the harness is where the behavior actually lives, but you’re now baking your harness’s quirks into the model weights… so you change your retry logic next quarter, you’ve silently shifted what the model was trained on. So now, you have to version your harness,” adds Albuquerque-based Banerjee.

Microsoft has released the complete workflow and training scripts to facilitate reproducible, harnessed agentic RL in Agent Lightning v1.0 on GitHub under the MIT license, including data cleaning and reward-hacking prevention.

The post Microsoft just released Agent Lightning v1.0. Here’s why it matters for platform engineers. appeared first on The New Stack.

  •