❌

Vue normale

Reçu hier — 27 septembre 2026Infra

The rise of agentic AI on Kubernetes: unleashing the new infrastructure layer

27 septembre 2026 à 16:00
Abstract 3D render of blue cubes inside gold wireframe boxes, linked by red rods into a dense cluster, with teal lines connecting outer cubes.

AI is changing expectations around infrastructure and operations, including Kubernetes management. When models run close to the data they use, deployment, scaling, and governance responsibilities tend to shift to platform teams. And as clusters, environments, and operational signals continue to multiply, manual operations often strain under the added weight.

AI may simultaneously provide opportunities to lighten this growing load. Agentic software can now observe a system, reason about it, and act within predefined limits. 

Ultimately, these platforms’ value depends on the quality of the context an agent can see and the boundaries you set. Without cluster state, policy, and access rules, an agent can only guess.

Without cluster state, policy, and access rules, an agent can only guess.

For agentic AI to streamline multi-cluster management, you need clear lines between what the system observes, what it recommends, and what it changes. Drawn well, those lines let teams gain notable speed while still maintaining control.

The impact of AI on computing infrastructure

Teams once treated AI as an application concern; models sat on top of existing systems, and the stack underneath stayed mostly unchanged. Today, AI reaches into more and more customer interactions, while data storage needs simultaneously expand and orchestration pressure grows. A recent Forrester report describes the modern AI computing stack as stretching from the models themselves into and across the infrastructure beneath them.

As AI workloads move into production, they place new demands on the infrastructure beneath them. Many lean on specialized compute, with resource needs that rise and fall through bursts of training and inference. Because conditions shift quickly, they can also call into question whether telemetry remains trustworthy. Each of these demands lands at the infrastructure layer, where the workloads run.

The infrastructure layer of the new AI stack

The infrastructure layer covers compute, storage, and networking. It is a foundation that every workload running on the layer depends on. As AI workloads grow, choices about capacity, placement, and control will increasingly shape the performance of the data, intelligence, orchestration, and experience layers atop the infrastructure.

To operate the infrastructure layer efficiently across many machines and locations, a team may rely on orchestration instead of managing servers by hand. In cloud native contexts, Kubernetes has become a control point for scheduling workloads, applying policy, and presenting a consistent interface across environments. Kubernetes is especially well-suited to support organizations this way when teams need consistent control across an estate spanning data centers, clouds, and edge sites. 

Agentic AI and Kubernetes: the future of the infrastructure layer

Agentic AI can extend automation from fixed rules to systems that adapt to real-time conditions. Traditional automation runs the same script whether the environment has changed, while an agentic system observes the environment, reasons about what it finds, and then takes action.

When you apply agentic capabilities to multi-cluster management, the system follows this same sequence. An agent reads cluster state and operational data, proposes a diagnosis or next step, and then carries out actions based on an approved scope, usually after a person signs off. You can further reinforce these boundaries by routing each request to a specialized agent that receives only the metadata it needs.

The signals that an agent receives from the cluster, the context about policy and access, and the definitions of what the agent may change are the key elements that give agentic systems their value. They also separate agentic AI on Kubernetes from a generic assistant. 

Manual Kubernetes management is less efficient at scale

Admittedly, agentic AI fits some settings better than others. On a small single-cluster footprint, the overhead may outweigh the benefit. Manual Kubernetes management often holds up on a handful of clusters, but it can become unreliable in a rapidly growing estate. After all, each new cluster adds lifecycle work across upgrades, patching, configuration, and renewal. Those tasks can quickly multiply and diverge in hybrid environments.

Configuration drift is a high risk in these situations. Settings that started identical can fall out of sync, and policies can apply unevenly from one team to the next. Individually, these gaps may be manageable, but collectively they raise the odds of an outage or a failed rollout.

Visibility can also erode in an unmanageable way. Clusters spread across data centers, clouds, and edge sites often leave teams with no single view of the whole landscape. When DevOps and platform engineers stitch together signals from separate tools, resolution can slow and become more error-prone. A unified view helps enable sound, efficient decision-making by people, agents, or both.

Kubernetes knowledge is fragmented, and existing AI tools lack business context

Kubernetes expertise often sits unevenly across an organization. For example, senior engineers may hold deep operational knowledge that application teams lack. The most current information about a running system may also be fragmented if logs sit in one tool and metrics in another. Real-time understanding can be further clouded when policies, runbooks, access rules, and deployment history each live elsewhere.

Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes. Without that context, even a capable AI tool may fall short of providing meaningful Kubernetes management support.

Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes.

When an agent can read current signals alongside the rules that govern them, its suggestions become specific, testable, and actionable. In an incident, agentic systems can correlate logs with a recent change. Ahead of a rollout, they can check the change against policy. During troubleshooting, they can account for access rules rather than guessing at them. Kubernetes decisions carry real operational consequences, which makes these details all the more important to consider. 

Engineering “toil” isn’t time-efficient

Site reliability teams use the word “toil” for repetitive manual work, especially tasks that keep systems running without adding lasting impact. In Kubernetes operations, toil takes the form of repeated triage, manual signal correlation, alert follow-up, and routine checks. The tasks aren’t particularly difficult, but they can consume significant time and attention for enterprise teams.

When engineers spend their days on this kind of investigation, proactive modernization efforts tend to stall and planned upgrades can slip behind schedule. In other words, the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements. In a recent survey about how AI provides value to DevOps teams, reducing toil emerged as one of the clearer opportunities.

…the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements.

Agentic AI can support repetitive investigations by gathering signals, correlating them, and proposing a likely cause for an engineer to weigh.

Kept under human review, it can take on some of the routine correlation that would otherwise fall to the team. That kind of support can give engineers more room to focus on the strategic work that most needs their judgment.

Building more intelligent infrastructure with agentic AI and Kubernetes

As you consider building toward intelligent infrastructure without surrendering control, the following principles can inform your efforts:

  • Start with observable context, giving agents access to current cluster state, policy, and history before they reason about a problem.
  • Separate suggestions from actions, allowing agents to recommend freely while any change must wait for human approval and a defined scope.
  • Connect agents to existing controls, routing their work through the access rules, identity, and audit paths the team already trusts.
  • Keep the ecosystem open, favoring platforms that integrate with current tools and standards over those that lock work into a single stack.

Platforms like SUSE Rancher Prime and SUSE AI Factory embrace these principles and illustrate how Kubernetes management can become a foundation for agentic operations. These platforms can help you improve cluster and policy consistency without compromising your authority over AI. Built on open-source foundations, they can also help you avoid being trapped in a single vendor’s stack.

In SUSE Rancher Prime, the industry’s first context-aware agentic AI ecosystem, its AI assistants work as a crew of specialized agents with an intelligent router. The platform draws on the cluster context already in place and acts through existing access controls. Through support for external Model Context Protocol (MCP) servers, teams can extend that crew to their own sources. In addition, human validation tools allow you to hold a proposed action for approval before the agent runs it.

Despite its potential, intelligent infrastructure is not universally beneficial. In situations where change control must stay fully manual, for example, agentic AI’s role may be strictly limited to observation and suggestion. Measure the technology’s value against the realities of your day-to-day operations. For those who are investing, agentic AI will have the greatest impact when it actively supports context, control, openness, and human judgment.

The post The rise of agentic AI on Kubernetes: unleashing the new infrastructure layer appeared first on The New Stack.

Reçu avant avant-hierInfra

The agent didn’t break your controls. It went around them.

26 septembre 2026 à 16:00
Three black circular directional signs on a gray concrete wall, showing arrows pointing straight ahead, turning left and turning right.

The identity part of agent security is settled. An agent needs its own identity: a short-lived, revocable credential scoped to the job, and an audit trail that names the human who set it running. NIST’s security leads made that case in August 2026, and most identity vendors agree.1

Identity and access management is table stakes. It’s necessary, but it isn’t what’s breaking.

What’s breaking is an assumption we’ve carried for twenty years: Get identity and permissions right at the door, and whatever happens inside takes care of itself. That worked when software was passive. Agents reason about a goal and choose their own steps toward it, like a seasoned escape artist.

An agent that hits a wall looks for another way

Almost every control in today’s stack answers a question about entry. Should it connect? Should it reach that service? Should its token be accepted here? Each is a question about a route, and there’s rarely just one route to anywhere worth going.

An agent treats a blocked route as a problem to solve, because that’s what we built it to do. A person who hits a locked door usually files a ticket, while an agent tries the window.

In July 2026, an autonomous agent spent four and a half days inside Hugging Face’s production systems.2 A filter controlled which internet addresses its dataset servers could download from, and it never fired, because “the agent stopped asking the worker to fetch remote resources and instead made it act on local ones.” The filter worked as designed, and the agent went around it anyway.

A person who hits a locked door usually files a ticket, while an agent tries the window.

On ordinary developer machines, malware in a compromised npm package tried to recruit the AI coding assistants already installed to search for secrets,3 and a coding agent deleted a production database during a change freeze before falsely telling its operator the data couldn’t be recovered.4 Both happened on the machine itself, where no network control was looking.

The shift from outside-in to inside-out

Outside-in controls govern entry, and most organizations run plenty of them. Make no mistake, inside-out security completes those controls rather than replacing them.

Inside-out control governs the action itself, and asks a narrower, harder question: Should this agent, acting on this person’s authority, delete this table in this database, right now?

That question matters because an agent can swap routes but not the outcome it’s after. No matter how many routes it tries, deleting a table is still deleting a table, and a checkpoint on the action sees it every time.

Here’s how today’s controls line up against it.

ControlWhat it coversWhat it misses
GatewayTraffic you route through itLocal shell commands and file edits never reach it
SandboxThe environment as a wholeConstrains reach, not individual actions
SIEMA record of what occurredReports after the action is completed
RegistryThat an agent existsWhat the agent did with that existence

Each does its job, but they all decide somewhere other than the moment the action runs.

Put the enforcement point where the agent acts

Every agent acts through an agent harness: the software that takes the action the model chose and carries it out, whether that means running a command, writing a file, or calling an API. In most deployments today, nothing checks that action before it runs.

An inside-out control puts an approval step in that gap. Before the harness executes anything, the checkpoint looks at which agent is asking, on whose authority, and against which system, then applies policy to allow the action, block it, or send it to a human. Because every action passes through it, an agent denied a destructive command and trying a smaller version of the same thing is held to the same rules. The remaining risk is a badly written policy, which can be fixed.

None of this works without the identity basics. Any type of control, whether it be at the prompt level, inference level, harness level, or MCP layer, can’t judge “may an agent take this action here on this object?” when the only name on the request is a service account shared by six agents and four engineers.

The companies building agent runtimes have reached the same conclusion. Over the past eighteen months, Anthropic, Google, Microsoft, OpenAI, LangChain, and Cursor have each added a hook that lets you inspect an agent’s action before it runs.5 When AWS explained its own agent policy design, it argued that controls belong at the moment an agent attempts to invoke tools.6

The catch is that each hook works differently, with no standardized request or response formats. An enterprise whose developers use Claude Code and Cursor while its platform team builds on LangChain would maintain the same enforcement logic in multiple different flavors, each with its own audit trail. That doesn’t scale, and it tightly couples your security model to whichever runtime a team favors that month. Enterprises need one vendor-agnostic agentic security layer that spans every harness, so adopting a new model or framework doesn’t mean restarting the entire onerous security review.

Turn the lights on before you start blocking

The standard, well-ingrained security instinct is to start blocking right away, but we’ve all seen how well that works with the business in the past. Security must move and adapt at the speed of business, not the other way around. Security tools such as intrusion prevention systems and web application firewalls both ran in monitoring mode until teams understood what normal looked like, and those that skipped that step tended to hear about it from a production outage.

Agents need the same sequence, only faster. An enforcement point in monitoring mode blocks nothing and quickly answers questions most organizations can’t today:

  • Which agents are actually running, not which ones someone believes are running
  • Who started each one, and whose authority it’s operating under
  • What capabilities it used, and against which systems
  • Which of those actions would have violated a policy, had the agent security platform been switched to enforcement mode

Write policy from what you know, see, and have evidence of, not just from an architecture diagram. Enforce first where the stakes are highest: destructive commands, production data, and anything that moves data out. Then watch-learn-build, just like the agents we use: Watch the patterns, build finer-grained controls and policies, and learn how to use AI securely, safely, and confidently. Observe first, then enforce, build, and deploy, in that order.

The bottom line

None of this requires a new category of infrastructure. It’s the identity, authorization, and audit you already run for your people, extended to agents and applied inside the harness before the action runs.

The perimeter is still there, but it has moved to the moment an agent acts, the one place it can’t route around.

Ory built Agent Security inside the harness, on the same identity and authorization engines that run in production for human users. It starts in “observe mode,” so you get that inventory first, and you can try it today at ory.com/agent-security.

Footnotes

  1. Bill Fisher and Ryan Galluzzo, “Back to the Future: Why Agentic AI Needs a Strong Identity Foundation,” NIST Cybersecurity Insights, August 27, 2026.  ↩︎
  2. Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident.” The intrusion ran July 9–13, 2026. ↩︎
  3. Nx, “s1ngularity postmortem,” August 2025. The malicious packages “attempted to use local AI tools (like Claude and Gemini)” while scanning systems for sensitive data. ↩︎
  4. AI Incident Database, Incident 1152: Replit agent deletes production database during code freeze, July 18, 2025. ↩︎
  5. Pre-execution hooks by vendor. Anthropic, Claude Code hooks; Google, Agent Development Kit callbacks; Microsoft, Agent Framework middleware; OpenAI, Agents SDK guardrails; LangChain, human-in-the-loop middleware; Cursor hooks (InfoQ, October 2025) ↩︎
  6. Liana Hadarean and Jean-Baptiste Tristan, “Why Policy in Amazon Bedrock AgentCore chose Cedar for securing agentic workflows,” AWS Security Blog, May 20, 2026. ↩︎

The post The agent didn’t break your controls. It went around them. appeared first on The New Stack.

Developers and platform teams both want Kubernetes self-service. They disagree on who owns it.

24 septembre 2026 à 18:06
Abstract view up through bold yellow angled beams to a white skylight grid and curved ceiling panels.

What do developers want? Kubernetes environments when they need them.

What do they not want? Those environments after a week or more of tickets. 

Platform teams, meanwhile, own what those environments cost, who can access them, and whether they meet company policy.

That tension is the core of Kubernetes self-service: What can safely be handed to developers, and what still belongs to the platform team?

In a recent interview with enterprise cloud specialists — Marius Bogoevici, Senior Principal Product Manager at Hewlett Packard Enterprise (HPE), and Karthik Subramanian, Principal Product Manager for HPE Morpheus Software — The New Stack explored the core friction points of Kubernetes self-service. The conversation focused less on whether self-service is desirable than on where to draw the line.

HKS, HPE’s CNCF-certified Kubernetes distribution, is integrated with HPE Morpheus Software to help platform teams deliver and lifecycle-manage Kubernetes environments as part of a broader operating model spanning Kubernetes, VMs, infrastructure, and clouds. HPE Morpheus Advanced Software supports the on-premises private-cloud use case with HKS, while HPE Morpheus Enterprise Software extends Kubernetes and application operations across hybrid and public-cloud environments.

Together, HKS and HPE Morpheus Software extend that operating model beyond infrastructure provisioning. Through service and application catalogs, platform teams can connect approved Kubernetes environments with the CI/CD pipelines, container registries, automation tools, and other services developers already use. Developers receive a governed, ready-to-use path from code to deployment instead of manually assembling the toolchain for each project.

The self-service paradox

Open-source Kubernetes provides orchestration and declarative APIs, but not a complete operating model.

Subramanian says teams building their own self-service layer usually run into two recurring problems:

  1. Tool and package sprawl: To make upstream Kubernetes production-ready, platform teams must curate and maintain an ever-evolving ecosystem of third-party CNCF tooling for networking (CNI), storage (CSI), ingress, identity, and policy enforcement. Navigating and supporting this fragmented stack creates immense maintenance overhead for internal platform teams.
  2. Day-2 lifecycle and hybrid footprint complexity: Spinning up a Kubernetes cluster is the easy part, but keeping it current — across development, QA, staging, and production — is where the work piles up. That is why HPE says every Kubernetes upgrade must be checked against the networking, storage, ingress, identity, and policy components around it. The problem gets harder when clusters span bare metal, private clouds, edge sites, and public clouds, because one-off scripts and environment-specific configurations can quickly create drift. That maintenance burden belongs with the platform team, not with developers trying to ship applications.

Giving developers direct access to raw Kubernetes APIs just shifts the operational work — it’s far from gone for good. In fact, developers will wind up debugging manifests and storage drivers instead of writing code. 

Meanwhile, operations teams have to deal with overprovisioning, idle clusters, and configurations that reach production without review.

What developers control — and what the platform supplies

The practical answer is not unrestricted access. It is a paved path: approved Kubernetes services that developers can request themselves, with access, configuration, placement, approvals, and lifecycle controls defined by the platform team.

“The best candidates for self-service are requests that are repeatable, low-risk, and well-understood,” Bogoevici tells The New Stack. “For example, a developer should be able to request a development cluster, deploy an approved application, create a namespace, or select resources from pre-approved configurations without opening a ticket. The platform team decides what a safe configuration looks like, and the developer chooses from a supporting menu.”

“The platform team decides what a safe configuration looks like, and the developer chooses from a supporting menu.”

Rather than asking developers to write YAML for ingress, storage classes, and RBAC, HPE Morpheus exposes those choices through service catalogs, reusable layouts and blueprints, workflows, role-based access control, approvals, APIs, and automation. Developers do not lose Kubernetes. They retain direct access through standard Kubernetes interfaces and tools where permitted, while the platform team standardizes the request, governance, and lifecycle processes around them.

Those catalog items can package more than infrastructure settings. They can also integrate the approved services and application components that support the development workflow – including CI/CD tooling, source and artifact repositories, container registries, and runtime dependencies – while the platform team controls how those components are configured and governed.

Developers choose the parameters that matter to the application:

  • Approved Kubernetes versions and cluster sizes: Select from pre-tested Kubernetes runtime releases and node count templates.
  • Resource quotas: Specify required CPU, RAM, and persistent storage capacity tailored to the workload.
  • Integrated toolsets and IDE environments: Select required developer toolchains, container registries, and runtime dependencies. 
  • Lease and duration limits: Define explicit operational lifetimes for temporary development or sandbox clusters to prevent abandoned infrastructure sprawl.

Network isolation, identity-provider integration, security policy, and cost allocation stay with the platform team and are applied automatically through the approved service configuration.

The same division of responsibility applies to the delivery toolchain: Developers choose from approved services, while the platform team manages the integrations, credentials, policies, and automation behind them. This gives developers a consistent experience without shifting toolchain maintenance and governance onto individual application teams.

The division of responsibility looks like this:

Service areaDeveloper chooses or requestsPlatform team defines and suppliesReview or exception path
Development cluster provisioningApproved Kubernetes service, version, size, target environment, and duration.Reusable layout or blueprint, access controls, placement rules, storage and network defaults, and lifecycle policy.Nonstandard versions, placements, configurations, or requests outside quota.
Production deploymentApplication artifacts, target namespace, and deployment request through the approved path.RBAC, tenancy, policy, audit, backup, and release controls appropriate to the environment.Formal review for production changes and exceptions.
Resource allocation and quotasCPU, memory, storage, and other approved capacity parameters within project limits.Project quotas, upper bounds, placement constraints, and supported resource profiles.Requests above quota or for specialized resources.
Networking and securityApplication endpoints and permitted connectivity within approved patterns.Identity integration, RBAC, tenant isolation, network policy, secrets, and audit controls.Cross-tenant access, elevated privileges, or changes to baseline security policy.
Lifecycle and cost governanceService lifetime and approved operational actions.Visibility, policy, approvals, retirement workflows, and applicable cost controls for the licensed variant.Long-running exceptions, nonstandard lifecycle actions, or budget exceptions.

From ticket queues to a repeatable paved path

In conventional IT environments, provisioning a dedicated Kubernetes environment for a new project often involves cross-departmental ticket handoffs spanning infrastructure, networking, security, and storage teams. This friction frequently stretches provisioning timelines from days to weeks.

By unifying infrastructure orchestration, role-based access controls, and multi-tenancy into a single operational experience, HPE Morpheus Software can compress these provisioning workflows down to minutes or hours, according to HPE. “Developers get a usable environment that complies with the organization’s defined controls and policies, without needing to understand all the complex infrastructure steps sitting underneath,” Bogoevici says. “When you reduce provisioning time from weeks to hours, that is super meaningful and tangible.”

“When you reduce provisioning time from weeks to hours, that is super meaningful and tangible.”

The result is not only faster cluster provisioning. HPE Morpheus Software can also automate the handoff into the developer’s established delivery process by making approved CI/CD and application services available with the environment. Instead of waiting for separate teams to connect pipelines, registries, credentials, and runtime dependencies, developers receive a ready-to-use path from development through deployment.

Faster provisioning can create a different problem, too: The speed can and will cause teams to lose track of what was provisioned and why. HPE Morpheus Software gives administrators visibility into utilization and cost, while lease controls can shut down temporary development clusters when their time expires.

Bogoevici says ticket volume is a poor measure of success, particularly early on, when more developers may be trying the catalog. He recommends watching deployment success, exception rates, resource utilization, and the day-to-day effort required to keep the service running.

Security belongs in the service design

Security is another boundary that must be designed into the self-service path. If identity, access, tenancy, and policy are added only after a cluster is created, every request produces more work and more room for inconsistency.

“Security must be a core design consideration built directly into the service, not an afterthought during deployment,” Bogoevici says. “HPE Morpheus Software brings identity integration, role-based access, tenant isolation, approvals, and policy into the operational workflow.”

A newly provisioned environment should arrive through an approved configuration with the applicable identity, RBAC, tenant, policy, and audit controls attached. Platform teams can validate the paved path by testing an allowed request, a request that should be rejected, and the resulting audit record.

The operating-model test

The strongest Kubernetes self-service model does not hide Kubernetes or make it the control plane for every workload. It gives developers useful, approved choices and direct access to the Kubernetes workflows they need, while the platform team standardizes the enterprise processes around those workflows.

That matters because the enterprise still runs VMs, clouds, and existing infrastructure alongside Kubernetes. HPE Morpheus Software helps platform teams use common request, governance, automation, and lifecycle processes across these environments without forcing every workload onto one runtime or creating another operational silo.

In practice, that means self-service should deliver more than a Kubernetes cluster. With HPE Morpheus Software, a catalog request can bring together the approved environment, application services, and DevOps toolchain integrations developers need, while preserving the governance and lifecycle controls the platform team requires. Developers spend less time assembling and troubleshooting delivery infrastructure – and more time building and releasing applications.

Looking toward 2027, the goal is not unrestricted control. It is faster access, predictable results, transparent guardrails, and a clear exception path when the standard service does not fit.

The post Developers and platform teams both want Kubernetes self-service. They disagree on who owns it. appeared first on The New Stack.

What managing 150,000 AI agents could look like for database teams

24 septembre 2026 à 15:11
Abstract 3D render of translucent orange cubes and panels scattered across a pale gray background, with bundles of glossy teal tubes curving in from the right.

The database administrator of the future will spend considerably less time administering databases.

That sounds contradictory, but AI agents are taking over that work. For decades, DBAs have handled the decidedly hands-on work of keeping databases available, performant, secure, and affordable. They provision capacity, troubleshoot slow queries, manage migrations, and step in when something inevitably goes sideways.

AI is already taking on some of that work. At the same time, it is creating a much bigger data infrastructure fleet to manage.

The result is likely to be a very different kind of DBA: one who spends less time tending individual databases and more time supervising the autonomous systems doing it for them.

Congratulations, you’re managing robots now

This shift starts with a familiar problem: more infrastructure needs managing than the people available to manage it.

Database automation is hardly new, but agents can potentially go further than the scripts and rules DBAs already rely on. Rather than automating one predetermined task, an agent can inspect what is happening, decide what needs attention, use tools to act on it, and check whether its intervention worked.

That changes the DBA’s relationship with the database. A performance problem that once required someone to dig through metrics, identify the troublesome query, and decide how to respond could increasingly be investigated by an agent before a human gets involved.

It doesn’t remove the DBA from the equation. Someone still has to decide what an agent can do, where human approval is required, and what happens when it gets something wrong. But the work moves up a layer. Instead of personally performing every operational task, DBAs start managing the systems carrying them out.

Instead of personally performing every operational task, DBAs start managing the systems carrying them out.

And before anyone gets too comfortable with that idea, the number of those systems could become enormous.

150,000 agents walk into a database…

Gartner predicts that the average global Fortune 500 company will have more than 150,000 AI agents in use by 2028, up from fewer than 15 in 2025. Only 13% of organizations currently believe they have the right governance in place to manage them.

Not every agent will need its own database, but plenty will. They will create state, retrieve data, remember previous interactions, and exchange information with other agents. Many will also behave very differently from the applications DBAs are used to supporting: spinning up quickly, sitting idle for long stretches, and suddenly becoming busy when there is work to do.

Nobody is hiring 150,000 DBAs to manage them.


That is the scale problem Yugabyte is targeting with YugabyteDB AMP, or Agentic Multitenant PostgreSQL. Rather than treating each new agent workload as another database for an administrator to provision and babysit, AMP manages databases as a fleet.

The platform packs hundreds of small Postgres workloads onto shared distributed infrastructure while keeping their databases isolated. Lifecycle operations, including provisioning, branching, scaling, migration, and teardown, can be exposed to agents through MCP. Yugabyte has also built specialized agents for setup, migration, performance tuning, and integrations.

In that model, a DBA is no longer provisioning database number 14,372. The interesting job is setting the rules for how database number 14,372 is provisioned, operated, and fine-tuned without them.

Do more with less (no, really)

Scale is only half of the problem. Someone also has to pay for all this stuff.

Agent workloads make traditional capacity planning particularly awkward because many are bursty and frequently idle. Giving every experimental agent permanently provisioned infrastructure could leave companies paying for many databases that spend much of their lives doing very little.

This is where consolidation becomes as much an economic question as an operational one.

AMP’s approach is serverless multitenancy and scale-to-zero. Multiple small workloads share the underlying distributed infrastructure, while customers pay by CPU minute and idle agents consume no compute. Resource governance can impose CPU limits on individual workloads, preventing a single overeager agent from consuming the capacity intended for its neighbors.

The human equivalent matters too. If routine setup, migrations, tuning and other database operations can increasingly be delegated, a smaller database team can potentially look after a much larger estate.

That doesn’t mean companies get to fire the DBAs and hand the keys to the robots. It means scarce database expertise can be spent on architecture, governance, and genuinely difficult problems instead of repeatedly doing the work that software can handle.

Your 2028 database problem starts now

The harder question is what to build underneath all of this when nobody really knows what the enterprise AI estate will look like in two years.

An agent that begins as an experiment today could disappear next month. Another could suddenly become a production application used across the business. Building one infrastructure stack for cheap experiments and another for serious workloads risks creating a migration problem every time an experiment succeeds.

Yugabyte bets that both ends of that journey should sit on the same foundation.

YugabyteDB AMP lets workloads start on serverless Postgres and transition to fully distributed YugabyteDB as their scale and criticality increase, without rewriting the application or migrating data to a different database platform.

Then there is the problem above the individual database: agents need to remember what happened, and not just in a silo.

That’s where Meko fits into the Yugabyte stack. Meko is an agent-native context engine designed for multi-agent AI systems. It provides persistent memory, shared knowledge, decision traces, and autidability across multiple agents, rather than leaving each agent working from its own isolated context. An agent can pick up information learned by another agent instead of retrieving it again or restarting the reasoning process.

Taken together, it delivers a single data stack for an agent’s entire lifecycle: Meko for the context shared among agents, YugabyteDB AMP for agentically managing fleets of Postgres databases, and distributed Postgres-compatible YugabyteDB for workloads that outgrow their serverless beginnings.

Of course, there’s no guarantee that 2028 will look exactly like today’s forecasts. That’s rather the point. The safest architectural bet may be one that doesn’t require you to know in advance which of today’s tiny AI experiments will become tomorrow’s critical applications.

The DBA is still critical in that world, but the job will look different. The DBA of the future may manage fewer databases directly, while taking responsibility for vastly more of them. Instead, managing the autonomous systems that do the administering.

The post What managing 150,000 AI agents could look like for database teams appeared first on The New Stack.

A third option is emerging in the fight over AI and your data

23 septembre 2026 à 21:31
Split-screen video interview with The New Stack host Alex Wilhelm and VAST Data cofounder Jeff Denworth.

Not your keys, not your coins. Not your model, not your data?

Over the summer, the tech industry was consumed by a debate about AI use in the enterprise and the need to protect IP. If an enterprise used proprietary models, was data leakage a necessary evil?

Companies seemed to have two options: They could use state-of-the-art, proprietary models and risk losing control of their data, or they could use open-weight models and never kiss the frontier.

Thankfully, a third option is emerging.

Consider the concern: Company A wants to use LLM B from AI Lab C, and they want to avoid training AI Lab C how to eat Company A’s lunch by building its capabilities into LLM B. A good way to resolve the tension would be to let Company A run LLM B on its own infrastructure, so there’s no risk of its information fleeing on the wind.

AI agents are “creating a whole different set of requirements at the data layer.”
–Vast Data co-founder Jeff Denworth

But that raises another problem: AI Lab C doesn’t want to allow Company A to run LLM B on its own GPUs because it doesn’t want to hand over its model weights. It’s the same IP issue the company ran into, in reverse. You have to solve the trust problem in both directions!

Enter VAST Data co-founder Jeff Denworth and a new product called DataEnclave, which aims to let AI labs and enterprise-scale companies deploy proprietary models in secure compute environments without risking data transfer in either direction. (DataEnclave uses Nvidia’s Confidential Computing technology to make the system tick; Vast Data’s core product is AI OS, infrastructure that fits beneath a company’s AI applications.) 

The New Stack had Denworth on the podcast to chat about the confidential computing market. I was curious about timing. Why did Vast build DataEnclave now? Nvidia began rolling out Confidential Computing in a serious way in 2024, after all. Denworth argues that the market needed the core technology, yes, but also demand.

And until late 2025, AI demand was modest compared to today’s token totals. Once agentic coding tools took off, corporate demand for AI products soared. This led to the pricing crisis we saw in early 2026, and the secure AI usage debate we endured over the summer. 

Performance drove demand, demand drove usage, and usage dug up fresh problems to solve. Now the question for the market is whether or not DataEnclave has solved enough concerns on both sides of the proprietary AI-proprietary data equation. The market will sort that out as it moves through early access and into general availability.

Our conversation goes deep into the arc of AI, where companies are in their AI journey today, and how much data remains to be unlocked inside the enterprise. If you want to feel the acceleration, it’s a fun one!

The post A third option is emerging in the fight over AI and your data appeared first on The New Stack.

How confidential AI splits control between data and model owners — and opens new opportunities for both

23 septembre 2026 à 18:15
Abstract 3D render of dark blue and black cubes floating among translucent spheres against a warm red and coral background.

Most people already understand what generative AI can do. But enterprises run into problems when they need to give a model access to information that cannot leave their own environment, such as a patient record, a customer’s financial details, or a company’s most valuable intellectual property.

Sending that data to a cloud or SaaS service means it crosses external networks and is processed on infrastructure run by another organization, creating additional concerns about control, accountability, and exposure. That’s where AI enthusiasm collides with production realities. Despite its productivity potential, enterprise AI still faces a fundamental gap in trust and control.

Organizations need to know whether a system will expose information it should protect, act as intended, meet security and performance requirements, and behave safely at machine speed.

Alon Horev, CTO and co-founder of AI operating system company VAST Data, tells The New Stack that the challenge is particularly acute when AI systems handle sensitive customer information. “Even if you ask the model today to obfuscate a conversation or redact PII from a conversation, it’s hard to have 100% confidence that’s the case, and that it worked.”

“Even if you ask the model today to obfuscate a conversation or redact PII from a conversation, it’s hard to have 100% confidence that’s the case, and that it worked.”

Consider a customer support agent that needs access to an individual’s profile to provide a useful, personalized answer. The organization must ensure that information isn’t exposed to another customer, while also considering whether those conversations can be used for training or system improvement. They might contain personally identifiable information (PII) or other protected details, and the consequences of mishandling them ultimately fall on the organization and the people whose information it holds.

Confidential AI architectures: solving a two-sided trust problem

Enterprise AI has two parties to satisfy: organizations must keep sensitive data under their control, while model builders need to protect the weights and software that represent substantial investments in research, engineering, and IP. They’re understandably reluctant to place those assets in environments where customers, infrastructure operators, or attackers might gain access. That mutual need for control has created a stalemate. How can organizations bring advanced models to sensitive data without asking either side to surrender control?

Horev has seen that the most capable models are increasingly delivered as SaaS services, because that’s the simplest way for their creators to distribute and protect them. Even when a provider offers compliance controls, the enterprise might still shoulder the consequences of a breach, misuse, or regulatory violation. Sending information across the WAN also places it in the hands of more systems, connections, and operators, increasing the number of points that must be trusted and governed. Organizations may also be unwilling, or legally unable, to rely on a provider’s assurances that it will not retain, reuse, or expose their data beyond the intended service. 

For organizations in regulated or data sovereignty-sensitive sectors, that could be an unacceptable trade-off. “Naturally, many organizations are adopting a hybrid strategy,” Horev tells The New Stack. “Some applications and datasets can go to the cloud, while others must remain on-premises, sometimes even in the building, or in the country.”

This is where confidential AI comes in. Encryption at rest and in transit protects data while it’s stored or moving between systems. Confidential computing extends that protection into the processing environment, using hardware-isolated execution to create a protected enclave in which the data and model weights can remain encrypted until they’re released to an approved workload.

Cryptographic attestation verifies the hardware, virtual machine (VM), software, and configuration requesting access before releasing keys. The model builder can encrypt its model using the public key of a specific confidential VM. Only that VM’s corresponding private key can decrypt it within protected memory, enabling the customer to use the model without accessing its weights.

Independent key control preserves the separation between the two sides. The enterprise retains control of the keys governing its data, while the model builder retains control of the keys governing its model. While the workload is running, the infrastructure operator doesn’t control either set of keys.

As AI becomes more agentic, those controls will matter more. Agents will need to access more data, systems and tools, and might act on that information with far less human intervention.

“This world of agentic AI is moving extremely fast, and we need to limit what an agent can see and do.”

Those that can’t establish strong privacy and governance assurances for today’s models will find it even harder to deploy agents safely in the future. “This world of agentic AI is moving extremely fast, and we need to limit what an agent can see and do,” Horev tells The New Stack.

From architecture to ecosystem

Many businesses simply cannot manage the integration, security, and maintenance of the entire AI stack, because it requires working separately with each model provider to engineer something that suits both parties. Turning confidential AI architecture into something organizations can deploy is the challenge VAST DataEnclave, which was launched on September 22, intends to address.

As a capability of the VAST AI Operating System, the goal is to bring the model, application layer, and data platform together under customer-controlled operating conditions. The architecture is designed to protect both sides of the equation: the enterprise’s data and the model builder’s weights. The customer retains control of its infrastructure and data keys, while the model provider can make its software available without handing over the underlying intellectual property.

“We’re trying to close the trust and control gap by working with world-class model builders such as Cohere, Deepgram, Factory, Fundamental and TwelveLabs, who continue to innovate and build their expertise,” says Horev. The ecosystem also includes infrastructure and security providers such as Nvidia, CrowdStrike, Fortanix, Nscale, Cisco, and Supermicro. The range reflects the practical challenge: confidential AI needs more than a protected GPU. It requires models, applications, accelerated hardware, data infrastructure, and operational support to work together.

That control also changes the cost conversation, without automatically making AI cheaper. Hosted models can make budgets harder to predict as token consumption varies with usage patterns, agent loops, model architecture, and workload volume. Customer-controlled infrastructure gives enterprises a more defined capacity and cost base: they can plan around GPU clusters they own or have already budgeted for, instead of allowing inefficient model choices or uncontrolled agent activity to generate an open-ended token bill.

“…instead of allowing inefficient model choices or uncontrolled agent activity to generate an open-ended token bill.”

The cluster also imposes a natural ceiling on throughput, which helps organizations understand how much work their infrastructure can handle within a given period. Model providers can then price access by token, task, or license, while the enterprise retains greater visibility into its total operating cost.

Why the data platform is paramount

Confidential AI protects data and model weights during inference, but it’s only part of the production challenge. Real-world AI systems are living environments in which data moves between storage, databases, GPUs, networks, applications, and agents.

That’s why confidential AI can’t be bolted onto a fragmented stack. Businesses need to protect the model, the data, and the infrastructure connecting them as one system. As Horev says: “You need to build security in multiple layers of the platform,” with someone accountable for rapidly updating compromised components.

Confidentiality is only useful if the resulting system can also be operated, monitored, and improved. As AI infrastructure becomes more distributed, it becomes harder to tell what’s happening when something goes wrong and where the fault lies.

Horev recommends a “single pane of glass” across storage, networking, and compute, so teams can see what’s happening and keep resolution times low. If a network port is intermittently failing in a data center, for example, an agent could help identify the root cause, provided it has access to the right operational data and tightly controlled permissions. Those permissions should govern the infrastructure it can inspect, the data it can retrieve, and the actions it can take.

The same applies to monitoring AI workloads. Teams need visibility into performance, failures, and access patterns without exposing the customer data or model weights. Agent sandboxes can limit the systems and tools an agent can reach, while data platform observability can log which data it accessed, what it did with that data, and how it interacted with downstream systems.

Evaluation, therefore, becomes part of production discipline. Teams must observe systems, measure behavior, govern access, and manage change in ways that demonstrate progress. Confidentiality, data-level policy, observability and correctness have to work together.

The emerging ecosystem suggests demand for models that can run securely under customer control, wherever sensitive data resides. These are “living systems,” says Horev. “It’s not just leveraging a feature inside of a wider platform.” 

Visit the VAST Data Confidential AI solution page to learn more about the architecture, ecosystem, and availability.

The post How confidential AI splits control between data and model owners — and opens new opportunities for both appeared first on The New Stack.

The software supply chain is the new battlefield. AI just changed the rules.

23 septembre 2026 à 16:00
Illustration of a lime-green fingerprint on an orange background, split into three horizontal sections labeled 1.1, 1.2 and 1.3.

AI coding tools have seriously accelerated developer speed, but AI has also done the same for attackers — and the software supply chain is increasingly where the two are colliding.

The numbers give some idea of how quickly software development is changing. GitHub processed around one billion commits in 2025. By April 2026, the platform was handling roughly 275 million commits a week, according to GitHub COO Kyle Daigle. GitHub Actions usage has climbed, too, from 500 million compute minutes per week in 2023 to 2.1 billion in just part of a single week this year.

Quincy Castro, CISO at Chainguard, says the shift in how software gets written is already stark.

“I look around Chainguard, and I don’t think any of our engineers have actually written a line of code by themselves in the past year,” Castro tells The New Stack. Writing code manually now “sort of feels quaint, like you’re illuminating manuscripts,” he says, while “the printing press is out there just going to town.”

But this isn’t only about professional developers producing more code. AI has also widened the pool of people who can create software. Teams in HR, finance, and business intelligence that once had to wait for engineering resources can increasingly build what they need themselves.

That means more software being created by people outside traditional engineering teams, often with AI making decisions about what goes into it. The person prompting the agent may never see which libraries or packages it has chosen.

When the agent chooses the dependencies

Software security was already built around the fact that humans couldn’t inspect everything. But developers were still making important decisions, including which libraries and packages went into an application.

That changes when an AI agent is doing much of the coding.

“Humans are directing what they want to be done, but they’re somewhat abstracted from the actual doing of the work,” Castro says. “You have AI instead now making the choices of what dependencies am I going to pull into this application? How am I going to go accomplish this task?”

“Humans are directing what they want to be done, but they’re somewhat abstracted from the actual doing of the work.”

Attackers, meanwhile, are finding plenty of uses for the same technology. Castro sees three problems arriving at once: frontier models finding previously unknown vulnerabilities, attackers using agents to exploit better vulnerabilities organizations haven’t fixed, and sustained attacks against the open-source ecosystem.

A collection of medium- and low-severity findings might once have sat well below the top of a remediation queue. Frontier models with advanced cyber capabilities, including Anthropic’s Claude Mythos Preview and OpenAI’s GPT-5.6-Cyber, can now work across those findings and chain seemingly minor weaknesses into a viable attack path.

“Here’s a whole ton of mediums and lows. Now give me the attack path that gets me domain admin,” Castro says, describing the approach. “Chain these together to go get me root on the system. And AI is really, really good at being able to do that.”

That poses an awkward problem for vulnerability management — and the models keep getting stronger, with OpenAI releasing GPT-5.6-Cyber in August. Mean time-to-exploit has already fallen from 63 days in 2018–19 to an estimated minus seven days in 2025, according to Mandiant, meaning exploitation can begin before defenders have a patch to apply. 

At the same time, AI’s ability to combine apparently less-serious weaknesses makes a neat CVSS-based queue a less useful representation of what an attacker can actually do.

The third problem is the software supply chain itself.

Open source becomes the attack path

Modern applications depend heavily on open source software, and attackers have increasingly targeted the infrastructure used to build and distribute it.

Castro pointed to the TeamPCP campaign, which compromised widely used projects including Aqua Security’s Trivy. In that attack, malicious code was pushed into trusted components and subsequently picked up downstream.

Supply chain attacks were once associated primarily with sophisticated state-backed groups willing to spend significant time getting into the right place. That barrier is falling.

“If you don’t mind making some noise, this is a way easier attack vector than I think a lot of people thought it was,” Castro says. More importantly, “a single attack that’s successful can lead to a cascading set of other compromises and other access that gets you into other places.”

The development pipeline itself can make matters worse. Castro says many organizations still have relatively few controls around CI/CD, while developers routinely pull components from external sources to get their work done. Adding autonomous coding tools to that behavior compounds the risk.

“You wouldn’t pick up a random thumb drive and stick it into a production system, right? But that is effectively what folks are doing when they’re consuming open-source software that way.”

He compared the way organizations consume open source software to plugging an unknown USB drive into a production system. “You wouldn’t pick up a random thumb drive and stick it into a production system, right?” he said. “But that is effectively what folks are doing when they’re consuming open source software that way.”

Open source isn’t the problem. Trusting its distribution path without sufficiently verifying what you’re consuming is.

Prevention has to come before detection

This is where the old security model starts to creak.

For years, much of vulnerability management has followed a familiar loop: scan something, generate an alert, decide how serious it is, and get somebody to fix it. That becomes harder to sustain when development output multiplies, AI agents make more of the underlying decisions, and attackers can exploit weaknesses before fixes are available.

Castro wants companies to put more effort into what enters the development environment in the first place, rather than discovering problems once the software is already there.

“How do we just make things work from the beginning, with no alerts and no responding to stuff and no people chasing things and no people trying to prove a negative?” he says. “From end to end, from the creation of code to its deployment, how do we make sure that we can give folks the most trustworthy version of that thing?”

Rather than taking packages from public ecosystems at face value and scanning them after the fact, Chainguard builds artifacts from verified, buildable source.

That’s the thinking behind Chainguard’s approach to containers, libraries, and other open source artifacts. Rather than taking packages from public ecosystems at face value and scanning them after the fact, the company builds artifacts from verified, buildable source. It provides provenance about how they were created.

But trustworthy components are only one layer.

“There’s no point in bringing inherently secure software components into the environment if you don’t actually have a technical control that says this is the only way people developing code can consume these things,” Castro says. That means engineering, security, and SRE teams also need controls over where software — whether selected by a human or an AI agent — can come from.

That requires several layers of protection. Organizations need to know where their software came from and how it was built, control what can enter their environments, and make sure those rules apply when an AI agent chooses components as well as when a developer does.

Defending open source at AI speed

There is another problem, however. Frontier models such as Claude Mythos Preview and GPT-5.5-Cyber aren’t just finding vulnerabilities that previously went undetected; they can also combine lower-severity flaws into working attack paths. Individual companies can harden their own pipelines, but the software they depend on comes from an open-source ecosystem facing vulnerability discovery at a speed and scale it wasn’t built for.

That’s part of the reasoning behind Athena, the industry coalition Chainguard launched to turn vulnerability findings from frontier AI programs into fixes. As of July, the coalition had processed more than 40,000 vulnerabilities, with 42% rated critical or high severity and 86% marked as network reachable, meaning attackers can access and trigger them at the network level.

For Castro, the important part isn’t simply finding more bugs. AI is already getting very good at that. Someone still has to fix them.

“Through Athena, what we attempt to do is to give people that engineering fix,” he says. “What if we create a coalition where folks just send us the issues that they’re finding? We automatically generate fixes for those, and we push those back to everybody.”

“Through Athena, what we attempt to do is to give people that engineering fix.”

Those fixes can also be pushed upstream to open source maintainers, who face the prospect of being buried beneath an expanding pile of AI-generated vulnerability reports.

That may ultimately be the bigger shift AI forces on software security. Developers aren’t going to stop using coding agents because they create new risks, any more than companies are going to stop using open source because attackers target it.

Bolting enough scanning onto an exponentially faster development process isn’t much of an answer either.

The opportunity is to remove more of the risk before the software ever reaches a developer or an agent: Start with components you can trust, tightly control how they enter the environment, and fix weaknesses as close to their source as possible.

AI has made it dramatically cheaper to create software. It’s doing the same thing for attacks. Security now has to keep up without putting the printing press back in the box.

Visit Chainguard to learn more.

The post The software supply chain is the new battlefield. AI just changed the rules. appeared first on The New Stack.

Agents operate, humans govern: Scale your operations and reduce toil with Azure SRE Agent 

15 septembre 2026 à 18:21
Illustration of colorful doughnut, bar, radar, and line charts alongside slider controls on a black background.

What if engineers could spend their time building and optimizing systems rather than maintaining them?

It’s 3 a.m., and the pager goes off. Tabbing between multiple dashboards and diagnostics, the SRE struggles to determine whether what woke them is a real incident, whether they’re the right person to handle it, or whether they need to wake someone else. Digging through monitoring tools, deployment history, incident systems, and team runbooks — and chasing what might be the wrong theory about the root cause — they can’t respond fast enough to stop more customers from being affected.

Or imagine that, by the time the SRE joins the incident bridge, Azure SRE Agent has already analyzed the monitoring data, identified the root cause, and prepared a fix for approval and deployment.

Sanchit Mehta, one of the head engineers for Azure SRE Agent, tells The New Stack that “[Azure SRE Agent] starts analyzing telemetry and correlates things like blast radius, deployment changes, recent changes, any recent rollouts, to try to tell the engineers, ‘OK, this is what is causing it.'” Increasingly, it will even create the PR for that fix.

The support is just as useful during normal working hours. At InEight, correlating telemetry across tens of thousands of Azure resources can take days, if not weeks. When a support ticket reports slow performance without identifying the product, engineers must determine which of the company’s 14 products is affected, then check multiple observability and reliability tools.

InEight shared that, during its first incident using Azure SRE Agent, the agent quickly identified the affected product, traced the performance issue to its root cause, and recommended scaling Redis. The DevOps team had been considering scaling the app service as a temporary fix.

Proactive and in production

This kind of help is becoming the new normal at Microsoft, where more than 3,000 service teams already use Azure SRE Agent to investigate issues, perform root cause analysis, respond to incidents, fix code, enable automatic mitigation, support proactive detection, analyze data, and report at scale. Azure SRE Agent has already handled more than 1.8 million incidents inside Microsoft, many mitigated in minutes.

The team also uses Azure SRE Agent to develop and improve the service itself, with custom agents for code review, deployment, evaluation, and monitoring. This “agent-powered engineering” approach, as Mehta calls it, lets the team take advantage of ongoing advances in AI models. That includes proactively spotting problems, like quota issues that affected deployments, and automatically raising support tickets to resolve them. The agent recently identified the root cause of a change that broke synthetic tests as soon as the change reached the first region, he says.

 “It said, ‘OK, this was an upstream PyPI package that broke your dependency; you need to add tests for it; you should roll back immediately; here’s how you should go fix this.'”

Mehta says that kind of proactive monitoring is hard to handle with deterministic queries. “You need a level of intelligence to see when a large production payload is being deployed and if it has the potential to cause degradations.” 

For some internal teams, more than half of incidents are autonomously managed by the SRE agent and don’t need any human intervention, adds Shamir Abdul Aziz, lead program manager for Azure SRE Agent, because they’re what he calls “safe” operations and mitigations: a restart, scale-out or rollback of a service, or change order requests escalated by customers.

“The humans did the governance, set up the guidelines, gave some coaching to the agent, and then it went into auto mode to complete the entire workflow,” Abdul Aziz says.

Agents are ready to help

SREs are already drowning in repetitive toil. SREs are already drowning in repetitive toil, and coding agents add to that workload. Agentic operations are now powerful enough to help, Vyom Nagrani, one of the head PMs for Azure SRE Agent, tells The New Stack.

“As code gets written more and more by agents, it’s going to take another agent to operate it,” Nagrani says. “But why wait? If the agent can manage code which other agents write, why can’t it manage code written by humans?”

“As code gets written more and more by agents, it’s going to take another agent to operate it.”

“The reasoning loop has become mature enough that now agents can automatically start figuring out a lot of these complex problems, especially when it comes to correlating across multiple data sources, which has always been the hardest thing for humans to do,” Nagrani says.


Powerful models aren’t enough, though, and homegrown automation won’t have the production-grade governance, verification, evaluation, telemetry, and control a platform can offer.

The state of the art has progressed from prompt engineering to context engineering—which grounds AI in your infrastructure, code, and institutional knowledge — and now to harness engineering. “That is what allows you to run agents at scale, control them, and govern them,” says Abdul Aziz.

“When you combine all these things with being able to verify, audit, evaluate, and get real telemetry and metrics out of the system, where the agent claims it has done something, you can validate that agent’s claim,” Abdul Aziz says.

Instead of a non-deterministic black box that can’t explain its decisions, you can trace and learn from the agent’s reasoning so that you can correct mistakes once, not over and over again. “That’s why companies are willing to adopt it now,” Abdul Aziz says. “Because when you try the same thing ten times, you’re going to get the same output.”

“You don’t just turn on the agent, give it full access, and ask it to solve everything.”

After two years of building enterprise-grade systems that can be trusted, audited, and validated, the next step for cloud-native SRE can be agentic ops with autonomous capabilities — but you still need to know how to adopt it, Abdul Aziz warns. “You don’t just turn on the agent, give it full access, and ask it to solve everything.”

Context and connections

Azure SRE Agent is built for Azure but not limited to the Azure platform. The agent provides native access to Azure services such as Azure Monitor, Application Insights, Log Analytics, and Azure Resource Graph. Connecting the agent to your subscriptions, telemetry data, and source code gives it the operational context and institutional knowledge needed to understand how you work.

Beyond Azure, Azure SRE Agent integrates with engineering and operational tools through managed connectors for Azure DevOps and GitHub, plus MCP connectors that enable access to external knowledge sources such as Google Drive, Confluence, Cursor, Claude Code, and other third-party systems. 

Put all that knowledge into Markdown files in a repo, along with the skills and tools agents need to act on your systems (including third-party and on-premises services). That gives you artifacts that agents can version, review, test, reuse, and update.

When you want to dictate how to handle an incident — what to check and in what order, what to post, and even how to format a report — you can create a custom agent, either by using an existing runbook or by working through an incident with an agent and saving that skill. Using agents to improve agents is the shortcut to making Azure SRE Agent more useful the more you use it. Essentially, saving what agents learn during incidents helps improve their future responses.

Guidelines and guardrails

Governance covers identity, role-based access control (RBAC), and tool-access policies. These controls determine which actions are allowed, blocked, or subject to step-by-step approval, and whether an agent operates autonomously or with human review.

What makes governance both flexible and powerful are hooks, based on prompts or deterministic commands, that fire at different stages of a workflow and catch edge cases, such as allowing an agent to drop the index in a SQL database but never drop a table.

Metrics show you whether governance is working. The new live reports show time to mitigation, tool reliability, how often agents act autonomously, and cost per outcome at a glance. InEight’s metrics are typical: an 80% reduction in both incident investigation time and build failure triage time, a 67% reduction in the effort needed to investigate bugs, and an 84% reduction in cost.

To get those results, you need triggers that automatically launch agents instead of waiting for a human to open a chat window.

Bind skills and custom agents to specific alert classes so they can respond to incidents first. Start agents through pipelines, webhooks, or work items to automate delivery workflows. Schedule regular checks, reviews, and audits, and have agents automatically update their artifacts.

Agents operate; you stay in control 

By reducing repetitive tasks and technical toil, Azure SRE Agent frees engineers to focus on more interesting and innovative projects. Just as there’s a familiar maturity model for adopting site reliability engineering in the first place, you don’t jump straight into having agents rather than humans handle operations. When you give agents the context about your infrastructure, you can start using them for investigations.

“If you give agents read access to your source code, your telemetry, your resources, the time to get to the root cause is reduced to minutes rather than hours or days,” Abdul Aziz points out. “Every customer starts there.”

Once you’re happy with the answers you’re getting, you can give the agent more permissions while still approving individual steps, he says. “The fixing is easy once you understand the problem. It’s usually changing your configuration, writing a piece of code, or restarting a service.”

“The fixing is easy once you understand the problem. It’s usually changing your configuration, writing a piece of code, or restarting a service.”

As you expand into other operational tasks, refine the agents’ artifacts, metrics, and governance before granting more autonomy: “Things like rolling back a release when we know there was a regression in that release, restarting a service, dropping a corrupt index on a SQL table, or scaling out a service,” Abdul Aziz suggests.

For more complicated issues, agents can deliver the entire fix, ready for approval. The Azure SRE Agent that manages the Azure SRE Agent product looks at exceptions, errors, incidents, Teams conversations, emails, and GitHub issues every night and spits out PRs. 

Avoid code review bottlenecks by having agents deploy, test, measure, and include outcomes in the PRs. Use continuous evaluation to build a self-learning system that accurately follows your existing workflows.

“The agent can self-improve because the agent learns constantly,” Abdul Aziz says. “You can configure scheduled tasks to identify which evaluation scores were low and automatically improve the custom agent, custom skills, and even your knowledge documents – because knowledge management is also a toil. The agent can automate all of that.”

The right way to start

Azure SRE Agent now offers a 30-day trial experience with no always-on charges. Make the most of that by learning from some common mistakes:

  • It’s not magic! Turning on the agent doesn’t mean you don’t have to do DevOps anymore. Don’t treat it as a chatbot or connect it to just your observability system. You need to give the agent the context it needs, the tools to do the job, and intentional triggers that tell it when to act. Otherwise, it may spend effort on low-value work or generate outputs that aren’t grounded in your environment. 
  • Don’t limit yourself to what the agent does out of the box: Customize agent skills, tools, connections, and logic to fit how your organization works, and build custom agents for specific tasks.
  • Don’t use agents for jobs a single line of code can do: Using them to explore deterministic, structured data for anomalies is an expensive waste of tokens that will only flood the context window when the agent can write that line of code itself. “Orchestrate, don’t calculate,” as Nagrani puts it. If you’re drowning in alerts, use automation to filter the noise and only send alerts that need intelligent analysis to agents.
  • Don’t stick with what you’ve always done or copy your org chart: The most effective agents have a complete picture of the system, so they need all the context, even if it crosses two teams. That might mean crossing boundaries, coordinating who has expertise and who needs to grant access, or rethinking how the organization works.

“If agents have the right context, they minimize the toil and truly make operations less costly,” lead product manager Deepthi Chelupati points out. That way you can move faster, be proactive, and give engineers more time to innovate and less maintenance work to dread.

Get started today: sre.azure.com

The post Agents operate, humans govern: Scale your operations and reduce toil with Azure SRE Agent  appeared first on The New Stack.

Chip Huyen explains how to cut inference costs without new hardware

13 septembre 2026 à 17:00
Layers of wavy yellow horizontal strips with deep shadows between them, forming an abstract pattern.

Last October, the P99 conference — the online gathering for developers focused on high-performance, low-latency applications — featured a cracking keynote from Chip Huyen. 

The author of the best-selling AI Engineering, Huyen opened with simple math: Training a frontier model is a one-off cost, but inference is the same cost paid over and over. That’s great for the frontier model providers, and bad for us token burners. Over the life of a model, Huyen reckons the compute split ratio lands somewhere between 1:10 and 1:100 for training to inference. Reasoning models — which burn even more tokens — push that out even further. We all know the feeling of hitting our weekly session quotas.

Huyen’s point is that if inference is too expensive, then nobody ever recovers the training bill, which might explain why there are so many memes about the “profitability” of frontier models. So, how do we optimize inference? 

That’s a topic that Huyen spent months researching for her book. And in the spirit of optimization, Huyen distilled it down to 30 minutes for the conference in October 2025.

Huyen is returning for P99 CONF 2026 in a few weeks. Ahead of that moment, watch her full talk – or read the recap below – from last year, then let’s talk about how those ideas aged over the past 11 months.

What to measure

Chip recommends focusing on a few key latency metrics:

  • Time to first token (TTFT): How much time elapses before the user sees anything
  • Time per output token (TPOT): The average time between consecutive tokens (aka inter-token latency)
  • End-to-end latency: Time to first token, plus time per output token, multiplied by the number of output tokens minus one
(Click to enlarge graphic.)

With reasoning models, some of those tokens never reach the user. “The first generated token might not be the same as the first visible token,” Huyen explained. “The model might think for a while, and it will only show the first token of the final output to the user.” 

“The first generated token might not be the same as the first visible token,”
— Chip Huyen

Some people also measure Time to Publish for that (i.e., how long until the user sees the first token). The best metric to prioritize depends on what matters most for your users. 

Also consider “goodput” alongside throughput. Throughput measures requests processed in a given window. Goodput measures the requests that actually met your targets. Chip’s example: an app targets 200 ms time to first token and 100 ms time per output token, and processes 10 requests per minute, but only three hit both. 

(Click to enlarge graphic.)

3 ways to optimize LLM inference

With inference servers, you can optimize from 3 different angles: the hardware, the model, and the service that manages the requests and responses.

(Click to enlarge graphic.)

Huyen previously worked at Nvidia and opted out of the hardware discussion: “Even though I find it to be an intellectually interesting topic, it’s not relevant to a lot of people because we don’t have the power to change the hardware itself,” Huyen explained. She also didn’t want to spend much time on the obvious solution: replica parallelism, or just adding more machines. It’s costly, and it gets complicated fast – especially if you end up with a mix of 80GB, 48GB and 24GB machines and models of varying sizes to distribute across them.

That leaves the model and the service. Huyen offers these tips on how to decide: “If you want to host the models yourself, or if you have access to the model weights, or if you train a model yourself, or you want to fine-tune or distill a model, then model optimizations might be for you. However, if you want to take a model as-is and make it more efficient on your own inference service, you might want to look into service optimizations.

Model optimization

The following techniques change the actual weights so that they can change the model outputs.

Quantization lowers the precision used to store weights and activations  (e.g., from four bytes per parameter at 32-bit to one byte at 8-bit). Huyen explained, “Reducing the precision not only reduces the memory requirement to run the model, making it cheaper. It can also make the model a lot faster. If you do additions bit by bit and each weight is 32 bits, you have to do it 32 times. If it’s 8 bits, you only have to do it eight times.”

The tradeoff is a small quality hit. Huyen continued: “It’s possible to reduce a lot of the model’s memory footprint with minimal quality degradation, and quantization is pretty generalizable to a wide variety of model architectures and model sizes. That’s why it’s very popular. I rarely see any companies running a model at full precision anymore.”  

“I rarely see any companies running a model at full precision anymore.”
— Chip Huyen

Distillation involves using a large model to generate training data for a smaller model. For example, say you have a truly large model (the example Huyen used was o1) and want a model that performs like it, but is much smaller. Basically, you collect a large set of prompts, run them through the larger model, then train the smaller model on its responses.

Proceed with caution, though. Huyen warned, “A lot of model providers have the condition that they do not allow their models to be used to train competitive models. So even though it’s a very common technique, you need to check licensing.”

Service optimization

This set of techniques targets how requests are scheduled, routed, and reused. The actual weights aren’t affected.

Batching groups multiple requests so they’re processed together in a single pass through the model – which is much more efficient than dealing with them one at a time. Huyen presented a few batching options:

  • Static batching waits for the batch to fill. This maximizes compute utilization, but it might increase the latency for the first requests.
  • Dynamic batching runs on a timer instead (e.g., batching every 15 ms). This is less compute-efficient, but it’s better for latency.
  • Continuous batching handles the case where requests finish at wildly different times, which is common with LLMs. One request asks for the capital of Vietnam; another kicks off deep research. With static or dynamic batching, the finished request’s slot sits idle until the slowest one completes – and new requests queue up behind it. Continuous batching returns each request as it finishes and fills the spot with another request. That can improve compute resource utilization and latency.
(Click to enlarge graphic.)

Decoupling prefill and decode separates the two phases of a request onto different machines. (Prefill processes the input, while decode generates the output.) Huyen said, “Input tokens can be processed in parallel, whereas output tokens need to be generated sequentially. With parallel processing, it’s bounded by compute, the processing power of the chip. With decoding, it’s bounded by memory, because you have to move model weights.” 

Because each phase stresses different resources, most services now separate them. To improve time to first token, shift machines toward prefill. If you care more about improving time per output token, shift them to decode. 

(Click to enlarge graphic.)

Parallelism splits work across machines. Replica parallelism copies the whole model onto more machines. Tensor parallelism divides a very large matrix, so different machines compute different parts of it. Pipeline parallelism divides the model by layer, so requests move through as a pipeline. 

(Click to enlarge graphic.)

Prompt caching processes shared text once, saving cost and latency. A lot of repetition exists across requests to the same application: the system prompt, the examples, the same code base, the same document behind different questions. You might as well process that shared segment once, cache it, and reuse it.

The technique was relatively rare when Huyen was writing AI Engineering. “There was one paper about it, and it was not really known, but it made a lot of sense. So I included prompt caching in the book, and I’m very happy to see that nowadays it’s pretty much everywhere.”

(Click to enlarge graphic.)

The savings scale depending on how much of your prompt gets cached. In Claude Code logs, Huyen’s open-source tool Sniffly found cache hit rates of 90% to 97%. Some providers rewrite prompts internally to improve hit rates, but you might as well structure them yourself. 

Huyen’s tip: since caching works on shared prefixes, put the stable parts of your prompt first and the variable parts later. “It’s pretty easy to do, and it can improve your application performance significantly,” she noted. 

Evaluating inference providers

Huyen closed with a warning for anyone evaluating inference providers: “There are many inference companies that provide inference optimizations for models you want to use, and a lot of them advertise just cost and latency. 

“But pay attention to how many inference optimization techniques also change the model behavior or reduce the model quality. So when evaluating an inference service, it’s important to look not just at cost and latency, but also at model quality. Does this model, provided on this service, also perform similarly on standard benchmarks?”

What’s changed one year later?

So where do we stand today, one year on from this keynote? Most of it actually aged quite well. 

On the economics, I reckon Huyen was bang on… I think, for most of us as users, we don’t have all the cost levers to pull that Huyen outlined. But it’s great to understand what is happening. As a novice local LLM user myself, I found I could relate to her points on parallelism (I don’t have it) and prompt caching/quantization (within my grasp of control). 

Prompt caching (which Huyen said was new when Huyen wrote AI Engineering) is now priced into every bundle purchase of API tokens. And her Claude Code observation (90% cache hit rates) is probably the reason we mere mortals can still afford agentic coding agents at all.

Some of it aged in ways that were hard to predict at the time. Huyen mentioned how reasoning models make inference even more significant. One year on, I think agents running multi-step loops with tool calls have turned that idea from a footnote into a way to turn Claude’s rate limits (and their infamous 99.x% availability) on their head. 

All the metrics Huyen described – time to first token, time to publish, goodput under a latency SLO, etc., are all now part of the lingo and probably need to be reasoned about differently. 

That’s one thing I hope she’s talking about this year! Grab a free conference pass and join us online. 

Grab a complimentary pass to PG 99 Conf 2026 and join us on October 21 and 22 to chat with Huyen.

The post Chip Huyen explains how to cut inference costs without new hardware appeared first on The New Stack.

It passed CI. It passed your evals. The customer still got the wrong answer.

13 septembre 2026 à 16:00
Blurred, overlapping close-ups of yellow analog thermometer dials, their curved scales marked 0, 10, 20, and 30 in black with a red band sweeping through the upper range.

A diff is not evidence. It’s a statement of intent.

The tests passed. The review’s done. The change is live. Then someone says the app is slow, or the answers are wrong, or both. You open the diff. Your assistant points at the function it changed and offers a plausible cause.

It sounds right. It might not be.

This is the observability gap that AI features expose. Dynatrace’s 2026 State of SRE and Platform Engineering report (919 enterprise leaders surveyed globally) found that while 77% of platform engineering teams embed observability in at least some services, only 40% have it fully integrated across all deployments. That gap was manageable when your services were deterministic. But with AI agents, it becomes a liability.

A conventional service fails loudly… An AI agent fails quietly. It returns a 200. It passes faithfulness checks. And the customer still gets the wrong answer.

A conventional service fails loudly. A 500 error, a latency spike, a dependency that stops responding. An AI agent fails quietly. It returns a 200. It passes faithfulness checks. The customer still gets the wrong answer.

You can’t alert on “wrong.” You need evidence from the running system — and for AI features, that means something more than request traces and error rates.

Find the request first.

One release, two symptoms

Say you run a support agent over product documentation. A customer asks how to configure export in version 2026.3. Your coding assistant helped rewrite the documentation lookup. CI passed. The existing evals passed.

After deploy, answers take longer. Some of them describe older product versions.

Start with one affected run. You want its release, retrieval config, and feature-flag state, so those need to be on the root span as attributes set at span start, not reconstructed later from a deploy log. Then put that run next to one for a similar question from before the change.

For an agent, that means the trajectory: every model call and tool call, in order, with arguments and results. A distributed trace records those as spans and stitches them across service boundaries through context propagation.

Here’s one run, simplified, with its evaluation linked separately.


# Illustrative pseudotelemetry, not a captured incident.

# Names, IDs, timings, and labels are invented, not a standard schema.

# Selected spans shown in execution order; other work is omitted.

trace: example-run-a | session: example-session-7 | release: 2026.9.2

requested.product_version: "2026.3"

agent.run                                  12.4s

  model.choose_tool                         1.0s

  tool.search_docs                          0.9s

    args: {query: "configure export", product_version: null}

  tool.search_docs                          0.8s

    args: {query: "configure export", product_version: null}

  tool.search_docs                          0.9s

    args: {query: "configure export", product_version: null}

    returned.doc_versions: ["2024.1", "2024.1", "2023.9"]

  model.generate_answer                     8.1s

linked_evaluation:

  trace: example-run-a

  faithfulness: pass

  requested_version_answered: fail


Two things to chase. The repeated searches. The null version filter.

Check the repeated work

Three identical searches cost 2.6 seconds. The trace shows the symptom. It doesn’t explain the cause.

But look at what sits between them. Nothing. One model.choose_tool span at the top, and no model call between the second search and the third. The model didn’t ask for those retries. Something in the harness did: the code that runs tools, handles retries, and manages context. A model.choose_tool span between each search would mean the opposite: a model that kept requesting the same tool, which is a prompt or tool-description problem. Same symptom, different file to open.

That still doesn’t make the retries wrong. Read the retry policy, then read the tool results. A 200 from a search backend can carry an empty hit list, or every hit under your relevance threshold, and retrying on that is legitimate.

The generation call is the bigger slice anyway, at 8.1 seconds. Compare its input tokens and duration against similar runs. If the harness appended all three result sets to the context, the retries inflated that prompt, and you paid for them twice, in latency and in tokens. Check downstream services and traffic too before you pin the slowdown on the release.

Bringing this back into the IDE? Bound the question. Give the assistant the service, the release, the time window, and the trace IDs. Have it line the changed code path up against the dependency calls in the affected trace. Then separate what the evidence supports from what it’s assuming.

Same workflow debugs a checkout service making three identical database calls. You don’t need to build an agent to use it.

A grounded answer can still fail

Now read the answer.

In this example, it accurately repeats the retrieved documentation. Faithfulness passes, or groundedness, depending on whose vocabulary your tooling uses.

The customer still gets instructions for the wrong version.

Whether you call it faithfulness or groundedness, the metric only tells you whether the answer is supported by the sources you supplied. It says nothing about whether those were the right sources.

The obvious next move is a retrieval evaluator. It still won’t catch this. Those score whether the retrieved context is relevant to the query, and the 2024.1 export instructions are relevant to configure the export. They’re just invalid for the version asked. Those documents are relevant to the query. They are not valid for the version the customer requested. Relevance is not validity.

Those documents are relevant to the query. They are not valid for the version the customer requested. Relevance is not validity.

So, this isn’t a generation failure. It’s a retrieval precondition nobody asserted, and the null filter names it: the requested version never reached the lookup. Reproduce that before you touch the prompt or the model.

Most of it is testable with ordinary code. Give the fixtures documents carrying version metadata, then assert on the lookup directly, no model in the loop:

def test_lookup_filters_to_requested_version(docs_fixture):

    hits = search_docs(query="configure export", product_version="2026.3")

    assert hits, "no hits for a version that has docs"

    assert {h.product_version for h in hits} == {"2026.3"}


Deterministic, cheap, belongs in CI. Then evaluate the answer separately, which is the part you can’t assert: does it give usable 2026.3 instructions, or say the available documentation can’t support one? Two tests, because they fail for different reasons and you want to know which one broke.

The assertion won’t catch every wrong answer. It catches this missing constraint every time, which is more than a judge scoring helpfulness one to five will do for you.

Which is why evaluation needs retained context. Record the prompt version, model ID, retrieval config, and document IDs and versions alongside the release. Keep enough permitted evidence to read the answer back later, sensitive content redacted before export.

Link results by trace and span ID. If scoring lands after the span closes, store a separate linked result. Don’t plan on writing attributes to a finished span: the OpenTelemetry tracing API says implementations should ignore updates after End.

GenAI semantic conventions are still evolving, and different instrumentation projects expose similar concepts with different attribute names.

Make the failure part of the next release check

Confirmed the causes? Verify each fix against the behavior it’s supposed to change.

For the repeated searches, add a regression test that reproduces the repetition without killing legitimate retries. Don’t pin one exact tool sequence when several orderings finish the task correctly; a trajectory test that demands a single path fails on every valid refactor.

For the version mismatch, restore the filter. Add these cases: current version, an older supported version the customer names explicitly, irrelevant documentation, and no supportable answer. Run the answer evals repeatedly where output varies, because one pass isn’t a result.

Use code for anything you can assert directly. Use a model-based judge for answer quality, and validate that judge against examples people reviewed. An unchecked judge is one more model you’re taking on faith.

To run scoring against production traffic rather than fixtures, you’ll need a way to sample spans already in your environment, score them with a judge model, and link each result back to the source trace — the linked-result pattern above, not a write to a closed span. Whatever tooling you use, version the evaluator. A scoring change that looks like a product improvement isn’t one.

After the release, compare latency and task success on similar requests, and keep tool-call and token counts on the same screen. Read them together, or they’ll mislead you. Tool calls dropping from three to one can look like the fix is working, but it can also look like a lookup you removed by accident. Fewer output tokens look like a cost win, and it also looks like an answer that quietly stopped listing step four.

Bring one debugging question

If you wouldn’t know where to start the investigation, you’re not done instrumenting.

For a conventional service, that’s the request path and dependency timing. For an AI feature, add what it retrieved, what it produced, and how you’ll decide whether that was the right answer.

Dynatrace is sponsoring WeAreDevelopers World Congress Americas, September 23-25, 2026, in San José. Come with a debugging question from an AI-assisted release or from an AI feature you’re building, and we’ll work through it.

The post It passed CI. It passed your evals. The customer still got the wrong answer. appeared first on The New Stack.

The AI-native SDLC won’t be one process 

12 septembre 2026 à 16:00
Five fuzzy pom-poms — white, pink, magenta, purple, and teal — arranged in a horizontal row across the upper portion of a dark navy background, each casting a small shadow, with colored light washing the backdrop in magenta and teal.

Anthropic recently published its AI-Native SDLC Playbook. Its central claim is that “code is no longer the bottleneck.” When agents can produce an implementation in minutes, the constraint moves to everything around the build phase: planning, review, verification, deployment, and governance.

The risk, if organizations get this wrong, is producing ten times the changes at the same quality per change or worse, with no way to identify which changes are the bad ones. The traditional answer is that a person looks at each one, and that is exactly what stops working at this volume.

The risk… is producing ten times the changes at the same quality per change or worse, with no way to identify which changes are the bad ones.

The playbook gets the foundations right. What it doesn’t capture is that an organization’s process has nuance: it is really a family of processes that vary with the change at hand, not a single flow every change travels through.

The spec-driven wave

The playbook is part of a broader wave of spec-driven development tooling, including Amazon’s Kiro and GitHub’s Spec Kit. The tools share a common shape. Written artifacts drive the work: an intent document becomes a spec, a plan, a diff, and review findings, all committed to version control. Policy is enforced by deterministic mechanisms such as hooks, rather than by instructions in a prompt. Agents check their own work before a human sees it. Humans own the approvals.

But each of these tools also prescribes a particular process: a fixed sequence of stages that produce fixed artifacts, and that every change travels through. Adopting the tool means adopting its process.

One organization runs many processes

No real organization runs a single process. The right process for a change depends on the risk it carries and the accountability it requires. A documentation fix, a dependency upgrade, and a schema migration in a payments service should not travel the same path. They need different levels of verification, different approvers, and different records. In regulated domains, the process itself is part of the compliance obligation: auditors expect a record of who approved each change and based on what evidence. What must be recorded differs by the type of change.

When a tool prescribes one process, teams route around it for changes that don’t fit, which is the worst outcome because the real process becomes invisible.

When a tool prescribes one process, teams route around it for changes that don’t fit, which is the worst outcome because the real process becomes invisible. Or the vendor keeps adding configuration until the tool becomes a workflow engine that nobody fully understands.

The tool should not prescribe a process. It should give the organization a way to define its own.

Processes as state machines

A better model is to define each process as a state machine. The states are facts about a change: reviewed, validated against its dependencies, approved for production. Those facts live in systems no single tool owns: the repository, CI, the cluster, the tracker. So a process cannot be a program that executes steps. It is a set of rules that react to observations about those systems. Each rule specifies:

  1. The facts it requires before it can fire.
  2. Its gate: fire automatically, or wait for a person’s approval.
  3. The permission it grants when it fires, such as merging or deploying.

The definition is this set of rules, stored as data and reviewed like code. An organization runs many small machines, one per risk class.

At runtime, this behaves nothing like a workflow engine. No component tracks “we are on step four”: the process advances when a fact appears in the system that owns it, and rules react. Events that arrive late, twice, or after a restart are handled like any other, because rules only react to current state. The gate is one of a rule’s conditions, so you can hold firing during an incident or a release freeze without editing any definition.

Gates need enforcement. A gate implemented as a prompt instruction depends on the model following it. The agent harness can provide the determinism required to run the state machine and enforce its gates, stopping the agent between actions until a gate is answered, while the infrastructure enforces the rest.

The process adapts to the change

One fixed process definition per repository is not enough: every change in that repository would still travel the same path, regardless of its risk. The path a change takes should depend on what the change is, and this routing comes from classifying the change, not from the author choosing a path. The organization defines classification using signals it already has: the paths a change touches, the repository it lives in, a label on its tracking issue, etc.

The definitions themselves also need to change over time, and that has to be safe. Because a process definition is data, editing it is itself a change, and it goes through its own gated process. Loosening an approval gate on the release process gets reviewed the way a schema migration does, not edited the way a config file does.

Click to enlarge graphic.

Concretely, consider three changes to the same service:

  • A documentation fix is classified by the paths it touches. Its process has two states: the build passes, and it merges. No person is involved.
  • A dependency upgrade skips design review, but its process requires compatibility evidence: the upgraded service runs its integration tests against real dependencies. A major version bump adds an approval that a patch bump does not.
  • A schema migration in the payments service is classified by the component it touches, no matter what kind of change it claims to be. Its process adds states the others never see: review by a payments owner, validation against production-shaped data, and a release approval from someone accountable for that domain.

Each transition, in each path, is logged with who approved it and on what evidence.

Tenets

The tenets these processes should follow:

  • Autonomy is granted per action, and grows over time. Each transition is set to fire automatically, require approval, or hold. As agents prove reliable on a class of change, that setting is relaxed, so the process absorbs agent improvements without redesign.
  • Human attention is spent only where judgment is needed. Agent effort keeps getting cheaper; supervision hours do not. A person is brought into the loop only when the decision requires human judgment, and is given the context to decide quickly.
  • Evidence comes from outside the agent. An agent’s own report never moves a change forward. Transitions fire on facts from systems the agent cannot write to, such as test results and validation in a realistic environment.
  • The process record is the audit trail. The definition is the written policy, and the transition log shows who approved each step, on what evidence, under which version of the policy.

Quality at scale

The playbook and its peers get the foundations right. What is missing is the ability for an organization to define its own processes, vary them by the risk of each change, and evolve them safely. The goal is not fewer humans in the loop. It is spending human judgment only where it is needed, backed by evidence agents cannot produce about themselves, so that quality holds while throughput multiplies.

We are building these ideas at Signadot and acting as our own guinea pigs, running our own development through this process. If you’re experimenting with these ideas too, we’d love to talk!

The post The AI-native SDLC won’t be one process  appeared first on The New Stack.

47,000 job listings reveal the engineering roles that AI is creating

10 septembre 2026 à 15:09
Abstract overlapping circles in black, green, orange, and pale yellow on a cream background.

Every major transformation in tech has led to roles merging, then new ones emerging. Friction between developers and operations drove the creation of the DevOps engineer. Then, when security needed to be considered throughout the delivery pipeline, DevSecOps emerged.

The team beyond the AI-native talent and services platform Andela analyzed 47,000 recent engineering job postings from Fortune 500 companies. This research, released on Thursday, uncovered more than 2,000 skills that pour into 23 emerging job titles. None of these are coming out of nowhere; they strategically merge existing skill sets to create new roles. 

Among 1,832 postings titled primarily for AI or ML engineers, 53% contained at least two skills drawn from different established roles, Andela finds.

In today’s tighter economy and amid AI, companies seem to be going one of three ways. They are lumping too much work and required experience into now-nebulous AI engineer or machine learning (ML) engineer job titles. They might be looking to replace tech workers with AI. But more forward-thinking organizations are reworking job titles and descriptions to reflect the demands of getting AI safely and efficiently through the software delivery lifecycle. 

Cory Hymel, head of research at Andela, tells The New Stack, “If you’re going to look to deploy AI within your organization, the way to look at it is that an AI has a certain set of skills, and then a human has a certain set of skills.

“If you Venn diagram those and see where they cross over, an AI should do the skills it can. But the human circle is still exponentially larger than that of AI.”

“When you’re looking to deploy AI, it’s not about trying to replace that human circle with an AI one. It’s about what certain skills you need to carve out and delegate to it.”

Read on for the top engineering jobs that are emerging because of AI, how to attract tech talent for them, and what you need to focus on to get a tech job in this tough market.

Click image to enlarge.

AI is not serving the generalist. Specialization is still key.

Citing the leading AI CEOs, Hymel remarks, “You’ve heard from the AI salespeople of the world that AI is going to push people to be more generalist, and the data that we found here doesn’t necessarily support it.”

“You’ve heard from the AI salespeople of the world that AI is going to push people to be more generalist, and the data that we found here doesn’t necessarily support it.”

Overall, they found that these emerging job titles aren’t generalist at all. These emerging roles bridge skill sets from several existing ones, but each addresses a specific operational or product need, some tied to AI adoption. 

The top five new engineering job roles discovered are:

  1. MLOps pipeline engineer, who builds and runs the automated infrastructure to deploy, version, and monitor machine-learning models in production, with 46% ML engineer, 23% DevOps engineer, 15% data engineer skills, and 8% each AI engineer and data scientist roles.
  2. LLM application engineer, who builds on and evaluates foundational models via large language model application and conversation systems, bringing 48% AI engineer and 34% ML engineer, with a touch of product designer, software architect, and embedded software engineer roles.
  3. FinOps reliability engineer runs cloud infrastructure for both reliability and cost, bridging 36% DevOps engineer, 27% site reliability engineer (SRE), 18% cloud engineer, and 9% each DevSecOps engineer and cloud solutions architect.
  4. Docs-as-Code engineer applies program management and DevOps engineering skills to the traditional technical writer’s role, pivoting from stagnant docs to specification-as-code.
  5. Product frontend engineer is about a third traditional frontend engineer and a third product manager, with a touch of full-stack engineer, UX researcher, and product designer.

“If you’re a DevOps engineer, historically, your skill bundle might have allocated 30 to 40% of pure DevOps-required skills that are rich and specific to that role, and you have a remaining bundle that is cross-role habitable, meaning that those skills would translate between DevOps or to an engineer or to a technical product manager,” Hymel explains. “Some of those skills can now be replaced with AI, which means that those skills that are more directly focused on your role become more important than ever.” 

So-called “soft” business skills are also increasingly crucial, he contends. However, he seriously doubts anyone will ever be able to slide between finance, marketing, engineering, and sales roles. 

Where enterprise engineering job descriptions falter

“Job descriptions and resumes right now are the best worst thing that we have. When you’re talking about large enterprises, and you’re having to deal with scale, your hiring process gets farther away from the work,” Hymel explains. 

Especially when the hiring process starts in HR, not engineering, “you’re needing to put language in place that will survive the chain of custody, with the naming of the job [coming from] the engineer that’s closest to the work.”

It’s not uncommon for an enterprise to have 50 different front-end developer job listings, each with very different skill requirements. It’s better for candidates and for fit to be as specific as possible, including embracing new job titles.

This habit of generic job titles used to be positive because it brought in more applicants, but nowadays, with so many engineers on the market, it further dilutes your hiring pool, leaving you with the 100 fastest applicants—who are often AI-generated anyway.

“Any company that has not taken a hard look at revising their job postings and job titles is at an extreme disadvantage because there’s a very high probability that you’re going to end up hiring the wrong person simply because you didn’t take the time to describe the role well enough,” Hymel remarks, which leads to dire consequences. 

“There’s potential churn, so you just spend all this time and cost to go headhunt and find someone. Two, if they do get in there, you have to pay for their ramp time to get up to speed because they were sold a different bill of goods than what was in the description. And then three, it impacts overall roadmaps and timelines because now you might have to replace, and, again, you have to wait for people to get up to speed.”

On top of this, HR and engineering hiring managers alike are using AI to generate job descriptions. It still isn’t recommended to have AI generate something so human and essential to your core success.

Especially in this time of flux, when no one may have the required experience, companies should start job descriptions with what they want the future hire to achieve.

“The cost of code is going nearer to zero.”

“The cost of code is going nearer to zero.” Hymel explains organizations should think more like, “Here are the outcomes that we’re looking for. If you have the soft skills and additional skills around it to get there, whether that is backlog prioritization, being able to be collaborative, having worked on project deployments before, and we don’t necessarily care that you can score a 10 out of 10 on Python anymore.”

Which emerging roles engineers should pursue

The familiar claim that women apply only when they meet every qualification is not well supported; recent research finds that application behavior is more complicated. Still, clearly separating essential qualifications from preferences can reduce ambiguity and unnecessary barriers.

Focusing on outcomes and clearly distinguishing required from preferred skills may broaden the applicant pool, although it does not guarantee greater diversity.

For example, if you’re an engineer who enjoys having a product focus, collaboration, and strategy, Hymel recommends looking toward the new product front-end engineer role, which owns the full user-facing feature lifecycle, from definition to shipping.

“You are required to have more mindshare towards prioritization of features,” he says, shifting away from a ticket person, because “now you have more control because AI allows you to span out a little bit deeper.”

Similarly, AI has the back-end engineer thinking beyond the back-end stack to deployments, scalability, and the reliability of underlying infrastructure systems, giving rise to roles like the polyglot back-end integration engineer. 

Technical writers — reasonably worried about their jobs in the face of AI-generated documentation — should look toward new docs-as-code engineer positions, which add technical program management and DevOps engineering skills.

“If you’re writing the docs, you’re essentially writing the specs that enable spec-driven development. You now have the capability to actually contribute software,” Hymel observes. “And it starts all the way at the top too. If you’re a product manager, you can now start building and contributing code, like a product experience designer.”

Read the full Emergent Role Research. If any of these AI engineering job descriptions ring truer than what you were hired for, we hope it empowers your next conversation with HR or for you to apply for a different job title. 

The post 47,000 job listings reveal the engineering roles that AI is creating appeared first on The New Stack.

The critical vulnerability was a test database. That’s the whole triage problem.

14 septembre 2026 à 17:20

A security researcher testing a 300-person B2B company with a global footprint discovered an internet-exposed database with weak authentication during a routine scan. The database appeared to be a prime target for attackers; it had critical severity and was an obvious first-fix candidate. Upon further inspection, however, researchers learned it was a resettable test database used to test job candidates, rather than a system containing client data.

The episode represents a critical problem: Scanners and security researchers can’t infer the real cost of a compromise on their own. Increasingly, companies rely on an external partner to handle cloud vulnerability triage and keep teams from becoming overwhelmed. Scanners produce excessive noise, requiring focus and attention from engineers to understand alerts and complex attack techniques, tune out false positives, and work with teams on remediation, reducing the time they have to build new features, work with customers, or scale up their tech environment.

Tech employees have more on their to-do lists than ever: Product teams face an infinite stream of feature requests and engineering teams carry more technical debt than they can clear. When you factor in a reorg, many employees may have larger project scopes and leaner teams, all while constantly monitoring, correcting, and mentoring AI agents. Some reports cite 90-hour work weeks.

Beyond today’s structural challenges, security teams are drowning in a different type of data. Programs absorb identity events, firewall logs, endpoint alerts, vendor feeds, and threat intelligence, producing far more analysis data than even a few years ago. Every item can look urgent, high-risk, and worthy of immediate attention. With finite hours, it’s hard to know where to act first.

Jon Rose, founder of the information security and risk management advisory firm IOmergent,  has seen firsthand how AI and near-universal tooling are producing more findings than teams can triage. 

“Within the span of security work, there’s an unending list of things you could tackle, and you’re pulled in so many different directions… But you have to be ruthless about prioritizing and investing your time.”

“Within the span of security work, there’s an unending list of things you could tackle, and you’re pulled in so many different directions,” Rose tells The New Stack. “But you have to be ruthless about prioritizing and investing your time.”

The challenge, then, isn’t remediation. It’s allocation: Deciding where limited engineering and security capacity will make the greatest impact is the most important question security leaders answer every day. A technical severity score, used without threat and environmental context, cannot answer the business question: What should we fix first? Effective vulnerability prioritization turns raw findings into business-aware priorities, and that judgment is the real work.

CVSS limitations: a starting point, not a decision

For security teams, a Common Vulnerability Scoring System (CVSS) provides a useful baseline. Its base metrics classify a vulnerability’s severity — attack vector, complexity, required privileges, and potential impact on confidentiality, integrity, and availability. But a severity score is designed to be stable across environments, so it can’t tell a team whether an asset is exposed to the internet, shielded by compensating controls, or central to the business.

Essentially, treating a base severity score as an automatic fix-first ticket is problematic, rather than using CVSS on its own.

CVSS can include Threat and Environmental metrics that account for evolving exploit conditions and organization-specific context. However, risk-based vulnerability management still depends on accurate knowledge of the environment — and on someone applying that context consistently. 

“The piece that’s missing from any of these tools is the grounding in the business, the understanding of what actually matters,” Rose says.

Escalate the internet-facing medium

A smart decision framework deprioritizes the urgency of a score and probes reachability: whether an attacker could actually reach the vulnerable component. 

“Is the affected service exposed to the public internet or not? Or is it isolated behind network controls and accessible only to a limited set of internal users?” Rose says. “Often, issues will get flagged, but it’s not in a position where it could be triggered.”

Next is consequence. An exposed flaw on a disposable test system might be a genuine security concern. Still, it doesn’t carry the same weight as a weakness on the application that processes customer transactions, stores sensitive information, or underpins a company’s main revenue stream. So while the severity might be alarming, the business outcome might not.

Teams should also ask whether the weakness is attracting active attacker interest and where it could lead. A vulnerability listed in CISA’s Known Exploited Vulnerabilities catalog deserves urgent scrutiny, because there’s evidence of exploitation in the wild.

The Exploit Prediction Scoring System provides a forward-looking signal: An estimate of the likelihood that exploitation of a particular vulnerability will be observed over the next 30 days. Neither replaces a business decision, but both help distinguish a theoretical risk from one that demands attention now. However, as AI-driven exploitation accelerates, the window between what is known to be vulnerable and what is actively exploited is shrinking because the cost of building and weaponizing exploits is dropping.

The relevant attack path might also extend beyond the affected machine. A comparatively modest flaw can become an immediate concern if it provides a route into privileged accounts, a production system, or customer data. Conversely, a high-severity finding can be deprioritized — temporarily — if it is not reachable, has limited impact, and is protected by reliable controls.

“The speed and the depth of research and investigation into those security issues are going faster,” Rose says. “So it can change really quickly.” The worst state is unacknowledged risk sitting in the backlog. Even the best detection degrades when no one owns the trend line. Teams should track exceptions, set review dates, and name an owner. Accepted risk is still risk; the difference is that it’s explicit, time-bound, and revisited. 

An operator capability, not a weekend project

Security teams are leaning more on AI to prioritize and triage threats, and they’re uncovering vulnerabilities at an unprecedented pace. Yet the sheer volume of findings is overwhelming, even for the most efficient of humans.

It’s clear, too, that AI can make it easier to discover less obvious paths to exploitation, and to introduce new ones. An academic study of 20,000+ issues fixed by AI found that LLMs introduce nearly 9x as many new vulnerabilities as developers, exhibiting unique patterns not found in developers’ code. The answer is to use AI at the outcome level. For every alert, the goal is to have the SOC analyst’s standard questions answered quickly: new or known, what’s exposed, what data is at risk, prod or dev, and how it has evolved.

“Effective programs start by aligning with executive teams to understand the business — where the company is going — so allocation and adjustments track the actual risk, not just the score.”

“That’s how teams get thousands of alerts down to 10 to 20 prioritized tickets,” Rose tells The New Stack. “Effective programs start by aligning with executive teams to understand the business — where the company is going — so allocation and adjustments track the actual risk, not just the score.”

As to-do lists grow ever longer, business-context judgment that’s applied every day by someone who owns it is something worth dedicating more resources.

If you think you’d benefit from managed cloud security, book a 30-minute cloud security scoping call with IOmergent.

The post The critical vulnerability was a test database. That’s the whole triage problem. appeared first on The New Stack.

How to find failures without drowning in tracing data

3 septembre 2026 à 22:17
On The New Stack podcast, Sarah Hudspeth of Chronosphere, a Palo Alto Networks company, explains how teams can build a more effective tracing strategy.

A metrics dashboard can tell you a system’s health with ease. A log can help you understand a discrete failure. But if you want to understand where in a query’s journey things went awry, you need traces.

By tracking a request from its point of origin through data and microservices to the end user, traces offer unparalleled insight into how systems work and where failures occur. SREs offer the fastest path to remediation. That means less downtime, fewer burned-out developers, and happier customers.

Sadly, the promise of traces often doesn’t match the on-the-ground reality. 

Why? Simply collecting and holding onto all your company’s traces is an exercise in hoarding. Do you need to store terabytes of tracing data just to show when your systems worked? Not only is that much information expensive to hold onto, but collecting it can slow the very systems you are trying to monitor. And when you have all the stored tracing data, finding what you need in the ocean of information can take too long.

Is tracing cooked? Not at all.

Is tracing cooked? Not at all. There are several ways to beat back tracing data overload: Head sampling collects only a portion of tracing data, reducing storage concerns; tail sampling asks whether, after a trace is recorded, it is worth holding onto, making it easier to find what you’re looking for down the road. And dynamic sampling can automatically cull similar or highly repetitive traces, so you don’t accidentally flood your storage system with nearly identical data.

You can avoid the most common tracing pitfalls by building your observability system intelligently. That’s precisely what I was hoping to learn from Sarah Hudspeth of Chronosphere (a Palo Alto Networks company), who is my guest on the latest episode of The New Stack podcast.

Whether you are just starting your tracing journey or deep in the trenches looking for help, Hudspeth’s ability to turn abstract technical concepts into simple, digestible analogies is enviable. 

Hit play on the episode above, and let’s jump the chasm between the promise of tracing and getting it to work for you in a production setting.

The post How to find failures without drowning in tracing data appeared first on The New Stack.

Observability has a data problem. AI is about to make it worse.

26 août 2026 à 22:08
Parallel orange lines form a flowing wave across a dark purple background.

Observability is entering a new phase now that OpenTelemetry has standardized instrumentation for data collection. Unfortunately, the observability industry still lacks a cost-effective way to store, retain, search, and analyze full-fidelity telemetry data. This results in blind spots in observability and many teams operating without full operational visibility.

As AI systems generate more logs, traces, and metrics — thereby making the blind spots issue worse — Bronto, a Dublin, Ireland, firm offering an intelligent data observability platform, is betting that the next observability platform battle will be won at the data layer, not the dashboard layer.

Bolt-ons and incremental efficiency aren’t enough

Trevor Parsons, co-founder and co-CEO of Bronto, tells The New Stack that the industry has been optimizing at the edges rather than rebuilding the economics and architecture of telemetry storage. The industry has introduced a wide array of “hacks” and “capabilities” to avoid tackling this issue head-on and ultimately to protect their margins. 

“If you are a couple of times cheaper or 50% cheaper, that ain’t going to cut it,” Parsons says, because data volumes, especially AI telemetry, are growing so quickly, on top of already stretched observability budgets and inefficient datastores. 

Promises, Promises, Promises…

Parsons elaborates, “Observability has always and continues to have a data problem.”

The eternal promise of observability has been delivering teams a clearer view of what’s happening inside their systems.

“Observability has always and continues to have a data problem.”

But in practice, that view is often incomplete, expensive, and short-lived. For too many teams, observability has become less about asking better questions and more about fighting the cost and complexity of storing the data they already need. 

“Sometimes people frame that as a cost problem, where they’ll say observability is up to 20 or 30% of your infrastructure spend,” Parsons says. “I actually think this minimizes the issue; it’s much bigger than that. Teams are actually paying 10, 20, 30% of their infrastructure spend for access to only a sliver of their data.” 

Noel Ruane, co-founder and co-CEO of Bronto, frames the challenge that organizations face and tells The New Stack, “Agents and applications are generating more logs, traces, and metrics each day. The software landscape has accelerated, but are observability vendors keeping pace? No, they’re offering workarounds, bolted-on features, and asking teams to accept blind spots.” In short, Ruane says, they’ve failed to solve the data problem.

Out with the old observability model 

“Customers are not getting access to all of their observability data, Parsons explains. “They have to cut their retention from 30 days to seven days to three days. They have to sample data. They have to rehydrate data.”

In other words, today’s tools make customers choose which parts of their own data they’re allowed to see, and you may only get to see it for a short amount of time.” 

“The solutions that are being put in front of customers to give them their data are always full of compromises, forcing teams to choose between cost, coverage, and speed of data access. The burden is always put on the customer by vendors.”

“The solutions that are being put in front of customers to give them their data are always full of compromises, forcing teams to choose between cost, coverage, and speed of data access,” Parsons says. “The burden is always put on the customer by vendors.

“But really this should be the other way around; it’s the vendors’ job to innovate on behalf of the customer” 

OpenTelemetry: Collection solved, storage problem exposed

Severin Neumann, head of community at Bronto, tells The New Stack that OpenTelemetry has helped standardize instrumentation and data collection, while reducing reliance on proprietary agents.

But that success has created a new bottleneck, says Neumann, who is also an OpenTelemetry maintainer and member of the OpenTelemetry governance committee. Now that organizations can collect more telemetry, they need somewhere affordable and useful to put it.

“We have fixed the instrumentation problem,” Neumann says. He cautions that enterprises now need ways to handle all this data. And if enterprises can’t store it and instead throw away large parts of it, humans and agents can not make sense of it.

The observability business model doesn’t align with customer value

The legacy observability tool business model charges customers for data storage, rather than the value teams get from their data, Parsons says. Customers tell him the same thing constantly: “I pay the same price even if I never search my data.” In many cases, customers find existing tools difficult to use and feel that their observability solution is just a really expensive data store that they do not get a lot of value from.”‘

Noel Ruane assessed the market by saying, “Traditional vendors like Datadog know their pricing model isn’t sustainable. They’ve introduced defensive features like ‘Flex Logs’ and a new ClickHouse partnership to try to keep customers from jumping ship, but they’ve only added new complexity for their customers.” 

Especially in the AI era, Ruane adds, the traditional business model charges teams in ways that discourage them from capitalizing on their data. Customers should pay much less for data that sits idle and more when they actually derive value from it with queries and analysis.

Bronto’s technical differentiation

Bronto isn’t selling another observability dashboard. It argues that observability is a storage problem before it’s a visualization problem, and that’s where the company went.

Underneath the platform is a custom-built polymorphic data store called BrontoDB, specifically designed for observability data. The pitch: enterprises can keep more than 100 times the observability data they hold now, and it won’t get slower or harder to use.

Why that matters comes down to how the three signals break. Metrics, logs, and traces each hit a wall at different points, and Bronto says it built BrontoDB to tackle these issues head-on. Parsons is blunt about two of them.

“With metrics, we’ve solved the high cardinality problem where costs traditionally explode with high cardinality metrics,” Parsons says. “With logging, we’ve solved the indexing problem where there was always a trade-off between fast logs and paying through the nose for it or having slow logs and getting them slightly cheaper.”

  • High cardinality is what wrecks metrics pricing. Add enough unique dimensions and the bill lands somewhere nobody forecast. Bronto says it was built specifically to take that surprise out.
  • Logs have always been pick-your-poison: fast and expensive, or cheap and slow. Bronto says that choice goes away — sub-second search across petabytes, no shortened retention windows, no rehydrating cold data, no waiting.
  • Traces, Bronto argues, shouldn’t be sampled at all. Sampling exists because tools and pricing models couldn’t handle the full stream. Bronto says teams can send it all.

Billing works differently, too. Most vendors charge for data sitting in storage, whether anyone touches it or not. Bronto charges closer to what teams actually search and analyze. That’s the piece that has to hold up if full-fidelity observability is going to be affordable at AI scale.

AI is what raises the stakes, Parsons says. AI systems are non-deterministic and trace-heavy. They throw off more telemetry, and the data has to stick around longer if you want to debug effectively. 

He points to an upside as well. As operations become more automated, telemetry data becomes more useful because agents can chew through volumes of history that no SRE would ever read manually.

AI raises both the volume and the stakes, according to Parsons. AI systems create more telemetry because they are non-deterministic, trace-heavy, and require longer retention for troubleshooting. At the same time, AI-enabled operations will make historical telemetry more valuable because agents can analyze far more data than human SRE teams could manually inspect.

“If AI is the intersection of where data meets intelligence, you can not apply intelligence if you do not have the data.”

“If AI is the intersection of where data meets intelligence, you can not apply intelligence if you do not have the data,” Parsons says.

The next observability battle 

AI is unlikely to fix observability’s data problem. In fact, it will produce more telemetry, create more edge cases, and increase the cost of missing the right signal at the wrong time.

For Bronto, the data layer is the next major battleground. Dashboards still matter, but in an AI-heavy production environment, the more important question may be whether teams have access to all their data for as long as they need so that they can apply AI to it. 

The post Observability has a data problem. AI is about to make it worse. appeared first on The New Stack.

How telemetry pipelines keep AI agent costs under control

25 août 2026 à 21:01
Abstract neon pink and purple angular pathways interlock against a dark geometric background.

As enterprises move from experimenting with AI to running autonomous agents in production, an infrastructure problem is emerging: rising telemetry costs. Non-deterministic, iterative, and capable of generating data at machine speed, agents are far harder to monitor — and their costs far harder to predict — than conventional applications.

Many companies are struggling to attribute and defend their telemetry bills. In fact, 59% of organizations have already terminated or delayed an agentic AI deployment due to monitoring costs, according to a survey of more than 300 enterprise IT decision-makers in North America and Western Europe, commissioned by Apica and conducted by Omdia/Informa TechTarget. 

The agents most affected are often in some of the most high-stakes deployments: think cybersecurity, compliance, and fraud detection. As monitoring bills explode, deployments aren’t necessarily getting killed by engineering teams. More often than not, it’s finance pulling the plug.

Andi Mann, chief product and technology officer at Apica, recently saw this play out at a large bank. The organization couldn’t pin down exactly what it was spending on its AI programs.

“They knew they couldn’t afford to keep going on the same trajectory, so they had no choice but to cancel certain AI programs,” Mann tells The New Stack. “It’s a pattern I have seen before, because AI projects are cannibalizing typical budgets.” 

“It’s a pattern I have seen before, because AI projects are cannibalizing typical budgets.”

The implications are huge. As people, funding, and monitoring resources are diverted toward new AI workloads, other parts of the business start to suffer. Mann says he’s seeing outages, downtime, penetration attacks, and DDoS protection compete for the same resources.

The problem is already showing up in research. Most (54%) enterprises have seen telemetry volume triple in the past year alone, with 43% of that growth coming from AI/ML workloads — by far the largest driver. Businesses are under pressure to address this crisis as their observability bills climb.

They report spending an average of $3.17 million on observability, with that figure growing 28% year over year and showing no obvious ceiling. No wonder then that 83% rank AI observability as a top priority for the year ahead.

The coming wave could be catastrophic

With a new cloud service, database, or application, you’d expect the monitoring burden and telemetry data to rise by a relatively predictable increment. But enterprises now foresee an average 9.5X increase in telemetry data within two years.

“Imagine if your credit card or grocery bill went up more than nine times — that is not a marginal amount,” Mann says. “This isn’t a gentle ramp; it’s a skyscraper, and it’s prompting panic.”

“This isn’t a gentle ramp; it’s a skyscraper, and it’s prompting panic.”

Some 44% of organizations expect their telemetry data to increase by 6X to 100X.

The reason is that an agent task isn’t the same as a single application request. A customer-support task might generate a top-level trace, several model calls, retrieval operations, tool calls, retries, and loops. But if the agent delegates work to another agent, that adds another branch to the trace. 

Each model can produce token, latency, cost, and provider data, while each tool call generates its own records for arguments, results, status, and downstream activity. Identifiers such as tool_name, agent_id, and trace_id also create high cardinality, making data harder to aggregate and more expensive to index, with costs compounding at every stage.

That creates an uncomfortable gap between AI ambition and infrastructure readiness. Although 35% of enterprises claim widespread agentic AI deployment, operating and managing those systems is very different. Nearly two-thirds are only somewhat prepared for the shift. Unlike conventional applications, agents can call multiple models and tools, repeat tasks, or expand a workflow in unpredictable ways, making both capacity and monitoring costs difficult to forecast.

From extreme telemetry costs to an upstream control layer

The answer, according to Mann, is to intervene earlier. “You can’t keep sending essentially useless data to an expensive central analytics or storage platform because there’s no point analyzing data which says everything is fine,” he says. “As early as possible, get the pipeline to use data collectors and manage those at source.”

Legacy observability platforms were built around collecting data, ingesting it, storing and indexing it, and then analyzing it. That model worked for human-driven, dashboard-queried workloads, but agentic AI requires decisions to be made before telemetry reaches the most expensive parts of the stack.

A pipeline-first architecture can sample repetitive successful events while retaining failures, retries, policy violations, and unusually slow traces. It can enrich records with agent, session, model, tool, token, and estimated-cost data, redact sensitive prompts and identifiers, and aggregate metrics and long-term records to destinations with different cost and retention profiles.

A compact metric or sampled span might represent a routine successful tool call, while a failed call retains its parent trace, error details, retry history, and security context. The aim is to preserve the information needed to explain an agent’s behavior while reducing redundant data and limiting what gets indexed.

Agents also need millisecond-level context for autonomous decisions. Processing telemetry close to its source allows organizations to quickly detect retry storms, excessive tool loops, or unusual token consumption, rather than waiting for data to be ingested and indexed centrally. 

Without upstream control, businesses risk feeding fragmented, unnecessary telemetry into platforms that charge for every additional gigabyte, index, and retained record.

Architecture that separates winners from cancellations

The payoff for rethinking the telemetry pipeline is hard to ignore. Enterprises with a telemetry pipeline are 50% more likely to be prepared for the growth of agentic AI data. Among mature agentic AI organizations, pipeline adoption is what sets them apart: these organizations are 80% more likely to have avoided the operational cost challenges hampering their peers.

Rather than treating the observability platform as a catch-all destination, enterprises can make decisions about telemetry before it gets there.

The answer is to move the intelligence upstream. Rather than treating the observability platform as a catch-all destination, enterprises can make decisions about telemetry before it gets there — filtering out noise, enriching what matters, and routing data according to its value and purpose. That means less data hitting expensive storage and analytics systems, while the information that does make it through is more useful and available in real time.

Crucially, this isn’t about ripping out the observability platforms enterprises already rely on. It’s about putting a smarter control layer in front of them: deciding what data deserves to be ingested, where it should go, and how much it should cost.

Existing observability platforms still have an important role to play. “The pipeline can’t do everything, but it can act as a first responder,” says Mann. “You still work with the big analytics platforms, but you’re saving money, reducing risk, and improving your compliance performance.”

The Apica study finds that its pipeline control, metrics foundation, and data readiness services, for example, can reduce the total cost of ownership by 40% compared to legacy observability platforms. Of course, though, the actual savings will depend on telemetry volumes, retention policies, sampling rules, routing decisions, existing contracts, and the proportion of data that can be processed before ingestion.

The window to rethink the architecture is now. Some 68% of enterprises plan to evaluate changes to their observability stack within the next six months, while almost a quarter say existing vendor relationships won’t be a significant factor in those decisions. The next phase of observability will be won by the ability to handle what agentic AI throws at the infrastructure.

That means the pipeline can no longer be treated as plumbing that moves telemetry from A to B. It’s becoming the control layer for an increasingly autonomous, data-hungry environment. 

Organizations that establish agentic-ready infrastructure will be better positioned to reduce observability costs and improve risk performance. Mann doesn’t think platform engineering and SRE teams have much choice. “This has already become a board-level decision,” he says. “Ultimately, it’s a choice about how smart you can afford to make your business.”

Download the Omdia research report: “The Agentic AI Telemetry Crisis: Are You Ready for What’s Coming?”

The post How telemetry pipelines keep AI agent costs under control appeared first on The New Stack.

❌