❌

Vue lecture

Vulnerability alert fatigue nearly swamped WHOOP. But its fix still keeps a human in charge.

Abstract image of thin vertical ribs against a bright orange background. In the center, a blurred rectangular glow shifts from green on the left, through dark red, to pink on the right, as if seen through ribbed glass.

With often hundreds of thousands of alerts a day, many tech organizations are buried in vulnerabilities and worn down by alert fatigue. The rise of AI has only made it harder to cut through the noise and to find actionable alerts. Manual security and site reliability engineering is not an option. 

The engineering team behind WHOOP‘s health and fitness tracker felt this pain, relying on multi-day, all-hands triage sessions to stay on top of the alert deluge. But, as a high-growth consumer health company handling sensitive user data, it couldn’t afford to miss anything. Which is why the team at WHOOP built an automated vulnerability-response workflow based on the company’s specific technical, operational, and trust considerations.

Join The New Stack on Wednesday, October 7 to learn from WHOOP staff engineer Vinay Raghu and Datadog senior product engineer Amber Tunnell how WHOOP built and implemented this workflow using Datadog Bits AI and Workflow Automation for faster, at-scale response. 

Join us on October 7, 2026, for a live Datadog x TNS event

REGISTER NOW FOR THIS WEBINAR
By registering, you consent to The New Stack’s Privacy Policy, Terms of Use and to receiving email communication from The New Stack and our event partner. You may opt out at any time.

DevSecOps, security, and cloud/platform engineers should bring their questions to this live demo-slash-case study to learn how to reduce friction between developer velocity and security requirements without increasing headcount. 

What you’ll take away from our live webinar

Raghu and Tunnell engineers will share how they were able to:

  1. Focus on real exposure vs. scanner noise. Not everything is critical. You’ll learn how WHOOP used Datadog’s Software Composition Analysis (SCA) to analyze runtime code execution and prioritize active threats.
  2. Route the right vulnerability to the right engineer. WHOOP automated vulnerability mapping to microservice owners, so developers received tickets with full context attached.
  3. Build automated guardrails for devs to self-resolve. This let security engineers pivot from frustrating gatekeeping and ticket-pushing to more proactive, systemic work that adds value. 
  4. Maintain a human in the loop. With such sensitive data and a demand to be always-on, WHOOP isn’t ready to automate the engineer out. Learn how they decided their team’s response had to change.

And then, of course, we will end the live discussion with how to measure it all. Don’t miss out and register to attend on October 7.

The post Vulnerability alert fatigue nearly swamped WHOOP. But its fix still keeps a human in charge. appeared first on The New Stack.

  •  

Your agent is only as good as your infrastructure

Dark abstract 3D glass ribbon rendering symbolizing complex AI agent infrastructure and bursty data workflows.

You built a great agent, but something happened when it moved into production.

In testing, your agent reviewed pull requests efficiently on its own. It read the diff, grepped the codebase for related usages, ran the test suite, checked whether CI was still red from an earlier commit, and drafted a comment—all before you’d finished reading the diff yourself.

In production, however, imagine the same five steps ran behind every other PR review your team’s agents performed that hour. Some reviews landed in seconds; others sat for minutes because the test-suite step landed on a node that was mid-burst from someone else’s agent. 

The agent didn’t change. The execution environment did, and that’s what decided whether review time held steady or crept up.

Agentic applications introduce a different execution pattern than traditional chat applications. As those workflows become longer and more dynamic, infrastructure has a much larger influence on latency, reliability, and cost than it does for a simple chatbot.

That’s why your agent is only as good as your infrastructure.

Agents aren’t chatbots with more steps

The difference between serving inference for a chatbot vs. an AI agent isn’t simply that one is “more capable.” They execute work differently:

  • A chatbot usually makes one inference call per user message. The model receives a prompt, generates a response, and waits for the next user input before proceeding.
  • An agent executes the entire workflow, turning one user message into a chain of inference calls. It might decide to search documentation, retrieve data from a database, call an API, execute code, evaluate the result, and then repeat that process before producing an answer. Each of those decisions may trigger another inference call, and every result becomes additional context for the next step.

“Infrastructure has a much larger influence on latency, reliability, and cost than it does for a simple chatbot.”

That execution model changes the infrastructure requirements for AI agents. 

One question, many steps behind it

Instead of optimizing for individual inference requests, the system has to support long-running workflows whose latency and reliability depend on every component in the chain.

A single user request often expands into a sequence of inference and tool execution steps, sometimes called multi-turn tool calls or the agentic loop. Rather than generating one response, the model alternates between reasoning and interacting with external systems.

For example, you ask an agent why checkout latency spiked overnight. The agent pulls the deploy log, queries the monitoring system, runs a diagnostic against the connection pool, weighs whether the culprit is a bad deploy or a capacity issue, and then folds that into another inference call before finally producing a full-fledged response.

Each reasoning step becomes another inference request, and every tool result is added to the model’s context before the next step.

This workflow changes what reliability means

Multi-turn workflows are inherently sequential, which is why even low latency can quickly add up to a significant amount. Every inference step waits for the previous one to finish. If a database query takes two seconds, the model can’t begin the next reasoning step until that result returns. The model may generate tokens quickly, but the other steps slow it down.

“In this agentic workflow, every step in the chain has to hold, because the chain is only as strong as its slowest link.”

In this agentic workflow, every step in the chain has to hold, because the chain is only as strong as its slowest link. Instead of processing isolated inference requests, the inference stack has to orchestrate a chain of dependent model invocations and external tool calls. As those workflows become longer, the stack increasingly determines how quickly, reliably, and cost-effectively the application performs.

But the user doesn’t see an orchestration hiccup. They see an agent that hung or gave up.

That’s why end-to-end agent reliability depends on much more than model quality. Infrastructure determines whether each step has the resources it needs to execute predictably under load.

Why the bill and the performance both feel unpredictable

A second difference in agentic workflows catches teams off guard: demand patterns and their impact on your inference bill. 

Most inference services, and the pricing built on top of them, assume traffic arrives at a predictable pace. A typical inference solution knows the predictable demand pattern: User traffic increases, request volume increases, and capacity scales accordingly. Cloud infrastructure is typically optimized for these steady request patterns using mechanisms such as autoscaling, load balancing, and capacity planning.

Agent workloads don’t behave that way. Individual workflows pause while waiting on external systems, then resume as soon as new information becomes available.

  • The pause: The agent waits on an external API or database, so the GPU serving that workflow has no inference work to perform, and its accumulated context may be evicted from GPU memory while it waits
  • The burst: As soon as external systems return results, inference resumes simultaneously across many workflows, creating short and sharp spikes in GPU demand, each re-processing its full accumulated context

If you’re watching GPU utilization and it looks less like steady load and more like a heartbeat — flat, then a spike every time tool results come back — that’s the signature. It means you’re provisioning for the average when you should be provisioning for the peak, and it’s usually the first place p99 latency quietly blows out.

“If you’re watching GPU utilization and it looks less like steady load and more like a heartbeat, that’s the signature.”

Inference services designed around steady or predictable request streams can struggle to allocate resources efficiently under these conditions. This leads to inconsistent latency, GPU underutilization, or higher operating costs.

When performance becomes unpredictable, or a bill doesn’t match what you expected, it’s evidence your infrastructure was built for a different workload than the one you’re actually running.

What infrastructure built for agents actually looks like

Agentic applications place different demands on infrastructure than traditional AI workloads: long dependency chains, bursty demand, and continuous evolution. Because of this, the infrastructure is deciding whether the chain holds, and whether the bill holds too.

For an agent to run, it needs infrastructure purpose-built to support:

Performance that holds across the whole chain. The infrastructure must keep latency consistent across multi-step and multi-tool workflows.

Scalability that responds to bursty demand. Infrastructure should scale quickly as inference demand fluctuates, without requiring capacity to remain provisioned during idle periods.

  • Predictable economics even for dynamic workloads. The infrastructure bill should reflect actual usage.
  • None of this makes bursty demand disappear, but it changes how the system absorbs it. A large enough simultaneous burst, or a workflow that accumulates enough context before pausing, still costs something. The goal isn’t zero cost or zero limit; it’s making both predictable.

Your agent is only as good as your infrastructure. Get it right, and your agent’s responsiveness, reliability, and cost-effectiveness will improve your work. 

Learn more about how infrastructure can be purpose-built for agentic workflows: Check out the documentation to get started.

The post Your agent is only as good as your infrastructure appeared first on The New Stack.

  •  

Study: Developers are addicted to AI, and managers are making it worse

AI is addictive, and managers are rewarding those who use it the most (even though they are shipping stuff they don’t understand). What a tangled mess. Let’s unpack findings from the AI Coding Addiction Report, what it says about the wayward use of AI coding tools, and how behavior differs based on whether you use Claude Code, Google Gemini, OpenAI Codex, or GitHub Copilot.

The survey targeted developers who use AI at least once a week, and the report is based on over 300 responses from developers of varying seniority levels and using different tools.

Animal mistreatment, applied to humans

A branch of psychology called behaviorism is based on trial-and-error learning. Edward Thorndike noticed that cats could learn skills to escape from puzzle boxes, but B.F. Skinner took it further. Skinner invented the “operant conditioning chamber”, a box that could be used to manipulate animal behavior by forcing it to respond to signals to earn rewards. If a rat or pigeon in the chamber hits a button when the light comes on, it gets food. If they fail to press the button, the electric grid in the floor delivers a punishment.

Skinner Box diagram by AndreasJS.

The ultimate result of this line of research was the creation of push-button slot machines, arranged in long rows, with flashing lights and occasional rewards for humans who drop coins in and press the button. In a business context, it shows up in gamification techniques, where managers treat employees like lab animals and are surprised when employees respond by acting like they are.

Intentionally or otherwise, AI coding tools seem wired like a Skinner Box. With coding loops delivering frequent flashing lights and rewards, 43% of developers hit home time but can’t tear themselves away from the reward-generation process. Overall, 80% of developers describe their relationship with AI as “more like a dependence than an advantage.” In fact, it’s harder to give up AI coding tools than to give up social media or video games, which are often colloquially described as “addictive.” This is why developers report skipping or delaying breaks, meals, and even bedtime.

What developers skip or delay during AI coding sessions. Source Coddy.

This leads us to ask some serious questions, like how much of the “AI productivity gain” comes from the behaviorism of keeping developers at the keyboard for longer hours with fewer breaks and less rest?

Differences by tooling

One interesting finding is that AI coding tool choice may influence at least some behaviors. Codex had the highest after-hours coding rate (62%), compared with lower rates for Gemini (45%), Claude Code (40%), and GitHub Copilot (36%). This could indicate that Codex most deeply embodies the stimulus/reward cycle of the Skinner box.

It could be illuminating to identify which specific aspects of these tools increase the likelihood of addictive coding loops spilling into the workday so that we can moderate their impact on personal time and crucial rest. As senior developers are most likely to work longer hours, organizations will end up with tired people making important decisions. This hustle has serious consequences, and none of them are good for the organization.

You’ll get more of what you reward

Meanwhile, managers quickly reward those who are most addicted. The heaviest AI users were more likely to get raises and promotions, even though 71% of developers shipped code they didn’t fully understand. Before AI, a developer copying and pasting code from Stack Overflow into their codebase was expected to understand and adapt the examples as part of the work. But now, managers are throwing cash at the developers glued so hard to the slot machine they can’t pause to visit the restroom, let alone assess the code they are committing.

With these heavy AI users getting the rewards, other developers are left to mimic the dysfunctional behavior, or at least give the impression of it. There’s a famous scene in the movie Shaun of the Dead where the band of misfit protagonists must cross a road infested with zombies. They achieve this by emulating the jerky movements and slurred speech of the infected. Developers who want to understand the code they commit will feel pressure to compromise when they see rewards going to the flippant.

Organizations must realize that software value comes from a series of good decisions. Hustle mode rapidly diminishes the rate at which decisions cross the threshold of “good”. The value is in gracefully solving user problems, but too many managers value software by the volume of features, code, or hours spent with their hands on the keyboard.

Look around you. The world is in a state of excess. There’s an avalanche of content, code, and crunch. We don’t need more; we need better.

The post Study: Developers are addicted to AI, and managers are making it worse appeared first on The New Stack.

  •  

Agents operate, humans govern: Scale your operations and reduce toil with Azure SRE Agent 

Illustration of colorful doughnut, bar, radar, and line charts alongside slider controls on a black background.

What if engineers could spend their time building and optimizing systems rather than maintaining them?

It’s 3 a.m., and the pager goes off. Tabbing between multiple dashboards and diagnostics, the SRE struggles to determine whether what woke them is a real incident, whether they’re the right person to handle it, or whether they need to wake someone else. Digging through monitoring tools, deployment history, incident systems, and team runbooks — and chasing what might be the wrong theory about the root cause — they can’t respond fast enough to stop more customers from being affected.

Or imagine that, by the time the SRE joins the incident bridge, Azure SRE Agent has already analyzed the monitoring data, identified the root cause, and prepared a fix for approval and deployment.

Sanchit Mehta, one of the head engineers for Azure SRE Agent, tells The New Stack that “[Azure SRE Agent] starts analyzing telemetry and correlates things like blast radius, deployment changes, recent changes, any recent rollouts, to try to tell the engineers, ‘OK, this is what is causing it.'” Increasingly, it will even create the PR for that fix.

The support is just as useful during normal working hours. At InEight, correlating telemetry across tens of thousands of Azure resources can take days, if not weeks. When a support ticket reports slow performance without identifying the product, engineers must determine which of the company’s 14 products is affected, then check multiple observability and reliability tools.

InEight shared that, during its first incident using Azure SRE Agent, the agent quickly identified the affected product, traced the performance issue to its root cause, and recommended scaling Redis. The DevOps team had been considering scaling the app service as a temporary fix.

Proactive and in production

This kind of help is becoming the new normal at Microsoft, where more than 3,000 service teams already use Azure SRE Agent to investigate issues, perform root cause analysis, respond to incidents, fix code, enable automatic mitigation, support proactive detection, analyze data, and report at scale. Azure SRE Agent has already handled more than 1.8 million incidents inside Microsoft, many mitigated in minutes.

The team also uses Azure SRE Agent to develop and improve the service itself, with custom agents for code review, deployment, evaluation, and monitoring. This “agent-powered engineering” approach, as Mehta calls it, lets the team take advantage of ongoing advances in AI models. That includes proactively spotting problems, like quota issues that affected deployments, and automatically raising support tickets to resolve them. The agent recently identified the root cause of a change that broke synthetic tests as soon as the change reached the first region, he says.

 “It said, ‘OK, this was an upstream PyPI package that broke your dependency; you need to add tests for it; you should roll back immediately; here’s how you should go fix this.'”

Mehta says that kind of proactive monitoring is hard to handle with deterministic queries. “You need a level of intelligence to see when a large production payload is being deployed and if it has the potential to cause degradations.” 

For some internal teams, more than half of incidents are autonomously managed by the SRE agent and don’t need any human intervention, adds Shamir Abdul Aziz, lead program manager for Azure SRE Agent, because they’re what he calls “safe” operations and mitigations: a restart, scale-out or rollback of a service, or change order requests escalated by customers.

“The humans did the governance, set up the guidelines, gave some coaching to the agent, and then it went into auto mode to complete the entire workflow,” Abdul Aziz says.

Agents are ready to help

SREs are already drowning in repetitive toil. SREs are already drowning in repetitive toil, and coding agents add to that workload. Agentic operations are now powerful enough to help, Vyom Nagrani, one of the head PMs for Azure SRE Agent, tells The New Stack.

“As code gets written more and more by agents, it’s going to take another agent to operate it,” Nagrani says. “But why wait? If the agent can manage code which other agents write, why can’t it manage code written by humans?”

“As code gets written more and more by agents, it’s going to take another agent to operate it.”

“The reasoning loop has become mature enough that now agents can automatically start figuring out a lot of these complex problems, especially when it comes to correlating across multiple data sources, which has always been the hardest thing for humans to do,” Nagrani says.


Powerful models aren’t enough, though, and homegrown automation won’t have the production-grade governance, verification, evaluation, telemetry, and control a platform can offer.

The state of the art has progressed from prompt engineering to context engineering—which grounds AI in your infrastructure, code, and institutional knowledge — and now to harness engineering. “That is what allows you to run agents at scale, control them, and govern them,” says Abdul Aziz.

“When you combine all these things with being able to verify, audit, evaluate, and get real telemetry and metrics out of the system, where the agent claims it has done something, you can validate that agent’s claim,” Abdul Aziz says.

Instead of a non-deterministic black box that can’t explain its decisions, you can trace and learn from the agent’s reasoning so that you can correct mistakes once, not over and over again. “That’s why companies are willing to adopt it now,” Abdul Aziz says. “Because when you try the same thing ten times, you’re going to get the same output.”

“You don’t just turn on the agent, give it full access, and ask it to solve everything.”

After two years of building enterprise-grade systems that can be trusted, audited, and validated, the next step for cloud-native SRE can be agentic ops with autonomous capabilities — but you still need to know how to adopt it, Abdul Aziz warns. “You don’t just turn on the agent, give it full access, and ask it to solve everything.”

Context and connections

Azure SRE Agent is built for Azure but not limited to the Azure platform. The agent provides native access to Azure services such as Azure Monitor, Application Insights, Log Analytics, and Azure Resource Graph. Connecting the agent to your subscriptions, telemetry data, and source code gives it the operational context and institutional knowledge needed to understand how you work.

Beyond Azure, Azure SRE Agent integrates with engineering and operational tools through managed connectors for Azure DevOps and GitHub, plus MCP connectors that enable access to external knowledge sources such as Google Drive, Confluence, Cursor, Claude Code, and other third-party systems. 

Put all that knowledge into Markdown files in a repo, along with the skills and tools agents need to act on your systems (including third-party and on-premises services). That gives you artifacts that agents can version, review, test, reuse, and update.

When you want to dictate how to handle an incident — what to check and in what order, what to post, and even how to format a report — you can create a custom agent, either by using an existing runbook or by working through an incident with an agent and saving that skill. Using agents to improve agents is the shortcut to making Azure SRE Agent more useful the more you use it. Essentially, saving what agents learn during incidents helps improve their future responses.

Guidelines and guardrails

Governance covers identity, role-based access control (RBAC), and tool-access policies. These controls determine which actions are allowed, blocked, or subject to step-by-step approval, and whether an agent operates autonomously or with human review.

What makes governance both flexible and powerful are hooks, based on prompts or deterministic commands, that fire at different stages of a workflow and catch edge cases, such as allowing an agent to drop the index in a SQL database but never drop a table.

Metrics show you whether governance is working. The new live reports show time to mitigation, tool reliability, how often agents act autonomously, and cost per outcome at a glance. InEight’s metrics are typical: an 80% reduction in both incident investigation time and build failure triage time, a 67% reduction in the effort needed to investigate bugs, and an 84% reduction in cost.

To get those results, you need triggers that automatically launch agents instead of waiting for a human to open a chat window.

Bind skills and custom agents to specific alert classes so they can respond to incidents first. Start agents through pipelines, webhooks, or work items to automate delivery workflows. Schedule regular checks, reviews, and audits, and have agents automatically update their artifacts.

Agents operate; you stay in control 

By reducing repetitive tasks and technical toil, Azure SRE Agent frees engineers to focus on more interesting and innovative projects. Just as there’s a familiar maturity model for adopting site reliability engineering in the first place, you don’t jump straight into having agents rather than humans handle operations. When you give agents the context about your infrastructure, you can start using them for investigations.

“If you give agents read access to your source code, your telemetry, your resources, the time to get to the root cause is reduced to minutes rather than hours or days,” Abdul Aziz points out. “Every customer starts there.”

Once you’re happy with the answers you’re getting, you can give the agent more permissions while still approving individual steps, he says. “The fixing is easy once you understand the problem. It’s usually changing your configuration, writing a piece of code, or restarting a service.”

“The fixing is easy once you understand the problem. It’s usually changing your configuration, writing a piece of code, or restarting a service.”

As you expand into other operational tasks, refine the agents’ artifacts, metrics, and governance before granting more autonomy: “Things like rolling back a release when we know there was a regression in that release, restarting a service, dropping a corrupt index on a SQL table, or scaling out a service,” Abdul Aziz suggests.

For more complicated issues, agents can deliver the entire fix, ready for approval. The Azure SRE Agent that manages the Azure SRE Agent product looks at exceptions, errors, incidents, Teams conversations, emails, and GitHub issues every night and spits out PRs. 

Avoid code review bottlenecks by having agents deploy, test, measure, and include outcomes in the PRs. Use continuous evaluation to build a self-learning system that accurately follows your existing workflows.

“The agent can self-improve because the agent learns constantly,” Abdul Aziz says. “You can configure scheduled tasks to identify which evaluation scores were low and automatically improve the custom agent, custom skills, and even your knowledge documents – because knowledge management is also a toil. The agent can automate all of that.”

The right way to start

Azure SRE Agent now offers a 30-day trial experience with no always-on charges. Make the most of that by learning from some common mistakes:

  • It’s not magic! Turning on the agent doesn’t mean you don’t have to do DevOps anymore. Don’t treat it as a chatbot or connect it to just your observability system. You need to give the agent the context it needs, the tools to do the job, and intentional triggers that tell it when to act. Otherwise, it may spend effort on low-value work or generate outputs that aren’t grounded in your environment. 
  • Don’t limit yourself to what the agent does out of the box: Customize agent skills, tools, connections, and logic to fit how your organization works, and build custom agents for specific tasks.
  • Don’t use agents for jobs a single line of code can do: Using them to explore deterministic, structured data for anomalies is an expensive waste of tokens that will only flood the context window when the agent can write that line of code itself. “Orchestrate, don’t calculate,” as Nagrani puts it. If you’re drowning in alerts, use automation to filter the noise and only send alerts that need intelligent analysis to agents.
  • Don’t stick with what you’ve always done or copy your org chart: The most effective agents have a complete picture of the system, so they need all the context, even if it crosses two teams. That might mean crossing boundaries, coordinating who has expertise and who needs to grant access, or rethinking how the organization works.

“If agents have the right context, they minimize the toil and truly make operations less costly,” lead product manager Deepthi Chelupati points out. That way you can move faster, be proactive, and give engineers more time to innovate and less maintenance work to dread.

Get started today: sre.azure.com

The post Agents operate, humans govern: Scale your operations and reduce toil with Azure SRE Agent  appeared first on The New Stack.

  •  

Your AI coding spend bought 25% more output. Duplication rose 81%.

Since they arrived on the scene, a great swathe of the software industry has pinned its hopes on AI tools, whether that’s early chat interfaces or modern agentic swarms. But the tone has shifted over the past few weeks, with HR software provider Rippling adding an anti-tokenmaxxing AI spend console to give CFOs and CTOs visibility into spend, and IBM Vice Chairman Gary Cohn saying last week that the ROI has “not been nearly as high as people might think.” As the northern hemisphere feels the Fall cooldown, it seems that Winter is coming for AI tool budgets.

As organizations balance the books, teams will start feeling pressure on their Claude Code and Cursor budgets. That means they may face harder usage limits where returns are unclear, or budgets that can’t sustain the usage levels.

For organizations that have figured out how to measure AI impact, there’s a growing realization that generating code volume at pace doesn’t guarantee movement in the metrics that matter. If you count lines of code, the number of pull requests, or even the number of features delivered, you’ll see no clear relationship to value. Not every line of code or feature matters equally to the business or its customers. This distribution is galaxy-wide.

Even at the output level, many organizations haven’t worked out how to turn siloed gains into end-to-end improvements. Gains in coding speed transfer to new tasks introduced by AI or get absorbed by downstream changes. If you haven’t worked out the inherent properties of value streams before you bought AI, you’ll be getting painful lessons when you try to track your ROI.

If the only problem were translating the cost of AI tools into end-to-end value, it would be serious enough. But something far worse is happening.

Productivity in terms of output

Let’s look at the data, which GitClear collected and analyzed for the Maintainability Gap report. The report, published in June, covers 623 million analyzed changes from 2023 to 2026. This is a substantial dataset, with millions of change operations included across three and a half years. As teams rapidly adopt AI developer environments and tools such as Cursor and Claude Code, GitClear’s code-change-operation database allows them to detect and classify code duplication, hotspots, and signals of good or poor code factoring.

Heavy AI users gained 25% on their own prior velocity, far from the claims of 10x increases. The same report shows those heavy users out-producing non-AI users by 4 to 10x, which sounds like the opposite finding until you look at who they are. Teams that outperformed their peers in output were doing so before AI tooling arrived. And remember, there’s no guarantee this output will accrue to the value stream, or provide meaningful value to the organization or its customers.

The first part of the ROI calculation is to determine whether these increases are worth the cost. For many organizations, I would be surprised if they were.

Perhaps because much of the discussion of AI tools has focused on speed, other factors have received little attention. The software industry may have found a different kind of value if it had focused on the tools as a forklift truck, rather than a racing car, because the straight-line speed doesn’t seem so impressive. Yet they can perform heavy lifts that are tricky for us mere humans, like large-scale changes across a codebase, such as replacing an unmaintained library with a replacement.

For those who pass this first gate, we can look at the next factor.

Productivity in terms of code quality

The shift to AI has brought about a giant behavioral change in the software industry. For several decades, the importance of code maintainability has been emphasized repeatedly. More than half the programming books on my shelf focus on architecture, code design, coupling, and cleanliness. The idea of refactoring, supported by automated tests, appears across many of these books.

Yet the signals GitClear is getting from the data are a complete reversal: a return to the code-and-fix era of software development. Across the dataset, block duplication rose 81% over 2023, from 40.3 to 73.0 per million changed lines. Those multiple expressions of the same concept drift apart and create whack-a-mole bugs. Moved code, the signature of refactoring, fell from 21% of changed lines in 2022 to 3.8% in 2026, which means code is becoming harder to understand, and that will hit maintainers with or without AI.

Chart showing a dramatic drop over four years in refactoring changes and a steep rise in duplication over the same time period.

Source: GitClear

When you make these changes, you get away with it initially because you’re early in the maintenance cost curve. Over time, however, the rising costs will become unbearable. Rework rates will rise, stealing time from new feature development. Seemingly minor issues will take far too long to pinpoint and resolve, with many simply becoming part of how it works because the fix is economically unviable. The accumulation of tightly coupled, incomprehensible code units will reach the point where the software stops being valuable.

We’ve been trying to validate the claims of 10x boosts with AI coding assistants. The data shows the opposite. Before AI, developers chose refactoring over copy-and-paste about two to one. Now they’re roughly five times likelier to copy and paste.

Technical practices are the mission, not a side quest

When I’ve presented at conferences and user groups on what great software delivery looks like, I mention, among other things, test automation and refactoring. In the Q&A that follows, this question will inevitably come up in one form or another: “How do I get permission from my boss to do these things?”

Developers are whipped hard for fast progress, so they are trained to avoid what they see as the side quest. If they need to increase output, they streamline coding tasks, leaving no time to write tests or improve the code’s design, which would delay the feature. Inevitably, this makes all feature development vastly slower over time.

The premise of lightweight software delivery processes is that they rely on technical practices that control the cost of maintaining software over time. The wisdom is that working more deliberately today lets us maintain the pace of change indefinitely. If we skip these practices, change becomes increasingly slow and expensive.

Chart showing a traditional software project with costs rising superlinearly over time and an XP project with cost growth subdued.”
Based on figures in Extreme Programming Explained (Beck, 1999)

Those technical practices, like test automation and refactoring, aren’t side quests; they are the work. Technical discipline is a fundamental requirement of commercial software delivery, and these practices stopped being optional some time ago.

When asked for techniques to convince managers to allow these practices, I’m confused. I’ve never asked for permission to do what is right for me, the software, its users, and the organization. No compromise can be reached, because omitting technical practices harms everyone involved.

This “side quest” thinking was unresolved in many organizations, and adding AI into the mix has made things far worse. When teams are given AI tools, they come with the expectation of a big return. When teams are, in reality, seeing a 25% increase in their rate of change against an industry misperception of some 10x boost, they will feel even more pressure to deliver.

Under these dysfunctional circumstances, it’s no wonder those who treat good practice as a side quest are skipping crucial steps.

Real high performance is well known

High-performing teams have worked out that a set of software delivery practices is no longer optional. They worked it out because they were scaling long before the new tools arrived.

For software that matters, that people depend on, and that still needs to exist in a year, in five years, and beyond, we’ve moved from the pick-and-mix of the past, and there’s a new bar for professional software delivery.

There is a glimmer of hope here. The teams doing well with AI are the same teams that outpaced the industry before AI. They maintain rigorous technical disciplines, monitor code health indicators, and prioritize the craft of keeping code maintainable for the long haul.

The post Your AI coding spend bought 25% more output. Duplication rose 81%. appeared first on The New Stack.

  •  

Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots

Abstract dark red digital grid texture representing AI code verification loops and developer workflow checks.

Red Hat released Red Hat AI 3.5 this week, a move designed to let software engineering teams run AI with the same operational rigor as enterprise apps on mission-critical infrastructure.

Echoing both the “pilots-to-production” and “single control plane for infrastructure, models, and agents” narratives playing out across much of the tech industry, Red Hat’s key move here seems to be its expansion of platform capabilities to run enhanced multi-tenancy for AI service providers. 

Crucially, the new AI platform release is built to run AI use cases that require complete hardware-to-software isolation (needed when AI workloads have to wrangle sensitive data, proprietary models, and regulated information), as well as handle priority-aware service requests (where mission-critical workloads execute in favor of lower-grade tasks) with native multi-tenancy on a shared GPU infrastructure. 

Red Hat’s Senior Director of Product for Red Hat AI is Tushar Katarki. He tells The New Stack that running enterprise AI without safety controls is “like driving a supercar blindfolded” in real terms.

“With Red Hat AI 3.5, we are delivering the operational guardrails, verifiable trust, and multi-tenant controls needed to run AI as a mission-critical service rather than an unpredictable experiment,” Katarki says. “You can’t scale what you can’t measure, and you certainly shouldn’t deploy what you can’t verify. By unifying pre-deployment safety benchmarking, real-time observability, and GPU resource management, we are giving platform teams the power to turn isolated AI pilots into a fully governed enterprise architecture.”

Every GPU request now becomes a priority decision

Applied mathematician, data scientist, and fractional CMO Joshua Estrin, Ph.D., tells The New Stack that Red Hat’s work is of the time and of the moment; primarily because “every GPU request now becomes a priority decision”, so one developer’s internal experiment cannot be treated with the same urgency as a financial close.

“Looking at the state of AI infrastructure players out there now, Red Hat has clearly seen that priority-aware multi-tenancy lets companies use expensive compute more efficiently, but efficiency without isolation is just a faster way to create a security and reliability crisis,” Estrin says.

“Every GPU request now becomes a priority decision.”

He thinks that the winners in this market (he tags Nvidia, Nutanix, Suse with Rancher, HPE Ezmeral and VMware Cloud Foundation under Broadcom as usual suspects) will be the organizations that can “share capacity while still proving what happened where” in live production.

“That means proving whose workload actually executed and ran, who had access, what it cost, and what happens when demand spikes. Those answers rarely come from the infrastructure diagram; they get settled in the boardroom, usually after someone’s critical workflow has slowed down, but regardless, this sums up where AI infrastructure is now,” adds Estrin.

What are the elevated multi-tenancy pain points?

To unpack what’s happening here, let’s remind ourselves that GPUs are expensive, obviously. As organizations move to live production use cases of agentic AI, they will want to maximize their GPU state’s ability to serve multiple workloads across multiple teams, multiple customers, multiple apps, and so on. 

This all means that the breadth of AI infrastructure efficiency becomes the new agentic bottleneck.

The priority-aware services above for native multi-tenancy on shared GPU infrastructure are important right now; this function dynamically allocates GPU capacity based on workload priority. When lower-priority workloads can be run on spare (or cheaper) capacity (rather than separately provisioned GPU resources having to be spun up for lesser jobs), then everyone gets to go home earlier on Friday.

“The breadth of AI infrastructure efficiency becomes the new agentic bottleneck.”

But that’s not all the balls being juggled here; Red Hat mentioned isolation too, and that’s a concurrently complex AI infrastructure discipline challenge. GPU compute resources managed through isolation techniques enable AI services to run without accessing or interfering with another service’s data, models, or compute environment. 

In other words, this combines hardware consolidation with strong tenant isolation.

What new Red Hat technologies are on offer?

Red Hat says this release lets developers verify models before deployment through EvalHub, enabling risk-focused safety benchmarking and regulatory compliance certifications. 

New observability dashboards give platform teams metrics for a real-time view of inference health, GPU utilization, and AI model performance. Non-admin users can access dashboards for per-user token consumption showback (another term for token tracking) and distributed inference workloads.

Also new is shared GPU control for multi-tenant inference. So-called “fair-share GPU scheduling” manages resource allocation across tenants, while priority-aware serving provides admission control and priority-based request routing to protect real-time inference. As suggested above. it also allows background workloads to use available capacity.

VP of product management at Nutanix, Anindo Sengupta, tells The New Stack that running multi-tenant AI at scale does indeed require secure tenant isolation.

“The essential isolation is best achieved through virtualization,” Sengupta says. “For specialized at-scale AI workloads, the choice could be to run Kubernetes on bare metal. On top of that, to create real value, agents need access to both LLMs that run on containers and enterprise systems (databases, business systems, etc.) that run on traditional infrastructure. For hybrid AI to run efficiently, the platform must manage both these environments in a performant way, with a common operating model.”

Observability & model-as-a-service showback

Built-in observability and MaaS showback in Red Hat’s latest release are present to provide per-user token metering, performance dashboards for models and agents, MLflow visual agentic tracing, and GPU utilization dashboards for operational and usage transparency.

For efficient GPU memory management, the general availability of CPU offloading and the developer preview of storage offloading allow models to handle longer conversations and larger documents without additional GPU hardware. 

Red Hat AI Hub also introduces agent templates and starter kits with pre-configured reference implementations for common enterprise patterns, including code review, document processing, and research workflows.

Red Hat is hoping its Red Hat AI 3.5 version release gets us to a state where GPU-based AI resources are viewed as a policy-controlled infrastructure pool.

Power without interconnect bandwidth is botched.

As enterprise AI pilots succeed and initial results show some returns, IT teams must then address the need to deliver at scale. But scaling AI across the business demands the same operational rigor as any mission-critical infrastructure: verified safety before deployment, precise resource controls across shared GPU environments, governed agent behavior and transparent usage metrics.

Yoram Novick, CEO of sovereign AI edge cloud provider Zadara, has previously been on the record on this exact topic. He has said that when teams need to scale AI, “Simply adding more GPUs without ensuring adequate interconnect bandwidth can lead to diminishing returns” in the modern AI era. 

Overall, with its ability to direct priority-aware inference, tenant isolation, capacity sharing, and observability, Red Hat hopes its Red Hat AI 3.5 release gets us to a state where GPU-based AI resources are viewed as a policy-controlled infrastructure pool.

The post Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots appeared first on The New Stack.

  •  

Nvidia and Palantir fine-tune a 30B Nemotron model for Nvidia’s supply chain. It beats a model 18 times its size.

Illustration of two yellow robotic arms on an automated assembly line, reaching toward a conveyor belt beside a server rack with glowing amber cooling fins, depicting AI and the supply chain.

Nvidia and Palantir announced Thursday that they’re working together to bring “sovereign AI to critical supply chains,” kicking off initially with Nvidia’s own sprawling supply chain.

The news builds on a partnership that kicked off last October, when the duo said they would combine Nvidia’s AI computing and models with Palantir’s software to help companies use AI to make complex operational decisions. Then in June, they expanded that effort into sovereign AI, allowing organizations to run and customize Nvidia’s AI models inside tightly controlled environments while keeping sensitive data and model weights under their own control.

Now, they’re applying that technology inside Nvidia itself, where they say a smaller, fine-tuned model is already outperforming a far larger one.

A proving ground for sovereign AI

The companies have fine-tuned Nvidia’s 30-billion-parameter Nemotron 3.5 Lightning model on decisions made by Nvidia’s supply-chain operations team. Palantir’s Foundry and Artificial Intelligence Platform (AIP) bring together the data behind those decisions, while its Ontology acts as a live map connecting components, factories, capacity and production commitments. Nvidia’s cuOpt software, meanwhile, works out how to distribute scarce parts, with Nemotron weighing the wider context and recommending what planners should do.

They then plan to “extend the learnings from Nvidia’s deployment” to companies in other sectors, including manufacturing, energy, healthcare, automotive and aerospace. Palantir’s own customers will be able to build versions tailored to their own supply chains by training Nemotron on their proprietary data using Foundry and AIP, then run the resulting system on-premises or through cloud and colocation providers.

So, in effect, Nvidia and Palantir are putting the sovereign AI partnership they outlined in June into practice inside Nvidia, while using that deployment as a proving ground for an architecture other companies can adapt to their own use cases.

Nvidia as a test case

As the world’s most valuable public company at a $5.4 trillion market cap, there’s good reason for Nvidia to start close to home. Its supply chain spans millions of parts, thousands of suppliers, and a global network of manufacturing partners, with the company saying a single Vera Rubin rack alone contains some 1.3 million parts. Those components have to arrive in the right place at the right time: if one part is missing, assembly can stall while everything else that arrived sits waiting.

“Supply chains are the operating system of the physical economy, and AI factories are among the most complex systems ever built.

Jensen Huang

And that complexity is what Nvidia founder and CEO Jensen Huang says makes supply chains a natural target for the technology. From chips and memory to manufacturing, networking, power and cooling, he argues that building modern AI systems increasingly depends on coordinating an enormous web of companies and components.

“Supply chains are the operating system of the physical economy, and AI factories are among the most complex systems ever built,” Huang says in a statement.

Palantir co-founder and CEO Alex Karp goes further, arguing that Nvidia’s operations provide an unusually demanding environment in which to put the companies’ approach to the test.

“Nvidia has arguably the most valuable, intricate, and complex supply chain in the world.”

“Nvidia has arguably the most valuable, intricate, and complex supply chain in the world,” Karp adds in a separate statement.

The sovereignty selling point

Nvidia has long been positioning itself at the center of the open-model debate. In July, Huang even used his first-ever post on X to promote an industry letter lobbying Washington to support frontier open-weight models, arguing that they give companies and countries more control over their AI infrastructure.

Then in early September, Nvidia swooped in with a $12.9 billion deal for Hugging Face, the so-called “GitHub for AI models.” Amid concerns that ownership by the world’s dominant AI chipmaker could undermine Hugging Face’s neutrality, Huang pledged that it would remain open, continue hosting models from across the industry and support hardware beyond Nvidia’s own.

Nemotron is central to Nvidia’s own open-model push. The name dates back to 2023, when Nvidia released its first Nemotron-3 8B models for enterprises to customize and fine-tune. Those early models were downloadable through Hugging Face and Nvidia’s NGC catalog, although access was gated and governed by Nvidia’s own community license. So they were customizable, and their weights were available, but the much broader “open model” positioning Nvidia uses today came later.

The current Nemotron 3 series arrived back in December, initially spanning Nano, Super and Ultra models aimed at different agentic AI jobs. Nvidia now publishes weights and, for many of the models, training data and recipes so developers can customize themselves. Nemotron 3.5 Lightning, released in August, is the 30B model Nvidia and Palantir have fine-tuned for this supply-chain deployment.

That openness is also at the heart of the whole sovereignty pitch: companies can adapt Nemotron using proprietary data while keeping that data, the model weights, and inference inside their own environment.

Specialization over size

Nvidia’s own deployment gives outsiders a result to chew on. It says the fine-tuned 30B Lightning scored 86.7% accuracy on its supply-allocation task, versus 55.5% for the 550B Nemotron 3 Ultra—a model roughly 18 times its size.

Accuracy scores of post-trained Nemotron Lightning compared against Nemotron Ultra
Accuracy scores of post-trained Nemotron Lightning compared vs Nemotron Ultra (Source: Nvidia)

In a technical blog post published on Thursday alongside the main announcement, Nvidia solutions architects Nell Barber, Rana Haber, and Aastha Jhunjhunwala note that the result shows how far specialization can go. On a tightly defined allocation task, the 30B model outperformed a general-purpose model more than an order of magnitude larger.

“This doesn’t mean the smaller model is more capable overall. Its gains are concentrated in the domain it was post-trained on.”

“This doesn’t mean the smaller model is more capable overall,” they add. “Its gains are concentrated in the domain it was post-trained on. Future production risk forecasting remained difficult despite fine-tuning. Specialization improved the decision task but failed to solve every prediction problem attached to it.

For companies considering Nvidia’s blueprint, the more interesting takeaway may be this: a smaller open model, trained on business specifics, can sometimes be more useful than reaching for the biggest model available.

The post Nvidia and Palantir fine-tune a 30B Nemotron model for Nvidia’s supply chain. It beats a model 18 times its size. appeared first on The New Stack.

  •  

How much control should AI get? A CISO roundtable takes on SOC autonomy

Illustration of a robot with a headset and antenna sitting between empty office cubicles, with a phone handset and looping white cables running around it, depicting an AI agent standing in for a human customer support worker.

Security operations centers have struggled with alerts for years, and AI agents offer a new way to tackle it: Enable machines investigate some of those alerts themselves.

That’s already starting to happen: Security teams are experimenting with AI that can pull together signals from different systems, investigate suspicious activity, and recommend next steps to humans in the loop.

There’s an obvious appeal: SOC analysts have finite time and attention, while the volume of potential threats does not come with the same constraint. Attackers are also getting access to AI tools that can accelerate parts of their own operations. Simply giving analysts better ways to work through an ever-growing queue may only get security teams so far.

But moving from AI-assisted security to increasingly autonomous security creates a new problem: How much control are organizations actually prepared to hand over?

That question is at the heart of The New Stack’s AI-Speed SOC CISO Roundtable on September 15, where security leaders will discuss how far to let AI agents go — and when humans need the final say.

There is a big difference between asking an AI agent to investigate a suspicious login and allowing it to disable the account responsible for it. The same goes for isolating an endpoint, blocking network traffic, or making other changes that could immediately impact the business. An autonomous agent could potentially make those decisions much faster and at much greater scale than a human analyst.

The model is only part of the trust equation. Security teams also need to know what an agent is doing, when a human gets the final say and, crucially, whether they can undo a bad decision. That could mean putting some fairly hard limits on autonomy,  including a way to shut the whole thing down if an agent goes off course.

Giving agents more responsibility also changes the role of the people working alongside them. If AI handles a large chunk of routine investigation, analysts could spend less time working through queues and more time threat hunting, making judgment calls and overseeing the agents doing the repetitive work. The SOC analyst starts to look less like an investigator and more like an orchestrator.

Eventually, the bigger change may be to the SOC itself. “Continuous detection and response” has become familiar security language, but AI agents could make it something more literal. Instead of detection, investigation, and response being separate steps, an agent could move between them, with what it learns during one investigation feeding directly into how the next threat is detected.

That starts to look less like AI bolted onto the existing SOC and more like a different operating model altogether. It also presents CISOs with a familiar problem: tooling. Security teams already have sprawling stacks, and vendors are racing to add agents and AI capabilities to them. Organizations risk ending up with another collection of products to manage rather than the continuous system they were promised.

Join us on September 15, 2026

On September 15, I’ll be joined by Jami Hughes, deputy CISO at Zions Bancorporation, and Oren Saban, co-founder and CPO of Mate Security and former Microsoft Defender XDR and Security Copilot product lead, to discuss alert overload, autonomous agents, the future of the SOC analyst, and what continuous security actually looks like.

The session is limited to 20–25 security leaders, with applications reviewed to keep the group small and relevant. This isn’t a traditional webinar with hundreds of people listening in: everyone in the room will be expected to take part. Chatham House Rule will apply throughout, so participants can speak candidly about what’s working, what isn’t, and where they still have concerns.

Apply for a seat at the table

Because if attackers increasingly operate at AI speed, security teams need to work out how much of the response they’re willing to hand to AI, too.

The post How much control should AI get? A CISO roundtable takes on SOC autonomy appeared first on The New Stack.

  •  

DeepSeek is hiring 150 engineers, and none of them will touch a model

abstract bubbles

Hundreds of thousands of AI agent sandboxes can already run concurrently on a single DeepSeek cluster. Now the company is staffing up to handle what happens as that number — along with its training, evaluation, and other backend workloads — keeps climbing.

Cui Tianyi, who joined DeepSeek in March and works on its Harness team, the group responsible for the infrastructure and environments used to run and evaluate agents, announced in an X post that roughly 150 engineering positions on September 7, with the hiring concentrated in server-side engineering and Agent Elastic Compute rather than AI research. The work spans operating systems, virtualization, networking, storage, scheduling, and the control-plane services that coordinate those resources.

Cui said DeepSeek’s existing backend systems will need upgrades, maintenance, and rewrites as workloads grow. One such system at the center of that scaling challenge is DeepSeek Elastic Compute, or DSec, the sandbox infrastructure DeepSeek built to execute agent workloads during post-training and evaluation.

Cui said DeepSeek’s existing backend systems will need upgrades, maintenance, and rewrites as workloads grow.

Four sandboxes, one SDK

Agent workloads require more than GPUs for inference, with each agent also needing an isolated environment to run code, call tools, change files, and collect the results.

DSec supports four types of those environments through the same Python SDK. Simple function calls go to pre-warmed containers, while Docker-compatible containers handle jobs that need a persistent environment. DeepSeek uses Firecracker microVMs when stronger isolation is needed and QEMU virtual machines for workloads that require a full guest operating system.

That range means the same infrastructure can handle anything from a simple tool call to a software-engineering task that needs an entire OS. It’s a similar challenge to the one the rest of the industry is bumping into as agents move from demos to production. OpenAI, for instance, recently designed custom silicon specifically to address the compute pressure that agent workloads create, and DeepSeek open sourced its own agent harness in August.

Lazy loading agent environments

Every sandbox needs its own environment, but copying complete container or VM images onto every host would consume enormous amounts of storage and network bandwidth while adding to startup time. DeepSeek gets around that by tying DSec into 3FS, the distributed filesystem it originally built for its AI infrastructure, and keeping container base images and filesystem commits as read-only layers backed by 3FS.

The metadata stays local, but the underlying data blocks are fetched only when they’re actually needed. MicroVMs use a similar setup, sharing their read-only base layer through 3FS while writes from individual sandboxes are kept in local copy-on-write layers.

DeepSeek says DSec reduces duplicate page-cache usage across virtualized environments and reclaims memory to allow safe overcommitment, while changes to the container runtime cut the CPU overhead of each sandbox.

The team also had to deal with spinlock contention inside the container runtime. At small scale, the CPU time spent there barely registers. At scale, it limits how densely those environments can be packed onto each host.

DeepSeek says DSec reduces duplicate page-cache usage across virtualized environments and reclaims memory to allow safe overcommitment, while changes to the container runtime cut the CPU overhead of each sandbox.

When replay breaks training

During reinforcement learning and other post-training workloads, large numbers of agent rollouts can be running at once, and jobs may be interrupted as compute gets reassigned. Starting over wastes everything the agent has already done, but picking up where it left off isn’t as simple as replaying its previous commands.

Some of those commands may have changed a file or otherwise altered the environment, so running them again could produce a different result or leave the training trajectory in the wrong state. DSec avoids that with a globally ordered trajectory log that records commands along with their results.

When a rollout resumes, DSec can fast-forward through the completed work using those recorded results rather than executing the commands a second time. That reduces the cost of interruptions across thousands of training and evaluation runs, while the same logs preserve a history of how each sandbox changed and allow earlier sessions to be replayed.

Engineers, not researchers, wanted

The roughly 150 openings reach across DeepSeek’s backend, including the lower-level systems work behind Agent Elastic Compute as well as the services that support its models and agents.

DeepSeek said in June that it planned to at least double the size of every department, but this round of hiring leans heavily toward the systems underneath its models rather than the models themselves. DSec is part of that work, with hundreds of thousands of sandboxes running concurrently and putting pressure on everything from how jobs are scheduled to how they recover after an interruption.

The roughly 150 openings reach across DeepSeek’s backend, including the lower-level systems work behind Agent Elastic Compute as well as the services that support its models and agents.

The post DeepSeek is hiring 150 engineers, and none of them will touch a model appeared first on The New Stack.

  •  

AI agents are creating more work, not less — and OpenAI’s own numbers back it up

abstract bot

OpenAI says it hit a goal it set last fall, stating researchers are now using what the company calls an “automated research intern,” which is an agent that can handle well-defined tasks that would normally take a researcher several days.

The data shows coding-agent use climbing throughout 2026, and by mid-August its agents were logging 3.1 agent-workdays for every human workday across the research organization. The median researcher was spending more than $600 on inference per day at API prices, while those in the 90th percentile spent more than $7,000.

The median researcher was spending more than $600 on inference per day at API prices, while those in the 90th percentile spent more than $7,000.

Agent hours versus useful output

But everyone knows that an agent-workday and a human workday aren’t the same. The company converts the time agents spend working on tasks into standard eight-hour workdays. Because researchers can run several agents at once, the figure tells us how long the agents are working, but not necessarily what they’re completing.

For engineering teams, that leaves plenty of work on the human side, which means running more agents can increase the amount of work happening at once, but it can also increase the amount of work a human needs to keep track of.

OpenAI’s very specific definition of a research intern highlights that it must be able to complete well-defined research tasks that would take a skilled person several days, but a human is still in charge. The company’s next goal, an automated AI researcher, is one they hope to reach by March 2028.

The company’s next goal, an automated AI researcher, is one they hope to reach by March 2028.

Supervision becomes the constraint

Using a taxonomy from Epoch AI, OpenAI broke the agents’ work into six areas — Decide, Design, Build, Run, Analyze, and Communicate — and found activity increased across all six between January and August, although agents still did relatively little of the work involved in deciding what research to pursue.

Much of the work is practical, with agents writing research and infrastructure code, monitoring experiments, and providing enough technical support that OpenAI says attendance at debugging office hours has fallen, prompting one team to stop holding the sessions altogether.

And yet, more agent hours don’t automatically mean more useful research. OpenAI says code output and experiment counts are relatively easy to track, but neither shows how much progress those agents actually made. Compute also increased significantly as the number of experiments rose.

OpenAI used another model to judge how well agents performed on tasks of varying difficulty and found that, despite improving success rates between January and July, humans still had to step in on more than half of successful tasks that would have taken a person four to eight hours.

Security incidents limit Astra deployment

Once engineers can run several agents at once, with those agents launching subagents of their own, the challenge shifts to keeping up with what they produce — catching runs that go off track, reviewing code diffs, and deciding what is ready to ship or feed into a training run.

⁠Astra’s persistent-agent capabilities already let researchers hand off multi-day assignments⁠, which makes this supervisory strain worse, not better.

The company acknowledges that as agents take over more of the execution, the parts of research that are hardest to automate will consume more of an engineer’s time, putting a practical limit on how much agent output one person can realistically review.

On July 20, a series of outages caused by agents disrupted OpenAI’s research infrastructure badly enough that the company took its training container service offline and later brought it back with tighter restrictions.

Nearly a month later on August 7, OpenAI tightened access again after early evidence suggested Astra could reach the “Critical” cybersecurity threshold in its Preparedness Framework, restricting the model to higher-security research areas and adding safeguards that developers may already be encountering as unexpected API interruptions⁠

Workloads shift between models fast

Astra-class GPU allocation fell 59.2% the following week, but that compute didn’t sit idle for long. Researchers moved much of the work to other models, which saw GPU allocation rise 17.2% and made up for roughly 85% of the drop in Astra usage. Instead of reducing the amount of work being run, the restrictions pushed it to other models, showing how easily workloads can move when one part of the system is locked down.

OpenAI’s researchers are handing off larger jobs to agents, running more of them at once and launching more experiments, but whether that translates into faster research is harder to measure — and OpenAI is still figuring out how to price it.⁠

OpenAI’s researchers are handing off larger jobs to agents, running more of them at once and launching more experiments, but whether that translates into faster research is harder to measure — and OpenAI is still figuring out how to price it.⁠

The post AI agents are creating more work, not less — and OpenAI’s own numbers back it up appeared first on The New Stack.

  •  

The systems guide to production token optimization

Collage of a woman climbing progressively taller stacks of coins, with arrows tracing her upward path.

When enterprise AI applications scale, they inevitably hit a wall. For many engineering teams, this wall is initially diagnosed as a billing issue, a monthly API invoice that has grown out of control. However, viewing token consumption purely as a financial metric fundamentally misunderstands how LLMs operate in production. Token optimization, unlike its other optimization cousins, is not an accounting exercise; it’s a distributed systems and hardware utilization challenge.

“Token optimization, unlike its other optimization cousins, is not an accounting exercise; it’s a distributed systems and hardware utilization challenge.”

In this guide, we explore, through the lens of Concierge (a latency-sensitive, synchronous customer support agent) and Pathfinder (an asynchronous, multi-step autonomous CI debugging agent), how these systems fell victim to autoregressive bottlenecks as they grew, and how we fixed these issues.

What you’re actually paying for

A token is not a word: treating it like one will break your budgeting “models.” Every major LLM provider tokenizes text using byte-pair encoding (BPE), breaking words into subword units. While common words stay intact, rarer words or punctuation split into fragments. As a rule of thumb, 1 token = 4 characters, or 0.75 words in standard English prose.

When budgeting for production, you must account for the structural pricing spread; providers bill input tokens and output tokens at different rates. Output tokens are typically 4-5X more expensive than input tokens.

As a baseline, assume a mid-tier frontier model runs roughly $3 per million input tokens and $15 per million output tokens.

The quadratic history tax

LLM provider APIs are completely stateless; to make an LLM behave as if it remembers past events, you must resend the entire history of the session and input with every single API call.

This means that a model’s own previous outputs are continuously re-billed to you as inputs on subsequent steps. This triggers a compounding cost that impacts both Concierge and Pathfinder, though their curves scale differently.

Let S be the static system context (instructions and schemas), u be the incoming data per step, and r be the model’s response payload. The input cost for every given turn k is

Formula defining the input cost for every given turn.

When you sum this across a complete execution run of N steps, the total input token volume compounds quadratically. 

Formula for the total input cost, i.e. the sum of input costs of every given turn in an execution run of N steps.

This O(N^2) accumulation of history is the exact mechanism that causes the explosion in cost and latency.

VariableConcierge (Chat system)Pathfinder (Autonomous agent)
Static context (S)3,100 tokens (Full returns/shipping policies and brand guidelines)1,200 tokens (Tool definitions, system constraints, CI environment data)
Incoming data (u)80 tokens (Short customer chat replies)900 tokens (Massive raw text payloads: log excerpts, file reads, shell outputs)
Response payload (r)220 tokens (Polite customer-facing answers)300 tokens (Internal monologue + JSON Tool Arguments)
Step multiplier (N)10 turns (Average support thread length)15 steps (Average agent troubleshooting loop length)

When we calculate the total input tokens consumed by a single session using the quadratic formula

  • Concierge: Consumed 45,300 tokens per 10-turn ticket
  • Pathfinder: Consumed 150,000 tokens per 15-step turn

Because Pathfinder’s step increment was 4X larger than Concierge, its token cost curve was drastically steeper. If Pathfinder were to get stuck in an infinite tool-use loop and hit 30 steps, a single run could consume 570,000 tokens.

The solution

Fixing the individual call

Prompt hygiene: Hardcoding static reference documentation into the system prompt means you pay to parse identical text on every turn. So we stripped static text from the prompt and switched to dynamic injection. 

For Concierge, we implemented a RAG step to fetch only the 2-3 policy snippets relevant to the ticket. The prompt dropped from 3,100 tokens to 380. A 60% reduction for a 10-turn thread.

For Pathfinder, we applied automated prompt compression using LLMLingua-2 to compress verbose CI log files before sending them to the model. By filtering out non-essential log lines, we reduced the size of incoming tool observations by 3X without sacrificing debugging accuracy.

from llmlingua import PromptCompressor

compressor = PromptCompressor(
model_name="microsoft/llmlingua-2-xlm-roberta-large-meetingbank",
use_llmlingua2=True
)

try:
compressed_result = compressor.compress_prompt(
raw_ci_log_text,
rate=0.33,
force_tokens=["Error", "Exception", "Failed", "Traceback", "FATAL"]
)
# Pass high-density payload to the frontier model
compact_prompt = compressed_result["compressed_prompt"]
except Exception as e:
print(f"Compression failed, falling back to raw log text: {e}")
# Graceful degradation: pass the raw (or truncated) log if compression fails
compact_prompt = raw_ci_log_text

Eliminating the retry: Relying on open-ended prose instructions “return JSON” caused malformation. When parsing failed, the system initiated a synchronous retry, sending the entire accumulated context as if it were a new attempt. We replaced the entire natural language formatting request with strict structural contracts via forced schema validation.

In both Concierge and Pathfinder, we converted the output format to a strict pydantic schema for tool-calling mode and tool-execution payloads. Malformed outputs across both systems dropped to under 0.5%, eliminating tail latency spikes caused by cascading queues.

# Unified Schema Enforcement for Concierge Responses & Pathfinder Tool Execution
from pydantic import BaseModel
from typing import Literal

class TicketResponse(BaseModel):
    reply: str
    category: Literal["shipping", "returns", "billing", "product", "other"]
    escalate: bool
    confidence: float

# The API is structurally locked into emitting validated JSON matching the schema
response = client.messages.create(
    model="claude-opus-4",
    system=SYSTEM_PROMPT,
    messages=messages,
    tools=[
    {
    "name": "respond_to_ticket",
    "description": "Formulate a response and classify the support ticket.",
    "input_schema": TicketResponse.model_json_schema()
    }
    ],
    tool_choice={"type": "tool", "name": "respond_to_ticket"},
)

Output token bounding: Models naturally generate verbose reasoning chains and conversational filler, inflating expensive output tokens. Where the LLM provider exposes logit bias, you can directly suppress every token outside the valid set at decode time; where it doesn’t, constrained decoding libraries (Outlines, Guidance) or a forced tool call with an enum-typed schema will get you the same guarantee.

class ClassifyOnly(BaseModel):
    category: Literal["shipping", "returns", "billing", "product", "other"]
    priority: Literal["low", "medium", "high", "urgent"]

State management 

The stateless nature of the models meant we had to parse the static prompt prefix and historical steps on every turn. We introduced explicit cache breakpoints to allow the inference engine to reuse the states of static blocks. We altered both Concierge and Pathfinder to flag stable, historical segments for caching. Under standard vendor pricing, cache reads are discounted by 90%. It is important to check with your vendor on whether caching is enabled. 

# Caching the stable history prefix for a multi-turn session
response = client.messages.create(
    model="claude-sonnet-4",
    max_tokens=4096,
    system=[{
        "type": "text",
        "text": SYSTEM_PROMPT,
        "cache_control": {"type": "ephemeral"} # Cache hits drop prefix costs by 90%
    }],
    tools=TOOL_SCHEMAS,
    messages=session_history + [{"role": "user", "content": current_step_input}],
)

For a 10-turn Concierge chat, this dropped input costs by ~70%. For a 15-step Pathfinder trajectory, it resulted in a 76% cost reduction.

Semantic caching 

Duplicate queries across separate sessions were triggering redundant frontier model invocations. We implemented a vector similarity cache layer upstream of the LLM using Redis. Our Concierge service analysis showed that 34% of customer support tickets were semantic duplicates of common FAQs.

Intercepting these requests reduced latency to sub-50ms for hits. Because of the nature of CI pipeline logs, we have not yet found a suitable cache for Pathfinder’s inputs. 

import os
import json
import redis
from redis.commands.search.query import Query

# Configure connection via environment variable for environment portability
redis_url = os.environ.get("REDIS_URL", "redis://localhost:6379")
r = redis.Redis.from_url(redis_url)

def get_cached_response(tenant_id, query_text, threshold=0.92):
    try:
results = r.ft(f"cache_idx:{tenant_id}").search( # scoped by tenant -- see below
Query("*=>[KNN 1 @vector $vec AS score]").sort_by("score").dialect(2),
query_params={"vec": query_vec.tobytes()},
)
except redis.RedisError as e:
print(f"Redis cache error: {e}")
return None # Fail-open: gracefully fall back to a cache miss)
  query_vec = embed(query_text) # small, fast bi-encoder -- not the frontier model

try:
results = r.ft(f"cache_idx:{tenant_id}").search( # scoped by tenant -- see below
Query("*=>[KNN 1 @vector $vec AS score]").sort_by("score").dialect(2),
query_params={"vec": query_vec.tobytes()},
)
except redis.RedisError as e:
print(f"Redis cache error: {e}")
return None # Fail-open: gracefully fall back to a cache miss

Semantic caching could be a security problem, because if you choose a global cache, Customer A’s account-specific answer could get served to Customer B because their phrasing embeddings are close enough. To mitigate this, we split the cache into two tiers: a global cache for tenant-agnostic content, and a per-tenant, per-user namespace keyed with the tenant ID baked into the prefix itself for anything touching account state.

Cache poisoning is another risk: we only write to the cache from responses that passed schema validation and the injection-pattern classifier, we stamp every cache entry with its source traceId, and we encourage routine purging of unknown caches.

Context compaction

Uncapped conversation or agent trajectories allowed N to grow continuously, expanding the cost curve and causing latency degradation. We capped N by implementing a sliding window that summarizes historical context via a small, ultra-cheap model. For Concierge, we kept the last 3 turns verbatim while condensing older turns into a rolling metadata block.

For Pathfinder, when the debugging steps exceeded 4 runs, we trimmed and summarized the oldest tool execution outputs into a compact chronological timeline, transforming the open-ended quadratic cost explosion into a predictable, bounded window.

def compact_session_history(history_steps: List[Dict[str, Any]], keep_recent: int = 3) ->     List[Dict[str, Any]];
    """Flattens older history into a cheap summary block, preserving recent context."""
    if len(history_steps) <= keep_recent:
        return history_steps
    old_steps = history_steps[:-keep_recent]
    recent_steps = history_steps[-keep_recent:]
    
# Compress the old history using a fast, low-cost utility model
try:
historical_summary = summarize_with_utility_model(old_steps)
# Note: Anthropic prohibits 'system' roles in the messages array.
# Using 'assistant' ensures cross-provider compatibility.
return [{"role": "assistant", "content": f"[System Context: Summary of prior steps: {historical_summary}]"}] + recent_steps
except Exception as e:
print(f"History compression failed: {e}")
# Fallback: Return the uncompressed history to gracefully degrade
return history_steps

Model cascading

Directing every single operation to an expensive frontier model represents massive overprovisioning for mundane tasks. We integrated LiteLLM as an internal routing gateway to implement model cascading, routing every request to the lowest-cost model capable of completing the task.

“Directing every single operation to an expensive frontier model represents massive overprovisioning for mundane tasks.”

# litellm_config.yaml
model_list:
  - model_name: fast-path
    litellm_params:
      model: openai/mistral-support-ft
      api_base: http://vllm-internal:8000/v1
  - model_name: frontier-path
    litellm_params:
      model: anthropic/claude-opus-4

Simple, repetitive tasks are routed to a lower model, which offloads 70% of Concierge chats from the frontier model. For Pathfinder, we broke the agent loop down into separate sub-tasks: high-level planning, tool selection, and code-patch synthesis remained with the frontier model, while mechanical, text-heavy operations, such as log parsing, regex extraction, and error-string formatting, were offloaded to the lower models. This hybrid orchestration reduced Pathfinder’s token costs by more than 50%.

What’s next

The transformation of Concierge and Pathfinder proves a fundamental truth about production AI. You cannot achieve scale by simply relying on the natural language capabilities of a frontier model. You must engineer the system around it. By shifting your focus from naive token reduction to maximizing system resource utilization, we reclaimed absolute control over the infrastructure.

“Efficiency in the era of gen AI is not defined by how cheaply you can operate but by how densely you can pack information.”

Efficiency in the era of gen AI is not defined by how cheaply you can operate but by how densely you can pack information, how quickly you can serve it, and how reliably you can parse the output. The architectural decisions detailed here represent more than just a token optimization strategy; they are a required foundation for building high-throughput, battle-tested, and resilient AI systems at scale.

The post The systems guide to production token optimization appeared first on The New Stack.

  •  

Your organization prioritized AI adoption, but you actually need AI fluency.

Abstract digital neural network visualization representing central hub-and-spoke enterprise AI infrastructure.

Thanks to increasingly capable models, some parts of your business are getting faster, more capable, and more productive every month. These teams are using artificial intelligence to compress timelines, surface insights, and automate work that has historically been time-consuming and tedious.

Meanwhile, other functions just down the hall are still waiting for a formal rollout, a governance approval, or someone to tell them what to do and how to start. The gap between the AI haves and have-nots in your organization is widening, and addressing it requires a new operating model.

Your teams need more support

When leaders notice the uneven distribution of capability across their business, the instinct is to treat it as a tooling problem. They push to get everyone access to the same platforms, provide general-use training, and hire some specialists to slot into IT.

But access is table stakes. It’s a good start, but it won’t get you to strong organizational adoption.

Harvard Business School reports that workers using these tools completed tasks 25% faster and produced results rated more than 40% higher in quality. But the same study also found that performance declined when people used the tools without understanding where they applied and where they didn’t. Fluency, not just access, drives results.

Departmental leaders need guidance on how to apply capabilities in the context of their day-to-day work. Without that knowledge, they can’t ask the right questions.

“Performance declined when people used the tools without understanding where they applied and where they didn’t. Fluency, not just access, drives results.”

Teams playing catch-up tend to focus on how to inject new tools into existing workflows, when they should be thinking about re-engineering processes entirely. They’re focused on evolution in a world undergoing revolution.

Reimagining a process also requires stepping back from it, which is easier said than done. Here’s how it plays out in practice:

An SDR team comes to IT with a specific, bounded ask: “improve our sales lead routing.” Completely reasonable. But only when someone from IT, with visibility across the broader system, dives into the problem does the real opportunity surface. The data pipeline supporting lead routing is unnecessarily complex. With the right support, the conversation shifts to overhauling the entire pipeline and opens the door to fully agentic lead follow-ups.

Departmental leaders don’t lack ambition but throwing a software license and Slack channel at them won’t build the right kind of adoption. Technical support and strategic guidance are required to reimagine work from first principles.

AI fluency must be a structural consideration

The typical pattern puts a centralized team in charge of taking requirements, interpreting them in isolation, and delivering capabilities to departments months later. This model can’t keep pace when AI capabilities launch weekly.

A more effective approach pairs a central “hub” that owns platform strategy, governance, and reusable patterns with AI engineers embedded directly inside business departments. AI engineers serve as “spokes” inside departments, helping them identify vertical use cases day-to-day and delivering the cross-functional visibility needed to make a real impact. The AI engineer who solved a problem for finance can share the pattern with someone facing the same challenge in operations.

“A more effective approach pairs a central “hub” with AI engineers embedded directly inside business departments.”

In a department just getting started, the embedded AI engineer is the primary technical capability: scouting, prototyping, building. In a more mature department, they shift toward enablement, feeding patterns back to the “hub” and helping teams navigate AI without getting buried in process. Over time, departments will organically become AI-fluent as they learn from the engineers.

Make fluency your advantage

The right operating model drives how a function actually works, and strong fluency strengthens processes and institutional knowledge, so outcomes improve over time. As the flywheel builds, each problem solved raises the ceiling of what your team can do independently. 

McKinsey finds that the right workflow redesign is the single biggest factor in whether an enterprise sees meaningful bottom-line impact. Knowing what to redesign depends on how your teams understand and work with AI.

Everyone is adopting AI capabilities. The question now is whether your operating model helps your teams see the best path forward for applying them. If it doesn’t, that’s the gap to close first.

The post Your organization prioritized AI adoption, but you actually need AI fluency. appeared first on The New Stack.

  •  

OpenAI wants to charge only when AI gets it right — here’s the catch

abstract yellow grade

AI companies have always charged customers for the tokens they use, whether the model gives them exactly what they need or completely misses the mark. Now, OpenAI is experimenting with a different approach by allowing customers to pay only when the AI gets the job done right.

First reported by The Information, the company has already started testing that approach with some enterprise customers, billing them based on successful outcomes rather than simply the amount of compute consumed along the way.

OpenAI has not publicly disclosed pricing or exactly how it determines when a particular task counts as a success. Still, the approach presents an interesting technical challenge for developers building AI agents. If a customer only pays when an agent completes a task, someone needs to determine exactly when that task is complete. As we know, determining if an AI agent succeeded is more complicated than counting tokens.

Counting tokens, not success

Some outcomes are easy for software to verify, such as a support ticket closing without human involvement, but other jobs leave room for interpretation. For example, take a coding agent asking to fix an authentication bug. It might rewrite the code and pass every test, only for the patch to cause another problem once it reaches production. In that case, the agent technically completed the task, but the customer probably wouldn’t consider it a successful outcome. 

With outcome pricing, getting through 90% of a task may not be enough for the run to count as billable.

The same issue arises with longer-running agents that interact with browsers, databases, APIs, and other systems. An agent may carry out nine of the ten steps before failing at the last one. With token pricing, all of that activity can still be charged for. With outcome pricing, completing 90% of a task may not be enough for the run to be billable.

Evals become billing infrastructure

Developers already use evals to catch problems with models and agents. OpenAI’s hosted tools, for example, can check responses against expected results and grade a model’s performance on a given task.

Braintrust adds visibility into what happens during an agent run. It records model calls, retrievals, and tool calls in a trace, then scores the run on factors such as task completion, factual accuracy, and correct tool use. Developers can also turn those traces into datasets for future testing.

Some results are easy: a unit test either passed or it didn’t, an API returned the expected response, or a database contains the record it was supposed to create. There isn’t much room for debate. 

It records model calls, retrievals, and tool calls within a trace, then scores the run for things like task completion, factual accuracy, and correct tool use.

Semantic evals are different because they require a judgment rather than a pass/fail. An LLM-as-a-judge can help developers compare two versions of an agent, but using that judgment to trigger a charge is another matter. A false positive could leave a customer paying for unfinished work, while a false negative could leave the vendor covering the cost of a successful run.

Grading your own work

When an AI company runs the agent and sets the criteria for success, it is effectively grading its own work and then billing the customer for the result.

That gets trickier with more subjective work. Asking an agent to generate a monthly sales report gives you something you can check. But asking it to generate a good one is different because someone still has to decide whether the result is actually any good. And if the goal is to improve conversion rates, figuring out how much of that improvement is attributable to the agent is even harder. 

There’s also the question of who gets blamed when something outside the agent fails. An agent might handle a support request correctly only to hit a timeout in the customer’s CRM. A coding agent could finish its work but fail because a separate service is unavailable. If those runs don’t count as successful, the vendor could end up paying for failures it didn’t cause.

Failed agents shrink margins

Outcome pricing changes who pays when an agent fails. Under token pricing, an agent can burn through tokens and retry failed steps without ever finishing the job. The customer still pays for that usage.

If it never finishes, the provider has spent money on compute without anything to bill.

If the customer pays only for successful outcomes, failed runs become the provider’s expense. An agent that completes a task on its first attempt is more profitable than one that requires 20 model calls and several retries. If it never finishes, the provider has spent money on compute resources without anything to bill for.

The post OpenAI wants to charge only when AI gets it right — here’s the catch appeared first on The New Stack.

  •  

The 3 roles AI agents play in your developer platform

Dark abstract 3D metallic waves illustrating complex AI agent roles in developer platform architecture.

Engineering organizations are trying to deliver as fast as technology allows, bringing agentic AI into their developer platforms and working out how to use AI agents to maximize engineering productivity.

But when it comes to AI agents in an agentic developer platform, every team we talk with sees the role of those agents a little differently. Following hundreds of calls with our customers, we captured three types of roles.

Role 1: AI agents as platform consumers

In this role, an agent is basically a user of the platform. It uses the platform as part of its task to read context and to run actions. When agents first showed up, a lot of companies said they treat their AI agents like employees, and in a development platform, that makes the agent just another engineering resource consuming it.

Diagram showing an AI agent reading context and running actions from an agentic SDLC platform.

A typical case: an engineer asks Claude Code to add an endpoint to the payments service. Before it writes any code, the agent pulls the service owner, dependencies, and the standards it must meet from the platform, then spins up a preview environment via a self-service action and runs the tests.

“A lot of companies said they treat their AI agents like employees, and in a development platform that makes the agent just another engineering resource consuming it.”

For this to work, the agent has to reason over real, current information about your systems, starting with the service catalog and extending to ownership, dependencies, standards, and current state. If you get the context wrong, there’s a good chance an agent will get overconfident and do the wrong thing. Most teams solve this one agent at a time by providing local context. But then the same information ends up connected to agents in fragile ways, not to mention that none of them are governed. Compare that to a context lake that provides every agent with a single governed source of truth. The platform also has to be reachable the way an agent works, which means being API- and MCP-first.

What does it require from the platform? An API and MCP-first interface, a governed context layer the agent reads from, and a set of self-service actions it can call.

Role 2: AI agents as internal platform components

Platforms that can register agents and run them within workflows are using AI agents as part of a full business process. The agent runs within the platform, is triggered by an event rather than requested by a person, and sits in the orchestration engine alongside the deterministic steps.

Diagram showing an AI agent as an internal platform component

Run a nightly scan that flags vulnerable dependencies across 40 services. The platform pulls the remediation agent from the registry and runs it once per service, so every owning team wakes up to an open PR awaiting review.

What does it require from the platform? An orchestration layer to run the agents, a registry to pull the right agent from, an identity per agent so the action is logged against the agent rather than a borrowed human credential, and a human-in-the-loop step where the risk is significant.

Role 3: AI as a resource with its own lifecycle (a.k.a AgenticOps)

In this role, the agent is a resource like any other, as are the LLMs, MCP servers, and the skills that come with it. The platform provisions them, governs them, and hands them back, just as it does with a service, a database, or an environment.

Diagram showing AI agents as a resource with a lifecycle, being requested by an engineer.

For example, say an engineer needs an on-call triage agent. They pick the model, the tools, and the environment it runs in, either through a form or by describing what they need, and the platform provisions everything needed, a bit like a vending machine, with the addition of a well-governed agent, in the right standards. 

“That makes it a golden path problem. A golden path is the route that, by default, gets a team a resource the right way.”

That makes it a golden path problem. A golden path is the route that, by default, gets a team the right resource, and the agent lifecycle needs one: request it, get it provisioned and registered, and publish it for the next team. It is also what customers ask us for most, with an agent and skill registry raised by 47% of the organizations we spoke with through early 2026. I went into this in more depth in our golden paths post.

What does it require from the platform? A self-service path that provisions the runtime, issues the identity and scoped credentials, wires in the approved context, and registers the agent on the way out, plus a route to publish it for the next team.

The three roles at a glance

RoleWhat the agent isExampleWhat it requires from the platform
Role 1: AI agents as platform consumersA user of the platform, reading context and running actions as part of its taskClaude Code pulls the service owner, dependencies, and standards, then spins up a preview environment and runs the tests.A governed context layer, such as a context lake, and self-service actions it can call.l
Role 2: AI agents as internal platform componentsA step inside a workflow, triggered by an event rather than requested by a personA nightly scan flags a vulnerable dependency across 40 services, and the remediation agent opens a PR for each one.An orchestration layer to run agents in, a registry to pull the right agent from, an identity per agent, and a human in the loop where the risk is real
Role 3: AI as reusable building blocks (AgenticOps)A resource the platform provisions, governs, and hands backAn engineer requests an on-call triage agent, picks the model and tools, and gets one back already registeredA self-service path that provisions the runtime, issues the identity and scoped credentials, wires in the approved context, and registers the agent

Sometimes the three roles connect

The chained case is an interesting one. An engineer requests a triage agent through role 3; it’s added to the agent registry as part of the creation workflow, and a week later, an incident workflow calls it as a component (role 2). When it runs, it reads service ownership and recent deploys out of the same context lake, which is role 1. It is the same agent throughout, and which role it is in depends on when you look at it.

“It is the same agent throughout, and which role it is in depends on when you look at it.”

What does a platform that covers all three look like?

That is what we built Port for. An agent can use Port as a user through a service account, reading the context lake and running self-service actions. Agents run within Port workflows as part of a business process. And AgenticOps runs as self-service workflows, so a team can request an agent and receive a registered one. All three sit on the same catalog, the same context, and the same audit trail.

If you want the full picture, it is in our playbook, From Agentic Chaos to an AI-Native SDLC.

The post The 3 roles AI agents play in your developer platform appeared first on The New Stack.

  •  

Observability has a data problem. AI is about to make it worse.

Parallel orange lines form a flowing wave across a dark purple background.

Observability is entering a new phase now that OpenTelemetry has standardized instrumentation for data collection. Unfortunately, the observability industry still lacks a cost-effective way to store, retain, search, and analyze full-fidelity telemetry data. This results in blind spots in observability and many teams operating without full operational visibility.

As AI systems generate more logs, traces, and metrics — thereby making the blind spots issue worse — Bronto, a Dublin, Ireland, firm offering an intelligent data observability platform, is betting that the next observability platform battle will be won at the data layer, not the dashboard layer.

Bolt-ons and incremental efficiency aren’t enough

Trevor Parsons, co-founder and co-CEO of Bronto, tells The New Stack that the industry has been optimizing at the edges rather than rebuilding the economics and architecture of telemetry storage. The industry has introduced a wide array of “hacks” and “capabilities” to avoid tackling this issue head-on and ultimately to protect their margins. 

“If you are a couple of times cheaper or 50% cheaper, that ain’t going to cut it,” Parsons says, because data volumes, especially AI telemetry, are growing so quickly, on top of already stretched observability budgets and inefficient datastores. 

Promises, Promises, Promises…

Parsons elaborates, “Observability has always and continues to have a data problem.”

The eternal promise of observability has been delivering teams a clearer view of what’s happening inside their systems.

“Observability has always and continues to have a data problem.”

But in practice, that view is often incomplete, expensive, and short-lived. For too many teams, observability has become less about asking better questions and more about fighting the cost and complexity of storing the data they already need. 

“Sometimes people frame that as a cost problem, where they’ll say observability is up to 20 or 30% of your infrastructure spend,” Parsons says. “I actually think this minimizes the issue; it’s much bigger than that. Teams are actually paying 10, 20, 30% of their infrastructure spend for access to only a sliver of their data.” 

Noel Ruane, co-founder and co-CEO of Bronto, frames the challenge that organizations face and tells The New Stack, “Agents and applications are generating more logs, traces, and metrics each day. The software landscape has accelerated, but are observability vendors keeping pace? No, they’re offering workarounds, bolted-on features, and asking teams to accept blind spots.” In short, Ruane says, they’ve failed to solve the data problem.

Out with the old observability model 

“Customers are not getting access to all of their observability data, Parsons explains. “They have to cut their retention from 30 days to seven days to three days. They have to sample data. They have to rehydrate data.”

In other words, today’s tools make customers choose which parts of their own data they’re allowed to see, and you may only get to see it for a short amount of time.” 

“The solutions that are being put in front of customers to give them their data are always full of compromises, forcing teams to choose between cost, coverage, and speed of data access. The burden is always put on the customer by vendors.”

“The solutions that are being put in front of customers to give them their data are always full of compromises, forcing teams to choose between cost, coverage, and speed of data access,” Parsons says. “The burden is always put on the customer by vendors.

“But really this should be the other way around; it’s the vendors’ job to innovate on behalf of the customer” 

OpenTelemetry: Collection solved, storage problem exposed

Severin Neumann, head of community at Bronto, tells The New Stack that OpenTelemetry has helped standardize instrumentation and data collection, while reducing reliance on proprietary agents.

But that success has created a new bottleneck, says Neumann, who is also an OpenTelemetry maintainer and member of the OpenTelemetry governance committee. Now that organizations can collect more telemetry, they need somewhere affordable and useful to put it.

“We have fixed the instrumentation problem,” Neumann says. He cautions that enterprises now need ways to handle all this data. And if enterprises can’t store it and instead throw away large parts of it, humans and agents can not make sense of it.

The observability business model doesn’t align with customer value

The legacy observability tool business model charges customers for data storage, rather than the value teams get from their data, Parsons says. Customers tell him the same thing constantly: “I pay the same price even if I never search my data.” In many cases, customers find existing tools difficult to use and feel that their observability solution is just a really expensive data store that they do not get a lot of value from.”‘

Noel Ruane assessed the market by saying, “Traditional vendors like Datadog know their pricing model isn’t sustainable. They’ve introduced defensive features like ‘Flex Logs’ and a new ClickHouse partnership to try to keep customers from jumping ship, but they’ve only added new complexity for their customers.” 

Especially in the AI era, Ruane adds, the traditional business model charges teams in ways that discourage them from capitalizing on their data. Customers should pay much less for data that sits idle and more when they actually derive value from it with queries and analysis.

Bronto’s technical differentiation

Bronto isn’t selling another observability dashboard. It argues that observability is a storage problem before it’s a visualization problem, and that’s where the company went.

Underneath the platform is a custom-built polymorphic data store called BrontoDB, specifically designed for observability data. The pitch: enterprises can keep more than 100 times the observability data they hold now, and it won’t get slower or harder to use.

Why that matters comes down to how the three signals break. Metrics, logs, and traces each hit a wall at different points, and Bronto says it built BrontoDB to tackle these issues head-on. Parsons is blunt about two of them.

“With metrics, we’ve solved the high cardinality problem where costs traditionally explode with high cardinality metrics,” Parsons says. “With logging, we’ve solved the indexing problem where there was always a trade-off between fast logs and paying through the nose for it or having slow logs and getting them slightly cheaper.”

  • High cardinality is what wrecks metrics pricing. Add enough unique dimensions and the bill lands somewhere nobody forecast. Bronto says it was built specifically to take that surprise out.
  • Logs have always been pick-your-poison: fast and expensive, or cheap and slow. Bronto says that choice goes away — sub-second search across petabytes, no shortened retention windows, no rehydrating cold data, no waiting.
  • Traces, Bronto argues, shouldn’t be sampled at all. Sampling exists because tools and pricing models couldn’t handle the full stream. Bronto says teams can send it all.

Billing works differently, too. Most vendors charge for data sitting in storage, whether anyone touches it or not. Bronto charges closer to what teams actually search and analyze. That’s the piece that has to hold up if full-fidelity observability is going to be affordable at AI scale.

AI is what raises the stakes, Parsons says. AI systems are non-deterministic and trace-heavy. They throw off more telemetry, and the data has to stick around longer if you want to debug effectively. 

He points to an upside as well. As operations become more automated, telemetry data becomes more useful because agents can chew through volumes of history that no SRE would ever read manually.

AI raises both the volume and the stakes, according to Parsons. AI systems create more telemetry because they are non-deterministic, trace-heavy, and require longer retention for troubleshooting. At the same time, AI-enabled operations will make historical telemetry more valuable because agents can analyze far more data than human SRE teams could manually inspect.

“If AI is the intersection of where data meets intelligence, you can not apply intelligence if you do not have the data.”

“If AI is the intersection of where data meets intelligence, you can not apply intelligence if you do not have the data,” Parsons says.

The next observability battle 

AI is unlikely to fix observability’s data problem. In fact, it will produce more telemetry, create more edge cases, and increase the cost of missing the right signal at the wrong time.

For Bronto, the data layer is the next major battleground. Dashboards still matter, but in an AI-heavy production environment, the more important question may be whether teams have access to all their data for as long as they need so that they can apply AI to it. 

The post Observability has a data problem. AI is about to make it worse. appeared first on The New Stack.

  •  

Tokenmaxxing is out. How to minimize AI spend without sacrificing security capability.

Dark abstract 3D digital wireframe grid symbolizing AI security detection architecture and data filtering.

Security teams are discovering that the most capable AI models cost too much to run on routine, high-volume work, and they’re finding that out after the first invoice arrives. I lead a security operations team, and I’ve brought detection costs down to roughly $1 per day for trust and safety work. 

That number surprises people, because enterprise AI is supposed to be expensive. It isn’t, if you design the work correctly. The teams burning through budget are skipping the design work that determines which cases should reach a model at all.

The funnel is the cost lever

Detection funneling isn’t new. Before LLMs, security teams built layered filters to narrow high-volume event streams down to the cases worth a human analyst’s time. The same logic applies to AI spend. The narrower and more precise the funnel feeding your models, the lower your cost per accurate outcome.

“The narrower and more precise the funnel feeding your models, the lower your cost per accurate outcome.”

In trust and safety, most abuse is identifiable before any model runs. Deterministic, rules-based pattern detection captures a significant portion of volume upfront. Account age, email provider, and behavioral signals all feed automated filters that resolve the obvious cases and narrow what remains. Only the subset that clears those filters reaches an LLM. That’s where the dollar-a-day figure comes from. A well-designed funnel keeps expensive work to a minimum.

Think of it like a home security system. You don’t need a camera monitoring everyone who walks past your house. You care once someone’s actually inside, and that’s when you bring in the higher capability response.

Tiered models, tiered cost

For cases that clear the initial funnel, a lightweight model handles the first pass. The output is structured — a determination of malicious or benign at high or low confidence. High-confidence outcomes resolve automatically, while low-confidence cases escalate to a more capable model with broader context and stronger reasoning.

Only a fraction of events reach that second tier. We ran structured efficacy testing across model options and found only a 1-2% difference in accuracy between lightweight and frontier models for our use cases. Frontier models cost roughly five times more per token. That math only works if the cases reaching them need that capability.

Prompt engineering matters as much as model selection. One prompt I wrote for agentic detections runs over 1,900 words, covering every scenario the agent is likely to encounter, including when to escalate and when to act autonomously. Not every case needs that depth. Some trust and safety prompts are two or three sentences, but the scope of what you’re asking an agent to handle determines the precision the prompt requires.

“Frontier models cost roughly five times more per token. That math only works if the cases reaching them need that capability.”

Context is what separates accurate AI analysis from hallucination. Give a model an abuse report and ask whether the user is abusive, and it may take the report at face value. Give it the actual artifact being reported, along with the report, and it can independently assess whether the claim holds up.

Where humans still belong

Automation handles the clear cases. It’s the ambiguous ones that need judgment that agents don’t yet have.

Distinguishing a legitimate security researcher who hosts malware samples for analysis from a malicious actor who hosts the same content to target others requires discernment that AI still struggles with. So does a dispute in an issue thread where the terms of service could be interpreted differently. These cases reach a human because the question itself requires contextual reasoning the current generation of agents can’t reliably provide.

“Automation handles the clear cases. It’s the ambiguous ones that need judgment that agents don’t yet have.”

The goal is to give human reviewers the time to spend on the cases that actually need them. Instead of clicking through individual events, engineers on my team are building systems, writing prompts, and defining patterns that orchestrate detection at scale.

What this means in practice

Attackers try new obfuscation techniques, and we adapt prompts and models to catch them. It’s iterative work, closer in spirit to detection engineering than a one-time deployment. Prompt engineering is just another form of that: write the rule, test the output, tune when accuracy slips.

The cost question is solvable if you treat it as a design problem from the start. How much reaches a model, in what form, and with what context determines the bill. Most teams that find AI expensive haven’t made those decisions deliberately.

The post Tokenmaxxing is out. How to minimize AI spend without sacrificing security capability. appeared first on The New Stack.

  •