❌

Vue normale

Reçu hier — 27 septembre 2026

Performance engineering from kernel analysis to AI: Adrian Cockcroft’s take

27 septembre 2026 à 17:00
Abstract dark digital landscape with glowing contour lines representing multidimensional performance data and response time distributions.

Over its five-year history, P99 CONF has hosted quite a few speakers who’ve offered pointed takedowns of the namesake metric. At last year’s conference, Adrian Cockcroft didn’t explicitly state that P99s are BS… but he did allude to it.   

If you don’t know Cockcroft, he’s spent decades architecting, scaling, and optimizing resilient, high-performance systems at giants like Sun Microsystems, Netflix, eBay, and Amazon. We could probably dedicate an entire day of P99 CONF to discussing the lessons learned from just some of his projects (Solaris kernel performance, multi-processor optimization, Netflix’s on-prem to cloud migration, Chaos Monkey…) 

Fortunately, RedMonk analyst Rachel Stephens proved the perfect host for a conference that’s all about making things fast. She sat down with Adrian and led us on a whirlwind tour of how AI has impacted performance engineering. Here are some highlights from the chat (full video below).

Note: P99 CONF 2026 – a free + virtual conference on all things performance – is going live October 21-22. Grab a complimentary pass and join us!

From kernel analysis to vibe coding perf tools 

As a performance specialist at Sun in its heyday, getting to the root of performance problems involved lots of digging and divination. Cockcroft recalls, “Back in the old days with Sun, people would look at the output of system metrics in vmstat or whatever, and they’d be guessing what the numbers meant. There was a very vague understanding of what these things meant. The manual page wasn’t very clear.” 

Cockcroft ended up going to the source, literally. “I went and read all the kernel source code and figured out exactly where these numbers came from, exactly what they meant, which ones were approximating what, and wrote all that down.” That led to two performance books: Sun Performance and Tuning and Resource Management.

“My speedup is infinite, because this code would never exist without these tools. I wouldn’t have the time to build them.”

Four decades later, there’s now a wealth of helpful tools for end-to-end tracing, but Cockcroft’s curiosity still lies in what the tools are not showing. He continued, “Everything sort of looks okay in the tools – but the system isn’t behaving well. I usually come in and try to find a new way of looking at the data. A new type of analysis, or go a little bit deeper or finer grain, or stop looking at averages and start looking at distributions, and find all kinds of interesting things that nobody knew were happening.”

Currently, he’s vibe coding tools to better analyze the anomalies he finds. Saved from having to brush up on Python or hunt down graphics library fragments on Stack Overflow, Cockcroft can now stand up custom tooling in minutes. “My speedup is infinite, because this code would never exist without these tools. I wouldn’t have the time to build them.”

Peaks not percentiles

One specific vibe coding project: Cockcroft built (and open-sourced) tooling to get a better understanding of response time distributions. 

Response time distributions have been on Cockcroft’s mind for over a decade. While most people obsess over percentiles – yes, P99 CONF included – Cockcroft is most intrigued by the distribution of response time peaks in a histogram. He believes percentiles don’t work when trying to understand the latency and performance of modern web services. A single number like P99 can’t tell you whether the underlying distribution has one peak or several. And when there’s more than one peak (as is often the case in the real world), the mean, the standard deviation, and even the P99 itself lose most of their meaning.

“Percentiles don’t work when trying to understand the latency and performance of modern web services.”

Image showing what people think response time distributions looks like vs what they really look like
(source: A Tale of Two Histograms)

For example, assume you have a histogram with two response time peaks: a fast one from a cache hit and a slow one for misses that require actual work. As the cache hit rate shifts, each peak’s position remains the same (i.e., the latency values of the fast-response mode and the slow-response mode don’t change), but the peak heights rise and fall. “Your averages and your P99 are changing all over the place, but all that’s really happening is your cache hit rate is changing,” Cockcroft said.

So how do you go beyond measuring P99s and averages? Cockcroft did what he’s done for decades: dive in and build a custom tool. But these days, it’s much simpler thanks to LLMs.

“Your averages and your P99 are changing all over the place, but all that’s really happening is your cache hit rate is changing.”

He had already worked out the statistical approach for analyzing the distribution. Once ChatGPT came out, he quickly used it to build a tool that automated it. Instead of collapsing everything into an average, it identifies an arbitrary number of peaks in a distribution and tracks how they fluctuate over time. It’s implemented in R – a language Cockcroft hadn’t used in a while, but ChatGPT knew quite well – and it’s open source. If you’re curious, learn more in his Percentiles Don’t Work article and “A Tale of Two Histograms” talk and deck (“It was the best of response times, it was the worst of response times…”)

Where do we go from here?

To close, Stephens asked Cockcroft what advice he’d share with teams working on high-performance systems today. His top tip was to start with the macro view to find what’s interesting, then keep digging deeper until you’re inspecting individual slow requests end-to-end.

“Remember the microscope that you got when you were a kid,” Cockcroft said. “First, you have to focus it using the lowest resolution, at 10x, and then you can click it to 100x and adjust that, looking at just one speck now. Once you get that in focus, you click it to 1,000x.”

Cockcroft has spent his career building tools that bring obscure performance issues into focus. We look forward to seeing what others have cooked up with agentic tooling to help identify and solve performance problems this year at P99 CONF. 

Learn about the latest performance optimization techniques, tooling, and case studies at P99 CONF – free and virtual, October 21-22. Grab a complimentary pass and join us!

The post Performance engineering from kernel analysis to AI: Adrian Cockcroft’s take appeared first on The New Stack.

Reçu avant avant-hier

OpenTelemetry and Prometheus are getting along. What’s still missing?

25 septembre 2026 à 18:56
Abstract 3D illustration of metallic blue spoked hubs connected by purple tubes against a pink background.

Welcome to another edition of Road to KubeCon, where we’re tracking the major movements in the Kubernetes and cloud native ecosystem on the path to KubeCon + CloudNativeCon NA 2026, happening Nov. 9–12 in Salt Lake City, Utah.

This week, we look at how cloud-native teams are putting observability to work. There’s progress on OpenTelemetry and Prometheus interoperability, a migration spanning 100,000 hosts, and new data on the costs and benefits of monitoring AI systems. Plus, HPE’s latest Gartner recognition, agent governance updates, and a father-and-son story from KubeCon India.

HPE GreenLake named a Leader in Gartner quadrant

On Wednesday, HPE announced it had been named a Leader in Gartner’s Magic Quadrant for Infrastructure Platform Consumption Services for the second consecutive year.

Hewlett Packard Enterprise (HPE) is a presenting sponsor of Road to KubeCon. HPE Software helps IT organizations modernize infrastructure, streamline operations, and accelerate AI initiatives across hybrid, multi-vendor environments.

GreenLake, HPE’s cloud operations platform, helps teams monitor resource consumption, secure data, and manage infrastructure across data centers and private and public clouds. HPE points to recent additions, including agentic AI-powered operations, as part of the platform’s development.

Varma Kunaparaju, senior vice president and general manager of cloudops software and platform at HPE, says in the announcement: “We are building the operating model and platform for the agentic enterprise, giving customers the ability to simplify operations, govern intelligently, continuously optimize, and modernize without sacrificing choice.”

OpenTelemetry and Prometheus work better together

OpenTelemetry (OTel) and Prometheus are widely used for cloud-native monitoring and observability, often side by side. A new survey looks at how well that combination works.

Published Tuesday, the 2026 survey on Prometheus and OpenTelemetry interoperability found that nearly half of respondents mix Prometheus- and OTel-style instrumentation for infrastructure metrics. For application metrics, 30.7% use both.

While the two ecosystems haven’t always worked well together, the 2026 survey shows improvement: the average ease-of-use rating rose 0.5 points, from 3.1 to 3.6, while the share of those who find the two hard to use together fell from 29% to 10%.

As OTel contributors Dhruv Ahuja of SigNoz, Grafana Labs‘ Andrej Kiripolsky and Arthur Sens, and Ana Muenz share: “Two years of work on interoperability is paying off.” 

There’s still work to do. Respondents want better alignment between the projects’ data models, better handling of resource attributes and metadata, and fewer naming and formatting issues.

Atlassian moves metrics from 100,000 hosts to OpenTelemetry

A case study on the Cloud Native Computing Foundation (CNCF) blog details how Atlassian migrated its metrics collection to OpenTelemetry from gostatsd, its open-source Go implementation of Etsy’s StatsD.

The original pipeline had worked for years, handling metrics from roughly 100,000 hosts across 14 regions, but was increasingly out of sync with the shift to OTel. “It became the thing everyone standardized on, and more and more of what fed our pipeline was emitting OTel data we simply didn’t support,” write Atlassian’s Iris Grace Endozo, Farzad Vazirnia and Albert Kerr.

To maintain continuity throughout the migration to OTel, Atlassian swapped the collection and pipeline mechanics underneath while keeping the service-facing interface unchanged. This turned an organization-wide overhaul into what the authors call a “platform-team migration.”

According to the authors, aggregation now uses about half the CPU for the same traffic. Operations are more unified through the OTel Collector, CPU usage is more evenly distributed across ingest shards, and sidecar costs are down roughly 30% at fleet scale.

New Relic finds observability gains — and gaps

On Tuesday, New Relic released its 2026 Observability Forecast, based on a survey of 2,575 IT and engineering leaders and practitioners. The report found that 73% are standardized on OTel, actively migrating to it, or testing it.

The report also looks at observability’s role in AI adoption. It found that 83% of respondents consider observability essential for AI-generated code. Organizations monitoring AI agents are twice as likely to report a threefold return on observability investment as those running agents without monitoring.

The study also paints a picture of the impact of outages. Engineers now report spending 37% of their time addressing disruptions, while 42% of organizations learn about disruptions through inefficient channels, like manual checks or customer complaints.

Outages take a business toll. New Relic found that organizations lose $74 million a year on average due to high-impact outages. That’s $1.85 million per hour, or over $30,000 for every minute a system is down.

The findings show why teams are looking for ways to detect and resolve problems faster as their systems grow more complex.

As Kubernetes evolves, so do the demands on the teams running it. Presenting sponsor HPE helps teams address that complexity with software spanning virtualization, cloud management, observability, and automation.

Observability Day returns to KubeCon

If you’re into observability and attending KubeCon NA, definitely check out the agenda for Observability Day, happening during the co-located events in Salt Lake on November 9.  

OpenTelemetry’s graduation in May and growing production use give teams more experience to draw on as they adopt the standard.

New AI workloads, inference monitoring, and interoperability with other projects still present challenges. Those issues give practitioners plenty to compare notes on.

According to the Observability Day schedule, the agenda includes project updates and lessons from Capital One, Cisco, Nubank, and other organizations.

“Observability Day provides a vendor-neutral place for maintainers and practitioners to compare approaches and learn how the wider ecosystem is responding,” write Austin Parker, Iris Dyrmishi, Eduardo Silva Pereira, and Juraci Paixão Kröhling on the CNCF blog.

Komodor adds controls for agentic operations

Technically one we skipped last week, but potentially interesting vendor news nonetheless: Komodor, the site reliability engineering platform, announced its Komodor Agentic Operations Platform on Wednesday, September 16.

The additions let engineers deploy autonomous workflows and build or import agents under shared governance and context. Komodor says the release responds to the growing use of agents, including coding agents, and concerns about governance, return on investment, and costs.

“The hardest part of running agentic operations in production is not building the agents,” shares Itiel Shwartz, Komodor’s co-founder and CTO. Instead, the challenges lie in maintaining context, persistent memory, accuracy, security, and cost control — things the new Komodor release aims to address.

Spectro Cloud expands in the Middle East

The Middle East is an increasingly important technology market and a hotbed of data center construction. At the same time, data sovereignty and compliance requirements are driving interest in sovereign infrastructure.

This week, Spectro Cloud announced plans to expand in the Middle East, including new local partners and a dedicated regional office.

“The Middle East has bold ambitions for global AI leadership, from sovereign AI factories to AI-powered economies,” shared Tamer Riyal, Spectro Cloud’s sales director for the Middle East, in the announcement. “We’re investing in the regional expertise and partnerships to support that vision for the long term…”

The company aims to expand adoption of PaletteAI, its platform for managing AI infrastructure across sovereign clouds, enterprise data centers, and edge locations. The expansion reflects the region’s growing role in cloud-native infrastructure.

A father and son take the KubeCon stage

This week’s updates also have a personal side. Earlier this year, analyst, advisor, and TNS columnist Janakiram MSV co-presented a talk at KubeCon + CloudNativeCon India 2026 with his son, Shreyas Mocherla, a CNCF Kubestronaut and software engineer at Nirmata.

As Mocherla describes on the CNCF blog in a post published this week: “Presenting alongside him made this special on a level that goes beyond the conference itself. I grew up watching him speak at technology events. Standing next to him at the same podium, in front of the KubeCon audience, felt like a full-circle moment.”

You can watch the talk, “Run Your Own AI Cluster on a DGX Spark: Kubernetes, GPUs, and DRA,” below:

It’s a reminder that the Kubernetes community’s connections can span generations as well as organizations.

Other updates from the K8s universe

More updates for the platform engineers and cloud operators working in the Kubernetes ecosystem:

Follow the Road to KubeCon

Road to KubeCon is an eight-part series presented by HPE, which will be at KubeCon + CloudNativeCon North America in Salt Lake City. Before you go, explore how HPE Software helps IT teams do more with less complexity.

We’ll be here every Friday until KubeCon.

If you’d like to participate, Bill Doerrfeld, the writer of this series, is open to pitches — you can send release notes, quotes, reports, videos, case studies, or hot takes through his contact page.

If you missed the previous editions covering Kubernetes v1.37 and AI inference, you can catch up through those links. Visit the Road to KubeCon page for the complete archive.

The post OpenTelemetry and Prometheus are getting along. What’s still missing? appeared first on The New Stack.

Cursor acquired Firetiger. A month later, it launched a bot that tracks code changes from PR to production.

24 septembre 2026 à 14:56
Inspecting changes on a laptop screen document

We all know that producing code is easier than ever thanks to the abundance of AI coding tools and agents. The harder part undoubtedly comes after that code is written: making sure changes are safe to ship, spotting regressions in production, and figuring out what went wrong.

And that’s why Cursor is introducing Rollouts, a new agent that follows code changes into production and monitors whether they behave as intended.

The Firetiger effect

The announcement comes a little over a month after SpaceX closed its bumper $60 billion acquisition of Cursor, giving the AI coding company access to SpaceX’s vast GPU infrastructure as it develops its own models.

The day before that deal closed, however, Cursor quietly announced an acquisition of its own: it snapped up the team behind Firetiger, a three-year-old startup building AI agents that monitor software changes from pull request through deployment.

At the time, Firetiger co-founder and CEO Rustam Lalkaka argued that coding agents had dramatically reduced the effort involved in creating software changes, while doing little to reduce the risks involved in actually deploying them.

“Over the last two years, agentic coding has changed software dramatically,” Lalkaka wrote in a LinkedIn post following the deal’s announcement. “The cost of creating changes has dropped to near zero. The cost and risk of deploying them has stayed largely the same.”

“Writing code is no longer the slow part. What hasn’t sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.”

Rustam Lalkaka, Cursor

Fast forward to today, and Lalkaka, now at Cursor, has unveiled the first fruits from that acquisition — including Rollouts. In a blog post published on Wednesday, Lalkaka notes that the new agent, or “bot” as the company calls it, is all about helping developers “get safe, reliable code into production faster.”

“Writing code is no longer the slow part,” Lalkaka writes. “What hasn’t sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.”

Rollouts is effectively Firetiger’s Change Monitors reborn inside Cursor, rebuilt using a tool dubbed Bot Development Kit. This kit, too, appears to be new from Cursor: an early-stage framework for building and serving Cursor bots and agents, published as the @cursor/bdk package on npm. Its documentation says developers can define agents using Markdown and TypeScript, with support for tools, skills, subagents, webhooks and scheduled runs.

Like Change Monitors before it, Rollouts starts working when a pull request opens. It examines the proposed code change, works out which systems could be affected, and produces a monitoring plan covering what the change is supposed to do, the risks it sees, the signals it intends to watch, and any holes in the available instrumentation. Developers can review and edit that plan before the code reaches production.

Rollouts in action (1)
Rollouts generates a monitoring plan for a change

Once the change is deployed, Rollouts checks the resulting telemetry — including logs, metrics and traces — against that plan. Staging and production are assessed independently, with each deployment ultimately receiving one of three verdicts: verified healthy, regression detected or inconclusive.

That means a change could, for example, pass its checks in staging before Rollouts subsequently spots a problem when the same code reaches production.

Rollouts in action (2)
Rollouts reports deployment status as changes ship

If Rollouts does detect a regression, it can identify the change it suspects, alert the developer responsible and, depending on how it’s been configured, either open a revert pull request for review or hand the problem to a Cursor cloud agent to attempt a fix. There is still a human in the consequential part of that loop for now: Rollouts doesn’t merge fixes or roll back deployments by itself, though it can pause a progressive rollout.

Lalkaka notes that Rollouts is already capable of picking up problems limited to a particular endpoint or region before they trigger a broader alert, while it can also distinguish expected changes in behavior from genuine regressions.

Also “coming soon” to Rollouts, according to Cursor, is an integration with feature flags so it can directly adapt the traffic reaching a change, while support for release trains and deployment freezes is also in the works.

Enter Security Reviewer

Alongside Rollouts, Cursor is also introducing an upgraded Security Reviewer bot, which first appeared in beta back in April.

At launch, the bot could automatically inspect pull requests for security vulnerabilities, authentication regressions, privacy and data-handling risks, agent tool auto-approvals, and prompt-injection attacks, leaving findings alongside the relevant code.

As with Rollouts, the idea is that developers don’t have to remember to invoke it manually: Security Reviewer can be set to run whenever a new pull request is opened.

Security Reviewer in action
Security Reviewer runs automatically on new pull requests

In its current guise, Security Reviewer analyzes pull requests in the context of the wider codebase, with a focus on exploitable issues such as injection flaws and broken authentication, and returns a severity rating, attack path and proposed fix.

“Security Review reads code the way a security engineer does,” Lalkaka writes. “Where does user input enter, where does it end up, what does it pass through on the way.”

“Security Review reads code the way a security engineer does.”

He says that things have sped up considerably, too: average review time has fallen 21%, from 4.8 minutes to 3.8, while developer acceptance of its comments has risen from roughly 45–50% to 60–70%.

Both Rollouts and Security Reviewer are available through Cursor’s Automations tab for customers on its Teams and Enterprise plans.

The Origin story

Digging into the nuts and bolts of Rollouts reveals how it might serve as a boon for Cursor as it builds out Origin, the fledgling Git-compatible code hosting platform it launched back in August.

Origin is essentially an effort to build an alternative to GitHub for an agent-heavy software development world. It remains early, with limited functionality, but Cursor has been clear that tighter integration with its own agents is supposed to become one of the main reasons to use it.

When Cursor announced the Firetiger acquisition last month, Maxime Prades on the Cursor product team noted in a blog post that the deal was part of a “broader investment in long-running, autonomous, context-aware agents for teams.”

And he pointed to Origin and Change Monitors as two examples of that investment.

“Agents that write code should also be able to tell whether it works in production,” Prades wrote. “Today, those systems are mostly separate. Cursor and Firetiger bring them closer together so an agent can ship a change, see how it behaves, and respond when something goes wrong.”

Rollouts offers an early glimpse of that. It can connect to either Origin or GitHub for source control, pull deployment events from continuous delivery systems, and use signals from Datadog and other telemetry providers. If it spots a regression, it can then pass the problem back to a Cursor cloud agent to investigate or attempt a fix.

Origin potentially gives Cursor a native home for more of that loop: its cloud agents can already create branches, commit and push code, and open pull requests against Origin repositories. Rollouts then adds information about what happened after.

That could become increasingly important as more companies take aim at GitHub’s central role in software development. Zed, for example, put Delta into public beta last week, with its own ideas about how source control should change for teams working heavily with agents.

Cursor also faces competition further downstream. Datadog’s Bits Release, launched in preview in June, similarly follows changes from pull request into production and checks telemetry for regressions. Harness has long offered automated deployment verification and rollback based on logs and metrics, while LaunchDarkly’s Guarded Rollouts can monitor feature releases for regressions and automatically reverse them.

What Cursor can potentially bring to the table is proximity: the coding agent, repository, pull request, security checks, and production feedback can all sit much closer together. Rollouts doesn’t require Origin — GitHub remains supported — but owning the forge gives Cursor more room to integrate those pieces over time. And that may prove more compelling than simply recreating GitHub’s existing feature set.

The post Cursor acquired Firetiger. A month later, it launched a bot that tracks code changes from PR to production. appeared first on The New Stack.

It passed CI. It passed your evals. The customer still got the wrong answer.

13 septembre 2026 à 16:00
Blurred, overlapping close-ups of yellow analog thermometer dials, their curved scales marked 0, 10, 20, and 30 in black with a red band sweeping through the upper range.

A diff is not evidence. It’s a statement of intent.

The tests passed. The review’s done. The change is live. Then someone says the app is slow, or the answers are wrong, or both. You open the diff. Your assistant points at the function it changed and offers a plausible cause.

It sounds right. It might not be.

This is the observability gap that AI features expose. Dynatrace’s 2026 State of SRE and Platform Engineering report (919 enterprise leaders surveyed globally) found that while 77% of platform engineering teams embed observability in at least some services, only 40% have it fully integrated across all deployments. That gap was manageable when your services were deterministic. But with AI agents, it becomes a liability.

A conventional service fails loudly… An AI agent fails quietly. It returns a 200. It passes faithfulness checks. And the customer still gets the wrong answer.

A conventional service fails loudly. A 500 error, a latency spike, a dependency that stops responding. An AI agent fails quietly. It returns a 200. It passes faithfulness checks. The customer still gets the wrong answer.

You can’t alert on “wrong.” You need evidence from the running system — and for AI features, that means something more than request traces and error rates.

Find the request first.

One release, two symptoms

Say you run a support agent over product documentation. A customer asks how to configure export in version 2026.3. Your coding assistant helped rewrite the documentation lookup. CI passed. The existing evals passed.

After deploy, answers take longer. Some of them describe older product versions.

Start with one affected run. You want its release, retrieval config, and feature-flag state, so those need to be on the root span as attributes set at span start, not reconstructed later from a deploy log. Then put that run next to one for a similar question from before the change.

For an agent, that means the trajectory: every model call and tool call, in order, with arguments and results. A distributed trace records those as spans and stitches them across service boundaries through context propagation.

Here’s one run, simplified, with its evaluation linked separately.


# Illustrative pseudotelemetry, not a captured incident.

# Names, IDs, timings, and labels are invented, not a standard schema.

# Selected spans shown in execution order; other work is omitted.

trace: example-run-a | session: example-session-7 | release: 2026.9.2

requested.product_version: "2026.3"

agent.run                                  12.4s

  model.choose_tool                         1.0s

  tool.search_docs                          0.9s

    args: {query: "configure export", product_version: null}

  tool.search_docs                          0.8s

    args: {query: "configure export", product_version: null}

  tool.search_docs                          0.9s

    args: {query: "configure export", product_version: null}

    returned.doc_versions: ["2024.1", "2024.1", "2023.9"]

  model.generate_answer                     8.1s

linked_evaluation:

  trace: example-run-a

  faithfulness: pass

  requested_version_answered: fail


Two things to chase. The repeated searches. The null version filter.

Check the repeated work

Three identical searches cost 2.6 seconds. The trace shows the symptom. It doesn’t explain the cause.

But look at what sits between them. Nothing. One model.choose_tool span at the top, and no model call between the second search and the third. The model didn’t ask for those retries. Something in the harness did: the code that runs tools, handles retries, and manages context. A model.choose_tool span between each search would mean the opposite: a model that kept requesting the same tool, which is a prompt or tool-description problem. Same symptom, different file to open.

That still doesn’t make the retries wrong. Read the retry policy, then read the tool results. A 200 from a search backend can carry an empty hit list, or every hit under your relevance threshold, and retrying on that is legitimate.

The generation call is the bigger slice anyway, at 8.1 seconds. Compare its input tokens and duration against similar runs. If the harness appended all three result sets to the context, the retries inflated that prompt, and you paid for them twice, in latency and in tokens. Check downstream services and traffic too before you pin the slowdown on the release.

Bringing this back into the IDE? Bound the question. Give the assistant the service, the release, the time window, and the trace IDs. Have it line the changed code path up against the dependency calls in the affected trace. Then separate what the evidence supports from what it’s assuming.

Same workflow debugs a checkout service making three identical database calls. You don’t need to build an agent to use it.

A grounded answer can still fail

Now read the answer.

In this example, it accurately repeats the retrieved documentation. Faithfulness passes, or groundedness, depending on whose vocabulary your tooling uses.

The customer still gets instructions for the wrong version.

Whether you call it faithfulness or groundedness, the metric only tells you whether the answer is supported by the sources you supplied. It says nothing about whether those were the right sources.

The obvious next move is a retrieval evaluator. It still won’t catch this. Those score whether the retrieved context is relevant to the query, and the 2024.1 export instructions are relevant to configure the export. They’re just invalid for the version asked. Those documents are relevant to the query. They are not valid for the version the customer requested. Relevance is not validity.

Those documents are relevant to the query. They are not valid for the version the customer requested. Relevance is not validity.

So, this isn’t a generation failure. It’s a retrieval precondition nobody asserted, and the null filter names it: the requested version never reached the lookup. Reproduce that before you touch the prompt or the model.

Most of it is testable with ordinary code. Give the fixtures documents carrying version metadata, then assert on the lookup directly, no model in the loop:

def test_lookup_filters_to_requested_version(docs_fixture):

    hits = search_docs(query="configure export", product_version="2026.3")

    assert hits, "no hits for a version that has docs"

    assert {h.product_version for h in hits} == {"2026.3"}


Deterministic, cheap, belongs in CI. Then evaluate the answer separately, which is the part you can’t assert: does it give usable 2026.3 instructions, or say the available documentation can’t support one? Two tests, because they fail for different reasons and you want to know which one broke.

The assertion won’t catch every wrong answer. It catches this missing constraint every time, which is more than a judge scoring helpfulness one to five will do for you.

Which is why evaluation needs retained context. Record the prompt version, model ID, retrieval config, and document IDs and versions alongside the release. Keep enough permitted evidence to read the answer back later, sensitive content redacted before export.

Link results by trace and span ID. If scoring lands after the span closes, store a separate linked result. Don’t plan on writing attributes to a finished span: the OpenTelemetry tracing API says implementations should ignore updates after End.

GenAI semantic conventions are still evolving, and different instrumentation projects expose similar concepts with different attribute names.

Make the failure part of the next release check

Confirmed the causes? Verify each fix against the behavior it’s supposed to change.

For the repeated searches, add a regression test that reproduces the repetition without killing legitimate retries. Don’t pin one exact tool sequence when several orderings finish the task correctly; a trajectory test that demands a single path fails on every valid refactor.

For the version mismatch, restore the filter. Add these cases: current version, an older supported version the customer names explicitly, irrelevant documentation, and no supportable answer. Run the answer evals repeatedly where output varies, because one pass isn’t a result.

Use code for anything you can assert directly. Use a model-based judge for answer quality, and validate that judge against examples people reviewed. An unchecked judge is one more model you’re taking on faith.

To run scoring against production traffic rather than fixtures, you’ll need a way to sample spans already in your environment, score them with a judge model, and link each result back to the source trace — the linked-result pattern above, not a write to a closed span. Whatever tooling you use, version the evaluator. A scoring change that looks like a product improvement isn’t one.

After the release, compare latency and task success on similar requests, and keep tool-call and token counts on the same screen. Read them together, or they’ll mislead you. Tool calls dropping from three to one can look like the fix is working, but it can also look like a lookup you removed by accident. Fewer output tokens look like a cost win, and it also looks like an answer that quietly stopped listing step four.

Bring one debugging question

If you wouldn’t know where to start the investigation, you’re not done instrumenting.

For a conventional service, that’s the request path and dependency timing. For an AI feature, add what it retrieved, what it produced, and how you’ll decide whether that was the right answer.

Dynatrace is sponsoring WeAreDevelopers World Congress Americas, September 23-25, 2026, in San José. Come with a debugging question from an AI-assisted release or from an AI feature you’re building, and we’ll work through it.

The post It passed CI. It passed your evals. The customer still got the wrong answer. appeared first on The New Stack.

Jacob Coxon warns AI could kill us all. Anthropic’s own report exposes safety gaps. 

12 septembre 2026 à 13:00
Abstract teal composition with overlapping translucent geometric shapes and scattered flecks of red, yellow, white, and black paint.

I’m Matt Burns, Chief Content Officer at Insight Media Group. Each week, I round up the most important AI developments, explaining what they mean for people and organizations putting this technology to work. The thesis is simple: workers who learn to use AI will define the next era of their industries, and this newsletter is here to help you be one of them.


Former OpenAI and Anthropic researcher Jacob Coxon made a lot of noise this week with a viral thread on X. He has since talked to Wired and Axios. His warning deserves attention, but another revelation this week gives developers something concrete to investigate.

In retrospective testing of the Mythos 5 incident, Anthropic’s offline monitor flagged around 1% of actions. When researchers removed the model’s chain of thought and showed the monitor only user turns, tool calls, and tool results, that rose to about 50%. 

The model’s written reasoning is what talked the monitor out of flagging that harmful behavior. The offline monitor kept concluding it was operating in a simulation.

The lesson is to test whether your monitoring accepts an agent’s justification a little too easily. An AI’s explanation can help a safety monitor reveal a problem, or persuade it to essentially convey, “move along, nothing to see here.”

An AI’s explanation can help a safety monitor reveal a problem, or persuade it to essentially convey, “move along, nothing to see here.”

Coxon is warning about the pursuit of self-improving superintelligence. Developers are already pointing agents at real credentials during everyday work. Those problems differ substantially, but both deserve attention.

Forecasts aren’t checkable. Incident reports are.

Jacob Coxon’s thread is short and hard. He says neither of his former employers is acting responsibly, and that both are “racing straight to self-improving superintelligence and gambling with our lives.”

Maxwell Zeff pressed him on the end-of-humans mechanism twice in an interview with Wired. Coxon described shutdown avoidance, then biological and cyber threats. His answers explain the concern, but leave substantial questions about how AI will kill us all.

Annie Jacobsen, whose books on this stuff I wholeheartedly recommend, spelled out the scary biological scenario: A system that can hack anything, turned loose on the world’s 3,600-plus BSL-3 and BSL-4 labs, is how you get from a software problem to a biological one. For me, Jacobsen’s commentary is the most concrete, real-world implication anyone has given of what Coxon is actually worried about. 

Two incident reports published in the last few weeks are more useful anyway. OpenAI’s report on the Hugging Face incident, published August 26, traces the whole thing from May 12 to July 20. The agents found a hole in an internal package server and used it to reach the open internet that the same server was supposed to keep them off of. OpenAI rebuilt the server; they got back out. By July 11, they were far enough into Hugging Face to run their own commands on its machines.

The model doing most of that was never released. It’s a research version OpenAI runs in-house, with fewer safety checks than the ChatGPT everybody else uses. When the company ran the same work again with those checks switched on, it found that this kind of behavior could drop by more than 100 times.

That context matters. 

Amanda Caswell covered this in August. The incident demonstrates a serious failure under particular conditions. It also illustrates why the surrounding safeguards belong in any assessment of an agent’s behavior.

The monitor was talked out of it

Anthropic’s assessment, published on September 9, covers four incidents and analyzes the first three in a scan of roughly 141,000 transcripts, identified a fourth in August while assembling material for METR, then widened the search to roughly 481 million transcripts and reported no additional incidents of similar or greater severity. In each case, the model reached the open internet through a misconfiguration rather than by breaking out of a properly isolated sandbox. That’s an important detail. It makes the evaluation infrastructure part of the story. 

If you’re a developer, you should look at the monitoring results. Mythos 5 kept reasoning its way to the conclusion that it was operating in a simulation, and the offline monitor read and accepted that reasoning. The limits matter here too: These were cybersecurity evaluations running without released cyber safeguards, Anthropic reports no coordination between models and no concealment, and its independent METR review was planned rather than finished. This is not the same failure as the OpenAI incident. Don’t equate them as similar incidents

But yet the shape of it is familiar, and it’s what Adrian Bridgwater reported on last week in his piece on agent observability, which is a useful read if you’re producing agents (who isn’t). Jacob Krell of Suzu Labs put it about as well as it can be put: “The model can recite your constraints and reason past them in the same breath.” Liran Hason of Coralogix added the result: “System guardrails help, but a guardrail only stops what the developer already thought of.”

What developers can take from Steven Adler’s recommendations

Largely overshadowed this week by Coxon, Steven Adler’s op-ed in The New York Times asks AI companies for four things: disclose incidents including near misses, the way aviation does; keep tamper-evident records of model behavior; never let a model cut power to its own alarm systems; and formally swear off training techniques that undermine researchers’ ability to detect deception. 

Adler worked for OpenAI from 2020 to 2024 and later founded a nonprofit called Guidelight AI Standards. It has since graded six control practices at the frontier labs using public information and gave out a top mark of C-plus.

Guidelight also publishes a Control standard for frontier labs, recently revised in August. It wants monitors to access the chain of thought and scan both reasoning and actions. 

But those asks of frontier labs don’t apply to most developers. Here are five you can check on your own setup right now.

Five checks you can run on your own agent setup

Steven Adler’s asks of the frontier labs, applied one layer down.

Check What passes
Someone reads the near misses An agent crosses a line and nothing breaks. There’s still a record, a person who looks at it, and a copy the agent can’t reach.
The agent can’t edit its own history Logs are append-only and stored outside the agent’s reach. If its credentials can delete the log, the log is decorative.
The agent can’t touch its own oversight Separate tokens help. Then check whether it can shut off monitoring, rewrite an approval rule or get around the kill switch.
Rule-breaking gets blocked and logged In a test environment, hand it a job it can only finish by breaking a rule. Judge the monitor on what it does with the action, not with the agent’s account of it.
The explanation doesn’t change the answer Run the same rule-breaking action past your monitor twice. Change nothing but the explanation, one flat and one reassuring. If the reassuring version gets flagged less, you’ve found the problem.

Sources: Steven Adler, The New York Times; Anthropic alignment assessment, Sept. 9, 2026.

These checks do not establish that frontier alignment is solved. They give a team evidence about the system it operates on and a clearer account of where its protections need work.

Anthropic’s report is why I’d include that last test in a deployment review. Hold the action and its permissions constant, change the explanation, and see whether your monitor changes its judgment.

That won’t settle Coxon’s warning. Will AI kill us all? I don’t know. But I know developers need to test their agents’ permissions, protect their logs, and make sure a convincing explanation can’t talk a safety monitor into overlooking an unauthorized action.

The post Jacob Coxon warns AI could kill us all. Anthropic’s own report exposes safety gaps.  appeared first on The New Stack.

How to find failures without drowning in tracing data

3 septembre 2026 à 22:17
On The New Stack podcast, Sarah Hudspeth of Chronosphere, a Palo Alto Networks company, explains how teams can build a more effective tracing strategy.

A metrics dashboard can tell you a system’s health with ease. A log can help you understand a discrete failure. But if you want to understand where in a query’s journey things went awry, you need traces.

By tracking a request from its point of origin through data and microservices to the end user, traces offer unparalleled insight into how systems work and where failures occur. SREs offer the fastest path to remediation. That means less downtime, fewer burned-out developers, and happier customers.

Sadly, the promise of traces often doesn’t match the on-the-ground reality. 

Why? Simply collecting and holding onto all your company’s traces is an exercise in hoarding. Do you need to store terabytes of tracing data just to show when your systems worked? Not only is that much information expensive to hold onto, but collecting it can slow the very systems you are trying to monitor. And when you have all the stored tracing data, finding what you need in the ocean of information can take too long.

Is tracing cooked? Not at all.

Is tracing cooked? Not at all. There are several ways to beat back tracing data overload: Head sampling collects only a portion of tracing data, reducing storage concerns; tail sampling asks whether, after a trace is recorded, it is worth holding onto, making it easier to find what you’re looking for down the road. And dynamic sampling can automatically cull similar or highly repetitive traces, so you don’t accidentally flood your storage system with nearly identical data.

You can avoid the most common tracing pitfalls by building your observability system intelligently. That’s precisely what I was hoping to learn from Sarah Hudspeth of Chronosphere (a Palo Alto Networks company), who is my guest on the latest episode of The New Stack podcast.

Whether you are just starting your tracing journey or deep in the trenches looking for help, Hudspeth’s ability to turn abstract technical concepts into simple, digestible analogies is enviable. 

Hit play on the episode above, and let’s jump the chasm between the promise of tracing and getting it to work for you in a production setting.

The post How to find failures without drowning in tracing data appeared first on The New Stack.

Anthropic’s Claude failures have made agent observability a security priority

2 septembre 2026 à 22:12
Aerial view of a dark, futuristic digital city outlined in blue and purple neon.

Anthropic aimed to steer its ship into safer, more carefully charted waters this week. The company announced it was improving its alignment and security efforts, and the announcement read somewhat like an admission of responsibility and a mandate for tighter agent controls.

The company has recounted incidents in which its models “took a series of unauthorized actions” on the open web. However, they did so while “intentionally running without cyber safeguards” for evaluation purposes.

In a statement released on Monday, Anthropic attributes the July incidents in part to a third-party environment misconfiguration while saying it would approach the fixes as if responsibility were its alone.

Separately, on August 4, the UK AI Security Institute (AISI) reported that Claude Mythos 5 took a series of unauthorized actions during its cybersecurity testing. Both sets of incidents occurred during deliberately permissive capability evaluations, with normal cyber safeguards reduced or disabled. Anthropic identified six affected runs among 141,006 it reviewed. AISI found unauthorized behavior in 10 of 122 runs, said the attempts were unsuccessful, and found no evidence of resulting real-world harm. It also cautioned that the tested configurations were not commercially available.

“We believe the incidents reflect a failure of operational security, as well as two alignment issues: motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task (both of which we have described in previous system cards),” states Anthropic.

What isn’t Anthropic telling us?

Those AI engineers with an acute sense of dismal foreboding won’t be too surprised to see incidents like this surface, but the question of what happens next is ripe for discussion.

Anthropic’s account leaves several questions unresolved. How much of the risk arose from evaluation-environment failures, how much from model behavior, and what combination of containment, observability and alignment work is needed? Is this all about cybersecurity controls, or should we focus on architectural instabilities, cloud misconfiguration, lack of agent observability, or another missing piece of the jigsaw?

Senior director for secure AI solutions & cybersecurity at Suzu Labs, Jacob Krell, tells The New Stack that AI developers and systems engineers building agentic features need to “quit pretending their operational instructions are a security control” in real terms.

“Stop treating this like a malfunction,” Krell says. “Every developer shipping an AI agent is trusting the model to follow instructions. If you are building agentic features, stop pretending your instructions are a security control. The model can recite your constraints and reason past them in the same breath.”

“Stop treating this like a malfunction. If you are building agentic features, stop pretending your instructions are a security control. The model can recite your constraints and reason past them in the same breath.”

AI pursues its objective with persistence, past the rules

Anthropic’s post-mortem says Mythos 5 recognized evidence that it might be on the live internet but reasoned its way back to the conclusion that the environment was simulated. The other models behaved differently: Opus 4.7 continued after recognizing real systems, while Anthropic’s newest internal model eventually stopped. Krell thinks that’s not something we should define as a bug; for him, it’s AI pursuing its objective with creativity and persistence, including persistence past the rules.

“This means AI should be coded to treat every agent action the way we would treat input from an untrusted user, validated by something the model cannot override before it touches anything real. Claude hacked three real companies because that is what a capable, goal-directed system does when you point it at a target and leave a door open,” explains Krell.

Of course, that is Krell’s interpretation. Anthropic says it found no evidence that the models pursued self-generated goals: they were following assigned capture-the-flag objectives while operating under false or confused beliefs about their environments.

He calls for hardcoded scope checks, deterministic approval gates, action-level allow lists, and a human who signs off before anything high-risk fires. Saying that the industry is currently “automating judgment and calling it progress,” Krell bemoans the proposition that, right now, the AI trade is “automating accountability failures at machine speed” instead.

System constraints and control prompts are not enough

The lesson is not that system prompts are useless, but that they cannot serve as the only security boundary. Scope instructions need to be backed by network isolation, least-privilege access, deterministic approval gates and monitoring that can stop an unauthorized action before it executes.

As Krell points out, “Anthropic’s Claude breached three organizations and rationalized away evidence it was on the live internet. OpenAI’s agent recognized it was crossing a boundary on Hugging Face and did so anyway. When the UK AI Security Institute tested Anthropic’s Mythos 5, it caught the model creating fake identities to social-engineer a human maintainer into approving malicious code. Different models, different evaluators, same result.”

The incidents fall into the same broad category of failure, but their mechanisms and outcomes differed. OpenAI’s models exploited vulnerabilities to escape isolation; Anthropic’s July models followed an unintentionally open network path; and AISI deliberately enabled internet access. AISI reported no resulting real-world harm.

VP of AI at Coralogix, Liran Hason, tells The New Stack that he’s exasperated, for mostly the same reason.

“System guardrails help, but a guardrail only stops what the developer already thought of,” Hason says. “AI engineers still need to see the behavior and what the agent delivers. Every agent throws off decisions, tool calls, and outcomes that nobody was collecting even six months ago. That is the new observability problem, and it is a big one.”

“System guardrails help, but a guardrail only stops what the developer already thought of.”

“What Anthropic is seeing now, every enterprise will see within a year. An agent can be healthy by every metric we have and still be doing exactly the wrong thing. Fast, available, no errors, yet it just accessed a system it shouldn’t have, called the wrong tool, took an action nobody asked for. Uptime was never built to catch that,” Hason adds.

The next question is all about agent scope

So then, the question for developers now stops being whether an agent is running in live production in a successfully booted instance with correct configurations. The question now is what the agent has done so far, what it can reach, and what happened as a result.

Anthropic’s account shows that its scope-setting was incomplete in these evaluations. The July prompts told Claude that it had no internet access but did not explicitly limit where it could search for the flag. AISI similarly said its agent was not specifically told to avoid the public internet or social engineering. The incidents therefore demonstrate the danger of ambiguous or contradictory instructions, as well as the need for enforced network boundaries.

If that’s even part of the answer, it’s definitely not all of the answer. As we journey outwards from Earth into whichever part of the western spiral arm of the galaxy this takes us, we may still find agents that persist in satisfying technical objectives while violating their developers’ broader intentions.

Anthropic to analyze, review & improve security & alignment

In the aftermath of these developments, Anthropic confirms it is conducting an in-depth analysis of both incidents. It is also planning to work with METR (a research nonprofit that scientifically measures whether and when AI systems might threaten catastrophic harm to society) for an independent review.

In response, on security, the company has described the improvements it has made to its containment and monitoring systems, along with practices for third-party evaluators. The company has also explained how its early research on model alignment relates to these agentic errors.

Anthropic said its “internal security posture was not a contributing factor” to the three incidents it disclosed July 30. Those incidents involved mistakenly available internet access, whereas AISI deliberately enabled internet access for its separate evaluation.

“The July incidents have stressed that the urgency of improving our cybersecurity defenses is even higher than we previously believed. We are redoubling our efforts in this direction and will say more in our next risk report,” concluded Anthropic.

The post Anthropic’s Claude failures have made agent observability a security priority appeared first on The New Stack.

Observability has a data problem. AI is about to make it worse.

26 août 2026 à 22:08
Parallel orange lines form a flowing wave across a dark purple background.

Observability is entering a new phase now that OpenTelemetry has standardized instrumentation for data collection. Unfortunately, the observability industry still lacks a cost-effective way to store, retain, search, and analyze full-fidelity telemetry data. This results in blind spots in observability and many teams operating without full operational visibility.

As AI systems generate more logs, traces, and metrics — thereby making the blind spots issue worse — Bronto, a Dublin, Ireland, firm offering an intelligent data observability platform, is betting that the next observability platform battle will be won at the data layer, not the dashboard layer.

Bolt-ons and incremental efficiency aren’t enough

Trevor Parsons, co-founder and co-CEO of Bronto, tells The New Stack that the industry has been optimizing at the edges rather than rebuilding the economics and architecture of telemetry storage. The industry has introduced a wide array of “hacks” and “capabilities” to avoid tackling this issue head-on and ultimately to protect their margins. 

“If you are a couple of times cheaper or 50% cheaper, that ain’t going to cut it,” Parsons says, because data volumes, especially AI telemetry, are growing so quickly, on top of already stretched observability budgets and inefficient datastores. 

Promises, Promises, Promises…

Parsons elaborates, “Observability has always and continues to have a data problem.”

The eternal promise of observability has been delivering teams a clearer view of what’s happening inside their systems.

“Observability has always and continues to have a data problem.”

But in practice, that view is often incomplete, expensive, and short-lived. For too many teams, observability has become less about asking better questions and more about fighting the cost and complexity of storing the data they already need. 

“Sometimes people frame that as a cost problem, where they’ll say observability is up to 20 or 30% of your infrastructure spend,” Parsons says. “I actually think this minimizes the issue; it’s much bigger than that. Teams are actually paying 10, 20, 30% of their infrastructure spend for access to only a sliver of their data.” 

Noel Ruane, co-founder and co-CEO of Bronto, frames the challenge that organizations face and tells The New Stack, “Agents and applications are generating more logs, traces, and metrics each day. The software landscape has accelerated, but are observability vendors keeping pace? No, they’re offering workarounds, bolted-on features, and asking teams to accept blind spots.” In short, Ruane says, they’ve failed to solve the data problem.

Out with the old observability model 

“Customers are not getting access to all of their observability data, Parsons explains. “They have to cut their retention from 30 days to seven days to three days. They have to sample data. They have to rehydrate data.”

In other words, today’s tools make customers choose which parts of their own data they’re allowed to see, and you may only get to see it for a short amount of time.” 

“The solutions that are being put in front of customers to give them their data are always full of compromises, forcing teams to choose between cost, coverage, and speed of data access. The burden is always put on the customer by vendors.”

“The solutions that are being put in front of customers to give them their data are always full of compromises, forcing teams to choose between cost, coverage, and speed of data access,” Parsons says. “The burden is always put on the customer by vendors.

“But really this should be the other way around; it’s the vendors’ job to innovate on behalf of the customer” 

OpenTelemetry: Collection solved, storage problem exposed

Severin Neumann, head of community at Bronto, tells The New Stack that OpenTelemetry has helped standardize instrumentation and data collection, while reducing reliance on proprietary agents.

But that success has created a new bottleneck, says Neumann, who is also an OpenTelemetry maintainer and member of the OpenTelemetry governance committee. Now that organizations can collect more telemetry, they need somewhere affordable and useful to put it.

“We have fixed the instrumentation problem,” Neumann says. He cautions that enterprises now need ways to handle all this data. And if enterprises can’t store it and instead throw away large parts of it, humans and agents can not make sense of it.

The observability business model doesn’t align with customer value

The legacy observability tool business model charges customers for data storage, rather than the value teams get from their data, Parsons says. Customers tell him the same thing constantly: “I pay the same price even if I never search my data.” In many cases, customers find existing tools difficult to use and feel that their observability solution is just a really expensive data store that they do not get a lot of value from.”‘

Noel Ruane assessed the market by saying, “Traditional vendors like Datadog know their pricing model isn’t sustainable. They’ve introduced defensive features like ‘Flex Logs’ and a new ClickHouse partnership to try to keep customers from jumping ship, but they’ve only added new complexity for their customers.” 

Especially in the AI era, Ruane adds, the traditional business model charges teams in ways that discourage them from capitalizing on their data. Customers should pay much less for data that sits idle and more when they actually derive value from it with queries and analysis.

Bronto’s technical differentiation

Bronto isn’t selling another observability dashboard. It argues that observability is a storage problem before it’s a visualization problem, and that’s where the company went.

Underneath the platform is a custom-built polymorphic data store called BrontoDB, specifically designed for observability data. The pitch: enterprises can keep more than 100 times the observability data they hold now, and it won’t get slower or harder to use.

Why that matters comes down to how the three signals break. Metrics, logs, and traces each hit a wall at different points, and Bronto says it built BrontoDB to tackle these issues head-on. Parsons is blunt about two of them.

“With metrics, we’ve solved the high cardinality problem where costs traditionally explode with high cardinality metrics,” Parsons says. “With logging, we’ve solved the indexing problem where there was always a trade-off between fast logs and paying through the nose for it or having slow logs and getting them slightly cheaper.”

  • High cardinality is what wrecks metrics pricing. Add enough unique dimensions and the bill lands somewhere nobody forecast. Bronto says it was built specifically to take that surprise out.
  • Logs have always been pick-your-poison: fast and expensive, or cheap and slow. Bronto says that choice goes away — sub-second search across petabytes, no shortened retention windows, no rehydrating cold data, no waiting.
  • Traces, Bronto argues, shouldn’t be sampled at all. Sampling exists because tools and pricing models couldn’t handle the full stream. Bronto says teams can send it all.

Billing works differently, too. Most vendors charge for data sitting in storage, whether anyone touches it or not. Bronto charges closer to what teams actually search and analyze. That’s the piece that has to hold up if full-fidelity observability is going to be affordable at AI scale.

AI is what raises the stakes, Parsons says. AI systems are non-deterministic and trace-heavy. They throw off more telemetry, and the data has to stick around longer if you want to debug effectively. 

He points to an upside as well. As operations become more automated, telemetry data becomes more useful because agents can chew through volumes of history that no SRE would ever read manually.

AI raises both the volume and the stakes, according to Parsons. AI systems create more telemetry because they are non-deterministic, trace-heavy, and require longer retention for troubleshooting. At the same time, AI-enabled operations will make historical telemetry more valuable because agents can analyze far more data than human SRE teams could manually inspect.

“If AI is the intersection of where data meets intelligence, you can not apply intelligence if you do not have the data.”

“If AI is the intersection of where data meets intelligence, you can not apply intelligence if you do not have the data,” Parsons says.

The next observability battle 

AI is unlikely to fix observability’s data problem. In fact, it will produce more telemetry, create more edge cases, and increase the cost of missing the right signal at the wrong time.

For Bronto, the data layer is the next major battleground. Dashboards still matter, but in an AI-heavy production environment, the more important question may be whether teams have access to all their data for as long as they need so that they can apply AI to it. 

The post Observability has a data problem. AI is about to make it worse. appeared first on The New Stack.

How telemetry pipelines keep AI agent costs under control

25 août 2026 à 21:01
Abstract neon pink and purple angular pathways interlock against a dark geometric background.

As enterprises move from experimenting with AI to running autonomous agents in production, an infrastructure problem is emerging: rising telemetry costs. Non-deterministic, iterative, and capable of generating data at machine speed, agents are far harder to monitor — and their costs far harder to predict — than conventional applications.

Many companies are struggling to attribute and defend their telemetry bills. In fact, 59% of organizations have already terminated or delayed an agentic AI deployment due to monitoring costs, according to a survey of more than 300 enterprise IT decision-makers in North America and Western Europe, commissioned by Apica and conducted by Omdia/Informa TechTarget. 

The agents most affected are often in some of the most high-stakes deployments: think cybersecurity, compliance, and fraud detection. As monitoring bills explode, deployments aren’t necessarily getting killed by engineering teams. More often than not, it’s finance pulling the plug.

Andi Mann, chief product and technology officer at Apica, recently saw this play out at a large bank. The organization couldn’t pin down exactly what it was spending on its AI programs.

“They knew they couldn’t afford to keep going on the same trajectory, so they had no choice but to cancel certain AI programs,” Mann tells The New Stack. “It’s a pattern I have seen before, because AI projects are cannibalizing typical budgets.” 

“It’s a pattern I have seen before, because AI projects are cannibalizing typical budgets.”

The implications are huge. As people, funding, and monitoring resources are diverted toward new AI workloads, other parts of the business start to suffer. Mann says he’s seeing outages, downtime, penetration attacks, and DDoS protection compete for the same resources.

The problem is already showing up in research. Most (54%) enterprises have seen telemetry volume triple in the past year alone, with 43% of that growth coming from AI/ML workloads — by far the largest driver. Businesses are under pressure to address this crisis as their observability bills climb.

They report spending an average of $3.17 million on observability, with that figure growing 28% year over year and showing no obvious ceiling. No wonder then that 83% rank AI observability as a top priority for the year ahead.

The coming wave could be catastrophic

With a new cloud service, database, or application, you’d expect the monitoring burden and telemetry data to rise by a relatively predictable increment. But enterprises now foresee an average 9.5X increase in telemetry data within two years.

“Imagine if your credit card or grocery bill went up more than nine times — that is not a marginal amount,” Mann says. “This isn’t a gentle ramp; it’s a skyscraper, and it’s prompting panic.”

“This isn’t a gentle ramp; it’s a skyscraper, and it’s prompting panic.”

Some 44% of organizations expect their telemetry data to increase by 6X to 100X.

The reason is that an agent task isn’t the same as a single application request. A customer-support task might generate a top-level trace, several model calls, retrieval operations, tool calls, retries, and loops. But if the agent delegates work to another agent, that adds another branch to the trace. 

Each model can produce token, latency, cost, and provider data, while each tool call generates its own records for arguments, results, status, and downstream activity. Identifiers such as tool_name, agent_id, and trace_id also create high cardinality, making data harder to aggregate and more expensive to index, with costs compounding at every stage.

That creates an uncomfortable gap between AI ambition and infrastructure readiness. Although 35% of enterprises claim widespread agentic AI deployment, operating and managing those systems is very different. Nearly two-thirds are only somewhat prepared for the shift. Unlike conventional applications, agents can call multiple models and tools, repeat tasks, or expand a workflow in unpredictable ways, making both capacity and monitoring costs difficult to forecast.

From extreme telemetry costs to an upstream control layer

The answer, according to Mann, is to intervene earlier. “You can’t keep sending essentially useless data to an expensive central analytics or storage platform because there’s no point analyzing data which says everything is fine,” he says. “As early as possible, get the pipeline to use data collectors and manage those at source.”

Legacy observability platforms were built around collecting data, ingesting it, storing and indexing it, and then analyzing it. That model worked for human-driven, dashboard-queried workloads, but agentic AI requires decisions to be made before telemetry reaches the most expensive parts of the stack.

A pipeline-first architecture can sample repetitive successful events while retaining failures, retries, policy violations, and unusually slow traces. It can enrich records with agent, session, model, tool, token, and estimated-cost data, redact sensitive prompts and identifiers, and aggregate metrics and long-term records to destinations with different cost and retention profiles.

A compact metric or sampled span might represent a routine successful tool call, while a failed call retains its parent trace, error details, retry history, and security context. The aim is to preserve the information needed to explain an agent’s behavior while reducing redundant data and limiting what gets indexed.

Agents also need millisecond-level context for autonomous decisions. Processing telemetry close to its source allows organizations to quickly detect retry storms, excessive tool loops, or unusual token consumption, rather than waiting for data to be ingested and indexed centrally. 

Without upstream control, businesses risk feeding fragmented, unnecessary telemetry into platforms that charge for every additional gigabyte, index, and retained record.

Architecture that separates winners from cancellations

The payoff for rethinking the telemetry pipeline is hard to ignore. Enterprises with a telemetry pipeline are 50% more likely to be prepared for the growth of agentic AI data. Among mature agentic AI organizations, pipeline adoption is what sets them apart: these organizations are 80% more likely to have avoided the operational cost challenges hampering their peers.

Rather than treating the observability platform as a catch-all destination, enterprises can make decisions about telemetry before it gets there.

The answer is to move the intelligence upstream. Rather than treating the observability platform as a catch-all destination, enterprises can make decisions about telemetry before it gets there — filtering out noise, enriching what matters, and routing data according to its value and purpose. That means less data hitting expensive storage and analytics systems, while the information that does make it through is more useful and available in real time.

Crucially, this isn’t about ripping out the observability platforms enterprises already rely on. It’s about putting a smarter control layer in front of them: deciding what data deserves to be ingested, where it should go, and how much it should cost.

Existing observability platforms still have an important role to play. “The pipeline can’t do everything, but it can act as a first responder,” says Mann. “You still work with the big analytics platforms, but you’re saving money, reducing risk, and improving your compliance performance.”

The Apica study finds that its pipeline control, metrics foundation, and data readiness services, for example, can reduce the total cost of ownership by 40% compared to legacy observability platforms. Of course, though, the actual savings will depend on telemetry volumes, retention policies, sampling rules, routing decisions, existing contracts, and the proportion of data that can be processed before ingestion.

The window to rethink the architecture is now. Some 68% of enterprises plan to evaluate changes to their observability stack within the next six months, while almost a quarter say existing vendor relationships won’t be a significant factor in those decisions. The next phase of observability will be won by the ability to handle what agentic AI throws at the infrastructure.

That means the pipeline can no longer be treated as plumbing that moves telemetry from A to B. It’s becoming the control layer for an increasingly autonomous, data-hungry environment. 

Organizations that establish agentic-ready infrastructure will be better positioned to reduce observability costs and improve risk performance. Mann doesn’t think platform engineering and SRE teams have much choice. “This has already become a board-level decision,” he says. “Ultimately, it’s a choice about how smart you can afford to make your business.”

Download the Omdia research report: “The Agentic AI Telemetry Crisis: Are You Ready for What’s Coming?”

The post How telemetry pipelines keep AI agent costs under control appeared first on The New Stack.

❌