❌

Vue normale

Reçu avant avant-hierInfra

OpenTelemetry everywhere: Migrating a metrics platform at scale

Why we did this at all

For most of the last decade our metrics pipeline ran on gostatsd, the open-source StatsD implementation we maintain. It primarily did two jobs: as sidecar on every host and the aggregation tier at the other end. It took metrics from roughly 100k hosts across 14 regions at a 99.95% SLO and minimal latency and it was fine. Nobody thought about it much, which is usually the sign of good infrastructure.

Gostatsd served us for many, many years. However, the community continuously and consistently converged on OpenTelemetry over the past few years. It became the thing everyone standardized on and more and more of what fed our pipeline was emitting OTel data we simply didn’t support. Gostatsd was UDP-only, had no story for traces or logs and every clever thing the OTel Collector community shipped was one more thing we’d eventually rebuild by hand just to stay level. We will lose that race. It’s only a question of when.

So the why was easy. The how is what we discuss here. With observability wired into thousands of services and many different bespoke platforms, the obvious plan (tear out the old pipeline, get every team to re-instrument on the OTel SDK, flip the switch) is a pipe dream: a multi-year org-wide slog on a pipeline that can’t take an outage with a real chance of dropping the exact data alerts fire on.

The question we actually needed to answer was narrower. How do we replace the whole engine without huge impact across Atlassian services?

The bet: swap the collection and pipeline, leave the interface alone

A metrics pipeline is really a contract with two ends. One end is what service owners see: “send StatsD over UDP to this address → your metrics show up in the backend.” The other is everything between that packet and long-term storage. Teams care enormously about the first end and very little about that middle layer. So we kept the interface and rebuilt everything behind it, which turned an org-wide migration into a platform-team migration.

Two things followed. We put purpose-built OTel Collector distributions at each of the four stages (collection, ingest, aggregation, forward) so we could work on any one without touching the others. And we made the collection-side speak StatsD and OTLP at the same time: nobody had to swap StatsD clients for the OTel SDK before we could start. It helped that the OTel Collector wasn’t new within Atlassian: the tracing team had run it as their pipeline core and host-metrics sidecar for years, so “is it production-ready at our scale?” was already answered.

App - backends flow chart

What we actually built

The migration went step by step in place.

Collection. We replaced the gostatsd sidecar with our OTel Collector distribution, the same one the tracing team already shipped, keeping the app side identical: applications still fire StatsD over UDP as before. Day one, no team noticed. The payoff is that you stop running two sidecars (a StatsD one and a tracing one) on every host. Folding metrics into the tracing sidecar and killing the StatsD one saved about 3.9% CPU on average per service across our priciest Micros services, roughly a 30% cut in sidecar cost at fleet scale. At the same time, we enabled an OTLP receiver enabling our collection layer to receive and forward OTEL metrics natively.

A flow chart image of the observability-sidecar

Ingest. Aggregation is stateful: every datapoint for a time series has to hit the same aggregator, so you can’t just use traditional load-balancing strategies. For years an in-house proxy (called nomad) guaranteed that by hashing (service, environment) to a shard. But our service to metric load distribution is non-uniform and follows a long-tail, whichever shards owned the biggest services turned into hot shards. Our fix was a better routing strategy: the contrib loadbalancingexporter can hash by streamID (the identity of an individual time series) instead of by service, so one big service smears evenly across the pool while any given series still always lands on the same shard.

An image of the service-based routing and streamID routing workflows

After this change: per-shard CPU went from a couple of tall bars and idle replicas to a flat, even distribution. Even load means a tighter autoscaling band, real off-peak scale-down and no more hot-shard pages!

Aggregation. This is the stage that makes the numbers affordable: we take in ~4.8 billion datapoints a minute and land ~220 million, roughly a 96% reduction. Most of our metrics are delta temporality and nothing upstream aggregated deltas the way our users expect, so we wrote our own delta aggregation processor and open-sourced it under atlassian-labs. Same traffic and the aggregation tier now runs on about half the CPU (the aggregators no longer parse gostatsd, load is even and we inherit the community’s tuning).

Forward. The last hop: a bespoke internal forwarder, became a stateless Collector distribution (metrics-gateway) built on upstream exporters. Fan-out to multiple backends (SignalFx, S3, etc) with no custom backend integrations; by default support for retries, queuing and backpressure from the community. Adding a destination is a config change, not a project. This was the easy one.

Lambda. Serverless can’t run a sidecar, so we built an OTel Lambda extension to replace the gostatsd one, with the same StatsD address, same env vars and no code changes. This completes the collection layer that forwards metrics from the services over to our pipeline.

Where this sits in the bigger picture

Getting every stage onto the OTel Collector gives us one codebase and one way of operating. Adding to the pipeline means writing a component, not standing up a service anymore. It unblocks OTEL-based instrumentation without us giving up the aggregation and cost controls that make our scale payable and it lets us drop wasteful datapoints at ingest, the cheapest place to do it. The end state is worth it on cost alone: the gostatsd aggregators and nomad together are ~38% of CPU requests in our metrics clusters and Nomad on its own is ~13% of total resources. Removing them is real money and the last thing between us and a pipeline that’s OpenTelemetry end to end.

What we tell ourselves before starting

  • Pick the right early adopters. Find the teams who have the most to gain and will iterate with you. Leading with dev and staging workloads and the services that felt the pain most gave us real signal fast, from people who were forgiving while we found the rough edges.
  • Profile continuously in production. The real cost and behaviour of a component, ours or upstream only showed up under production load. Small tests and benchmarks were not enough; continuous profiling in prod is what actually told us where to optimise.
  • Match operational workflows. A migration this size runs for months even years and for most of that you’re operating the old and new systems side by side. Keep the overhead of running two systems as low as you can: carry the same operational primitives across and keep parity so nobody has to learn a second way of doing things.
  • Progressive rollouts. Start in the lower environments and lead with the less critical services, then ramp 1% → 10% → 50% → 100%. You want to find problems where they’re cheap; not on the tier-0 path.

What’s next

We have now unblocked our users on moving metrics instrumentation over to OpenTelemetry. The next move is to shift left: get the instrumentation itself onto the OpenTelemetry SDK and off the vendor and in-house clients (Datadog/DogStatsD, StatsD libraries) we’ve carried for years.

We’re also going to start exploring further into the OpenTelemetry ecosystem to solve more large-scale Observability problems we have that the community has solved and also start contributing back as we grow OpenTelemetry with our usage and scale.

Building a reliable cloud native foundation for distributed AI training

AI workloads are changing what platform teams need from infrastructure. Provisioning GPUs and standing up a cluster no longer makes a platform “AI-ready.” Once training spans more than one node, the bottlenecks show up in places application platforms rarely treat as first-class concerns: inter-node communication, shared storage, placement, topology, and validation.

Our internal ML platform supports training and inference workloads behind product experiences such as search and ranking. As these workloads grew, some training jobs outgrew the practical limits of a single machine. Models moved into the tens of billions of parameters, and a single node stopped being able to hold the model, its optimizer state, and a workable batch size at the same time. Atlassian therefore needed a platform that could make distributed training reliable and repeatable, without exposing ML teams to the underlying infrastructure complexity.

For distributed AI, performance is not just optimization. It is part of correctness.

The workload problem behind the infrastructure work

Adding GPUs was the easy part. The platform had to make three things predictable:

  • Communication performance. Workers need fast, low-overhead inter-node communication for synchronization traffic, including gradients, model state, and collective operations, so a distributed job keeps progressing together. The gap here is not marginal: on current-generation GPU hardware, a socket-based path can leave a multi-node job running at roughly half the speed the same hardware delivers over RDMA.
  • Shared data access. Workers need high-throughput shared storage to read training data, write checkpoints, and access intermediate artifacts concurrently without turning storage into a bottleneck or causing long pauses and uneven progress.
  • Operational predictability. Hardware placement, network topology, storage, and validation need to work together so a job does not silently run on a degraded path.

Two technologies cover the first two:

  • RDMA (Remote Direct Memory Access), which enables lower-overhead, high-throughput communication between GPU nodes.
  • Lustre, a parallel distributed filesystem designed for high-throughput shared access to training data and checkpoints.

ML teams should not have to manage either one. That is the platform’s job.

Why this matters beyond one company: None of this is specific to us. As AI adoption grows, platform teams keep hitting the same wall: distributed systems, accelerators, storage and scheduling have to behave as one platform, not four separate layers.

What happened before the high-performance network and storage design

Before RDMA and Lustre, distributed jobs ran. How fast they ran was anyone’s guess. Communication and storage delays surfaced as low GPU utilization, uneven step times, and runs that took far longer than they should have.

The worst of it was silent RDMA fallback. A misconfiguration sent collective traffic over sockets, and the job carried on at a fraction of the speed it should have reached.

That made the degradation easy to miss, and expensive to ignore.

This was not limited to one environment. On our other cloud, the device plugin that advertises the RDMA fabric to Kubernetes has been in CrashLoopBackOff on every fabric-capable production node from the day it was deployed. It was not a regression; it had never worked. We found it 271 days later, by accident, while verifying an unrelated GPU operator upgrade. The cause was mundane: a mirrored container image resolved to a single-architecture manifest that did not match the nodes, so the plugin failed at exec. A stopgap had been applied to staging months earlier, and production never received the follow-up.

Two failure modes matter here. A job that explicitly requests the fabric stays pending indefinitely, which is loud and easy to diagnose. A job that does not request it runs over sockets with no error at all, which is both the default and the expensive case.

Three lessons generalize from that outage:

  • A capability with no consumers is invisible, however expensive it was to buy. No alert existed for “this DaemonSet has never been ready”, and that absence was the real defect.
  • Pod age and restart counts mislead when nodes autoscale. Every new node presents a fresh crash loop, which reads as a recent break when the underlying condition is months old.
  • A healthy environment beside a broken one is a trap. Staging passing told us nothing about production, because staging had quietly diverged onto a different fix.

Both failures share a cause that has nothing to do with networking. Multi-node demand is still low next to single-node work. Few jobs touch the fabric, so nothing generates the signal that would show it is broken. That argues for synthetic validation rather than better dashboards: a dashboard shows what your jobs did, not what a fabric nobody used would have done.

Why a new network and storage design was needed

We treated this as a platform design problem rather than a tuning exercise.

We integrated RDMA-capable networking and shared high-throughput storage into the platform, absorbing topology and cloud-specific complexity so ML teams could run distributed training reliably without managing the underlying infrastructure.

Platform concernBeforeAfter
Inter-node GPU communicationStandard socket/TCP path could become a hidden bottleneckRDMA-capable path for faster collective communication
Shared training storageLess efficient path for concurrent checkpoint and dataset accessLustre-based shared high-throughput storage
Operational confidenceJobs could appear healthy while underperformingValidation made transport and performance behavior visible
User experienceRisk of infrastructure details leaking to ML teamsPlatform-managed capability with consistent workflow

How we achieved it

That meant changes at every layer:

  • RDMA-capable cluster and node-pool setup. Clusters are created with the multi-networking support the fabric requires, and GPU node pools are provisioned against a specific reservation and location rather than a generic pool.
  • GPU-specific network interfaces and network mappings. Each GPU node carries dedicated RDMA interfaces alongside its ordinary one, and each node pool is mapped to the RDMA network and subnet matching where its hardware physically sits.
  • Node-level dependencies, drivers, and readiness controls. The RDMA userspace libraries, the collective communication runtime, and the network plugin are installed on the node image. A node that fails its readiness check never becomes schedulable, so jobs cannot land on a partially configured node.
  • Lustre integration for shared storage. The filesystem is exposed through a CSI driver as an ordinary ReadWriteMany PersistentVolumeClaim, mounted at the same path in every worker pod, with per-namespace and per-job subdirectories providing isolation.
  • Placement and scheduling constraints for the right hardware and topology. Jobs are pinned to node pools with the correct hardware, drivers and fabric wiring, and gang scheduling ensures a distributed job either receives all of its workers or waits, rather than half-starting and holding GPUs idle.
  • Workload-level configuration and validation. The networking and storage plumbing is injected into the pod spec at admission time so ML teams never hand-write it, and the transport is confirmed before a job is treated as healthy.
A flow chart image of the final workflow

The platform absorbs network, storage and topology complexity so that submitting a distributed training job stays ordinary.

One of the key lessons was that RDMA depended on physical topology, not just Kubernetes or workload configuration. GPU reservations could move across datacenters within the same zone, and the RDMA subnet mapping had to remain aligned with where the hardware actually lived. That meant the solution needed to support multiple RDMA network mappings and safe node-pool transitions instead of assuming one static configuration forever. How much of this you inherit depends on where you run. On one of our two clouds the managed fabric handles reservation and zone changes itself, and the platform team never sees them.

We also had to treat reservation changes as platform transitions: introduce a new node pool, move workloads safely, and then retire the old path. In practice that proved more reliable than trying to encode the whole problem as a one-time setup task.

What problems we faced along the way

None of this was a feature flag. We had to solve both the availability of suitable GPU nodes across two of the top hyperscalers and the challenge of introducing a network design that pushed beyond each provider’s usual patterns.

  • GPU node availability across clouds: securing enough compatible GPU capacity was a constraint in both clouds, affecting planning, placement, and the ability to move workloads between providers.
  • A new ask for cloud providers: this combination of GPU placement, high-performance networking, and shared storage was not a standard, one-size-fits-all request. There was no single approach that worked across providers, so the design had to adapt to each cloud’s capabilities and constraints.
  • Topology, reservation, and configuration timing: the correct network path depended on underlying physical placement, and some information needed to select the right RDMA mapping was not always available early enough. This is not universal, and the difference matters if you run in more than one place. On one of our two clouds the managed fabric absorbs reservation and zone changes, and none of it is visible to the platform team. On the other, the mapping is ours to maintain. Do not assume the behavior transfers.
  • Reservation topology is a provider constraint, not a platform choice: we have been allocated GPUs inside a single block, and we have been allocated them spread across several. Whether topology-aware scheduling can help once an allocation spans blocks is still an open question for us.
  • Shared storage carries its own operational cost: the parallel filesystem solved throughput, but onboarding a new region or cluster still means standing up a new filesystem instance by hand. We traded a throughput problem for an operational one.
  • Cross-layer, cross-team delivery: networking, node setup, storage, scheduling, and workload integration all had to line up, requiring close iteration across infrastructure, networking, and ML platform boundaries.
Important lesson: for distributed training, “the job ran” is not a sufficient success signal. You need validation that confirms it ran on the intended transport and storage path.

In practice our validation is coarser than we would like. There is no per-job telemetry that pins a slowdown on the transport. What we have is which storage path a job used, whether its transport was sockets or RDMA, and total training time. That combination is enough to catch the failures described here.

What changed after the new design

Treating RDMA, Lustre, placement and validation as one concern bought us more than a benchmark bump.

We ended up with a more reliable training foundation: fewer hidden performance failures, better GPU utilization, and multi-node workloads that were practical to run repeatedly.

The measured results were significant:

MeasurementTCP fallbackRDMAMethod
Peak bus bandwidth–355 GB/s2-node NCCL all-reduce, H200 141GB nodes
Median step time12.36 s6.07 sQwen2.5-14B FSDP supervised fine-tune, 16 H200 GPUs across 2 nodes
Training throughput1.0×2.04×Controlled A/B, identical ~360 s model-load phase
100-step benchmark~1,950 s~836 sSeparate fine-tune benchmark, 2 nodes

Speed is the least interesting part. What those numbers buy is GPU-hour efficiency, capacity planning we can trust, and confidence that the next model up will scale.

The broader cloud-native lesson

The lesson is not really about RDMA or Lustre. It is about what happens when you stop treating them as separate problems.

Cloud-native platforms already know how to orchestrate complex distributed systems. AI training pushes those platforms into a new set of constraints where network fabric, storage behavior, accelerator scheduling, and validation all become part of the product experience.

That is the job. Keep the user-facing story simple, and let the platform absorb the rest.

If you are building this: treat networking, storage, scheduling, readiness and validation as one system. That is what makes distributed training predictable as it scales.

Central takeaway: serious distributed AI training becomes viable when performance-critical infrastructure is designed as an integrated platform problem, not as a collection of independent features.

References

Automating root cause analysis at scale: Multi-signal correlation for cloud native incident response

The problem: Humans shouldn’t be correlation engines

At Atlassian’s scale, hundreds of interconnected microservices distributed across multiple regions mean a production incident generates an overwhelming volume of telemetry. The problem is that finding the causal factor in the vast amount of telemetry still relies heavily on human expertise, intuition, and manual cross-referencing.

A typical root cause analysis workflow today looks something like this: an on-call engineer gets paged, opens a metrics dashboard, spots an anomaly in error rate or latency, pivots to a logging tool to search for exceptions within that time window, then opens a tracing UI to inspect individual request paths. They visually correlate patterns across these three separate views, form a mental hypothesis about where the fault lies, and then work backward through the service dependency graph to validate it.

This is a serial, cognitively expensive process. It depends on the responder already knowing which dashboards to check, which log queries to run, and which services are upstream of the one that’s failing. Senior engineers with years of domain knowledge can do this in minutes. Everyone else takes significantly longer, and during a user-impacting incident, every minute matters.

We asked a simple question: what if we automated the hypothesis generation step entirely, so responders could skip straight to validation and resolution?

Our approach: Treat RCA as a multi-signal correlation problem

The insight behind our automated RCA system is that root cause analysis is fundamentally a correlation problem across three dimensions:

  1. Signal type (metrics, logs, traces)
  2. Time (anomalies that co-occur are likely related)
  3. Topology (faults propagate along service dependency edges)

If we can detect anomalies independently in each signal, align them on a shared timeline, and then trace them through the known service dependency graph, we can generate ranked hypotheses about where a fault originated and how it propagated to produce the user-visible symptoms.

The system is explicitly designed to be modular. Each anomaly detection method is a pluggable component, and the correlation engine operates on normalised anomaly events regardless of which detector produced them. This means we can iteratively improve individual components, swap a statistical model for an ML model, add a new signal type, without rebuilding the pipeline.

Architecture: From raw telemetry to ranked hypotheses

Step 1: Scope the search using service topology

When an incident is detected, the first thing we do is narrow the blast radius. Rather than analyzing every service in the platform, we query our service dependency graph to identify the set of services in the call path of the degraded user experience. This gives us a focused subgraph, typically tens of services rather than hundreds, where the fault is most likely to exist.

We use OpenTelemetry-derived service maps for this. The dependency graph is built from span-level parent-child relationships observed in production traffic, giving us a real-time picture of how services actually communicate rather than how they’re supposed to communicate according to documentation.

Step 2: Detect anomalies independently per signal

With the relevant services identified, we run specialised anomaly detection modules for each telemetry signal:

Metrics (RED signals): For each service endpoint, we monitor rate, error rate, and duration (RED) using a combination of statistical methods. Median absolute deviation handles spike detection, while percentile bands catch sustained deviations. When a metric crosses its dynamic threshold, we emit a normalised anomaly event with a severity score, the observed value, and the baseline it deviated from.

Distributed traces: We analyse traces flowing through the affected services for structural anomalies, including unexpected exceptions, novel error propagation patterns, and latency spikes at specific spans. The trace-based detector uses both statistical methods (latency percentile violations) and pattern analysis to identify exception types that correlate with the incident window. Each anomalous trace produces an event tied to the service and timestamp where the anomaly was observed.

Logs: We apply clustering techniques to log streams from affected services, surfacing new or rare error clusters that appeared during the incident window. The key challenge here is log volume. At scale, you cannot naively scan every line. We use embedding-based clustering to group semantically similar log entries and flag clusters that are statistically novel relative to the service’s normal error distribution.

Every detector produces events conforming to a common schema:

{

 "timestamp": "2025-07-24T15:24:00Z",

 "service": "payment-service",

 "signal_type": "metric",

 "severity_score": 0.85,

 "details": { ... }

}

This normalisation is critical. It allows the downstream correlation engine to reason across signal types without caring which detector produced the event.

Step 3: Temporal correlation: find anomalies that co-occur

The correlation engine’s first job is to identify clusters of anomalies that happened close together in time. The intuition: if a database starts throwing errors at 15:24, the service that calls it starts timing out at 15:24:30, and the frontend that calls that service starts returning 500s at 15:25, these are almost certainly related.

We use a sliding window approach (configurable, typically plus or minus 5 minutes) to group co-occurring anomalies into correlation bundles. Each bundle receives a temporal cohesion score:

S_temporal = (1 / N(N-1)) × Σ exp(-|ti - tj| / τ)

Where N is the number of events, ti and tj are event timestamps, and τ is a tunable decay constant. Tightly clustered anomalies score higher than dispersed ones.

A critical refinement we discovered in practice: the same causal chain often replays multiple times during an incident (the same upstream timeout propagates the same downstream failure every few seconds). Without deduplication, this produces redundant bundles that obscure the signal. We solve this with sequence fingerprinting, computing a fingerprint from the ordered list of services in each anomaly path and collapsing repeated sequences into a single bundle with a replay count. This lets us say “this failure pattern repeated 47 times in 5 minutes” rather than generating 47 identical hypotheses.

Step 4: Graph-based impact analysis: find the causal direction

Temporal correlation tells us which anomalies happened together. Graph-based analysis tells us which service is the cause and which are effects.

For each correlation bundle, we identify the sink node, the service with the highest anomaly severity, which typically represents the most impacted point visible to users. We then traverse upstream in the dependency graph (BFS, bounded to a configurable depth) looking for anomalous neighbours.

The key insight: if Service A calls Service B, and both are anomalous in the same time window, but Service B’s anomaly preceded Service A’s, then Service B is more likely to be the fault origin and Service A is experiencing a downstream effect.

We score each candidate causal path:

S_path = (1/m) × Σ S_anomaly(Ui) × w_edge(Ui → Ui+1) × exp(-α × Δt)

Where m is the path length, S_anomaly is the anomaly severity at each node, w_edge captures the strength of the dependency relationship, and the exponential decay penalises anomalies that are temporally distant from the sink. The path with the highest score represents our best guess at the fault propagation chain.

Step 5: Hypothesis generation and narrative

The final step combines the temporal cohesion score and path score into an overall confidence score for each correlation bundle:

S_overall = w1 × S_temporal + w2 × S_path

We rank bundles by this score and emit the top N as root cause hypotheses. Each hypothesis includes:

  • The suspected root cause service (the upstream origin of the fault)
  • The propagation path showing how the fault spread to produce user-visible symptoms
  • Evidence at each node (which metrics breached thresholds, which exceptions appeared, which trace IDs exhibit the failure)
  • A confidence score and breakdown of how it was calculated
  • A human-readable narrative explaining the hypothesis in plain language

This last point matters more than it might seem. A ranked list of services with scores is useful for machines, but responders need to quickly assess whether a hypothesis is worth pursuing. Our narrative templates produce explanations like:

“Between 15:24 and 15:28, the payment-service endpoint /charge exhibited a 4x increase in error rate (baseline: 0.2%, observed: 0.8%). This preceded a latency spike in checkout-service /complete (p99: 340ms to 2100ms), which propagated to the frontend as HTTP 500 errors. The fault likely originated in payment-service based on temporal precedence and graph position. Confidence: 0.87.”

Fitting into a broader reliability platform

Automated RCA does not exist in isolation. At Atlassian, we are building a cohesive incident response platform that integrates automated user-impact detection, faulty service identification, causal diagnosis, and an AI-powered incident copilot into a single responder experience.

Our RCA engine serves as the diagnostic brain of this platform. When a user-impacting incident is detected, whether automatically via real-time user experience signals or manually by support teams observing ticket surges, the RCA engine is triggered. It publishes its hypotheses into a shared incident context that other components consume:

  • Faulty service identification uses early-stage RCA results to page the right team, reducing time to engage.
  • An incident copilot uses the hypotheses and their evidence to explain what is happening to responders and recommend mitigation actions (rollbacks, feature flag disablement) grounded in the actual diagnosis rather than generic runbooks.
  • A feedback loop captures whether responders accepted, rejected, or refined each hypothesis, allowing us to tune weights and improve accuracy over time.

The shared incident context is the critical integration point. By anchoring all signals, hypotheses, and actions to a single per-incident context, regardless of which system produced them, we ensure responders see one consistent view rather than reconciling outputs from multiple disconnected tools.

Lessons learned and design trade-offs

Start with the simplest anomaly detection that is useful, not the most sophisticated. Our initial impulse was to build complex ML models for every signal. In practice, statistical methods (MAD, percentile bands) work surprisingly well for metrics anomaly detection and are far easier to debug when they produce false positives. We reserve ML approaches for signals where statistical methods genuinely struggle, such as log clustering and trace structural analysis.

Modularity pays compound interest. Because each anomaly detector is a pluggable module behind a normalised interface, we could ship a useful system with just metrics-based detection, then incrementally add trace and log detectors without touching the correlation engine. Each new module immediately improved hypothesis quality because the correlation engine had more evidence to work with.

Deduplication is not optional at scale. Without sequence fingerprinting and replay collapsing, a single fault pattern that replays 100 times during an incident produces 100 bundles. This overwhelms both the scoring pipeline and the human reading the results.

The dependency graph is your most powerful prior. Temporal correlation alone produces too many hypotheses. Many services are anomalous during an incident because they are affected, not because they are faulty. The graph provides causal direction and dramatically reduces the hypothesis space.

Narratives build trust. Engineers will not act on a confidence score alone. They need to see the evidence and the reasoning that connects it. Our templated narratives with embedded telemetry references (specific trace IDs, specific metric values) let responders validate a hypothesis in seconds rather than re-deriving it from scratch.

What’s next

We are actively exploring how LLM-based orchestration can make this system iterative rather than one-shot. Today, the RCA engine runs once per incident trigger. The next step is an agent that can request additional telemetry, refine its hypotheses based on what it finds, and adapt its investigation strategy based on what signals are available, much like an experienced human responder would. The challenge is doing this safely: with appropriate rate limits, sandboxed execution, and clear provenance so responders always know what evidence supports each conclusion.

We are also expanding the signal types the system can reason about, including infrastructure metrics, deployment events, feature flag changes, and synthetic check results, to move from “which service is broken” toward “what change caused it to break.”

Key takeaways

  • Automated RCA is a correlation problem across three dimensions: signal type, time, and service topology. Solving it requires normalised anomaly events, temporal alignment, and graph-based causal inference.
  • Modular, incrementally useful design beats big-bang delivery. Ship the simplest version that provides value, then layer in additional signal types and more sophisticated detection methods.
  • OpenTelemetry provides the foundation. Standardised traces give you the dependency graph for free and provide the structural data needed for trace-based anomaly detection. Without consistent, correlated telemetry, multi-signal RCA is impossible.
  • Invest in explainability from day one. Confidence scores are necessary but not sufficient. Human-readable narratives with evidence provenance are what build the trust needed for responders to actually act on automated diagnosis.
  • Build feedback loops early. The system improves only if you capture whether hypotheses were correct. Simple thumbs-up/thumbs-down on each hypothesis is enough to start tuning weights and identifying systematic blind spots.

❌