❌

Vue normale

Reçu avant avant-hierInfra

Distributed tracing for CI pipelines without touching a single workflow file

You’ve probably felt this one: GitHub Actions usage creeps up across your org, and your actual visibility into it doesn’t keep pace. Which workflows are slow? Which are flaky? How long are jobs sitting queued for a runner before they’ve even started doing anything? GitHub’s own insights are per-repo and shallow. There’s no cross-org view of CI health, no way to slice by team or workflow type, no way to alert when things quietly get worse. Someone eventually asks why CI took forty minutes yesterday, and the honest answer is “let me go check that one repo and get back to you.”

So let’s actually fix that properly, without asking a single team to touch a single workflow file.

The obvious fix doesn’t scale

The instinctive answer is to instrument each workflow: add a tracing step, wire up an SDK, sprinkle spans through the YAML. It works, technically. But it means every team has to opt in, every new repo starts blind until someone remembers to add it, and you end up maintaining instrumentation scattered across however many workflow files exist across the org. That approach doesn’t scale with your org, it scales with how diligent everyone stays about something that isn’t their actual job.

The actual insight

GitHub already knows almost everything you want. Every workflow run and every job inside it fires an event: workflow_run and workflow_job. You don’t need to ask each repo to report on itself. You just need to listen to what GitHub is already telling you, at the org level, once.

It’s like trying to track a whole apartment building’s water usage by asking every tenant to self-report their reading. Most will forget. New tenants won’t even know they’re supposed to. The easier answer is to read the one meter at the street, where every pipe in the building already converges, whether the tenants know it’s there or not.

How it actually works

An OpenTelemetry Collector with the githubreceiver component sits behind a single org-level GitHub webhook and converts incoming workflow_run and workflow_job events straight into OTLP spans. Worth knowing upfront: it’s a contrib component still at alpha stability, so the config surface can shift. Pin a specific collector version rather than tracking latest, and skim the changelog before bumping it, cheap insurance against a config field quietly changing shape under you. If tracing is new to you, the mapping is intuitive once you see it: a workflow becomes one outer span, each job inside it a child span, each step inside a job a child of that, so what you get is something you can actually drill into rather than a flat pile of events.

One nice detail: span and trace IDs are generated deterministically, hashed from the workflow’s run ID and each job’s check run ID. If you ever want to emit your own telemetry from inside a step, there’s tooling for this, it can compute the matching ID and attach directly to the same trace without any coordination with the collector.

receivers:
  github:
    webhook:
      endpoint: 0.0.0.0:19418
      path: /events
      secret: ${env:GITHUB_WEBHOOK_SECRET}
    scrapers: # required even if you only want tracing, a dummy entry is enough
      scraper:
        github_org: ${env:GITHUB_ORG}

exporters:
  otlp:
    endpoint: ${env:TRACE_BACKEND_ENDPOINT}
    headers:
      authorization: ${env:TRACE_BACKEND_API_KEY}

service:
  pipelines:
    traces:
      receivers: [github]
      exporters: [otlp]

That scrapers block looks unrelated to tracing, and it is, it belongs to a separate GraphQL/REST metrics feature the same receiver offers, but the config fails validation without at least a dummy entry, even if all you want is the webhook side. Easy to lose twenty minutes to that the first time.

Point the exporter at Tempo, Jaeger, Datadog, or whatever your team already pays for, and it just works, that’s the actual point of using OTLP rather than a vendor-specific format. The backend is genuinely the least interesting decision in this whole setup.

Before you deploy this

A few things worth deciding upfront, since the config alone won’t force you to think about them.

The collector needs a publicly reachable endpoint, GitHub has to deliver webhooks to it, so plan for IP allowlisting or a WAF restricted to GitHub’s webhook source ranges rather than leaning on the shared secret as your only line of defence. There’s also a GitHub App option if you’d rather not manage a shared secret directly, worth a look if secret rotation across many services is already a headache for your team.

Setting up an org-level webhook needs org admin access, worth confirming early rather than discovering it mid-rollout. And if you’re on GitHub Enterprise Server rather than github.com, I’d validate that webhook delivery behaves the same way in your setup before assuming this is a drop-in.

None of that is difficult, it’s just easy to skip past when you’re excited about the zero-instrumentation part. Get it sorted early and it’s a one-time cost: point one org-level webhook at the collector and every repo in the org is covered from that moment on, including repos that don’t exist yet. Nobody has to remember to switch anything on.

Sizing it before you build it

Worth walking through the actual method here, since “just turn on tracing for everything” is a good way to end up with an unwelcome bill or an unwelcome conversation with whoever owns your tracing budget.

Start by scanning the org for total repo count, then immediately throw that number away. It’s nearly meaningless on its own. Most orgs of any real size are carrying a long tail of dormant, forked, archived, or abandoned repos that inflate the headline count without generating any real CI traffic. What actually matters is the active slice: how many repos had genuine workflow activity in a real week, not how many exist.

From there, extrapolate outward. Active repo count times average runs per repo gives you expected workflow volume. Workflow volume times average steps per workflow gives you expected span volume. Span volume times typical payload size gives you an expected data volume per day. Compare that against whatever tracing volume your infrastructure already handles for application traces, and in most orgs, CI trace volume turns out to be a rounding error next to it.

That comparison is the actual point, not any specific number I could hand you. What generalises is the method: measure real activity instead of headline repo count, and walk in with a comparison rather than an assertion.

Getting buy-in

The technical build was the easy part. The harder part was justifying the data volume to whoever owns the tracing budget, especially with cost concerns already floating around about the backend in question. Showing up with an actual sizing exercise, not “trust me, it’s small,” turns that into a five-minute conversation instead of a drawn-out one. Nobody has to take your word for “it’s small” when they can see it sitting next to the tracing volume they’re already paying for without blinking.

It’s also worth remembering that standing up a new observability project is as much an ownership question as a technical one. Someone has to actually own the collector, the webhook, the alerting rules going forward. Sorting that out early saves the awkward moment three months later when something breaks and nobody’s sure whose pager it is.

Why this is worth doing

The zero-instrumentation part is the whole value here. New repos are observable the moment they’re created, not the moment someone remembers to add tracing to them. And because everything lands as proper OTel traces, CI health sits in the same tool as your application traces, so a slow deploy and a slow downstream service can be correlated instead of investigated in two different dashboards by two different people who don’t talk to each other until Thursday.

If you’re running self-hosted runners on the Actions Runner Controller, it’s worth being clear this doesn’t replace what ARC already gives you, it sits on a different layer entirely. ARC’s own metrics tell you about your runner fleet: how many pods exist, whether autoscaling is keeping up, how deep the queue is. This tells you about your workflows: why a specific run was slow, which ones are flaky, where the time actually went. Queue depth and queue time are practically the same question asked from two different vantage points. Worth running both, not picking one.


This covers the collection side. It doesn’t get into building good alerting on top of the trace data, that’s queue-time thresholds, flaky-test detection, and how noisy those alerts get before people start ignoring them, which is a genuinely separate problem and probably its own post.

Migrating a critical Kubernetes deployment from the default namespace without any downtime

Somewhere in your cluster there’s probably a deployment sitting in the default namespace that everyone knows shouldn’t be there. Nobody put it there maliciously, it just happened, early on, before anyone had opinions about namespace hygiene, and now half your other services quietly depend on it. Moving it is now a tricky problem.

That was the exact situation with a service I’ll call auth-svc: an authentication service that dozens of other services called constantly, sitting in default for years, and about to become a genuine problem the moment it needed namespace-scoped things, its own ingress rules, its own policies, that default structurally couldn’t give it. Moving it wasn’t optional forever. But it also couldn’t go down, not even for a few seconds. This wasn’t a vague “other services might complain” risk: auth-svc handled authentication for that entire region’s cluster, so if it went down, nobody in that region could log in. Full stop.

Before getting into why this is actually hard, it’s worth being precise about what “moving it” means. There are two completely separate paths into auth-svc, and both have to keep working throughout the move, or fixing one just creates an outage in the other. Everything inside the cluster reaches it the ordinary way: other services resolve auth-svc.default.svc.cluster.local through Kubernetes’ own internal DNS and get routed to a pod, the standard Service mechanism. Everything outside the cluster reaches it through an ingress instead, a completely separate mechanism that has nothing to do with that DNS name. Whatever the fix turned out to be, it had to solve for both paths, not just the one that’s easier to reason about.

Why “just move it” doesn’t work

The obvious plan, move the deployment, update the references, done, falls apart the moment you look at who’s actually calling this thing. As mentioned, dozens of other services reference auth-svc by its cluster-internal DNS name, owned by different teams, on different release cycles. There’s no atomic moment where you flip a switch and every one of them simultaneously starts using a new name. Some team’s service hasn’t been redeployed in months. You shouldn’t be coordinating that.

The tooling got in the way too. Our deploy pipeline only knew how to ship a service to one namespace. There was no “deploy this to two places at once” option, and modifying the shared pipeline logic every other team also depended on felt like exactly the kind of blast radius we didn’t want to introduce. Whatever the fix was, it had to fit inside a single-namespace deploy, not require rewriting shared infrastructure.

On top of that, we had an OPA policy which enforced that identical ingress rules couldn’t exist live in two namespaces at once, a sane rule that exists specifically to stop the kind of half-finished migration that leaves routing ambiguous.

The insight: a forwarding address

The piece that made this solvable: I didn’t need to migrate every consumer’s understanding of where auth-svc lives. I needed to migrate the service, and quietly redirect anyone still asking for the old address.

Kubernetes has exactly this mechanism, and it’s easy to forget it exists because you almost never need it: an ExternalName service. Instead of pointing at pods, it points at another DNS name, functioning essentially like a CNAME (see my DNS article if you want to learn more about how that works). Deploy the real thing at its new home, then convert the old Service object into a forwarding address:

apiVersion: v1
kind: Service
metadata:
  name: auth-svc
  namespace: default
spec:
  type: ExternalName
  externalName: auth-svc.authentication.svc.cluster.local

Every consumer still calling auth-svc.default.svc.cluster.local gets silently redirected to the real thing in its new namespace. Nobody changes a line of code on their end. It’s the same trick as a postal forwarding order: you don’t visit every person who might send you mail and update their address book, you tell the post office where you actually live now, and everything gets redirected until people eventually update it themselves, at their own pace, with zero coordination required on your part.

That forwarding trick only earns its keep if you actually confirm it’s working before you lean on it. Once the proxy was live, the next step wasn’t scaling anything down, it was watching the metrics through the crossover: checking that traffic hitting the old address was genuinely landing on the new deployment, not silently failing or looping somewhere. Only once that looked clean did the old pods get scaled to zero rather than deleted outright. Scaling to zero costs nothing and buys an instant rollback, just scale back up, if anything downstream looked wrong later. Deleting them outright would have meant rebuilding from scratch if something went sideways, so there was no reason to give up that safety net early. We could defer the cleanup to a later point in time.

The chicken-and-egg problem

The remaining wrinkle was the ingress. External traffic to auth-svc doesn’t come in through the DNS-based Service mechanism at all, it comes in through an ingress, and I needed a working ingress in the new namespace before I could safely remove the one in the old namespace. But the policy engine wouldn’t allow both to exist at once; identical ingress rules across two namespaces is exactly the ambiguous state it exists to prevent.

Classic chicken-and-egg: can’t create the new one without a policy exception, can’t delete the old one first without a traffic gap.

The fix was a temporary, explicit exception rather than fighting the policy itself: annotate the new namespace to bypass the duplicate-ingress check just for this migration, stand up the new ingress alongside the old one for a short overlap window, confirm traffic was flowing correctly to the new deployment, then delete the old ingress and let the exception age out. A brief, deliberate window where both existed, rather than a gap where neither did.

The patch would look something like this:

apiVersion: v1
kind: Namespace
metadata:
  name: authentication
  annotations:
    policy.example.com/allow-duplicate-ingress: "true"

How it actually shipped

None of this went straight to production. It ran in dev first, then a staging cutover, with a couple of weeks between the staging success and doing it for real, mostly just to sit with it and see whether anything subtle showed up under real traffic before betting a critical path on it.

The production cutover itself ended up being the boring part. All the actual difficulty was front-loaded into getting the design right. Once the plan was solid, executing it was closer to a formality than an event.

Why this matters beyond one service

Here’s the thing about default namespace sprawl: it’s rarely caused by carelessness. It’s caused by the complete absence of pressure to ever fix it. Someone creates a policy restricting new services from landing in default, good practice, but nobody sets a deadline or a plan for the services already there. They just sit. For years, in this case. Nothing forces the issue until a team needs something namespace-scoped that default structurally can’t give them, and only then does the debt come due.

If you’re staring at a similarly stuck service, the pattern generalises past this one migration. An ExternalName proxy buys you a zero-coordination path to move anything addressed by DNS, provided you’re willing to hold two versions in careful overlap for a short, deliberate window rather than trying to cut everything over at once. The next time someone tells you a service can’t be moved because too many things depend on it, that’s usually a sign nobody’s looked for the DNS-shaped seam it can be split along.


This only works cleanly because everything here talked to auth-svc through its DNS name rather than a hardcoded IP or ClusterIP. If you’ve got consumers that skip DNS entirely, and some legacy systems do, you’re solving a different, uglier problem.

❌