Security Slam 2026 – Fall Edition is a 30-day virtual event from October 5 through November 6, 2026.
What Is the Security Slam?
The Open Source Security Foundation (OpenSSF) is partnering with the Cloud Native Computing Foundation (CNCF) Security Technical Advisory Group (TAG Security) to support the 2026 Security Slam at KubeCon + CloudNativeCon North America.
The 30-day challenge runs from October 5 through November 6 and highlights OpenSSF projects as practical tools that help improve project security posture. Participants will use OpenSSF projects, among others, to achieve security hygiene milestones tailored to their project’s maturity level.
OpenSSF project leads, staff, and maintainers have assisted in the creation of the “Slam Library,” a set of web resources to guide participants through each challenge, and will continue to be available throughout the month via the official Security Slam website.
How to Participate
Register now to receive reminders and instructions before the event kicks off on October 5. Stop by the OpenSSF booth #313 in the KubeCon Solutions Showcase anytime during the week of November 10-12 to pick up participant achievement awards.
A Growing Community Effort
The Security Slam is a CNCF community activity that has taken many different shapes over the years. Now on its sixth iteration, the Slam is designed to help projects understand and improve their high level security posture.
Expanded Eligibility
Previously limited to CNCF projects due to the nature of the evaluation tools available, the Slam is now taking advantage of new tools to greatly broaden the qualifications for participation: Any open source project is invited to participate!
The event has had several permutations in its length. In the case of the Kubernetes Lightning Round, the slam was a day of onboarding new contributors to Kubernetes with a focus on security hygiene improvements to seven different subprojects. Taking it a step further, the 2025 event featured weeks of preparatory work with maintainers, and 45-minute live sessions with maintainers and anyone who wanted to join from the audience at KubeCon + CloudNativeCon Europe.
This year returns to the 30-day format that produced strong results in 2023. Then, projects were given their own iron-on badges and a framed plaque to highlight the milestones that they completed during the 30-day event. Not only were the plaques seen at project tables long after the event ended, but we received reports of significant project wins due to the efforts achieved during that event. The 2026 Fall Security Slam builds on the success of earlier events, including the Spring event, where projects achieved major security milestones.
What to Expect for the 2026 Fall Edition
Here are some key similarities you will see:
The project will last approximately one month, leading up to KubeCon
CNCF TAG Security & Compliance will publish a library of support resources to accelerate execution of the more complex goals
Advisors will be available via a dedicated CNCF slack channel all month, to offer clarifications and answer questions related to security hygiene
Participating projects will be given custom recognitions to demonstrate their success
Individual contributors will be given physical and digital badges corresponding to the project’s completed goals
Observability Day returns to KubeCon + CloudNativeCon North America on November 9, 2026, in Salt Lake City, Utah, bringing together maintainers, operators, and end users from across the CNCF observability community.
Observability has reached an important transition point. OpenTelemetry graduated within CNCF in May 2026, established projects continue to evolve, and teams are applying these technologies to new challenges including AI workloads, telemetry cost, scale, and data quality. Observability Day provides a vendor-neutral place for maintainers and practitioners to compare approaches and learn how the wider ecosystem is responding.
From FluentCon to Observability Day
Observability Day traces its roots to FluentCon, first held as a co-located event at KubeCon + CloudNativeCon Europe 2022 in Valencia, Spain, for the Fluentd and Fluent Bit communities. Later that year, Open Observability Day at KubeCon + CloudNativeCon North America broadened the gathering to include maintainers and community members from projects including Jaeger, Fluentd, Thanos, Prometheus, and OpenTelemetry, laying the foundation for Observability Day.
Cloud native observability already had strong individual project communities, but there was an opportunity to create a shared forum where those communities — and the practitioners combining their projects in production — could learn from one another.
The event grew from that opportunity, connecting maintainers, operators, platform teams, and end users around interoperability, real operational challenges, and the direction of open source observability as a whole.
Why Observability Day?
Observability Day concentrates on vendor-neutral, cross-project observability rather than centering the discussion on a single project.
The event brings together people building and maintaining CNCF observability projects with people operating observability systems in production. This creates space to examine not only individual technologies but also how instrumentation, telemetry pipelines, storage, correlation, and operations fit together across modern cloud native environments.
The official 2026 program spans project internals, cross-project architectures, the data engineering required to make telemetry useful and affordable, and emerging areas including AI and agent observability, eBPF-based instrumentation, and CI/CD and platform telemetry.
For the wider KubeCon + CloudNativeCon community, Observability Day provides an opportunity to explore these topics through real architectures, operational experience, technical trade-offs, and open source technologies.
What to expect
Observability Day takes place on November 9 as part of the CNCF-hosted co-located events at KubeCon + CloudNativeCon North America 2026.
The program brings together maintainers, operators, and end users from projects including Prometheus, Fluentd and Fluent Bit, OpenTelemetry, Jaeger, Thanos, Cortex, Perses, Pixie, Kepler, Inspektor Gadget, and the wider observability ecosystem.
The day is focused on experience-driven content and the practical realities of observing, debugging, and operating cloud native systems at scale.
Attendees will have an opportunity to hear how different communities and practitioners are approaching observability challenges, understand technical and architectural trade-offs, and connect with others working across the open source observability ecosystem.
Observability and Cloud Native AI
Unsurprisingly, Observability is the event’s primary cloud native theme. Sessions explore areas including project internals, production deployments, telemetry pipelines, cost and data quality, eBPF-based instrumentation, and cross-project architectures.
Cloud Native AI is also an increasingly relevant theme. As AI workloads and agents become part of cloud native environments, understanding their behavior introduces new observability questions alongside familiar concerns around performance, reliability, scale, and cost.
The official Observability Day program reflects this evolution, with AI and agent observability among the emerging areas being explored by the community.
What can attendees take away?
Attendees should leave with practical perspectives and trade-offs that they can apply to production systems, as well as a clearer understanding of how different parts of the open source observability ecosystem fit together.
The event provides opportunities to learn from real architectures and operational experience, including approaches to telemetry pipelines, data quality and cost, instrumentation, observability at scale, and emerging approaches to AI and agent observability.
Attendees should also gain a broader picture of how CNCF observability projects complement one another and how the ecosystem is evolving as requirements change.
Who should attend?
Observability Day is especially relevant to site reliability engineers, platform engineers, developers, and observability practitioners responsible for production systems, as well as maintainers and contributors who want to exchange knowledge directly with users.
Practitioners already operating observability systems at scale can explore deeper architectural and implementation trade-offs, while newcomers can use the event to better understand the projects, technologies, and communities that make up the cloud native observability ecosystem.
No specific preparation is required. Familiarity with Kubernetes, distributed systems, and the basic roles of metrics, logs, traces, and telemetry pipelines can help attendees follow deeper technical discussions.
Attendees should also review the final schedule in advance to identify the sessions and speakers most relevant to their work.
Strengthening the observability community
Observability Day brings project maintainers, contributors, operators, and end users together, creating opportunities for direct feedback between those who build open source observability technologies and those who depend on them.
By encouraging discussion across project boundaries, the event gives practitioners a place to share production experience, helps maintainers understand how their projects are being used together, and gives newcomers opportunities to discover projects and communities where they can participate.
That cross-project exchange is particularly valuable as observability continues to evolve. OpenTelemetry’s graduation, the continued development of established CNCF observability projects, and emerging challenges around AI, scale, cost, and data quality all reinforce the importance of collaboration across the ecosystem.
Join Observability Day on November 9, 2026, in Salt Lake City, Utah, as part of KubeCon + CloudNativeCon North America 2026.
Neutrality is quietly the hardest part of open source. It gets tricky the moment someone pays your salary — and staying honest about it takes more effort than anyone admits.
Here’s something we don’t say out loud often enough: most of us doing open source are paid to do it, at least part of the time. And this isn’t just a hunch. When Google’s Open Source Programs Office surveyed over a thousand contributors in 2023, they found that 82% of us do open source at least partly on paid time — and only 18% are pure hobbyists working purely on their own clock. More than half do both, blending personal passion with a paycheck. So the romantic picture of open source as a world of unpaid volunteers? It’s real for about one in five of us. For everyone else, a company is somewhere in the mix.
That’s not a bad thing, it’s how a huge amount of great work actually gets funded. But it does mean that conflicts of interest aren’t some rare edge case you might bump into one day. They’re there from the start, for basically all of us.
And it only gets more tangled the longer you stick around. Give it time and you’re rarely wearing just one label. Maybe you maintain a project, help run a working group, represent your employer somewhere, and volunteer for something on the side. None of that is unusual. But every one of those roles comes with its own responsibilities, its own audience, and its own quiet set of interests. Keeping them straight isn’t a nice-to-have. It’s part of the work.
So what holds all of it together? Neutrality.
The one question: who’s speaking?
The single most useful habit I’ve picked up is asking myself a small question, over and over: who’s actually speaking right now and in what capacity?
When you give a talk, sit for an interview, or drop a comment in a meeting, which version of you is talking? The maintainer? The person representing their team? The volunteer? The employee? Those aren’t the same voice, and quietly blurring them is how you end up nudging decisions you had no real standing to nudge.
That’s why, in a lot of meetings, people take a second to make it explicit: something as simple as “just to be clear, I’m saying this as a maintainer, not on behalf of my employer.” The first time you hear it, it can sound a little stiff. It isn’t. It’s a small honesty that helps everyone in the room weigh what you just said and, just as importantly, it forces you to notice which interest you’re really carrying in that moment. You don’t need a formal ritual for it. You just need the self-awareness to flag it when it matters.
Company belongs way in the back
If neutrality is the principle, here’s the practical version: your employer should sit as far in the back as you can manage.
Open source is vendor-neutral by design. Is the influence actually zero? Of course not. Priorities exist. Roadmaps get shaped by who’s paying whom to work on what — and if 82% of us are on someone’s clock, that shaping is happening constantly, whether we name it or not. That’s just reality, and pretending otherwise is naive. But the decisions should still come down to what’s best for the project and the community first and not what’s best for your employer, and not what’s best for your own career.
That last part is the uncomfortable one. Your personal goals have to take a back seat too. Not just the company’s agenda, also yours. That’s genuinely hard, and it takes ongoing, slightly awkward self-reflection. Anyone who claims they’ve got it perfectly figured out is either kidding you or not paying attention.
Give people the benefit of the doubt
Here’s the flip side, and it matters just as much. When you notice this stuff, not in yourself, but in someone else, start from the assumption that they meant well.
We’re all human. We all slip. Most of the time, when someone’s employer creeps into a conversation where it doesn’t belong, it isn’t some scheme. It’s a person who just didn’t catch themselves in the moment. Nearly always, a quiet conversation sorts it out. Sure, once in a while there’s a real agenda humming in the background. We’re people, that happens. But treating every slip as a conspiracy poisons the exact collaboration you’re trying to protect.
The whole point is that we’re working on this together. We want to move the projects, the community, and the wider ecosystem forward. That should sit behind every decision and every conversation. And if a hard conversation is what it takes to get there, have it, kindly, and in good faith.
The takeaway
Wearing a few different hats isn’t a status thing. It’s mostly a reminder to stay self-aware. Before you speak, take a second to notice which role you’re really in. Put your employer and your ego toward the back. And when someone else stumbles, assume good faith first.
Project first. Community first. Everything else after.
There is something surreal about your first KubeCon being one where you walk onto the stage as a speaker. Most people ease into this community by attending a few conferences, lurking in hallway tracks, and working up the courage to submit a CFP. I did it backward. KubeCon + CloudNativeCon India 2026 in Mumbai was my very first KubeCon + CloudNativeCon event, and I experienced it from both sides of the podium.
Here is how it went.
The talk: Running an AI cluster on the DGX Spark
I co-presented “Run Your Own AI Cluster on a DGX Spark: Kubernetes, GPUs, and DRA” with my co-speaker, who is also my dad, Janakiram MSV. Presenting alongside him made this special on a level that goes beyond the conference itself. I grew up watching him speak at technology events. Standing next to him at the same podium, in front of the KubeCon audience, felt like a full-circle moment.
The talk itself covered how we turned NVIDIA’s DGX Spark into a self-hosted AI cluster. We walked through building a Kubernetes cluster on the hardware, exposing GPUs to workloads, and using Dynamic Resource Allocation (DRA) to schedule them properly, ending with serving models on infrastructure you fully own and control.
Speaking at a conference of this scale for the first time taught me a few things quickly. The rehearsals matter. The AV check matters. And no amount of preparation fully prepares you for looking out at a hall that size. But once we got going, the nerves faded, and it just became a conversation about technology we genuinely love working with.
The Two Pins I Carried Home
Somewhere between the sessions and the hallway conversations, I picked up two small pins: the blue Kubestronaut pin and its golden counterpart.
The Golden Kubestronaut title means completing every CNCF certification there is. For me, that was less a trophy hunt and more a long, unglamorous grind of labs, practice environments, and a lot of weekends. I started it because I wanted my fundamentals to be real, not resume-deep.
As it happens, I am the youngest person in India to complete it. I did not think much about that fact until people at the event started reacting to it, and their reactions honestly meant more than the milestone itself. If there is anything worth taking from my path, it is not the record. It is that the entire journey ran on things this community built and gave away for free: open documentation, community-run study groups, and platforms like KodeKloud. The pins are just a small, physical reminder of that.
The People: Why KubeCon is really about the Hallway Track
Everyone tells you that the real value of KubeCon is the people. I can now confirm this firsthand.
Saiyam Pathak was one of the highlights of the entire event for me. After spending real time with him across the conference, I came away having made a genuine friend in the community. He is exactly as generous and energetic in person as his content suggests.
Mumshad Mannambeth, founder of KodeKloud, was another meeting I will not forget. KodeKloud’s labs were a core part of my certification journey, so getting to thank the person behind the platform in person and talk about where cloud native learning is headed meant a lot.
During a CXO meet hosted alongside the conference, I also met Yongkang He, founder of Kubestrong. When the Golden Kubestronaut milestone came up, he insisted on capturing the moment with a photo together. Moments like that are a reminder of how much this community celebrates its own.
And of course, spending the event with the Nirmata team, meeting community members at our booth, and putting faces to GitHub handles I have interacted with for over a year made the whole thing feel less like a conference and more like a reunion I had somehow never attended before.
What Surprised Me as a First-Timer
A few honest observations from someone who had never been to KubeCon + CloudNativeCon before:
The event scale is massive, but navigable. Thousands of attendees and a massive venue may sound overwhelming, but the community is unusually approachable. Speakers, maintainers, and founders all walk the same hallways, and almost everyone is happy to talk.
Being a speaker changes the experience. The speaker badge is a conversation starter. People come up to you after your talk with questions, ideas, and sometimes job leads. If you have been on the fence about submitting a CFP, this alone is worth it.
India’s cloud native community is enormous and hungry. The energy in the sessions, the depth of questions, and the sheer number of students and early-career engineers in attendance made it clear that this region will shape the next decade of this ecosystem.
What I Am Taking Home
Beyond the badge, the pins, and the photos, I am taking home three things:
Confidence. I now know I can stand on the KubeCon stage and deliver a technical talk. The next CFP will be easier to write.
Relationships. The friendships and connections from this week are the kind that compound over years in this community.
Momentum. Talking to people about GPUs, DRA, and AI on Kubernetes all week confirmed that this intersection is exactly where I want to be building.
If you are an early-career engineer wondering whether KubeCon + CloudNativeCon is worth it, or whether your CFP idea is good enough, take this as your sign. Write the talk, submit it, and show up. The experience changed how I see my place in this community.
See you at the next one.
Shreyas Mocherla is a Software Engineer at Nirmata, working on Kubernetes policy-as-code and AI agent tooling. He co-presented at KubeCon + CloudNativeCon India 2026 with Janakiram MSV.
This November, the cloud native community is heading to Salt Lake City.
From November 9–12, 2026, KubeCon + CloudNativeCon North America will bring together adopters and technologists from across open source and cloud native communities for four days of collaboration, learning, and innovation.
As CNCF’s flagship conference, KubeCon + CloudNativeCon is an opportunity to hear from the people building and using cloud native technologies, connect with communities across the CNCF ecosystem, and explore the technologies and ideas shaping the future of cloud native computing and AI infrastructure.
With the schedule now available, it’s time to start planning your week.
Monday: Connect with the cloud native community
Monday, November 9 is dedicated to pre-event programming, including CNCF-hosted co-located events and Project Lightning Talks.
The CNCF-hosted co-located program spans a wide range of communities and technologies, including Agentics Day: MCP + Agents, ArgoCon, BackstageCon, CiliumCon, Cloud Native AI + Inference Day, FluxCon, Kubernetes on Edge Day, Observability Day, Open Source SecurityCon, OpenTofu Day, and Platform Engineering Day.
It’s a chance to spend focused time with communities working on everything from GitOps, developer platforms, networking, observability, and security to AI, inference, edge computing, and infrastructure as code.
Project Lightning Talks are also on the Monday schedule, offering short sessions focused on projects and technologies from across the cloud native ecosystem.
If you want access to all CNCF-hosted co-located events, make sure you select an All-Access Pass when registering. The KubeCon + CloudNativeCon Only Pass does not include the CNCF-hosted co-located events, although it does include access to Monday’s Project Lightning Talks.
Tuesday: KubeCon + CloudNativeCon begins
The main conference gets underway on Tuesday, November 10 with keynotes, breakout sessions, the Solutions Showcase, Maintainer Track sessions, and opportunities to participate directly in project communities.
Maintainer Track sessions are an important part of the program. They give attendees opportunities to hear directly from the people working on Kubernetes and CNCF projects about current developments, technical challenges, roadmaps, and ways to get involved.
Tuesday’s schedule includes Maintainer Track sessions covering areas such as Kubernetes autoscaling, device management, security, authentication and authorization, developer experience, Helm, Cortex, Flux, Kyverno, KubeVirt, infrastructure, operational resilience, and more.
There are also ContribFest sessions throughout the conference. These are hands-on sessions where attendees can work alongside project maintainers and community members.
Tuesday includes ContribFests for Harbor, OpenTelemetry, Podman, and other communities, with some sessions explicitly welcoming new or first-time contributors.
If you’ve been looking for a way to move from using cloud native technology to contributing to it, this is a great place to start.
Tuesday evening also brings the community together for #KubeCrawl + #CloudNativeFest in the Solutions Showcase from 6:00–7:30 PM.
Wednesday: Learn from the ecosystem
Wednesday begins with keynotes before another full day of breakout sessions, Maintainer Tracks, ContribFests, and the Solutions Showcase.
Across the program, attendees can explore the breadth of challenges facing cloud native practitioners today.
That includes familiar areas such as Kubernetes operations, networking, security, observability, storage, scheduling, developer experience, multi-cluster management, and platform engineering, alongside a growing focus on AI workloads and infrastructure.
The schedule reflects how these areas increasingly intersect.
For example, Maintainer Track sessions explore topics including AI-aware networking in Kubernetes, infrastructure for AI workloads, GPU and other specialized device support, AI/ML workload orchestration, and how established cloud native projects are evolving as new AI use cases emerge.
At the same time, there is plenty for attendees focused on the core challenges of building and operating cloud native systems. Maintainer sessions cover projects and communities including OpenCost, Harbor, Vitess, Cilium, etcd, Linkerd, Karmada, Dragonfly, Kubernetes SIGs and many others.
This breadth is one of the reasons to build your agenda ahead of time: there are many different paths through KubeCon + CloudNativeCon depending on the technologies you use, the problems you are trying to solve, and the communities you want to meet.
Thursday: One more day to learn, and get involved
Thursday opens with another morning of keynotes, followed by breakout sessions, Maintainer Tracks, and ContribFests.
Maintainer Track sessions on Thursday cover communities and projects including Kubernetes SIG Storage, containerd, SPIFFE, Kubernetes WG Batch, Rook, SIG Multicluster, Open Policy Agent, minikube, Falco, and more.
There are also further opportunities for hands-on community participation. Thursday’s program includes ContribFest sessions for Argo CD and Meshery, among others.
That means there are opportunities right through the final day to meet maintainers, learn how projects work, share experiences, and contribute.
Why attend in person?
All breakout and keynote session recordings will be made publicly available on the CNCF YouTube channel after the event, and the published schedule states that recordings will be available within two weeks.
So why register and attend in Salt Lake City?
Because KubeCon + CloudNativeCon is about more than watching sessions.
Being there gives you the opportunity to participate in ContribFests, ask maintainers questions, meet contributors and end users, explore the Solutions Showcase, and have conversations with people facing many of the same technical and organizational challenges as you.
It’s the difference between watching a project update later and being in the room with its maintainers.
It’s the chance to hear an idea in a session and continue the conversation afterwards.
And it’s an opportunity to find a community, discover a project, or make a contribution that could continue long after the conference ends.
KubeCon + CloudNativeCon North America 2026 does not have a virtual component. While breakout and keynote recordings will be made available afterwards, the in-person community experiences happen in Salt Lake City.
AI meets cloud native
Another reason to pay attention to this year’s program is the growing intersection between cloud native and AI infrastructure.
The official event description itself highlights the future of both cloud native computing and AI infrastructure, and that connection can be seen throughout the schedule.
Monday includes both Agentics Day: MCP + Agents and Cloud Native AI + Inference Day. Elsewhere in the program, sessions explore topics such as AI-aware networking, GPU and specialized device management, AI/ML workload orchestration, agent observability, AI workloads on cloud native infrastructure, and the ways AI is affecting open source development and security.
But AI is one part of a much broader cloud native program.
Kubernetes, security, networking, observability, GitOps, platform engineering, storage, policy, developer experience, infrastructure, and open source project development all remain central to the week.
That combination makes Salt Lake City an opportunity to learn about what cloud native teams are working on today while also exploring the technologies and use cases that are beginning to shape what comes next.
Plan your week now
With four days of programming, it pays to plan ahead.
The online schedule lets you search sessions and speakers, filter the program, and build a personal agenda. Seating for sessions is first come, first served, so reviewing the schedule before arriving can help you make the most of your time.
When registering, it’s also worth considering which experience you want.
An All-Access Pass includes the CNCF-hosted co-located events on Monday, November 9, as well as the three days of the main conference, including keynotes, breakouts, and the Solutions Showcase.
A KubeCon + CloudNativeCon Only Pass includes the three main conference days and Monday’s Project Lightning Talks, but does not include access to the CNCF-hosted co-located events.
Whichever path you choose, KubeCon + CloudNativeCon North America is an opportunity to learn directly from the people building and operating cloud native technologies, discover what’s happening across the ecosystem, and participate in the communities moving cloud native forward.
Join us November 9–12, 2026 in Salt Lake City, Utah.
For most of the last decade our metrics pipeline ran on gostatsd, the open-source StatsD implementation we maintain. It primarily did two jobs: as sidecar on every host and the aggregation tier at the other end. It took metrics from roughly 100k hosts across 14 regions at a 99.95% SLO and minimal latency and it was fine. Nobody thought about it much, which is usually the sign of good infrastructure.
Gostatsd served us for many, many years. However, the community continuously and consistently converged on OpenTelemetry over the past few years. It became the thing everyone standardized on and more and more of what fed our pipeline was emitting OTel data we simply didn’t support. Gostatsd was UDP-only, had no story for traces or logs and every clever thing the OTel Collector community shipped was one more thing we’d eventually rebuild by hand just to stay level. We will lose that race. It’s only a question of when.
So the why was easy. The how is what we discuss here. With observability wired into thousands of services and many different bespoke platforms, the obvious plan (tear out the old pipeline, get every team to re-instrument on the OTel SDK, flip the switch) is a pipe dream: a multi-year org-wide slog on a pipeline that can’t take an outage with a real chance of dropping the exact data alerts fire on.
The question we actually needed to answer was narrower. How do we replace the whole engine without huge impact across Atlassian services?
The bet: swap the collection and pipeline, leave the interface alone
A metrics pipeline is really a contract with two ends. One end is what service owners see: “send StatsD over UDP to this address → your metrics show up in the backend.” The other is everything between that packet and long-term storage. Teams care enormously about the first end and very little about that middle layer. So we kept the interface and rebuilt everything behind it, which turned an org-wide migration into a platform-team migration.
Two things followed. We put purpose-built OTel Collector distributions at each of the four stages (collection, ingest, aggregation, forward) so we could work on any one without touching the others. And we made the collection-side speak StatsD and OTLP at the same time: nobody had to swap StatsD clients for the OTel SDK before we could start. It helped that the OTel Collector wasn’t new within Atlassian: the tracing team had run it as their pipeline core and host-metrics sidecar for years, so “is it production-ready at our scale?” was already answered.
What we actually built
The migration went step by step in place.
Collection. We replaced the gostatsd sidecar with our OTel Collector distribution, the same one the tracing team already shipped, keeping the app side identical: applications still fire StatsD over UDP as before. Day one, no team noticed. The payoff is that you stop running two sidecars (a StatsD one and a tracing one) on every host. Folding metrics into the tracing sidecar and killing the StatsD one saved about 3.9% CPU on average per service across our priciest Micros services, roughly a 30% cut in sidecar cost at fleet scale. At the same time, we enabled an OTLP receiver enabling our collection layer to receive and forward OTEL metrics natively.
Ingest. Aggregation is stateful: every datapoint for a time series has to hit the same aggregator, so you can’t just use traditional load-balancing strategies. For years an in-house proxy (called nomad) guaranteed that by hashing (service, environment) to a shard. But our service to metric load distribution is non-uniform and follows a long-tail, whichever shards owned the biggest services turned into hot shards. Our fix was a better routing strategy: the contrib loadbalancingexporter can hash by streamID (the identity of an individual time series) instead of by service, so one big service smears evenly across the pool while any given series still always lands on the same shard.
After this change: per-shard CPU went from a couple of tall bars and idle replicas to a flat, even distribution. Even load means a tighter autoscaling band, real off-peak scale-down and no more hot-shard pages!
Aggregation. This is the stage that makes the numbers affordable: we take in ~4.8 billion datapoints a minute and land ~220 million, roughly a 96% reduction. Most of our metrics are delta temporality and nothing upstream aggregated deltas the way our users expect, so we wrote our own delta aggregation processor and open-sourced it under atlassian-labs. Same traffic and the aggregation tier now runs on about half the CPU (the aggregators no longer parse gostatsd, load is even and we inherit the community’s tuning).
Forward. The last hop: a bespoke internal forwarder, became a stateless Collector distribution (metrics-gateway) built on upstream exporters. Fan-out to multiple backends (SignalFx, S3, etc) with no custom backend integrations; by default support for retries, queuing and backpressure from the community. Adding a destination is a config change, not a project. This was the easy one.
Lambda. Serverless can’t run a sidecar, so we built an OTel Lambda extension to replace the gostatsd one, with the same StatsD address, same env vars and no code changes. This completes the collection layer that forwards metrics from the services over to our pipeline.
Where this sits in the bigger picture
Getting every stage onto the OTel Collector gives us one codebase and one way of operating. Adding to the pipeline means writing a component, not standing up a service anymore. It unblocks OTEL-based instrumentation without us giving up the aggregation and cost controls that make our scale payable and it lets us drop wasteful datapoints at ingest, the cheapest place to do it. The end state is worth it on cost alone: the gostatsd aggregators and nomad together are ~38% of CPU requests in our metrics clusters and Nomad on its own is ~13% of total resources. Removing them is real money and the last thing between us and a pipeline that’s OpenTelemetry end to end.
What we tell ourselves before starting
Pick the right early adopters. Find the teams who have the most to gain and will iterate with you. Leading with dev and staging workloads and the services that felt the pain most gave us real signal fast, from people who were forgiving while we found the rough edges.
Profile continuously in production. The real cost and behaviour of a component, ours or upstream only showed up under production load. Small tests and benchmarks were not enough; continuous profiling in prod is what actually told us where to optimise.
Match operational workflows. A migration this size runs for months even years and for most of that you’re operating the old and new systems side by side. Keep the overhead of running two systems as low as you can: carry the same operational primitives across and keep parity so nobody has to learn a second way of doing things.
Progressive rollouts. Start in the lower environments and lead with the less critical services, then ramp 1% → 10% → 50% → 100%. You want to find problems where they’re cheap; not on the tier-0 path.
What’s next
We have now unblocked our users on moving metrics instrumentation over to OpenTelemetry. The next move is to shift left: get the instrumentation itself onto the OpenTelemetry SDK and off the vendor and in-house clients (Datadog/DogStatsD, StatsD libraries) we’ve carried for years.
We’re also going to start exploring further into the OpenTelemetry ecosystem to solve more large-scale Observability problems we have that the community has solved and also start contributing back as we grow OpenTelemetry with our usage and scale.
Managing infrastructure secrets on Kubernetes needs a backend that is self-healing and free of vendor lock-in, and that is exactly what OpenBao (the Linux Foundation’s open-source fork of HashiCorp Vault) and CloudNativePG give you: an entirely open-source stack built on two CNCF projects, Kubernetes, long since graduated, and CloudNativePG, a CNCF Sandbox project currently under evaluation for Incubation by the CNCF Technical Oversight Committee. OpenBao’s postgresql storage backend turns any PostgreSQL cluster into its encrypted key-value store, and CloudNativePG turns that cluster into a self-healing, synchronously replicated, certificate-authenticated Postgres instance with no cloud database dependency underneath it.
This recipe deploys a three-instance CNPG cluster as OpenBao’s storage backend and removes every password from the connection: the schema-owning role and the application role OpenBao itself uses both authenticate with a DatabaseRole-issued TLS client certificate, enforced by explicit pg_hba rules rather than by the absence of a password. pg_hba.conf is PostgreSQL’s client-authentication file, the thing that actually decides, per connection, whether a role needs a certificate, a password, or nothing at all.
Setting up a local test environment withcnpg-playground
Nothing about this recipe is specific to any one Kubernetes distribution: any conformant cluster with enough worker capacity will do. To follow along locally, though, the official cnpg-playground repository is the fastest path to one, since it is pre-configured with the CloudNativePG operator already. It is designed primarily around CNPG’s own demos, so it is worth knowing what it actually gives you: a single Kind cluster with six nodes, a control plane node, one node labelled for infrastructure workloads, one labelled for application workloads, and three carrying a node-role.kubernetes.io/postgres taint. That taint is exactly what our Cluster‘s tolerations in Step 1 target, and it is also what leaves OpenBao itself with only the two general-purpose nodes to schedule onto, which matters once pod anti-affinity enters the picture in Step 3. setup.sh provisions one Kind cluster per argument it is given, normally used to model separate regions; passing it a single, arbitrary label gives you one local cluster and skips the two-region disaster recovery demo entirely.
# Clone the CNPG Playground repository
git clone https://github.com/cloudnative-pg/cnpg-playground.git
cd cnpg-playground
# 1. Provision a single local cluster labelled "openbao"
./scripts/setup.sh openbao
# 2. Deploy CloudNativePG, cert-manager, the Barman Cloud plugin and a
# ClusterImageCatalog only, skipping the demo databases
REQUIREMENTS_ONLY=true ./demo/setup.sh
Architecture blueprint
Storage engine: OpenBao’s native postgresql storage backend, with ha_enabled = "true" for its HA lock table.
Database cluster: a 3-instance CNPG cluster with quorum-based synchronous replication (method: any, number: 1, the default dataDurability: required) for zero-data-loss failover.
Workload isolation: node selectors, tolerations and required zonal pod anti-affinity keep PostgreSQL on dedicated nodes across separate failure domains, following CNPG’s scheduling guidance.
Authentication: passwordless mTLS via the DatabaseRole CRD’s clientCertificate block, for both the schema owner and the application role, enforced by explicit pg_hba rules.
Step 1: deploy the CNPG cluster, roles and database
The Cluster below points imageCatalogRef at the postgresql-minimal-trixie ClusterImageCatalog that REQUIREMENTS_ONLY=true ./demo/setup.sh already deployed in the previous step, rather than pinning an image tag directly: CNPG resolves it to the latest minimal PostgreSQL 18 image in that catalog, so a kubectl apply against the same manifest keeps picking up new patch releases as the catalog is updated, no Cluster edit required. It also declares synchronous replication, workload isolation, and the two pg_hba rules that force certificate authentication for both roles OpenBao will use. Two DatabaseRole objects follow: role-openbao, the schema owner used once to run DDL, and role-openbao-rw, the restricted role OpenBao itself connects as at runtime. Both get a clientCertificate, because a one-shot DDL job is no more entitled to a password lying around than the application is.
{{apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: openbao-db
namespace: openbao
spec:
instances: 3
# Tracks the latest minimal PostgreSQL 18 image via the ClusterImageCatalog
# the playground's REQUIREMENTS_ONLY step already deploys.
# See https://cloudnative-pg.io/docs/current/image_catalog
imageCatalogRef:
apiGroup: postgresql.cnpg.io
kind: ClusterImageCatalog
name: postgresql-minimal-trixie
major: 18
# See https://cloudnative-pg.io/docs/current/scheduling
affinity:
nodeSelector:
node-role.kubernetes.io/postgres: ""
tolerations:
- key: node-role.kubernetes.io/postgres
operator: Exists
effect: NoSchedule
enablePodAntiAffinity: true
topologyKey: topology.kubernetes.io/zone
podAntiAffinityType: required
postgresql:
# Synchronous replication: dataDurability defaults to "required", giving
# RPO=0 at the cost of pausing writes if no standby is available.
# See https://cloudnative-pg.io/docs/current/replication
synchronous:
method: any
number: 1
# The operator does not add cert rules for DatabaseRole client
# certificates automatically: without these, "openbao" and "openbao-rw"
# would fall through to the default scram-sha-256 rule, and since
# neither role has a passwordSecret, every connection would simply fail.
pg_hba:
- hostssl openbao openbao all cert
- hostssl openbao openbao-rw all cert
- hostnossl openbao openbao all reject
- hostnossl openbao openbao-rw all reject
# See https://cloudnative-pg.io/docs/current/postgresql_conf
parameters:
max_connections: '100'
log_checkpoints: 'on'
log_lock_waits: 'on'
hot_standby_feedback: 'on'
shared_memory_type: 'sysv'
dynamic_shared_memory_type: 'sysv'
storage:
size: 10Gi
---
apiVersion: postgresql.cnpg.io/v1
kind: DatabaseRole
metadata:
name: role-openbao
namespace: openbao
spec:
cluster:
name: openbao-db
name: openbao
login: true
clientCertificate:
enabled: true
databaseRoleReclaimPolicy: retain
---
apiVersion: postgresql.cnpg.io/v1
kind: DatabaseRole
metadata:
name: role-openbao-rw
namespace: openbao
spec:
cluster:
name: openbao-db
name: openbao-rw
login: true
clientCertificate:
enabled: true
databaseRoleReclaimPolicy: retain
---
apiVersion: postgresql.cnpg.io/v1
kind: Database
metadata:
name: openbao-db
namespace: openbao
spec:
name: openbao
owner: openbao
cluster:
name: openbao-db
}}
Apply these resources:
kubectl create namespace openbao
kubectl apply -f cnpg-stack.yaml
Watch for all three instance pods to come up, which takes a couple of minutes on a fresh cluster:
kubectl get pods -w -n openbao
Once all three are Running and Ready, confirm the cluster itself has reached a healthy state:
kubectl cnpg -n openbao status openbao-db
Cluster Summary
Name openbao/openbao-db
System ID: 7674399793927340061
PostgreSQL Image: ghcr.io/cloudnative-pg/postgresql:18.6-202608131513-minimal-trixie@sha256:e488b1434919f455f2ee4e18a181ce9b33f34cdd8dfb821126855486bce6ad34
Primary instance: openbao-db-1
Primary promotion time: 2026-08-15 23:10:49 +0000 UTC (3m15s)
Status: Cluster in healthy state
Instances: 3
Ready instances: 3
Size: 135M
Current Write LSN: 0/6000060 (Timeline: 1 - WAL File: 000000010000000000000006)
Continuous Backup not configured
Streaming Replication status
Replication Slots Enabled
Name Sent LSN Write LSN Flush LSN Replay LSN Write Lag Flush Lag Replay Lag State Sync State Sync Priority Replication Slot
---- -------- --------- --------- ---------- --------- --------- ---------- ----- ---------- ------------- ----------------
openbao-db-2 0/6000060 0/6000060 0/6000060 0/6000060 00:00:00 00:00:00 00:00:00 streaming quorum 1 active
openbao-db-3 0/6000060 0/6000060 0/6000060 0/6000060 00:00:00 00:00:00 00:00:00 streaming quorum 1 active
Instances status
Name Current LSN Replication role Status QoS Manager Version Node
---- ----------- ---------------- ------ --- --------------- ----
openbao-db-1 0/6000060 Primary OK BestEffort 1.30.0 k8s-openbao-worker3
openbao-db-2 0/6000060 Standby (sync) OK BestEffort 1.30.0 k8s-openbao-worker4
openbao-db-3 0/6000060 Standby (sync) OK BestEffort 1.30.0 k8s-openbao-worker5
Note the PostgreSQL Image line: a SHA-pinned, dated minimal build resolved straight out of the postgresql-minimal-trixie catalog, not a floating tag we wrote by hand.
Both standbys show up as Standby (sync) with a Sync State of quorum at the same time, which is exactly the dynamic behaviour method: any is meant to give: with number: 1, either standby satisfies durability, and CNPG does not pin a fixed “the” synchronous standby.
Once reconciled, the operator has created two client certificate secrets, role-openbao-client-cert and role-openbao-rw-client-cert, following its <databaserole-name>-client-cert naming convention. openbao, as the database owner, already has CREATE on the public schema by default (PostgreSQL grants that to the owner even though it revoked it from PUBLIC in v15), so no extra schema grant is needed before the DDL step.
Every manifest that mounts one of these secrets sets defaultMode: 0640 on the volume. Kubernetes mounts Secret volumes at 0644 by default, which libpq refuses outright: it rejects a private key file that is group-or-world-readable, whether owned by root (0640 or less) or by the connecting user (0600 or less). Since the mounted files stay root-owned and only their group matches the pod’s fsGroup, 0640 is the setting that satisfies libpq here, and it applies to every pod in this recipe that reads a client certificate, the schema-init Job and the OpenBao pods alike.
Step 2: initialise the schema and grant table privileges
DatabaseRole does not yet manage table-level grants: the permissions stanza that would let a Database object express GRANT/REVOKE declaratively is still an open proposal (#10826), as I covered when DatabaseRole first shipped in Recipe 25. Until that lands, a one-time Job running the DDL as the schema owner is the correct way to create OpenBao’s tables and grant the restricted DML the openbao-rw role actually needs.
OpenBao’s postgresql storage backend expects two tables when ha_enabled = "true": openbao_kv_store, with a parent_path, path, key and value column and a primary key on (path, key), and openbao_ha_locks, holding its HA lock records. Getting the key column or the primary key wrong here is an easy mistake, since OpenBao would otherwise silently create the table itself on first connection using its own DDL, and that path only works if the connecting role already has CREATE, which openbao-rw deliberately does not. Pre-creating both tables under the owner role and setting skip_create_table on the OpenBao side (Step 3) keeps that DDL entirely off the restricted runtime role.
The same job also closes a gap PostgreSQL leaves open by default: every database grants CONNECT to PUBLIC, and the public schema grants USAGE to PUBLIC too, so any role that can log into the cluster at all can connect to openbao and see what is in its public schema unless told otherwise. Making REVOKE CONNECT … FROM PUBLIC the default posture across every database CloudNativePG manages is on the roadmap (#10831), but it is not there yet, so the schema-init job revokes it explicitly here and grants back only what openbao-rw actually needs:
{{apiVersion: batch/v1
kind: Job
metadata:
name: openbao-schema-init
namespace: openbao
spec:
ttlSecondsAfterFinished: 300 # Clean up job 5 minutes post-completion
template:
metadata:
name: openbao-schema-init
spec:
restartPolicy: OnFailure
securityContext:
runAsNonRoot: true
runAsUser: 26
fsGroup: 26
seccompProfile:
type: RuntimeDefault
containers:
- name: psql-init
image: ghcr.io/cloudnative-pg/postgresql:18-minimal-trixie
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop:
- ALL
resources:
requests:
cpu: "100m"
memory: "64Mi"
limits:
cpu: "500m"
memory: "128Mi"
command:
- psql
- "postgres://openbao@openbao-db-rw:5432/openbao?sslmode=verify-full&sslcert=/etc/certs/tls.crt&sslkey=/etc/certs/tls.key&sslrootcert=/etc/ca/ca.crt"
- -c
- |
CREATE TABLE IF NOT EXISTS openbao_kv_store (
parent_path TEXT NOT NULL,
path TEXT NOT NULL,
key TEXT NOT NULL,
value BYTEA,
CONSTRAINT openbao_kv_store_pkey PRIMARY KEY (path, key)
);
CREATE INDEX IF NOT EXISTS openbao_kv_store_idx
ON openbao_kv_store (parent_path);
CREATE TABLE IF NOT EXISTS openbao_ha_locks (
ha_key TEXT NOT NULL,
ha_identity TEXT NOT NULL,
ha_value TEXT,
valid_until TIMESTAMP WITH TIME ZONE NOT NULL,
CONSTRAINT openbao_ha_locks_pkey PRIMARY KEY (ha_key)
);
GRANT SELECT, INSERT, UPDATE, DELETE
ON TABLE openbao_kv_store, openbao_ha_locks
TO "openbao-rw";
REVOKE CONNECT ON DATABASE openbao FROM PUBLIC;
GRANT CONNECT ON DATABASE openbao TO "openbao-rw";
REVOKE ALL ON SCHEMA public FROM PUBLIC;
GRANT USAGE ON SCHEMA public TO "openbao-rw";
volumeMounts:
- name: certs
mountPath: /etc/certs
readOnly: true
- name: ca
mountPath: /etc/ca
readOnly: true
volumes:
- name: certs
secret:
secretName: role-openbao-client-cert
defaultMode: 0640
- name: ca
secret:
secretName: openbao-db-ca
defaultMode: 0640
}}
Apply the job:
kubectl apply -f schema-init-job.yaml
Wait for its pod to finish ContainerCreating and complete before reading its logs, otherwise kubectl logs fails outright rather than waiting:
kubectl wait --for=condition=complete -n openbao job/openbao-schema-init --timeout=60s
kubectl logs -n openbao job/openbao-schema-init
CREATE TABLE
CREATE INDEX
CREATE TABLE
GRANT
REVOKE
GRANT
REVOKE
GRANT
Eight statements in, eight confirmations out: both tables, the DML grant, and the three REVOKE/GRANT pairs that lock the database and the public schema down to openbao-rw.
Step 3: configure and deploy OpenBao via Helm
Configure the official OpenBao Helm chart. Mount the role-openbao-rw-client-cert secret into OpenBao and point the postgresql storage stanza at it with sslmode=verify-full. A few details that are easy to miss from the OpenBao side: skip_create_table must be set explicitly, since openbao-rw has no CREATE privilege and would otherwise fail on first connection when OpenBao tries to create the tables itself, and server.dataStorage needs disabling, since it defaults to a 10Gi PVC per pod that would otherwise sit there unused: the whole point of this stack is that OpenBao carries no local state at all. Mounting the certificate and CA secrets also needs the right chart field: server.extraVolumes looks like the obvious choice, but it uses its own simplified type/name/path schema rather than a raw Kubernetes volume, and there is no matching extraVolumeMounts field for the server StatefulSet at all. server.volumes and server.volumeMounts are the fields that pass straight through to the Pod spec, and are what the manifest below actually uses.
{{global:
enabled: true
server:
# No local persistence: all state lives in CNPG.
dataStorage:
enabled: false
ha:
enabled: true
replicas: 3
config: |
ui = true
listener "tcp" {
tls_disable = 1
address = "[::]:8200"
cluster_address = "[::]:8201"
}
storage "postgresql" {
connection_url = "postgres://openbao-rw@openbao-db-rw:5432/openbao?sslmode=verify-full&sslcert=/etc/openbao/certs/tls.crt&sslkey=/etc/openbao/certs/tls.key&sslrootcert=/etc/openbao/ca/ca.crt"
table = "openbao_kv_store"
ha_table = "openbao_ha_locks"
ha_enabled = "true"
skip_create_table = "true"
}
# cnpg-playground only: with the Postgres nodes tainted and off limits,
# only two general-purpose nodes are left, one short of what three
# required-anti-affinity replicas need. Tolerating the control-plane
# taint gives OpenBao a third node to land on. Drop this in a cluster
# with three or more untainted worker nodes, and never carry it into
# production: workloads should not run on the control plane there.
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
# Mount the openbao-rw DatabaseRole's client certificate and the
# cluster's client CA. "volumes"/"volumeMounts" are passed through to the
# Pod spec as-is; the chart's own "extraVolumes" field uses a different,
# simplified schema (type/name/path) that does not accept a raw Secret
# volume, and there is no "extraVolumeMounts" field for the server
# StatefulSet at all.
volumes:
- name: cnpg-client-cert
secret:
secretName: role-openbao-rw-client-cert
defaultMode: 0640
- name: cnpg-client-ca
secret:
secretName: openbao-db-ca
defaultMode: 0640
volumeMounts:
- name: cnpg-client-cert
mountPath: /etc/openbao/certs
readOnly: true
- name: cnpg-client-ca
mountPath: /etc/openbao/ca
readOnly: true
}}
Install OpenBao following the official OpenBao Kubernetes documentation:
helm repo add openbao https://openbao.github.io/openbao-helm
helm repo update
helm install openbao openbao/openbao \
--namespace openbao \
-f openbao-values.yaml
Pod anti-affinity for the OpenBao replicas is not something this recipe has to configure: the chart’s server.affinity default already renders a requiredDuringSchedulingIgnoredDuringExecution rule keyed on kubernetes.io/hostname, so the three server pods refuse to land on the same node. In cnpg-playground specifically, that is worth double checking rather than assuming: with the Postgres nodes tainted and off limits, only the infra- and app-labelled nodes are left for general workloads, one short of what three required-anti-affinity replicas need. The manifest above adds a server.tolerations entry for the node-role.kubernetes.io/control-plane taint so OpenBao can use that node as its third, which is a reasonable thing to do on a single-developer Kind cluster and not something to carry into a real cluster, where the control plane should stay clear of ordinary workloads. On a cluster with three or more untainted worker nodes, the toleration is unnecessary and the chart’s default anti-affinity just works on its own.
Only openbao-0 and openbao-1 exist so far, and openbao-1 shows 0/1: the StatefulSet's default OrderedReady policy will not even create openbao-2 until openbao-1 reports Ready, and readiness here is the unseal status, not just the process being up. That is exactly why the initialisation and unsealing below has to happen pod by pod, in order: openbao-0 first, then openbao-1, then openbao-2, each one only created once its predecessor is unsealed.
Initialise and unseal OpenBao
Initialise the cluster on openbao-0 and unseal all three pods:
kubectl exec -it -n openbao openbao-0 — bao operator init
Store the generated unseal keys and root token securely.
Each pod’s Shamir state is independent and in memory: unsealing openbao-0 does nothing for openbao-1 or openbao-2, so the same three keys have to be submitted again, against each pod by name, one pod at a time:
Until a pod’s three keys go in, kubectl describe pod on it shows Unseal Progress: 0/3 and a stream of Warning Unhealthy readiness-probe events. Neither is a problem: it is the probe correctly reporting that the pod is still sealed, and it clears as soon as that pod gets its keys.
The third key flips Sealed to false:
Key Value
--- -----
Seal Type shamir
Initialized true
Sealed false
Total Shares 5
Threshold 3
Version 2.6.1
Commit Date 2026-07-22T14:22:33Z
Storage Type postgresql
Cluster Name vault-cluster-97b6aa67
Cluster ID 634fcfcf-09f2-b719-f9bf-1f684cfbe88a
HA Enabled true
HA Cluster n/a
HA Mode standby
Storage Type: postgresql here is the whole point of this recipe, and it is not just reporting the config back: OpenBao writes its own bootstrap state, the keyring, the root key material, its seal configuration, before you ever create an application secret. Query the cluster directly and it is already there:
Every one of those rows lives in the openbao_kv_store table this recipe’s schema-init Job created, written by the openbao-rw role with nothing but a client certificate.
Read/write test
Confirm OpenBao can write an encrypted payload through to CNPG using the restricted openbao-rw role:
# Login
kubectl exec -it -n openbao openbao-0 -- bao login <ROOT_TOKEN>
# Enable KV v2 and write a secret
kubectl exec -it -n openbao openbao-0 -- bao secrets enable -path=secret kv-v2
kubectl exec -it -n openbao openbao-0 -- bao kv put secret/test-app username="admin" password="supersecretpassword"
# Read it back
kubectl exec -it -n openbao openbao-0 -- bao kv get secret/test-app
kubectl exec -it -n openbao openbao-0 -- bao kv get secret/test-app
Run the same SELECT key FROM openbao_kv_store; query again and a new row for secret/test-app shows up alongside the bootstrap keys, its value column holding the payload as an encrypted BYTEA blob, never plaintext, even to someone with direct PostgreSQL access to the table.
Now that all three pods are unsealed, -o wide on the whole namespace shows the full picture: the three CNPG instances on their three dedicated, tainted Postgres nodes, and the three OpenBao replicas spread across the control-plane node and the two general-purpose ones, required anti-affinity satisfied without a single pod colocated with another:
All seven workloads nicely distributed across the six available nodes on the playground, control plane included, exactly as intended.
Operational notes: certificate renewal
CNPG’s client certificates carry a 90-day validity period and are renewed automatically about a week before expiry, the same schedule the operator already applies to the streaming_replica certificate (see Certificates). Renewal replaces the contents of the role-openbao-rw-client-cert secret in place, no manifest change required. Both figures are inherited unconditionally from the operator’s own global settings today; a proposal to let clientCertificate override duration and renewBefore per DatabaseRole, mirroring cert-manager’s convention, is open in issue #11312, useful if a role like role-openbao-rw ever needs a renewal cadence different from the cluster-wide default.
OpenBao itself does not pick that renewal up on its own. Like Vault before it, OpenBao’s postgresql storage backend opens its connection pool once at process startup and never re-reads the certificate files afterwards, so a renewed certificate only takes effect after a rolling restart of the OpenBao pods. This is a property of the storage plugin’s own connection lifecycle, not something CNPG’s certificate reconciliation controls: CNPG’s job ends at keeping the secret current, and nothing on the CNPG side requires a restart. Until OpenBao’s storage backend gains a way to reload its connection pool’s TLS material on signal, budget for a scheduled rolling restart inside the 83-day renewal window, well before the old certificate actually expires.
Beyond this setup: backups and disaster recovery
This is a single-cluster deployment. High availability inside the Kubernetes cluster is covered by the three OpenBao replicas, CNPG’s synchronous replication and required zonal anti-affinity, but production workloads need more than that:
Automated PostgreSQL backups: the in-tree .spec.backup.barmanObjectStore stanza is deprecated as of CNPG 1.26 in favour of the Barman Cloud Plugin: define an ObjectStore pointing at AWS S3, Google Cloud Storage or Azure Blob Storage, reference it from the Cluster's .spec.plugins, and back it with Backup/ScheduledBackup resources using method: plugin for continuous WAL archiving and scheduled base backups, enabling Point-In-Time Recovery.
Disaster recovery: to meet real RTO/RPO targets, span the deployment across more than one Kubernetes cluster. CNPG’s distributed topology lets an asynchronous replica cluster in a second region or cluster promote to primary if the first one is lost entirely, the same capability I discussed for cloud-neutral portability.
Conclusion
Combining OpenBao with CloudNativePG gives you a fully open-source, enterprise-grade secrets engine running on native Kubernetes CRDs, with the DatabaseRole CRD covering every role in the stack rather than just the application-facing one. Synchronous replication gives zero-data-loss failover, explicit pg_hba rules turn the absence of a password into an actual enforced policy rather than an assumption, and dedicated scheduling rules keep the database’s failure domains separate from everything else running in the cluster.
If you have ideas, questions, or run into something this recipe does not cover, reach out to Rob and Gabrielle directly on LinkedIn, or find us on the CloudNativePG community’s CNCF Slack.
On July 18, 2026, we held the third edition of Kubernetes Community Days Lima at UTEC in Barranco. By now, we have already sent the Transparency Report to the CNCF, thanked our sponsors, and processed the surveys. With the administrative wrap-up behind us, I want to share a few reflections on what it meant to organize it.
Before I continue, a quick introduction of me. I’m Ronald Requena, CTO at Rumbo, professor at Universidad Ricardo Palma, co-organizer of KCD Lima and DevOpsDays Lima, and, as of August 2026, CNCF Ambassador for Peru. Feel free to connect with me or reach out on LinkedIn.
This is the third KCD I have had the opportunity to co-organize. I assumed that by the third edition many things would run on autopilot. They did not. Every edition brings new problems, and the old ones do not go away; they just change in scale.
An image I will not forget
I arrived at UTEC at 7:00 a.m., an hour and a half before the scheduled 8:30 opening. There was already a line of around 100 attendees outside the building. A Saturday morning, coffee in hand, waiting to get into the long-awaited Kubernetes Community Days Lima. We had to move registration forward to 8:00 because the line kept growing.
No survey can measure the kind of welcome we had that Saturday, from early risers eager to be the first ones through the door.
The numbers
I find it hard to write without data, so let me start there:
2,244 registrations and more than 900 attendees in a single day. Roughly 75% more attendance than in 2025 and the largest of our three editions.
60 speakers and 54 sessions across 5 parallel rooms, selected from 89 proposals submitted by 44 companies from 14 countries.
5 international keynotes: Jeffrey Sica, Viktor Farcic, Lin Sun, Mauricio Salatino and Daniel Oh.
11 sponsors across five tiers: Testkube, Red Hat, Valkey, and Interbank as Diamond; VMware by Broadcom as Platinum; SUSE | V LATAM, Upwind, and Orca Security as Gold; BCP and Crubyt as Silver; and Rumbo as Bronze.
4.7 out of 5 satisfaction in the post-event survey. 98% rated the event 4 or 5 stars, and no one rated it below 3.
That last number is the one that matters most. The others speak to size; this one speaks to whether it was truly worth it for the people who came. Every number here is public: each KCD organizing team submits a transparency report to the CNCF, and ours is no exception
Three editions, three stages
Looking back, each edition had a different goal, even if we never framed it that way at the time.
2024 was about proving that a KCD could happen in Peru. We had no track record to show sponsors and no idea how many people would show up. We closed with more than 500 attendees, 49 speakers, and 11 sponsors. That first edition was organized with a lot of will and very little certainty.
2025 was about proving it had not been a fluke. The second year is the hardest: you are no longer a novelty, but you are not yet a tradition. We reached 1,377 registrations and 520 attendees, with 59 speakers across 7 rooms. That year we learned to document: the sponsor guide, the contracts, the run-of-show, the cash flow. What we improvised in 2024 had to be written down in 2025 so it would not depend on anyone’s memory.
2026 was about proving the event now has a life of its own. We nearly doubled attendance, sponsors came back and new ones joined, and international speakers said yes because they already knew KCD Lima. This year I understood that the community no longer needs to be convinced to attend. What it needs is for us to live up to the expectations it created for itself.
With each edition, what the event represents to you changes. The first time, it was a personal challenge. The second, a commitment. The third starts to become a responsibility toward people you do not know: the student who came alone, the professional who carved out a Saturday to be there, the speaker who flew ten hours for a thirty-minute talk.
One thing does not change: the community cannot be delegated. You can delegate logistics, design, or the website. But deciding what kind of event we want to be, who is welcome, and what conversation we want to happen in Lima, that has to be done every year, deliberately.
What goes unseen
When you think about organizing an event, you picture the agenda, the speakers, and the venue. All of that exists. But a good part of the time goes into things that never make it into a photo.
This year we completed a variety of forms for US-based sponsors and went through several vendor onboarding processes, each with its own questionnaires and validations. We reviewed sponsorship agreement clauses with companies: liability, sector exclusivity, refund conditions.
We built a budget with cash flow in both soles and dollars, because revenue arrives in one currency and expenses go out in another. We drew and designed the event layout: five rooms, ten booths, four coffee stations, stairs, and elevators.
I am not saying this as a complaint. I am saying it because there is a notion that community “just happens” and it does not. Someone has to do the part that no one sees. Many KCD organizers around the world will recognize this immediately. If you have never organized one, you probably will not — and that is okay too. It is not meant to be seen.
What connects with my work
Three data points from the attendee profile made me think about what I do outside the event.
45% came from end-user companies, and within that group, banking and finance was the largest sector, accounting for a third of all attendees: BCP, Interbank, Scotiabank, Compartamos, Pacífico, and others. I spent several years in banking leading engineering and architecture teams, and seeing that industry fill a cloud native auditorium in Lima confirms that digital transformation in Peruvian banking is no longer a slide deck. These are real platform teams, with DevOps engineers and SREs who need a community to learn from.
15% were architects, the most common role, ahead of developers and DevOps. That tells me cloud native adoption in Peru is at a stage where design decisions carry as much weight as implementation. It is the conversation I have almost daily in consulting.
8% of attendees came from Universities. I teach at a university, and I know how hard it is for a student to feel that an event like this is meant for them. That is why KCD Lima is free and we reserve space for students: the barrier cannot be the price of a ticket. That people registered as students and showed up on a Saturday is one of the things I value most about this edition.
What should be done differently
The survey was positive, but it was also clear. Four things remain open:
The coffee break was not enough. We sized four stations for 600 people and 900+ showed up. Next time, we plan with a margin above expected capacity.
Punctuality between sessions. With five parallel rooms, a five-minute delay in one throws the other four off. We need room moderators with more control over the clock.
Session pre-registration. There were packed rooms and half-empty rooms for talks of equal quality. If we know demand in advance, we can allocate space better.
More women on stage. Out of 60 speakers, only 5 were women. Among attendees, female participation was 17%, so the gap on stage is wider than in the audience. It is a number we would rather publish than omit, because it is the starting point for doing something about it in a future CFP.
Thank you
To Josua Castro, who led the organization of this edition, and to the team I shared these months with: Andreé Cordero, Almendra Paz, Bianca Torres, Enzo Venturi, Jean Paul López, Pavel Puclla, and Michelle Luna. Nine people, all with full-time jobs, holding up an event with more than 900 attendees. What people saw on July 18 was the result of that team.And to our volunteers, who ran registration, guided attendees across five rooms, kept sessions on time, and solved a hundred small problems no one else noticed. An event of this size does not work without them.
To Eric Biagioli and Cindy Archenti, for helping us with the unexpected and everything that was not in the plan. And to all of UTEC for opening the entire campus to us. To the CNCF for backing the KCD program. To the 60 speakers who prepared, traveled, and gave up their Saturday. To the sponsors who bet on Lima before seeing the numbers, and who make it possible for an event of this size to be free for the community.
And to the more than 900 people who showed up on a Saturday, from 7:00 a.m. to 6:00 p.m.
Three editions later, in August, I was accepted into the CNCF Ambassadors program. It is an honor to join that global community, and I take it as motivation to keep working so that Lima keeps growing as a cloud native hub in the region.
What comes next, I do not know yet. But that 7 a.m. line is something I will not forget.
Cilium 1.20, the second major open source Cilium release of 2026 after Cilium 1.19, is finally here. Three themes stand out in this release:
Gateway API becomes a much broader traffic management layer. Cilium jumps from Gateway API v1.4 to v1.6 and adds ExternalAuth, CORS filters, ListenerSets, and TCPRoute and UDPRoute for non-HTTP traffic. Platform and security teams can manage more of their north-south traffic through the same API, with authentication and authorization handled before a request reaches the application. If you are still running Ingress NGINX, now is the time to let the CNI you already run take on that traffic management too.
Cilium becomes a platform that cloud providers can extend. Cilium is already a common networking foundation across hyperscalers and neoclouds. Datapath plugins, developed by Google, make Cilium less like a sealed networking appliance and more like a network operating system: a stable core that cloud providers can extend with their own eBPF programs, independently of the Cilium release cycle.
Innovation and standardization move forward together. Cilium continues to push the datapath with innovations such as netkit, while aligning more closely with the wider Kubernetes ecosystem through stable MCS API support and Kubernetes ClusterNetworkPolicy. As capabilities Cilium supported early, including multi-cluster services and cluster-wide policy, take shape as upstream Kubernetes APIs, Cilium supports those APIs alongside its own established CRDs. Users get a choice between Cilium-native features and portable Kubernetes resources.
choice between Cilium-native features and portable Kubernetes resources.
Thank you to every contributor, reviewer and maintainer who made Cilium 1.20 possible, including engineers from Datadog, Google, Microsoft and many other organisations across the Cilium community.
Networking
ENI IPAM mode with IPv6
One of the final gaps that stopped us from claiming feature parity across IPv4 and IPv6 in Cilium is gone. It even makes a footnote in Chapter 4 of Cilium: Up and Running redundant, and while we are a little sad that a note in the book is obsolete so soon after publication, we are glad that ENI mode on AWS finally supports IPv6, four years after the feature request was first raised.
Let’s recap. Up until now, running Cilium in AWS ENI IPAM mode (the AWS-specific mode that gives pods real VPC addresses) was only possible on IPv4. Now, as a beta feature in Cilium 1.20, ENI IPAM can allocate and hand out IPv6 as well. The operator attaches an IPv6 /80 prefix to each node’s ENI through Prefix Delegation (which we covered in detail in an earlier blog on how to overcome IP address exhaustion in Kubernetes), and the agent assigns pod addresses from it.
Enable ENI IPAM and IPv6 together:
ipam:
mode: eni
eni:
enabled: true
ipv6:
enabled: true
Pods on EKS then boot up dual-stack with VPC-routable IPv6:
NAME IPS
dualstack-demo-7f8876746b-5xvnk 192.168.128.71,2a05:d01c:38d:4c02:909e::e771
dualstack-demo-7f8876746b-8lj6z 192.168.177.118,2a05:d01c:38d:4c00:1cf5::57d1
dualstack-demo-7f8876746b-p84p7 192.168.169.190,2a05:d01c:38d:4c00:1cf5::f7ea
And here is the other side of it in the AWS console: the /80 IPv6 prefix delegated to a worker node’s ENI. The IPv6 address from the first pod above comes straight out of this 2a05:d01c:38d:4c02:909e::/80 range.
Kudos to the folks at Datadog, one of the most impactful organizations contributing to Cilium for years now, for building ENI IPv6 support and finally closing that gap. Watch the video for a short demo.
If you’d like to learn more about this feature, check the documentation for more information.
The TL;DR on netkit is this: it replaces the traditional veth pair and its overhead, bringing pod networking performance to host-level throughput.
When it was initially released in Cilium, it needed a recent kernel (>= 6.8), so on a fleet with mixed kernel versions you either gave up on netkit everywhere or carved your nodes into pools by kernel version and configured them separately.
Cilium 1.20 adds bpf.datapathMode=auto: each agent probes its own host at startup, uses netkit when the kernel supports it, and quietly falls back to veth when it does not. That means you only need to configure one setting across a mixed fleet, making netkit easier to adopt across your environment. You can see which mode each node actually landed on through cilium-dbg status and the cilium_feature_datapath_config metric:
Previously, extending or instrumenting Cilium’s eBPF datapath beyond what the project already supported left you with two hard options: upstream your change, or maintain a fork. Both are high-effort, and a fork means carrying datapath patches across every Cilium upgrade.
While most end-users just use the off-the-shelf Cilium, some cloud providers – like Google, the authors of this new feature – would manage their own Cilium fork, with the maintenance efforts it comes up with.
Now, Cilium 1.20 introduces datapath plugins (beta): third-party code can instrument Cilium’s eBPF datapath as its own plugin, running as a separate process that Cilium reaches out to, without patching or forking Cilium itself. Because a plugin is its own image, it is built, versioned and rolled out on its own cadence and can be upgraded independently of Cilium. And because it runs on its own, a plugin crash does not take the agent down with it. It opens the door to specialised observability, security and networking extensions that ship and evolve independently of the Cilium release cycle, with no fork required.
A plugin registers itself with Cilium through a deliberately small CiliumDatapathPlugin custom resource:
apiVersion: cilium.io/v2alpha1
kind: CiliumDatapathPlugin
metadata:
name: example
spec:
# Fail closed if Cilium cannot reach the plugin.
attachmentPolicy: Always
# Bump this to reinitialise the datapath for a new plugin release.
version: 0.0.
If you’d like to learn more about this feature, check the documentation for more information.
Per-pod disable source IP verification
By default, Cilium enforces source IP verification: a pod may only send packets whose source IP matches its own. This stops a compromised or misbehaving pod from spoofing another workload’s address. There are – admittedly rare – circumstances where some workloads legitimately need to emit packets with a different source IP: for example, NAT gateways, firewalls and other appliance-style pods (such as VPN/Tailscale) .
Until now, source IP verification was a single cluster-wide switch (enable-source-ip-verification, on by default) with no per-pod control – so letting one workload use a foreign source IP meant turning it off for the entire cluster, weakening anti-spoofing everywhere.
Cilium 1.20 makes it a per-pod, opt-in decision, gated at two levels. First, a cluster administrator permits it for a namespace by annotating the namespace with config.cilium.io/delegate-source-ip-verification: “true”. Only then can a pod owner opt a specific workload out of source IP verification by annotating it with config.cilium.io/disable-source-ip-verification: “true”. Both annotations are required, so a namespace owner cannot quietly turn anti-spoofing off on their own, and the protection stays on everywhere else. This feature solves a problem for the workloads that need it, but make sure to use Kubernetes RBAC to restrict who can modify namespace annotations, so you don’t accidentally reduce your security posture.
If you’d like to learn more about this feature, check the documentation for more information.
Migrate cluster-pool IPAM to multi-pool IPAM
Multi-pool IPAM is steadily becoming the preferred way to hand out pod IPs in Cilium. Instead of drawing every address from one flat cluster-wide range, you carve the space into named pools and assign them where you want them – per namespace, per workload, per tenant. What had been difficult until now was the migration from the default cluster scope mode. Until now, if your cluster was already running cluster-scope IPAM, moving to multi-pool meant standing up a fresh cluster and migrating workloads across, which is not something anyone wants to do on a live platform.
Cilium 1.20 adds an in-place migration path. You can switch an existing cluster from cluster-pool to multi-pool, with no rebuild and no re-IP’ing, by using the operator option enable-cluster-pool-to-multi-pool-migration. Check out the migration documentation for more info.
Traffic distribution PreferSameZone / PreferSameNode
Introduced in Cilium 1.16, Service Traffic Distribution became the preferred method for keeping Service traffic close to its client, replacing the older topology-aware hints method. Service Traffic Distribution enables users to reduce latency and avoid cross-zone data-transfer charges, as Dean covered in this excellent post.
Until now, Cilium only supported the PreferClose value, indicating a preference for routing traffic to endpoints that are topologically proximate to the client.
In Cilium 1.20, Cilium honours recent changes to the upstream Kubernetes trafficDistribution specification and now supports two newer values. PreferSameZone keeps traffic within the client’s zone while a healthy backend is available there, and PreferSameNode prefers a backend on the very same node.
Both values gracefully fall back to the rest of the cluster’s backends when no local one is healthy, so you get the locality benefit without trading away availability. Because this is the standard Kubernetes field rather than a Cilium-specific annotation, the same Service spec behaves consistently across conformant implementations.
Check the official Kubernetes documentation for more information on service traffic distribution.
EndpointSlice weights for Maglev backends
Cilium has supported Maglev as its load balancer hashing mechanism since Cilium 1.10. Originated at Google and named after the Japanese magnetic levitation train, Maglev provides consistent hashing to minimize disruption during node or backend changes (you can find more information about Maglev and other performance optimization features in Chapter 8 of Cilium Up and Running).
In prior releases, there was no supported way to give some backends of a Service a larger share of Maglev traffic than others, or to gracefully drain a group of backends. Traffic was evenly distributed across all available backends.
Now, Cilium 1.20 reads a service.cilium.io/weight annotation directly off the EndpointSlice. The weight applies to every backend in that slice, with valid values running from 0 to 65535 (note the weights are relative, so they do not need to add up to 100). This is aimed squarely at selectorless Services backed by several manually managed EndpointSlice objects, where each slice can carry a different weight: a classic weighted split such as 70/30 across two slices becomes a two line annotation change.
apiVersion: v1
kind: Service
metadata:
name: example-service
annotations:
service.cilium.io/lb-algorithm: maglev
spec:
type: ClusterIP
ports:
- name: http
port: 8080
protocol: TCP
targetPort: 8080
# no selector: endpoints come from the manual EndpointSlices below
---
apiVersion: discovery.k8s.io/v1
kind: EndpointSlice
metadata:
name: example-service-1
labels:
kubernetes.io/service-name: example-service
annotations:
service.cilium.io/weight: "70" # applies to every backend in this slice
addressType: IPv4
ports:
- name: http
protocol: TCP
port: 8080
endpoints:
- addresses: ["10.0.0.11"]
- addresses: ["10.0.0.12"]
---
apiVersion: discovery.k8s.io/v1
kind: EndpointSlice
metadata:
name: example-service-2
labels:
kubernetes.io/service-name: example-service
annotations:
service.cilium.io/weight: "30" # set to "0" to drain these backends
addressType: IPv4
ports:
- name: http
protocol: TCP
port: 8080
endpoints:
- addresses: ["10.0.0.21"]
- addresses: ["10.0.0.22"]
The special case is a weight of 0. Instead of removing the backends, it marks them as being in maintenance: existing connections continue to work, but the backends leave the Maglev lookup table so no new connections are steered to them. That gives you a clean, connection-preserving drain, which is exactly what you want when taking a group of backends out of rotation for maintenance. Invalid values are ignored and fall back to the default behaviour, so a typo never takes a Service down.
Cilium 1.19 introduced support for Gateway API v1.4. Cilium 1.20 moves from Gateway API v1.4 to v1.6 and implements several of the capabilities that graduated across those releases.
From v1.5, TLSRoute (a feature Cilium introduced back in 1.14), the HTTPRoute CORS filter and ListenerSets all reached the Standard channel as generally available, and Gateway (and HTTPRoute) level authentication landed as an experimental feature. From v1.6, TCPRoute and UDPRoute graduated to Standard as well.
Gateway API continues to be a thriving project. Its accelerated development is a boon to anyone looking at migrating off Ingress – and its most popular implementation, Ingress NGINX – and Cilium keeps up with its pace.
Gateway API ExternalAuth filter
Cilium’s security features have always stood out. Built-in traffic encryption within or between meshed clusters, network policy for cluster segmentation, mutual TLS (now with zTunnel, which we cover later) provide a compelling set of security capabilities. One gap remained at the north/south boundary: authenticating users, scripts, and agents arriving from outside wasn’t easy to accomplish natively and, in prior releases, required customizing Cilium’s built-in Envoy proxy.
Cilium 1.20 closes that gap, by supporting the ExternalAuth filter in Gateway API (GEP-1494).
The idea is simple. You attach the filter to an HTTPRoute, and for every matching request the gateway checks with an external authorization service before forwarding any traffic. If the caller hasn’t authenticated yet, the authorizer sends them through an authentication flow first – a 302 redirect to a login page for a human in a browser, or a 401 / 403 for a request that presents no valid credentials. Only once the caller has authenticated and the authorizer returns 200 OK does the gateway forward the request to your backend. On the way through, the authorizer can also inject identity headers (who the caller is and which groups they belong to) so your application makes the final authorization call without ever touching a password or a token.
The feature applies to very different callers on the same gateway. A human hitting the dashboard in a browser gets redirected to an SSO provider (such as Authelia or Duo) and comes back with a session cookie. A CI job or script presents a JWT that oauth2-proxy validates against an identity provider (such as Keycloak, Duo or Okta). And an AI agent talking MCP authenticates with a service-account client, with no interactive login at all.
The authorizer’s response determines what the gateway does:
200: forward the request to the backend.
302: redirect the caller, for example to an SSO login page.
401 or 403: block the request at the gateway.
If you’d like to learn more, check out our External Auth lab!
Gateway API TCPRoute and UDPRoute support
Ever since Gateway API arrived in the 1.13 release, Cilium has kept in lockstep with every Gateway API release, adding TLS termination, TLS Passthrough, GRPCRoute, GAMMA (for east-west use cases) and many more features along the way.
Cilium 1.20 carries on in this vein. On top of the ExternalAuth feature mentioned above, a big addition to Cilium’s Gateway API support is TCPRoute and UDPRoute. Prior to this release, to expose a plain TCP or UDP service (a database, a DNS server, a game server) you had to drop out of the Gateway API model and back to raw LoadBalancer or NodePort Services. With TCPRoute and UDPRoute support, you can now manage L4 services through the same Gateway API you already use for HTTP or gRPC.
UDPRoute is ideal for applications such as DNS, VoIP, gaming, streaming media, IoT and telemetry (GEP-2645), while TCPRoute is handy for databases, message brokers, mail protocols and other TCP-based services (GEP-2644).
You can attach HTTP, TCP and UDP routes to a single shared Gateway:
In this example, an HTTPRoute serves Dean’s eBee Pac-Man app, a TCPRoute connects to the MongoDB database holding its high scores, and a UDPRoute sends DNS queries to CoreDNS:
Each backend is reachable through the same Gateway address (172.18.255.200) on its own port. After playing a round of eBee Pac-Man over the HTTPRoute, the score lands in MongoDB and can be read back over the TCPRoute:
$ dig @172.18.255.200 -p 53 ebeepacman.cilium.rocks
; <<>> DiG 9.20.23 <<>> @172.18.255.200 -p 53 ebeepacman.cilium.rocks
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 15177
;; flags: qr aa rd; QUERY: 1, ANSWER: 1, AUTHORITY: 0, ADDITIONAL: 1
;; QUESTION SECTION:
;ebeepacman.cilium.rocks. IN A
;; ANSWER SECTION:
ebeepacman.cilium.rocks. 3600 IN A 192.0.2.42
;; Query time: 1 msec
;; SERVER: 172.18.255.200#53(172.18.255.200) (UDP)
Thanks to external contributor eminaktas for landing this long-standing request.
Gateway API ListenerSets
Shared Gateways create an awkward ownership problem: the platform team controls the Gateway object, but application teams own the endpoints behind it. Adding another hostname or port therefore requires someone with write access to modify that central resource, creating a potential bottleneck in multi-tenant clusters. Without a delegation mechanism, this risks reproducing the same centralized configuration problem found in the inflexible Ingress API, one of the limitations Gateway API was designed to address.
Cilium 1.20 implements Gateway API ListenerSets, allowing application teams to attach and manage additional Listeners from their own namespaces. These Listeners share the parent Gateway’s address and infrastructure, while the platform team retains ownership of the Gateway and controls which namespaces may extend it. The result is a clean delegation model with per-namespace ownership and correct hostname filtering for attached routes.
The following manifests show this delegation in practice: a platform-owned Gateway authorizes a labelled application namespace to add its own Listener, while an HTTPRoute in that namespace attaches directly to the delegated Listener.
For more details on ListenerSet, check the documentation for more information.
More Gateway API improvements
Alongside the move from Gateway API v1.4 to v1.6, Cilium 1.20 fills in several smaller but useful gaps in its HTTP traffic handling.
Cilium now supports the Gateway API CORS filter natively. You can declare allowed origins, methods and headers, exposed response headers, credential handling and cache duration directly on the route.
Redirect handling is more complete too. In addition to 301 and 302 redirect codes we covered in the deep dive tutorial on “Redirect, Rewrite, and Mirror HTTP with Cilium Gateway API”, Cilium now formally supports and tests the optional 303, 307 and 308 HTTPRoute redirect codes. This gives users more precise control over whether a redirect is temporary or permanent and whether the original HTTP method and body should be preserved. You can try out redirection in the Advanced Gateway API Use Cases lab.
Finally, CiliumGatewayClassConfig can now control the HTTP Server response header. Its OVERWRITE, APPEND_IF_ABSENT and PASS_THROUGH modes let operators replace the header, add one only when the backend has not supplied it, or preserve the backend value. This is useful for applying a consistent value or preventing backend implementation details from being exposed.
Security
ztunnel: sidecarless mTLS service mesh
Cilium 1.19 introduced ztunnel, a purpose-built per-node proxy that brings sidecarless mutual TLS to pod-to-pod traffic, and Cilium 1.20 keeps that work moving.
The longer-term direction is for ztunnel to supersede Cilium’s original mutual authentication feature, which has been Beta since 1.14. That feature intentionally separated authentication from encryption. Cilium agents performed an out-of-band mTLS handshake to authenticate each new identity pair, while Cilium’s existing IPsec or WireGuard transparent encryption protected the workload traffic. This allowed mutual authentication to build on encryption capabilities Cilium already provided. One consequence of this design was that the first packet could be dropped while the authentication handshake completed. In addition, because transparent encryption operates between nodes, traffic between pods on the same node was not encrypted.
ztunnel takes a different approach by providing authentication and encryption together in the traffic path. Once you enroll a namespace by applying the io.cilium/mtls-enabled=true label, ztunnel holds each new connection while it establishes an mTLS session using TLS 1.3 over an HBONE (HTTP-Based Overlay Network Environment) tunnel. This avoids the first-packet drop and encrypts every supported TCP flow between enrolled pods, including pods running on the same node.
ztunnel remains Beta, but Cilium 1.20 improves its operational readiness. The certificate authority is now configurable: the default internal mode runs ztunnel mTLS without requiring SPIRE, while spire mode uses SPIFFE workload identities, with SPIRE issuing and rotating certificates for teams that want the CNCF-graduated SPIFFE and SPIRE identity stack. There are also new Prometheus metrics for enrollment and connection health. The dashboard below combines those metrics with Kubernetes workload data to show protected workloads, enrollment activity, failures and the connection state of each node’s ztunnel.
These improvements continue the joint engineering effort between Microsoft and Isovalent and build on the upstream ztunnel project, originally created for Istio’s ambient mesh.
As ztunnel matures, it is becoming the clear successor to the original mutual authentication feature. In Cilium 1.20, the maintainers marked legacy Mutual Authentication as deprecated, with removal planned for a later release.
Ztunnel support in Cilium remains currently in Beta. Make sure to review the documentation for current limitations and configuration guidance before rolling this out in production. For more information, see the blog and the ztunnel introduction in the Cilium 1.19 release blog.
Kubernetes ClusterNetworkPolicy (KCNP) support
A Kubernetes NetworkPolicy (KNP) is namespaced and additive: it is written by application teams, and there is no native way for a cluster administrator to set a rule that applies cluster-wide or that takes precedence over what a namespace owner writes. The upstream network-policy-api subgroup has been working on filling that gap, and its latest iteration consolidates the earlier AdminNetworkPolicy and BaselineAdminNetworkPolicy into a single ClusterNetworkPolicy.
Cilium 1.20 implements ClusterNetworkPolicy. It is cluster-scoped and tiered: an Admin tier that takes precedence over namespaced NetworkPolicy, and a Baseline tier that acts as a default which namespaced policy can override.
A policy selects a subject (by namespace or pod), sets a priority, and lists ingress/egress rules with an action of Accept, Deny or Pass. Those selectors accept matchExpressions in addition to matchLabels so you can target sets of namespaces or pods with set-based rules like In, NotIn and Exists. This gives platform teams the guardrails they have wanted: enforce a cluster-wide baseline, or hard-deny some traffic that no namespace can re-open.
For example, an Admin-tier policy can deny all ingress to a selected namespace:
For users who prefer using upstream APIs like KNP and KCNP rather than Cilium-specific Cilium Network Policies and ClusterWide Cilium Network Policies, this is a welcome addition to your segmentation toolkit.
If you’d like to learn more about this feature, check the documentation for more information.
New cluster-mesh policy entity & policy map aggregation
In previous Cilium versions, allowing traffic from “anywhere in my Cluster Mesh” in Cilium Network Policies meant either enumerating remote clusters by name or falling back to label selectors that spanned the mesh, since the cluster entity only covered the local cluster. Such policy was verbose to write, easy to get wrong as clusters came and went, and costly at scale because each remote identity consumed its own BPF policy-map entry.
Now, Cilium 1.20 adds a cluster-mesh policy entity that selects every endpoint in every meshed cluster in a single word. You write fromEntities: cluster-mesh once and it keeps matching as clusters join or leave the mesh. Underneath, both cluster and cluster-mesh now aggregate their identities into far fewer policy-map entries, so the mesh-wide rule is also the cheaper one to enforce at scale.
The policy is as simple as using the new entity in fromEntities:
If you’d like to learn more about this feature, check the documentation for more information.
Hubble policy correlation for audit verdicts
Previously, when a flow was dropped by a default-deny (nothing matched it) or logged under Policy Audit Mode, Hubble showed the drop but left the policy fields empty. You could see that a flow was denied, but not which policy, or the absence of one, was responsible. That gap made audit mode harder to trust: the whole point of audit mode is to preview what a policy would block before you enforce it, but the verdict did not tell you what would have caught the flow.
Now, in Cilium 1.20, Hubble correlates policy for audit verdicts. Audit flows (Verdict_AUDIT) populate the ingress_allowed_by / egress_allowed_by and ingress_denied_by / egress_denied_by fields, so you can see exactly what would have denied the flow before you flip enforcement on.
Cilium Cluster Mesh provides cross-cluster service discovery and load balancing across up to 255 clusters by default, with built-in support for network policy (as highlighted earlier), observability and service affinity.ia Placeholder
While Cluster Mesh is a popular Cilium feature, it is coupled tightly with Cilium. not portable and vendor-neutral. The Kubernetes SIG multi-cluster working group drove the creation of the portable and vendor-neutral Multi-Cluster Services API standard. Cilium introduced beta MCS-API support a few releases ago and in 1.20, promoted its MCS-API implementation to a stable support level.
The MCS-API provides a standard way to make a Service discoverable across multiple clusters. Imagine that an application runs in two clusters called site-a and site-b, behind a Service named web in the default namespace. Normally, each Service is visible only inside its own cluster. To make web available across the mesh, you create a ServiceExport with the same name and namespace in both clusters. That is the only additional resource the application team needs to manage.
apiVersion: multicluster.x-k8s.io/v1beta1
kind: ServiceExport
metadata:
name: web
namespace: default
Cilium then combines the exported Services and their backends into a single multi-cluster Service. Workloads in any connected cluster can reach it through the standard MCS-API DNS name web.default.svc.clusterset.local. Because this naming convention and the ServiceExport resource are part of the Kubernetes MCS-API, the same application manifests can be used with any conformant implementation.
Under the hood, the Cilium operator creates the corresponding ServiceImport, allocates the clusterset virtual IP and reconciles the derived Kubernetes Service used to connect to backends across the mesh. Cilium 1.20 moves this implementation to the multicluster.x-k8s.io/v1beta1 CRDs and installs the required CRDs automatically.
Of course, Cilium Global Services remain available, but you now also have a stable, supported alternative for cross-cluster service discovery that uses a portable, standard Kubernetes API.
If you’d like to learn more about this feature, check the documentation for more information.
Shrink cilium-cni binary by 80%
Every Cilium agent ships the cilium-cni plugin and installs it onto the node at /opt/cni/bin/cilium-cni by default, where the kubelet calls it to wire up each pod. Over time what was originally a small binary had grown: the CNI plugin was pulling in large chunks of the agent codebase through Go imports, and by 1.19 it weighed 76 MB. In 1.20, the plugin is no longer importing packages it never needs and the same binary is 16 MB, resulting in a 79% cut.
# In a 1.19 cluster
root@cilium-119-worker:/# ls -la /opt/cni/bin/cilium-cni
-rwxr-xr-x 1 root root 76625528 Jun 16 12:06 /opt/cni/bin/cilium-cni
# In a 1.20 cluster
root@cilium-120-worker:/# ls -la /opt/cni/bin/cilium-cni
-rwxr-xr-x 1 root root 16210104 Jul 17 08:59 /opt/cni/bin/cilium-cni
A smaller binary on every node means a smaller image layer to pull and less to copy during each CNI install.
remote-node / world policy identity aggregation
In previous Cilium versions, a policy that selected a broad entity like world or remote-node expanded, under the hood, into one entry per matching identity in the BPF policy map. On a large or heavily meshed cluster that could mean thousands of entries for a single rule, and in the worst case the policy map could fill up.
Now, Cilium 1.20 teaches the policy map a new kind of wildcard for exactly these semantic groups. Selecting world or remote-node now inserts a single aggregated wildcard entry instead of one entry per matching identity, and the datapath does an extra wildcard lookup to match. The verdict is identical, but the footprint collapses from thousands of entries to a handful, so policy enforcement stays cheap as the cluster and the mesh grow.
BGP Tooling Improvements
Cilium’s BGP support is commonly used in self-managed Kubernetes clusters, to interconnect with an existing network fabric. Cilium BGP relies on the GoBGP platform, a robust and mature Go-based implementation of BGP. In previous versions, debugging a Cilium BGP session from the agent was sometimes challenging: the on-agent BGP commands gave limited output, and there was no way to see which route policies were actually being applied to the speaker.
Cilium 1.20 improves the on-agent BGP tooling, now surfaced through the hive-shell from inside the Cilium agent (cilium-dbg shell -- bgp/...). A brand-new bgp/route-policies command shows the route policies Cilium has actually programmed into GoBGP – something you previously had no way to inspect.
$ cilium-dbg shell -- bgp/route-policies
Instance Policy Name Type Statement Name Match Peers Match Families Match Prefixes (Min..Max Len) RIB Action
64513 allow-local import accept-local accept
peer-frr-export export PodCIDR-ipv4 172.18.0.5 10.0.0.0/24 (24..24) accept
In addition, bgp/peers gains a --format option (table, JSON or detailed); providing a detailed view of per-peer timers, negotiated capabilities and graceful-restart state.
Heads-up worth calling out: the on-agent cilium-dbg bgp subcommands and the local REST BGP API are deprecated in 1.20. Move to the hive-shell equivalents – cilium-dbg shell-- bgp/peers, bgp/routes, bgp/route-policies. The user-facing cilium bgp CLI is unchanged.
Conclusion
Cilium 1.20 is the product of a community that continues to grow well beyond the code itself. Asana, Etraveli, Michelin, OpenAI, Suse, Telefónica’s acens, and Zynga have all shared stories about running Cilium in production. If your organization has a story to tell, we would love to hear it.
The community also gathered at CiliumCon and the Cilium Developer Summit in Amsterdam, with the CiliumCon talks now available to watch online. The next opportunity to meet in person will be at KubeCon + CloudNativeCon North America in Salt Lake City, where CiliumCon and another Developer Summit will take place from November 9 to 12. We look forward to seeing you there!
AI workloads are changing what platform teams need from infrastructure. Provisioning GPUs and standing up a cluster no longer makes a platform “AI-ready.” Once training spans more than one node, the bottlenecks show up in places application platforms rarely treat as first-class concerns: inter-node communication, shared storage, placement, topology, and validation.
Our internal ML platform supports training and inference workloads behind product experiences such as search and ranking. As these workloads grew, some training jobs outgrew the practical limits of a single machine. Models moved into the tens of billions of parameters, and a single node stopped being able to hold the model, its optimizer state, and a workable batch size at the same time. Atlassian therefore needed a platform that could make distributed training reliable and repeatable, without exposing ML teams to the underlying infrastructure complexity.
For distributed AI, performance is not just optimization. It is part of correctness.
The workload problem behind the infrastructure work
Adding GPUs was the easy part. The platform had to make three things predictable:
Communication performance. Workers need fast, low-overhead inter-node communication for synchronization traffic, including gradients, model state, and collective operations, so a distributed job keeps progressing together. The gap here is not marginal: on current-generation GPU hardware, a socket-based path can leave a multi-node job running at roughly half the speed the same hardware delivers over RDMA.
Shared data access. Workers need high-throughput shared storage to read training data, write checkpoints, and access intermediate artifacts concurrently without turning storage into a bottleneck or causing long pauses and uneven progress.
Operational predictability. Hardware placement, network topology, storage, and validation need to work together so a job does not silently run on a degraded path.
Two technologies cover the first two:
RDMA (Remote Direct Memory Access), which enables lower-overhead, high-throughput communication between GPU nodes.
Lustre, a parallel distributed filesystem designed for high-throughput shared access to training data and checkpoints.
ML teams should not have to manage either one. That is the platform’s job.
Why this matters beyond one company: None of this is specific to us. As AI adoption grows, platform teams keep hitting the same wall: distributed systems, accelerators, storage and scheduling have to behave as one platform, not four separate layers.
What happened before the high-performance network and storage design
Before RDMA and Lustre, distributed jobs ran. How fast they ran was anyone’s guess. Communication and storage delays surfaced as low GPU utilization, uneven step times, and runs that took far longer than they should have.
The worst of it was silent RDMA fallback. A misconfiguration sent collective traffic over sockets, and the job carried on at a fraction of the speed it should have reached.
That made the degradation easy to miss, and expensive to ignore.
This was not limited to one environment. On our other cloud, the device plugin that advertises the RDMA fabric to Kubernetes has been in CrashLoopBackOff on every fabric-capable production node from the day it was deployed. It was not a regression; it had never worked. We found it 271 days later, by accident, while verifying an unrelated GPU operator upgrade. The cause was mundane: a mirrored container image resolved to a single-architecture manifest that did not match the nodes, so the plugin failed at exec. A stopgap had been applied to staging months earlier, and production never received the follow-up.
Two failure modes matter here. A job that explicitly requests the fabric stays pending indefinitely, which is loud and easy to diagnose. A job that does not request it runs over sockets with no error at all, which is both the default and the expensive case.
Three lessons generalize from that outage:
A capability with no consumers is invisible, however expensive it was to buy. No alert existed for “this DaemonSet has never been ready”, and that absence was the real defect.
Pod age and restart counts mislead when nodes autoscale. Every new node presents a fresh crash loop, which reads as a recent break when the underlying condition is months old.
A healthy environment beside a broken one is a trap. Staging passing told us nothing about production, because staging had quietly diverged onto a different fix.
Both failures share a cause that has nothing to do with networking. Multi-node demand is still low next to single-node work. Few jobs touch the fabric, so nothing generates the signal that would show it is broken. That argues for synthetic validation rather than better dashboards: a dashboard shows what your jobs did, not what a fabric nobody used would have done.
Why a new network and storage design was needed
We treated this as a platform design problem rather than a tuning exercise.
We integrated RDMA-capable networking and shared high-throughput storage into the platform, absorbing topology and cloud-specific complexity so ML teams could run distributed training reliably without managing the underlying infrastructure.
Platform concern
Before
After
Inter-node GPU communication
Standard socket/TCP path could become a hidden bottleneck
RDMA-capable path for faster collective communication
Shared training storage
Less efficient path for concurrent checkpoint and dataset access
Lustre-based shared high-throughput storage
Operational confidence
Jobs could appear healthy while underperforming
Validation made transport and performance behavior visible
User experience
Risk of infrastructure details leaking to ML teams
Platform-managed capability with consistent workflow
How we achieved it
That meant changes at every layer:
RDMA-capable cluster and node-pool setup. Clusters are created with the multi-networking support the fabric requires, and GPU node pools are provisioned against a specific reservation and location rather than a generic pool.
GPU-specific network interfaces and network mappings. Each GPU node carries dedicated RDMA interfaces alongside its ordinary one, and each node pool is mapped to the RDMA network and subnet matching where its hardware physically sits.
Node-level dependencies, drivers, and readiness controls. The RDMA userspace libraries, the collective communication runtime, and the network plugin are installed on the node image. A node that fails its readiness check never becomes schedulable, so jobs cannot land on a partially configured node.
Lustre integration for shared storage. The filesystem is exposed through a CSI driver as an ordinary ReadWriteMany PersistentVolumeClaim, mounted at the same path in every worker pod, with per-namespace and per-job subdirectories providing isolation.
Placement and scheduling constraints for the right hardware and topology. Jobs are pinned to node pools with the correct hardware, drivers and fabric wiring, and gang scheduling ensures a distributed job either receives all of its workers or waits, rather than half-starting and holding GPUs idle.
Workload-level configuration and validation. The networking and storage plumbing is injected into the pod spec at admission time so ML teams never hand-write it, and the transport is confirmed before a job is treated as healthy.
The platform absorbs network, storage and topology complexity so that submitting a distributed training job stays ordinary.
One of the key lessons was that RDMA depended on physical topology, not just Kubernetes or workload configuration. GPU reservations could move across datacenters within the same zone, and the RDMA subnet mapping had to remain aligned with where the hardware actually lived. That meant the solution needed to support multiple RDMA network mappings and safe node-pool transitions instead of assuming one static configuration forever. How much of this you inherit depends on where you run. On one of our two clouds the managed fabric handles reservation and zone changes itself, and the platform team never sees them.
We also had to treat reservation changes as platform transitions: introduce a new node pool, move workloads safely, and then retire the old path. In practice that proved more reliable than trying to encode the whole problem as a one-time setup task.
What problems we faced along the way
None of this was a feature flag. We had to solve both the availability of suitable GPU nodes across two of the top hyperscalers and the challenge of introducing a network design that pushed beyond each provider’s usual patterns.
GPU node availability across clouds: securing enough compatible GPU capacity was a constraint in both clouds, affecting planning, placement, and the ability to move workloads between providers.
A new ask for cloud providers: this combination of GPU placement, high-performance networking, and shared storage was not a standard, one-size-fits-all request. There was no single approach that worked across providers, so the design had to adapt to each cloud’s capabilities and constraints.
Topology, reservation, and configuration timing: the correct network path depended on underlying physical placement, and some information needed to select the right RDMA mapping was not always available early enough. This is not universal, and the difference matters if you run in more than one place. On one of our two clouds the managed fabric absorbs reservation and zone changes, and none of it is visible to the platform team. On the other, the mapping is ours to maintain. Do not assume the behavior transfers.
Reservation topology is a provider constraint, not a platform choice: we have been allocated GPUs inside a single block, and we have been allocated them spread across several. Whether topology-aware scheduling can help once an allocation spans blocks is still an open question for us.
Shared storage carries its own operational cost: the parallel filesystem solved throughput, but onboarding a new region or cluster still means standing up a new filesystem instance by hand. We traded a throughput problem for an operational one.
Cross-layer, cross-team delivery: networking, node setup, storage, scheduling, and workload integration all had to line up, requiring close iteration across infrastructure, networking, and ML platform boundaries.
Important lesson: for distributed training, “the job ran” is not a sufficient success signal. You need validation that confirms it ran on the intended transport and storage path.
In practice our validation is coarser than we would like. There is no per-job telemetry that pins a slowdown on the transport. What we have is which storage path a job used, whether its transport was sockets or RDMA, and total training time. That combination is enough to catch the failures described here.
What changed after the new design
Treating RDMA, Lustre, placement and validation as one concern bought us more than a benchmark bump.
We ended up with a more reliable training foundation: fewer hidden performance failures, better GPU utilization, and multi-node workloads that were practical to run repeatedly.
The measured results were significant:
Measurement
TCP fallback
RDMA
Method
Peak bus bandwidth
–
355 GB/s
2-node NCCL all-reduce, H200 141GB nodes
Median step time
12.36 s
6.07 s
Qwen2.5-14B FSDP supervised fine-tune, 16 H200 GPUs across 2 nodes
Training throughput
1.0×
2.04×
Controlled A/B, identical ~360 s model-load phase
100-step benchmark
~1,950 s
~836 s
Separate fine-tune benchmark, 2 nodes
Speed is the least interesting part. What those numbers buy is GPU-hour efficiency, capacity planning we can trust, and confidence that the next model up will scale.
The broader cloud-native lesson
The lesson is not really about RDMA or Lustre. It is about what happens when you stop treating them as separate problems.
Cloud-native platforms already know how to orchestrate complex distributed systems. AI training pushes those platforms into a new set of constraints where network fabric, storage behavior, accelerator scheduling, and validation all become part of the product experience.
That is the job. Keep the user-facing story simple, and let the platform absorb the rest.
If you are building this: treat networking, storage, scheduling, readiness and validation as one system. That is what makes distributed training predictable as it scales.
Central takeaway: serious distributed AI training becomes viable when performance-critical infrastructure is designed as an integrated platform problem, not as a collection of independent features.
This document describes three failure scenarios that separate having backups from being able to recover, and the guidance that follows from each. Every scenario is reproducible on a laptop from the lab repository above, and every terminal output shown is a real capture from that lab.
The document covers recovery of stateful applications running on Kubernetes: verifying that backups contain data, the split between declared state and stored state, and consistency across multi-volume applications. It does not cover compliance frameworks, product comparisons, or recovery of the underlying cloud or datacenter infrastructure, though it names where those responsibilities begin.
Specific tools appear where a scenario needs them (Velero, the CSI snapshot APIs). They are reference implementations used to make the scenarios concrete. The failure modes and the guidance apply to any tool occupying the same role.
Definitions
RPO and RTO on one timeline
The four recovery layers
For a recovery to count, four layers must come back:
The four recovery layers
Each layer has mature tooling, and each usually recovers fine in isolation. Recovery fails at the joins between the layers: a restored cluster with no data, restored data with no traffic path, an application definition that provisions an empty volume. The three scenarios below each break one join.
The lab
The whole lab in one picture
Two clusters and two services, all local:
Production: a Kubernetes cluster where each node is its own lightweight VM with its own kernel, so losing production means powering off a machine.
Recovery: a second cluster that exists before anything goes wrong.
Backup store: an S3-compatible object store outside both clusters, so losing either cluster cannot take the recovery points with it. Both clusters point their backup tool at the same bucket.
Git: a local Git service holding the application manifests, watched by a GitOps controller in the recovery cluster.
The workload is a PostgreSQL application with known contents (four rows), so every restore can be validated against an expected result rather than against a green dashboard.
Scenario 1: verifying that a backup contains data
Backup tool moves objects and volume bytes to an S3 store
A Kubernetes backup has two distinct parts: the resource definitions (YAML) and the persistent volume data. Backup tools protect volume data through provider or CSI snapshots, file system backup, or snapshot data movement to an external store. The lab uses the last of these, with Velero and its data mover.
Most verification stops at the backup’s Completed status. Go one step further and confirm that volume bytes actually moved:
$ kubectl -n velero get datauploads -l velero.io/backup-name=$BACKUP \
-o custom-columns='NAME:.metadata.name,PHASE:.status.phase,BYTES:.status.progress.bytesDone'
NAME PHASE BYTES
guestbook-rehearsal-20260727001126-q2j9m Completed 47989888
This is the data mover confirming that 47,989,888 bytes of volume data left the cluster and landed in the external store. A backup tool that cannot report this number for a given backup deserves scrutiny.
Deleting the namespace, PVC included, and restoring from this backup returned the same four rows in about two minutes. That is the happy path, and it hides three things no backup tool does automatically:
Protecting volume data does not make a database backup application consistent. Flush or quiesce hooks must be configured when the application requires them.
Restoring onto different infrastructure may require storage class mappings and other transformations. Tools provide the mechanisms; each team must design and test them.
A backup phase of Completed means the backup operation completed. It does not prove the application will start, contain the expected data, or serve traffic. Only an end to end recovery test provides that evidence.
Boundary. Backup tools restore resources into a cluster that already exists. They do not create the cluster, the nodes, the network, the load balancers, or DNS. Something else must recover Kubernetes itself, and that something is infrastructure as code or Cluster API. A DR plan that starts with “restore the backup” must state what the backup gets restored into.
Scenario 2: declared state is not stored state
The GitOps trap: the controller rebuilds the declarations, the store holds the data
The production cluster is powered off. The recovery cluster, which existed before the disaster, has a GitOps controller pointed at Git and a backup tool pointed at the shared store. It has never run the application.
Syncing the application from Git succeeds: the sync reports Synced, the StatefulSet rolls out, the database pod is Running and Ready, every dashboard is green. Querying the database then returns:
ERROR: relation "attendees" does not exist
The database is running and it is empty. Nothing malfunctioned. Git only ever contained the declarations, so Kubernetes did exactly what the YAML says: create a StatefulSet, create a Service, and provision a brand new, empty volume for the PVC. GitOps reconstructed the declared state perfectly and restored none of the stored state.
Both tools are required because there are two different things to bring back and each tool carries exactly one of them: Git stores intent, and backups store state. In the lab, the recovery that produced validated data was:
Remove the empty application the sync created.
Restore the application, volumes included, from the backup store.
Validate the data against the expected contents.
The restore also crossed infrastructure: the backup was taken on one node runtime and restored onto another. A disaster may force recovery onto different infrastructure, so restore portability is something to test, not assume.
Measurement. In the lab, powering off production to validated data in the recovery cluster took four minutes live and just under two minutes in a rehearsed rerun. Both figures measure only the scripted slice; a production RTO wraps detection, decision, traffic cutover, and failback around it. The general lesson: the moment the dashboards turned green was not the recovery. The moment the data came back and was checked was.
Scenario 3: multi-volume consistency
Two snapshots from different moments tear the data
Real stateful applications span multiple volumes: database data plus WAL, message broker partitions, replica sets. The lab stand-in writes matched pairs, order n to one PVC and payment n to another, five times a second, with one invariant: every payment must have its order.
Snapshotting the two volumes individually, five seconds apart, produced two snapshots that were each ReadyToUse and individually perfect. Restoring both and comparing the last committed sequence numbers:
last order committed : 108352 last payment committed : 108377
[FAIL] 25 payments have NO matching order. [FAIL] Each snapshot succeeded. The restore is still wrong.
Twenty five payments reference orders that do not exist. No component failed, every operation reported success, and the combined recovery point describes a moment in time that never existed. In production, that five second gap is a backup tool walking a list of a hundred PVCs one by one.
The API answer.VolumeGroupSnapshot reached GA in Kubernetes 1.36. One object selects PVCs by label, and the CSI driver receives one request for a coordinated, crash consistent recovery point across all of them:
One group snapshot cuts both volumes at the same moment
Restoring the group’s member snapshots and running the same verifier:
last order committed : 109169 last payment committed : 109169
[OK] Every payment has a matching order. Restore is consistent.
Caveats that apply beyond the lab:
Support is driver specific. A driver that supports ordinary VolumeSnapshots proves nothing about group snapshots; the CSI group RPCs are a separate implementation. As of mid 2026, most of the major cloud drivers checked for the lab do not implement them.
Setup is explicit: the CRDs and feature gates on the snapshot controller and CSI sidecar must be enabled by the operator.
Crash consistent is not application consistent. The API removes cross volume timing skew; it does not flush or quiesce the database.
The lab uses the CSI hostpath test driver, which implements the group RPCs but archives member volumes sequentially, so the writer is paused during the group snapshot to keep the demo deterministic. The point in time guarantee itself belongs to the storage backend of a production driver.
Recovery testing guidance
A recovery test is not deleting a pod and watching it return; that tests workload reconciliation. A recovery test:
Restores a complete stateful application into a clean target that has never run it.
Validates the data and the user path against expected contents, not against resource status.
Measures the whole thing with a clock.
The principles that transfer from the lab to production as-is:
Two independent failure domains
Open gaps in the ecosystem
The scenarios expose gaps that no single tool closes today:
No common cross-cluster failover contract. Data, workload, cluster, traffic, and identity each have tools, and every row is missing the same thing: a shared contract with the next row. Products answer this inside their own APIs; core Kubernetes does not define the sequence.
No standard recovery unit for an application. Core Kubernetes has no maintained Application resource that says which objects, operators, data services, and external dependencies must recover together. Backup tools use namespaces and labels, GitOps controllers have their own application objects, package managers have releases, and each draws the boundary differently.
Backup success is treated as recovery proof. Backup completion metrics are widely monitored; restore rehearsal results rarely are.
How to contribute
The Cloud Native Business Continuity initiative under CNCF TAG Operational Resilience is an open proposal seeking contributors, aiming at a landscape gap analysis, updated backup and DR guidance, and reference architectures: https://github.com/cncf/toc/issues/1779
It was a routine cost review. The slide showed the month’s GPU spend, the biggest line on the whole infrastructure bill, and someone asked a five-word question: “Are we using these things?” Nobody could answer. The most expensive hardware we owned was also the hardest to see.
The frustrating part is that the answer already existed. Every GPU’s utilization had been recorded, every second, for months, all of it flowing into one central, infrastructure-owned Prometheus that held every metric for every team across thousands of namespaces. The data was right there. It just sat somewhere no tenant was allowed to look, because a store that sees everyone’s metrics can’t safely be opened to any one of them. So visibility split in two:
When we finally went looking, we found a GPU that had sat at zero percent utilization for eleven straight days: allocated, powered on, doing nothing, and invisible to the team that owned it. You can’t fix what you’re not allowed to see. Multiply that one idle card across a fleet, all of it drawing power behind green health checks, and “we’re not sure” becomes real money every month.
This post is how we closed that gap: how we gave every team a safe, self-service view into their own metrics without handing them the keys to everyone else’s. No new metrics stack, no vendor platform, just CNCF-native pieces arranged so the people spending the GPU budget can finally see it.
Why “just share Prometheus” doesn’t work
The obvious fix is to give every team read access to the central Prometheus. We ran into two walls, each built from an individually correct decision.
The first is security. A Prometheus query endpoint isn’t namespace-aware: if a tenant can run one PromQL query, they can run any query, including one that reads another tenant’s request rates or capacity plans. “Everyone can read everything” isn’t a posture you can defend across thousands of namespaces.
The second is scale, and it has a name every platform engineer knows: the noisy neighbor. The central Prometheus is already scraping and storing series for the whole fleet. Point a few hundred engineers and ad-hoc queries at it and the store everyone depends on starts to buckle. One team’s expensive range query becomes everyone’s latency spike.
Both are reasonable on their own. Together they leave the data tenants need locked in a store you can’t safely open to them. We needed to give each tenant their own curated, isolated slice, served directly.
A metaphor that made it click
Picture the central Prometheus as one enormous reading room where every team’s private notebooks sit on open shelves. Hand out a room key and you break confidentiality in the same motion, and the moment a crowd arrives the shared room grinds to a halt. What you want is a librarian who takes your card and brings a copy of only your box. That librarian is the multi-tenant proxy.
The insight: a proxy in front, a contract in YAML
We didn’t need a new metrics stack. We needed a thin, tenant-aware layer in front of the one we already had. The design came down to three moves:
Identify: Authenticate the caller and establish which tenant they are.
Isolate: Restrict every query to that tenant’s namespace, enforced below the query language so it can’t be bypassed.
Deliver: Optionally copy a curated slice of each tenant’s metrics into their own small Prometheus, so their dashboards and alerts run against a store they own.
The other half is self-service. Platform teams can’t hand-curate metric lists for thousands of namespaces, so the contract is a small Kubernetes custom resource, a MetricAccess object, where a team declares which metrics it wants. The platform owns the mechanism; the tenant owns the policy. All of it sits on CNCF-native, open source pieces.
How the pieces fit
The infrastructure Prometheus keeps doing its job; everything tenant-facing sits behind the proxy:
On the read path, a tenant’s request enters through Nginx (load-balancing across proxy replicas), passes through kube-rbac-proxy for authentication and authorization, and reaches the proxy. The proxy discovers backend Prometheus instances via the Kubernetes API, fans the query across healthy backends, filters results to what the tenant may see, and returns the aggregate.
On the write path, the proxy periodically collects each tenant’s curated metrics and remote-writes them into that tenant’s own Prometheus. For HA tenants it resolves each replica’s pod DNS and writes to all of them, so every instance holds identical data. None of this is exotic: Prometheus, service discovery, and remote-write with a tenancy model on top.
The part that has to be airtight: isolation
Self-service is only safe if isolation isn’t optional. Two open source components do the work here.
kube-rbac-proxy handles authentication and authorization. It answers “who is this, and are they allowed?” using Kubernetes-native identity and RBAC, the same model you already trust for the API server, and carries the tenant’s identity through as a namespace assertion that anchors everything downstream.
Query-time isolation is handled by prom-label-proxy . It’s the piece that makes “everyone can read everything” impossible rather than merely discouraged: it rewrites every incoming query to inject a namespace matcher, so any query becomes query{namespace=”your-namespace”} before it reaches Prometheus. Enforcement happens below the query language, so no PromQL can escape it.
The proxy itself runs hardened: non-root (UID 65534), read-only root filesystem, all capabilities dropped, no privilege escalation, least-privilege service account. Defense in depth on the path that matters most.
Isolation that also cuts the bill
There’s a second, quieter isolation, where the cost story lives. Read-time filtering stops tenants seeing each other’s data, but a tenant’s own Prometheus can still store far more than it needs. The metricIsolation setting pushes the boundary to collection time: metrics are gathered through prom-label-proxy with the namespace filter already applied, so a tenant only ever ingests its own series.
The difference is not subtle:
Configuration
Series stored
Query speed
Isolation
metricIsolation: false
~10,000+ (all namespaces)
Slower (large dataset)
Query-time only
metricIsolation: true
~300 (this namespace)
Faster (focused dataset)
Collection + query time
For a typical tenant that’s roughly a 97% cut in stored series, from ten-thousand-plus down to a few hundred. A smaller store queries faster, costs less, and can’t leak data it never collected. The noisy-neighbor problem shrinks too, since the fleet-wide store leaves the critical path for everyday dashboards. Our original team went from empty dashboards to a Prometheus of its own.
What a tenant actually does
A tenant onboards with a single YAML file that declares the metrics and where to deliver them:
apiVersion: observability.ethos.io/v1alpha1
kind: MetricAccess
metadata:
name: gpu-team-metrics
namespace: gpu-team
spec:
source: gpu-team
metricIsolation: true # only collect this namespace's series
metrics:
- "DCGM_FI_DEV_GPU_UTIL" # exact match
- "container_(cpu|memory)_.*" # regex
- '{__name__=~"nginx_ingress_controller_.*"}' # PromQL selector
remoteWrite:
enabled: true
interval: "30s"
target:
type: "prometheus"
prometheus:
serviceName: "prometheus-operated"
servicePort: 9090
replicas: 2 # write to both HA replicas
statefulSetName: "prometheus-gpu-team"
extraLabels:
tenant: "gpu-team"
managed_by: "multi-tenant-proxy"
The metrics list mixes three styles freely: exact names, regexes, and PromQL selectors. Apply the file and the proxy starts collecting and delivering. Querying is a normal Prometheus API call with a tenant header:
One requirement: the tenant’s Prometheus must accept remote-write (start it with –web.enable-remote-write-receiver). Everything else is defaults.
Six PromQL queries that make idle GPUs visible
Once a team can see GPU metrics, a few queries do most of the work. These assume a DCGM-style exporter; adjust the names to yours. Queries 5 and 6 also use an ingress request-rate metric, so swap nginx_ingress_controller_requests for whatever your ingress exposes.
1. Average GPU utilization per namespace — the headline number the team was missing:
avg by (namespace) (DCGM_FI_DEV_GPU_UTIL)
2. Count GPUs that are effectively idle — under 5% utilization for the last hour (this one finds the money):
count by (namespace) (avg_over_time(DCGM_FI_DEV_GPU_UTIL[1h]) < 5)
3. GPU memory used vs. total per namespace — separates “busy and memory-bound” from “reserved but empty”:
sum by (namespace) (DCGM_FI_DEV_FB_USED)
/ sum by (namespace) (DCGM_FI_DEV_FB_USED + DCGM_FI_DEV_FB_FREE)
4. Power draw per namespace — a rough real-time proxy for cost:
sum by (namespace) (DCGM_FI_DEV_POWER_USAGE)
5. Hot cards, no traffic — GPUs busy while ingress is silent, often a stuck job:
avg by (namespace) (DCGM_FI_DEV_GPU_UTIL) > 70
and sum by (namespace) (rate(nginx_ingress_controller_requests[5m])) < 1
6. Traffic, no GPUs — requests arriving while GPUs sit idle points at a scheduling gap, not capacity:
sum by (namespace) (rate(nginx_ingress_controller_requests[5m])) > 10
and avg by (namespace) (DCGM_FI_DEV_GPU_UTIL) < 5
Query 2 closed the loop for our original team: it surfaced GPUs idle for hours, the invisible spend we opened with, capacity they could finally see and give back.
Holding up at fleet scale
Self-service without limits just moves the noisy-neighbor problem downstream. Two levers keep load bounded. Curated metric sets mean the proxy handles a fraction of the fleet’s cardinality, not all of it. And per-tenant collection intervals act as a quota: one team can run at 5-second resolution, a batch workload at five minutes, and neither pays for the other’s cadence.
Across many clusters it leans on the same primitives: dynamic backend discovery, independent per-tenant remote-write, and writes to all HA replicas so failover is a non-event. It’s still a system you operate, but its moving parts are ones a platform team already knows.
What we’d tell our past selves
Self-service still needs guardrails. A team asking for “all metrics” usually means “I don’t know which ones I need yet.” Curation is a conversation, and good defaults beat a big allowlist.
Cardinality is a cost decision. metricIsolation is the difference between a 300-series store and a 10,000-series one. Turn it on unless a team has a concrete cross-namespace reason not to.
Delivery is where the sharp edges are. Retries, backoff, and HA multi-replica writes aren’t optional extras. Budget for the failure modes up front.
Isolation can be overkill. Small teams do fine on filtered query access alone; reserve remote-write for teams that own real dashboards and alerts.
Names drift across clusters. Half our early “no data” tickets were a query naming a metric the exporter didn’t emit. Pin exporter versions and document the exact series names.
The portable idea
Strip away the specifics and the pattern fits any multi-tenant cluster: an authn/authz proxy, label-enforced isolation below the query language, and optional per-tenant remote-write, all CNCF-native and open source. No proprietary lock-in, no bespoke stack; if you already run Prometheus, you have the foundation. Other approaches exist and are worth a look. This is just the one that let us give thousands of namespaces their own view without opening the shared store to everyone.
The team that couldn’t see its own GPUs now runs its own dashboards and catches idle capacity within the hour. The reading room is quiet again, and everyone has their own desk.
Try it yourself
The proxy is open source under Apache 2.0; the MetricAccess CRD, manifests, and examples are in the repo. The KubeCon + CloudNativeCon India 2025 talk walks through a live demo.
Bingi Narasimha Karthik (Golden Kubestronaut, Adobe) is a Senior Cloud Engineer working on Kubernetes observability and GPU workload orchestration. He co-presented “Unlocking Kubernetes Observability: Secure, Tenant-Centric Metrics for GPU Workloads” at KubeCon + CloudNativeCon India 2025 and is based in Bengaluru, India.
Ramkumar Nagaraj (Golden Kubestronaut, Adobe) is a Senior Computer Scientist working on GPU infrastructure and Kubernetes platform engineering. He co-presented the same session and is based in Bengaluru, India.
“A sales guy writing code” used to be the lead-up to a joke. But now no one’s laughing.
Designers used to sit meekly waiting for the high priests of code to make their designs real. Now they don’t wait for anyone.
The tools behind this shift (Cursor, Claude, Lovable, Replit) were unknown names a few years ago; now they count their users in the millions, and a meaningful share of those users have never written a line of code by hand… and never will.
A lot of what gets built is toys. That’s fine. Playing with toys is how people learn, and basic hosting is all a toy needs.
But they aren’t all toys. A lot of the products vibe-coded into existence have real potential, and their creators have real ambition and real skill. They don’t understand what’s happening under the hood, but they understand what the market needs and what might fix it. If they could ever get to production. Which most won’t.
What they will do is go live on some sort of infrastructure, with some sort of database, and a few rudimentary nods to good architecture: one region, one credential that can do everything, a schema nobody reviewed. By and large, vibe-coded apps bypass the last 20 years of cloud native best practices and suffer the consequences.
Apps get built at 100 miles per hour, get deployed at 85, and screech to a halt at anything resembling production. I believe closing this velocity gap is the defining infrastructure problem of the next few years, and the engineers working in cloud native infrastructure can either help or hinder.
Starting was never the hard part
There is an old joke that the first 90 percent of a project takes 90 percent of the time, and the last 10 percent takes the other 90 percent. Software has always worked this way. Everyone knows how to start an app; the rare skill is finishing one and getting it into the world.
AI did not change that ratio. It made it more extreme. When starting an app costs real money and real engineers, the projects that died before production at least died as line items someone had approved. Now starting is nearly free, so vastly more gets started, and the fraction that finishes has fallen accordingly. If 80% of apps never shipped before, today’s number is probably a lot closer to 99%. Getting to production stops being 80 percent of the work and becomes nearly all of it.
But what exactly do we mean by “production”? Because the way an SRE uses the term is different from the way an AI agent uses it.
To an SRE, production is a set of testable claims. What is the p99 latency under peak load? Has the failover actually been rehearsed, or does it exist only in a diagram? What is the blast radius of a bad deploy, and how fast does it roll back? Who touched what, and when? To an agent, production is a URL that returns 200.
Agents gonna agent
Watch an AI agent stand up infrastructure and you will see the same terms appear again and again: Supabase, serverless functions, a managed one-click backend. Nothing is wrong with these tools. They are well built, and they are deliberately simple, which is exactly why agents choose them: an agent can hold the entire mental model in a few thousand tokens and produce a working demo without asking anyone for anything.
The agent is not choosing that stack because it evaluated the alternatives and found it superior. It is choosing the stack that is most convenient for the agent: the one that demands the least context to operate. Convenience for the agent is a heuristic for good infrastructure, and like most heuristics, it works until it doesn’t.
Agents gonna agent. They are relentlessly following their mandate, which is usually to produce the visible result the person asked for. Security posture, failover behavior, and audit trails are invisible in a demo, so why would the agent bother? So are liveness probes, resource limits, and the difference between a secret in a vault and a secret pasted into an environment variable. Bothering with them would mean bothering the user, and probably confusing them. Today’s AI defaults are optimized for prototypes and demos, not production.
Does cloud native get replaced?
For those of us in the cloud native community this is frustrating. Because the definition of “finished” already exists. It has been built in public for two decades, one project at a time. Kubernetes for orchestration. Prometheus and OpenTelemetry for knowing what your system is actually doing. Istio for service-to-service security, OPA for policy, and the CNCF’s graduated-project process for separating the proven from the promising. This body of practice encodes thousands of hard-learned lessons about what happens to software after the demo.
But AI often ignores it. Not because the practice is wrong, but because it is expensive. Expensive in context, expensive in tokens, expensive in the number of steps between “generate” and “visible result.” Expediency wins.
So the great irony of building infrastructure today is that the default “AI-native cloud infrastructure” is not cloud native infrastructure. It is a cheaper, quicker, less secure, less performant, dumbed-down version of it, cobbled together to give vibe coders a quick dopamine hit. And the parts that get dropped are exactly the parts a demo never exercises. Things like mutual TLS between services, least-privilege identity, autoscaling tuned against real load, and telemetry.
But surely as this space evolves, the definition of AI-native infrastructure must be no less secure, no less scalable, and no less observable than what a good platform team builds by hand. AI-generated infrastructure that drops the standards a human team would have held is a regression, not progress. And the way to clear the bar is not to abandon the cloud native stack for whatever an agent finds most convenient. It is to make two decades of accumulated practice as cheap for an agent to use as the shortcuts are: declarative interfaces an agent can operate deterministically, policy engines that reject a bad manifest before it ships, and the same reconciliation loop that keeps a human honest keeping the agent honest too.
Vibe coders are builders too
The other thing standing between where we are now and AI-native infrastructure that is also cloud native is not technical at all.
This is actually great news. Really, it is. Let me explain.
For my whole career, the people who understood the problem best (the operations manager, the salesperson, the designer) had to pass their knowledge through filters: requirements documents, tickets, a game of telephone that stripped out half the insight before an engineer ever saw it.
Now, those filters are gone. The person with the domain experience builds the workflow themselves, and the designer ships the interaction they intended instead of an approximation of it. This has produced a lot of AI slop, but it’s also produced some software that’s closer to the user, closer to solving the problem.
If we bring the same time, skill, and attention to detail to AI-built apps that we brought to hand-built ones, vibe-coded software should end up better for its users, because it is finally shaped by the people who understand them and work with them. Sometimes it’s built by the users themselves.
But our workflows have not caught up with this. Everything about the path to production (Git, YAML, CI gates, review queues) was designed by engineers for engineers.
We used to be able to treat non-developers like children. We kept them away from anything sharp, patted their heads when they had an idea, handed them a finished product, and complained that they didn’t use it.
But non-developers are building the software now, and we’ve got to start treating them like adults… like equal participants in the development process. Getting through this period requires a collaborative framework in which vibe coders, professional developers, and SREs work as one community of builders. The domain expert supplies the intent, the platform supplies guardrails instead of gates, and operational discipline lives in the paved road rather than at a tollbooth.
BYOD has become BYOApp
If this feels threatening, I understand. It felt threatening the last time, too.
Fifteen years ago, employees started carrying their own iPhones and laptops into corporate networks, and IT departments were alarmed. They had good reasons. They would be the ones on the hook if corporate data walked out the door with a user’s laptop.
The problem was, the teams that responded by banning everything got bypassed. People used their devices anyway, invisibly, which was far worse.
The teams that adapted, with device management, sensible policies, and clear boundaries, ended up somewhere better than the locked-down world they lost. They got a more dynamic workplace where people took real responsibility for their own technology.
AI and vibe coding are at the same stage BYOD was around 2010. It’s young, messy, and growing regardless of anyone’s permission.
The cloud native community gets the same choice IT got. We can hold the new builders at arm’s length and watch a generation of software default to whatever agents find convenient. Or we can do what this community has always done best: encode hard-won operational judgment into open, shared infrastructure, and this time make it legible to agents and non-developers too.
Doron Grinstein is CEO of Control Plane, which builds AI-native cloud infrastructure on the cloud native stack. Learn more atcontrolplane.com.
Access control belongs on the same day-zero checklist as networking and storage.
On most on-prem clusters, it never makes the list.
The Identity Gap
Managed cloud Kubernetes ships IAM or SSO integration out of the box. Self-hosted clusters don’t. Access defaults to a static client certificate or a long-lived token, issued once and rarely revisited. That certificate keeps working long after the person it was issued to has left, changed roles, or lost the device it lives on. Nothing in the cluster’s authentication path checks whether they should still have access. Revoking it means finding every copy of a file, and in practice, that doesn’t happen completely. The moment more than one or two people need different levels of access, managing that per person, per file, becomes its own ongoing job.
Put an identity provider (Keycloak or any OIDC-compliant provider) in front of the cluster instead. Access should follow an account and its group membership, not a certificate file. Configure it with a public OIDC client using PKCE, not a confidential client with a secret. Access changes become identity operations: add someone to a group, remove someone from a group. No file distribution required.
Architecture: three moving parts
The integration has three components that need to agree with each other:
kubectl, with the kubelogin exec plugin. Starts the login, gets a token from the identity provider, and attaches it to every API request.
The identity provider (Keycloak). Authenticates the user and issues an ID token carrying their username and group membership.
kube-apiserver, configured with --oidc-issuer-url, --oidc-client-id, and --oidc-groups-claim. Validates the token, extracts username and groups, and lets RBAC decide what that identity can do.
kubectl authenticates against the identity provider, then presents the resulting token to kube-apiserver, which validates it and hands it off to RBAC.
kubectl never talks to the API server first. A kubectl exec-credential plugin (kubelogin, also distributed as “kubectl oidc-login”) intercepts the request, drives the browser-based login against the IdP, and hands the resulting ID token back to kubectl as a bearer credential. The API server validates that token directly against the IdP’s public signing keys. It never needs network access to the IdP itself beyond fetching those keys once.
The client configuration decision that matters
Configure this client as public, not confidential. A confidential client issues a client secret, which then gets pasted into the kubelogin plugin config, and ships to every machine that needs cluster access.
A secret that has to be distributed to every client that uses it isn’t functioning as a secret. It’s a shared static credential with extra steps, and rotating it means a coordinated config push to every machine rather than disabling one compromised identity.
OAuth 2.1 already settles this for native and command-line applications. Make the client public. Issue no secret at all. Use PKCE (Proof Key for Code Exchange) instead, which stops anyone who intercepts the authorization code from redeeming it. PKCE works by having the client generate a random value locally, send a hash of it with the initial login request, then prove possession of the original value when exchanging the code for a token. An interceptor holding only the code can’t complete that proof.
The client, in Keycloak’s admin console, ends up configured as:
Client ID: kubernetes
Client authentication: Off (public client, no secret issued)
Standard flow: On
Direct access grants: Off
Require PKCE: On, method S256
Valid redirect URIs: http://127.0.0.1:* and http://localhost:* (loopback only, nothing external)
Web origins: http://127.0.0.1:* and http://localhost:*
Client scopes: openid, profile, email, groups
General settings: the Kubernetes client, registered as OpenID Connect
Access settings: redirect URIs and web origins locked to loopback only
Capability config: Client authentication Off, Standard flow, Require PKCE On (S256)
Deployment walkthrough
1. Add the groups claim mapper
Kubernetes has no concept of “users” as a first-class object. RBAC binds to usernames and groups asserted by the token, so the IdP needs to actually put group membership into the ID token. “In Keycloak”is a protocol mapper on the client scope, of type Group Membership, mapped to the claim name groups.
Group Membership mapper on the realm-level ‘groups’ client scope: Token Claim Name ‘groups’, Add to ID token On
If Keycloak’s certificate isn’t signed by a publicly trusted CA (the common case for a self-hosted) on-prem identity provider, add one more flag pointing at that CA’s certificate:
--oidc-ca-file=/etc/kubernetes/pki/oidc-ca.crt
The API server needs to trust this connection to fetch the issuer’s signing keys. Without it, OIDC authentication fails with a TLS verification error that has nothing to do with the login flow itself, which makes it a confusing one to debug the first time you hit it.
3. Configure the kubectl side
kubeconfig gets an exec-credential entry instead of embedded certs or a static token:
Access changes now happen entirely in the IdP. Add someone to the platform-viewer group, and the next token they mint carries that group. The binding above applies immediately. No cluster-side change. No new kubeconfig to distribute.
Try it out
Using kubelogin as the exec plugin, a first login looks like this from the terminal:
$ kubectl get pods
Opening in existing browser session.
NAME READY STATUS RESTARTS AGE
web-7f9c9c4d8-2xk9p 1/1 Running 0 3d
The first call opens a browser window against the IdP; every call after that reuses the cached token until it expires, at which point kubelogin silently uses the refresh token to get a new one without another browser round-trip.
To confirm what identity and groups actually landed in the token:
$ kubectl auth whoami
ATTRIBUTE VALUE
Username jane.doe@example.com
Groups [platform-viewer system:authenticated]
The bottom line
None of this requires reworking how the cluster runs. It is a public OIDC client, one group membership mapper, a handful of RBAC bindings, and a kubectl plugin many engineers already have installed for other clusters. The setup cost is a single afternoon, not a platform migration.
What changes is what the cluster gets in return. Access follows group membership in the identity provider instead of a certificate file, so granting or revoking a level of access becomes a group change, not a search for every copy of a file across every laptop. There is a deeper benefit too. Kubernetes can log every request that hits its API server, but that audit trail is only as useful as the identity attached to each entry. A shared kubeconfig authenticating everyone as the same generic identity, often literally ‘cluster-admin’ means every audit log entry says the same thing no matter who actually ran the command. Federate identity through an OIDC provider instead, and every request the API server logs carries the person who actually made it. The audit trail stops being a list of anonymous actions and becomes an actual record of who did what, and it costs far less to set up than most teams assume.
You’ve probably felt this one: GitHub Actions usage creeps up across your org, and your actual visibility into it doesn’t keep pace. Which workflows are slow? Which are flaky? How long are jobs sitting queued for a runner before they’ve even started doing anything? GitHub’s own insights are per-repo and shallow. There’s no cross-org view of CI health, no way to slice by team or workflow type, no way to alert when things quietly get worse. Someone eventually asks why CI took forty minutes yesterday, and the honest answer is “let me go check that one repo and get back to you.”
So let’s actually fix that properly, without asking a single team to touch a single workflow file.
The obvious fix doesn’t scale
The instinctive answer is to instrument each workflow: add a tracing step, wire up an SDK, sprinkle spans through the YAML. It works, technically. But it means every team has to opt in, every new repo starts blind until someone remembers to add it, and you end up maintaining instrumentation scattered across however many workflow files exist across the org. That approach doesn’t scale with your org, it scales with how diligent everyone stays about something that isn’t their actual job.
The actual insight
GitHub already knows almost everything you want. Every workflow run and every job inside it fires an event: workflow_run and workflow_job. You don’t need to ask each repo to report on itself. You just need to listen to what GitHub is already telling you, at the org level, once.
It’s like trying to track a whole apartment building’s water usage by asking every tenant to self-report their reading. Most will forget. New tenants won’t even know they’re supposed to. The easier answer is to read the one meter at the street, where every pipe in the building already converges, whether the tenants know it’s there or not.
How it actually works
An OpenTelemetry Collector with the githubreceiver component sits behind a single org-level GitHub webhook and converts incoming workflow_run and workflow_job events straight into OTLP spans. Worth knowing upfront: it’s a contrib component still at alpha stability, so the config surface can shift. Pin a specific collector version rather than tracking latest, and skim the changelog before bumping it, cheap insurance against a config field quietly changing shape under you. If tracing is new to you, the mapping is intuitive once you see it: a workflow becomes one outer span, each job inside it a child span, each step inside a job a child of that, so what you get is something you can actually drill into rather than a flat pile of events.
One nice detail: span and trace IDs are generated deterministically, hashed from the workflow’s run ID and each job’s check run ID. If you ever want to emit your own telemetry from inside a step, there’s tooling for this, it can compute the matching ID and attach directly to the same trace without any coordination with the collector.
receivers:
github:
webhook:
endpoint: 0.0.0.0:19418
path: /events
secret: ${env:GITHUB_WEBHOOK_SECRET}
scrapers: # required even if you only want tracing, a dummy entry is enough
scraper:
github_org: ${env:GITHUB_ORG}
exporters:
otlp:
endpoint: ${env:TRACE_BACKEND_ENDPOINT}
headers:
authorization: ${env:TRACE_BACKEND_API_KEY}
service:
pipelines:
traces:
receivers: [github]
exporters: [otlp]
That scrapers block looks unrelated to tracing, and it is, it belongs to a separate GraphQL/REST metrics feature the same receiver offers, but the config fails validation without at least a dummy entry, even if all you want is the webhook side. Easy to lose twenty minutes to that the first time.
Point the exporter at Tempo, Jaeger, Datadog, or whatever your team already pays for, and it just works, that’s the actual point of using OTLP rather than a vendor-specific format. The backend is genuinely the least interesting decision in this whole setup.
Before you deploy this
A few things worth deciding upfront, since the config alone won’t force you to think about them.
The collector needs a publicly reachable endpoint, GitHub has to deliver webhooks to it, so plan for IP allowlisting or a WAF restricted to GitHub’s webhook source ranges rather than leaning on the shared secret as your only line of defence. There’s also a GitHub App option if you’d rather not manage a shared secret directly, worth a look if secret rotation across many services is already a headache for your team.
Setting up an org-level webhook needs org admin access, worth confirming early rather than discovering it mid-rollout. And if you’re on GitHub Enterprise Server rather than github.com, I’d validate that webhook delivery behaves the same way in your setup before assuming this is a drop-in.
None of that is difficult, it’s just easy to skip past when you’re excited about the zero-instrumentation part. Get it sorted early and it’s a one-time cost: point one org-level webhook at the collector and every repo in the org is covered from that moment on, including repos that don’t exist yet. Nobody has to remember to switch anything on.
Sizing it before you build it
Worth walking through the actual method here, since “just turn on tracing for everything” is a good way to end up with an unwelcome bill or an unwelcome conversation with whoever owns your tracing budget.
Start by scanning the org for total repo count, then immediately throw that number away. It’s nearly meaningless on its own. Most orgs of any real size are carrying a long tail of dormant, forked, archived, or abandoned repos that inflate the headline count without generating any real CI traffic. What actually matters is the active slice: how many repos had genuine workflow activity in a real week, not how many exist.
From there, extrapolate outward. Active repo count times average runs per repo gives you expected workflow volume. Workflow volume times average steps per workflow gives you expected span volume. Span volume times typical payload size gives you an expected data volume per day. Compare that against whatever tracing volume your infrastructure already handles for application traces, and in most orgs, CI trace volume turns out to be a rounding error next to it.
That comparison is the actual point, not any specific number I could hand you. What generalises is the method: measure real activity instead of headline repo count, and walk in with a comparison rather than an assertion.
Getting buy-in
The technical build was the easy part. The harder part was justifying the data volume to whoever owns the tracing budget, especially with cost concerns already floating around about the backend in question. Showing up with an actual sizing exercise, not “trust me, it’s small,” turns that into a five-minute conversation instead of a drawn-out one. Nobody has to take your word for “it’s small” when they can see it sitting next to the tracing volume they’re already paying for without blinking.
It’s also worth remembering that standing up a new observability project is as much an ownership question as a technical one. Someone has to actually own the collector, the webhook, the alerting rules going forward. Sorting that out early saves the awkward moment three months later when something breaks and nobody’s sure whose pager it is.
Why this is worth doing
The zero-instrumentation part is the whole value here. New repos are observable the moment they’re created, not the moment someone remembers to add tracing to them. And because everything lands as proper OTel traces, CI health sits in the same tool as your application traces, so a slow deploy and a slow downstream service can be correlated instead of investigated in two different dashboards by two different people who don’t talk to each other until Thursday.
If you’re running self-hosted runners on the Actions Runner Controller, it’s worth being clear this doesn’t replace what ARC already gives you, it sits on a different layer entirely. ARC’s own metrics tell you about your runner fleet: how many pods exist, whether autoscaling is keeping up, how deep the queue is. This tells you about your workflows: why a specific run was slow, which ones are flaky, where the time actually went. Queue depth and queue time are practically the same question asked from two different vantage points. Worth running both, not picking one.
This covers the collection side. It doesn’t get into building good alerting on top of the trace data, that’s queue-time thresholds, flaky-test detection, and how noisy those alerts get before people start ignoring them, which is a genuinely separate problem and probably its own post.
This recipe is aimed at small and medium non-security focused projects. Maintainers of a high-risk security sensitive project, you probably need a more sophisticated approach (check out https://alpha-omega.dev/)
Scope (ingredients)
Defining SECURITY.md, managing reports and embargoes, and CVE disclosure.
Out of scope (do not add)
Large projects and those that are security-criticalVulnerabilities in dependenciesCoding best practices (ie how to avoid a vulnerability)
TLDR:
Set the Station: Define a clear reporting path in your README.md and SECURITY.md so reporters know where to deliver ingredients.
Taste Test First: Determine if a report is a real vulnerability or just a regular bug before starting the fire.
Plate for Everyone: Coordinate the patch release with a public CVE disclosure to serve your users safely.
The Recipe: Step-by-Step Instructions
This recipe card provides guidance for handling security vulnerabilities in open source projects. A security vulnerability is a flaw in a program that can be exploited by an attacker to compromise the confidentiality, integrity, or availability of a system. Security vulnerabilities often stem from bugs in a program, but not all bugs can be exploited. While many bugs appear in the kitchen, not all of them have the potential to compromise the confidentiality or integrity of your menu.
Spoiled ingredients can cause a lot of damage for patrons, in the same way that vulnerabilities can cause a lot of damage for end-users. So many users track & prioritize vulnerabilities to minimize impact. This process can generate a lot of work, so in this guide we aim to minimize the extra work needed for both project maintainers and end-users.
Preparing to receive a vulnerability (prep the kitchen)
As a head chef, you want to hear from your customers when something is wrong, but you want them to reach out to you directly with the information, instead of posting it as a public review. Similarly, as a maintainer, you don’t want vulnerabilities to be reported to you through a plain bug issue, publicly visible to the world. Make it as obvious and as easy as possible for people to find that path by including a security section directly in the README.md of the repository. This section can directly include the instructions, or point to the SECURITY.md file, which should also be at the root of the repository.
The instructions for reporting a vulnerability should include:
Your bug bounty policy: if one exists. For small/medium sized projects, it is ok to not offer a bounty in exchange for finding vulnerabilities on your project.
After you get a report (the taste test)
Once you discover a vulnerability or have one reported, you will want to respond to the vulnerability without talking publicly about the vulnerability. This involves an “embargo” where key individuals from the project, usually maintainers, are made aware of the vulnerability and work with the reporter to address it.
Just as the first step with a questionable ingredient is to determine if it is actually spoiled, the first step when you get the report is to determine if it’s actually a vulnerability. Some reports may turn out to be a non-exploitable bug (meaning it can’t lead to an exploit) or a misunderstanding of the project’s expected behavior. To decide, discuss the report with the reporter and relevant project experts. Try to minimize the number of people involved in the discussion at this stage, and make sure everyone agrees to maintain confidentiality until the report becomes public. If you decide the report is not a vulnerability, direct the reporter to file a public issue instead, and you may follow any of the procedures in “publishing” to inform users. Think of it as adding a note to the menu that specific dishes contain cilantro.
If you are not sure whether a report is a bug, you can reach out to TAG Security and Compliance or CNCF staff for guidance. When reaching out, ask generally for guidance, or for the TAG leads to start a DM. Make sure to keep the contents of the report out of public channels at this stage. You can also consider the following questions:
Is there a mitigation mechanism for users?
Can this bug lead to compromise, data leaks, or other malicious activity?
Is this a documentation error?
Use your judgement when adding someone to an embargo, and make sure they are aware that the vulnerability is not yet public, and that it needs to be handled appropriately and kept confidential, given the sensitivity of the matter.
Just as a chef modifying a recipe avoids having too many extra cooks in the kitchen who might speak too loudly and be overheard, you should carefully manage who is informed about an embargo. Be intentional with your disclosures, involving only those essential contributors whose direct assistance is required to resolve the vulnerability. The embargo ensures that attackers do not have access to the vulnerability before it is patched.
Fixing a Vulnerability (cooking)
When patching a vulnerability, keep your preparation out of the public dining room. This means keeping the patch out of public conversations and pull requests until after it is public. People can guess at the vulnerability from the patch, and you want to give users a chance to update before the vulnerability becomes public. To do:
If you are using the GitHub vulnerability reporting mechanism, you can create a private branch from the vulnerability report. If not, you can develop and review the patch over a private channel.
Test the fixes! Your existing CI may not work with private branches, so do your best to run tests locally. You are the maintainer, so use your best judgment about what testing needs to be done for the scope of the patch.
Merge fast, be coordinated! Once you are confident that the patch fixes the vulnerability and is adequately tested, work to merge it fast. This is especially important if you are not using private branches. Ensure that enough maintainers were involved in the patch creation and testing to ensure speedy response. You can coordinate this with the publishing process to ensure that the patch and vulnerability are public at the same time. Just remember to check the code before publicizing the fix. Make sure to re-apply the taste test to ensure that the vulnerability is fixed.
Publishing the Disclosure (plating and serving)
Once the patch is fully baked, you want to plate it up and tell your users about it. It’s generally considered best practice to do this within 90 days of receiving the report.
One strategy is to have a list of users to disclose to privately. These users can be added to your embargo and patch before the vulnerability is public. But this requires long-term maintenance of this list of users. If this list already exists, or if you have the resources to maintain it, this is great, but most small projects will not be able to keep a list of users and their contact information up to date.
Just as food needs to be plated just before service, a common strategy is to make the patch available and announce the vulnerability at the same time so that users can patch quickly. This gives everyone the patch at the same time, and is much easier to coordinate a safe meal. To do so:
Once the patch is landed in the codebase, you’ll need to build a new release. Use your normal process, but do so right away.
You also need to publish a CVE. If you used GitHub private reporting, you can do so from within the GitHub UI for the report.
To get a CVE number, it must be assigned by a CNA (CVE Numbering Authority). GitHub is a CNA and can handle this for you, or you and the reporter can report directly to one of the listed CNAs to get a number.
The CVE severity score is assigned based on a list of questions. Work with the reporter to ensure the answers are accurate and that you agree on the assigned severity score.
Once published, CVEs go into OSV and other vulnerability databases. It will then show up in users’ vulnerability scans, which will help them know to update to your latest release.
Consider publishing any proof of concept after the fix has been available for a while. This may be hard to do with private vulnerability reporting, but can give users extra time to patch.
AI infrastructure conversations often start with GPUs. Accelerators provide much of the compute behind model training and inference, so the focus is understandable.
But a production AI workload rarely starts and ends on a GPU. Data needs to be prepared and moved. Applications and orchestration services need to run. Models need to be loaded and served. Results may require additional processing.
Platform teams are therefore not simply managing GPU workloads. They are managing heterogeneous workloads that depend on CPU, GPU, memory, storage, and networking working together. For Kubernetes platform teams, the challenge is not just providing accelerators. It is matching the right resources to each stage of the workload.
Follow the workload, not the GPU
Consider a simplified AI inference pipeline:
Data → CPU preprocessing → GPU inference → CPU post-processing → application
The GPU may perform the most compute-intensive step, but overall performance depends on the complete path. If preprocessing cannot supply data quickly enough, the accelerator waits. If storage cannot deliver model artifacts efficiently, startup slows. If CPU, memory, or network capacity becomes constrained, adding more GPU capacity may do little to improve throughput.
Instead of asking: How many GPUs does this workload need?, platform teams should ask:
What resources does each stage need, and where are the dependencies between them?
That shift helps teams optimize the workload as a system rather than optimizing one expensive component in isolation.
Match resources to the work
Different stages of an AI workload have different infrastructure requirements.
CPU resources can handle data preparation, tokenization, retrieval, orchestration, application logic, and post-processing. GPUs and other accelerators are suited to highly parallel operations such as model training and inference. Memory, storage, and networking determine how efficiently data and model artifacts move between these stages.
Even inference itself is not necessarily one uniform workload. For large language models, prompt processing and token generation can have different compute and memory requirements. This creates an opportunity for platform teams to match resources to the work rather than forcing an entire AI pipeline onto a single infrastructure profile.
Kubernetes provides a common orchestration layer for doing this. Dynamic Resource Allocation (DRA), for example, extends Kubernetes’ resource model by providing a more flexible, declarative way for workloads to request specialized devices. The important point is not DRA itself. It is the direction: specialized compute is increasingly part of the same cloud-native resource model as the rest of the application.
Observe the handoffs
Heterogeneous infrastructure also changes what platform teams need to observe. GPU utilization alone does not tell you whether an AI workload is running efficiently. Low GPU utilization could indicate insufficient demand. But it could also mean the accelerator is waiting for CPU preprocessing, data access, scheduling, or another upstream dependency.
Platform teams, therefore, need visibility across the complete workload:
CPU → data → accelerator → application
Correlating infrastructure and application telemetry makes it easier to identify where time is being spent and which resource is limiting performance. The objective isn’t to keep every resource at 100% utilization. It is to understand whether those resources are working together efficiently enough to meet the workload’s performance requirements.
Design for the whole system
As AI workloads move into production, infrastructure is likely to become more heterogeneous, not less. Kubernetes provides platform teams with a common control plane across these resources, while capabilities such as DRA are expanding the ways specialized hardware can participate in that model.
The key shift is conceptual: AI infrastructure is not a collection of GPUs with supporting services around them. It is a system of interconnected compute, memory, storage, and network resources.
For platform engineers, designing around that complete system, not one component, is what turns accelerator capacity into useful AI infrastructure.
Kubernetes isn’t brand new anymore. Yet, for many teams, adopting it still feels intimidating. Even if you’ve watched Kubernetes become the default foundation for production software and AI workloads, it can still feel like a big leap when you’re the one making the call.
Recently, we’ve seen a wave of organizations making the jump, with AI now one of the primary drivers of Kubernetes usage and growth. While K8s has matured significantly over the years, stepping into it for the first time is still a major shift.
Today’s AI stacks add GPUs, bursty traffic, and stricter data boundaries, making Kubernetes start to feel like an entirely new operations discipline (even for teams already accustomed to deploying on K8s). Training requires massive bursts of compute. Inference demands clean scaling and automatic recovery. Data pipelines need a consistent control plane sitting right next to the rest of your application stack.
Most AI teams don’t start on Kubernetes, even though that’s usually where their infrastructure ends up. At some point, training jobs, inference services, and data pipelines need a real production environment, and the platform conversation comes up fast. For a lot of teams, that conversation is about ownership: who runs the cluster, who manages shared services, and who makes sure AI workloads don’t break everything else.
Getting a Kubernetes Cluster Running Is the Easy Part
Today, getting a basic cluster running is easier than ever. You can kick the tires with managed Kubernetes offerings like GKE, AKS, and EKS. Standing up a Kubernetes cluster isn’t the hardest part by any means.
The real test comes when that infrastructure has to carry production AI workloads without blowing through your GPU budget, starving other applications in the cluster, slowing down core services, or compromising security. You have to actively manage job placement, keep GPUs utilized rather than idling expensively, and enforce guardrails so platform stability doesn’t crumble when experiments go wrong.
It reminds me of what it was like moving to Linux for the first time. Linux is incredible once you get used to it. But if all you’ve ever known is Windows, it feels like an entirely different universe.
Live CDs were an easy way in. You could pop a CD into your drive, boot into a new Linux distribution on your actual hardware, and see how it all worked before committing to a full, risky installation (which, back then, always meant the stressful process of re-partitioning your hard drive).
Most teams want the same thing with Kubernetes and AI: a way to see how everything behaves on real infrastructure before they commit to owning the platform themselves.
Embracing the New Shift
Stepping into Kubernetes to meet today’s AI demands can feel overwhelming. But new and unfamiliar doesn’t have to mean risky. Once you adjust to the shift and have the right foundation in place, the power and flexibility it offers your infrastructure is hard to match.
If you’re planning AI workloads on Kubernetes or want an assessment of your current environment, reach out and we’ll walk through how we can support your business goals with a managed Kubernetes platform as you deploy workloads on Kubernetes.
There’s still time to join OSPOlogy + OSPO Summit China 2026, taking place on September 7, 2026, in Shanghai, China as part of KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China.
This event will focus on the evolution of Open Source Program Management (OSPO) functions and open source governance in the era of Agentic AI. It is an opportunity to explore some of the questions and topics emerging at the intersection of open source, AI governance, agentic AI adoption, and strategy in the enterprise
If you’re joining us in Shanghai for KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China, here’s what you can expect from OSPOlogy + OSPO Summit.
Explore OSPOs in the era of Agentic AI
As organizations navigate the Agentic AI era, OSPOlogy + OSPO Summit China will explore how Open Source Program Management functions and corporate open source governance are evolving.
The program will cover key topics including:
How Agentic AI can support practical OSPO activities
The role of OSPOs in AI and data governance: from open source AI policies and model usage to licensing, provenance, compliance, and responsible adoption in organizations
AI-driven software supply chains: security, transparency, and trust
Supporting cross-border open source strategies
Open source strategy as a driver of organizational innovation and business value
Together, these topics put the spotlight on the evolving relationship between open source program management, AI governance, agentic AI adoption, and enterprise strategy.
Plan your day
OSPOlogy + OSPO Summit China takes place on September 7 in Shanghai, with session times listed in China Standard Time (CST / UTC+8).
Attendees can explore the event schedule in advance and build a personal agenda for the day.
Please note that seating is first come, first served, so plan ahead for the sessions you’d most like to attend.
Can’t make every session? Sessions will be recorded and made available on the CNCF YouTube channel within two weeks of the event.
There’s still time to register
Late registration for OSPOlogy + OSPO Summit China is available through September 7 at:
$30 USD / ¥205
OSPOlogy + OSPO Summit China is an add-on option. All attendees must also be registered for KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China.
With Agentic AI creating new considerations for open source program management and corporate open source governance, OSPOlogy + OSPO Summit China offers a dedicated forum to explore these topics.
Join us on September 7, 2026, in Shanghai for OSPOlogy + OSPO Summit China.
Somewhere in your cluster there’s probably a deployment sitting in the default namespace that everyone knows shouldn’t be there. Nobody put it there maliciously, it just happened, early on, before anyone had opinions about namespace hygiene, and now half your other services quietly depend on it. Moving it is now a tricky problem.
That was the exact situation with a service I’ll call auth-svc: an authentication service that dozens of other services called constantly, sitting in default for years, and about to become a genuine problem the moment it needed namespace-scoped things, its own ingress rules, its own policies, that default structurally couldn’t give it. Moving it wasn’t optional forever. But it also couldn’t go down, not even for a few seconds. This wasn’t a vague “other services might complain” risk: auth-svc handled authentication for that entire region’s cluster, so if it went down, nobody in that region could log in. Full stop.
Before getting into why this is actually hard, it’s worth being precise about what “moving it” means. There are two completely separate paths into auth-svc, and both have to keep working throughout the move, or fixing one just creates an outage in the other. Everything inside the cluster reaches it the ordinary way: other services resolve auth-svc.default.svc.cluster.local through Kubernetes’ own internal DNS and get routed to a pod, the standard Service mechanism. Everything outside the cluster reaches it through an ingress instead, a completely separate mechanism that has nothing to do with that DNS name. Whatever the fix turned out to be, it had to solve for both paths, not just the one that’s easier to reason about.
Why “just move it” doesn’t work
The obvious plan, move the deployment, update the references, done, falls apart the moment you look at who’s actually calling this thing. As mentioned, dozens of other services reference auth-svc by its cluster-internal DNS name, owned by different teams, on different release cycles. There’s no atomic moment where you flip a switch and every one of them simultaneously starts using a new name. Some team’s service hasn’t been redeployed in months. You shouldn’t be coordinating that.
The tooling got in the way too. Our deploy pipeline only knew how to ship a service to one namespace. There was no “deploy this to two places at once” option, and modifying the shared pipeline logic every other team also depended on felt like exactly the kind of blast radius we didn’t want to introduce. Whatever the fix was, it had to fit inside a single-namespace deploy, not require rewriting shared infrastructure.
On top of that, we had an OPA policy which enforced that identical ingress rules couldn’t exist live in two namespaces at once, a sane rule that exists specifically to stop the kind of half-finished migration that leaves routing ambiguous.
The insight: a forwarding address
The piece that made this solvable: I didn’t need to migrate every consumer’s understanding of where auth-svc lives. I needed to migrate the service, and quietly redirect anyone still asking for the old address.
Kubernetes has exactly this mechanism, and it’s easy to forget it exists because you almost never need it: an ExternalName service. Instead of pointing at pods, it points at another DNS name, functioning essentially like a CNAME (see my DNS article if you want to learn more about how that works). Deploy the real thing at its new home, then convert the old Service object into a forwarding address:
Every consumer still calling auth-svc.default.svc.cluster.local gets silently redirected to the real thing in its new namespace. Nobody changes a line of code on their end. It’s the same trick as a postal forwarding order: you don’t visit every person who might send you mail and update their address book, you tell the post office where you actually live now, and everything gets redirected until people eventually update it themselves, at their own pace, with zero coordination required on your part.
That forwarding trick only earns its keep if you actually confirm it’s working before you lean on it. Once the proxy was live, the next step wasn’t scaling anything down, it was watching the metrics through the crossover: checking that traffic hitting the old address was genuinely landing on the new deployment, not silently failing or looping somewhere. Only once that looked clean did the old pods get scaled to zero rather than deleted outright. Scaling to zero costs nothing and buys an instant rollback, just scale back up, if anything downstream looked wrong later. Deleting them outright would have meant rebuilding from scratch if something went sideways, so there was no reason to give up that safety net early. We could defer the cleanup to a later point in time.
The chicken-and-egg problem
The remaining wrinkle was the ingress. External traffic to auth-svc doesn’t come in through the DNS-based Service mechanism at all, it comes in through an ingress, and I needed a working ingress in the new namespace before I could safely remove the one in the old namespace. But the policy engine wouldn’t allow both to exist at once; identical ingress rules across two namespaces is exactly the ambiguous state it exists to prevent.
Classic chicken-and-egg: can’t create the new one without a policy exception, can’t delete the old one first without a traffic gap.
The fix was a temporary, explicit exception rather than fighting the policy itself: annotate the new namespace to bypass the duplicate-ingress check just for this migration, stand up the new ingress alongside the old one for a short overlap window, confirm traffic was flowing correctly to the new deployment, then delete the old ingress and let the exception age out. A brief, deliberate window where both existed, rather than a gap where neither did.
None of this went straight to production. It ran in dev first, then a staging cutover, with a couple of weeks between the staging success and doing it for real, mostly just to sit with it and see whether anything subtle showed up under real traffic before betting a critical path on it.
The production cutover itself ended up being the boring part. All the actual difficulty was front-loaded into getting the design right. Once the plan was solid, executing it was closer to a formality than an event.
Why this matters beyond one service
Here’s the thing about default namespace sprawl: it’s rarely caused by carelessness. It’s caused by the complete absence of pressure to ever fix it. Someone creates a policy restricting new services from landing in default, good practice, but nobody sets a deadline or a plan for the services already there. They just sit. For years, in this case. Nothing forces the issue until a team needs something namespace-scoped that default structurally can’t give them, and only then does the debt come due.
If you’re staring at a similarly stuck service, the pattern generalises past this one migration. An ExternalName proxy buys you a zero-coordination path to move anything addressed by DNS, provided you’re willing to hold two versions in careful overlap for a short, deliberate window rather than trying to cut everything over at once. The next time someone tells you a service can’t be moved because too many things depend on it, that’s usually a sign nobody’s looked for the DNS-shaped seam it can be split along.
This only works cleanly because everything here talked to auth-svc through its DNS name rather than a hardcoded IP or ClusterIP. If you’ve got consumers that skip DNS entirely, and some legacy systems do, you’re solving a different, uglier problem.