❌

Vue lecture

Kubernetes access via an identity provider: Public client, not confidential

Access control belongs on the same day-zero checklist as networking and storage.

On most on-prem clusters, it never makes the list.

The Identity Gap

Managed cloud Kubernetes ships IAM or SSO integration out of the box. Self-hosted clusters don’t. Access defaults to a static client certificate or a long-lived token, issued once and rarely revisited. That certificate keeps working long after the person it was issued to has left, changed roles, or lost the device it lives on. Nothing in the cluster’s authentication path checks whether they should still have access. Revoking it means finding every copy of a file, and in practice, that doesn’t happen completely. The moment more than one or two people need different levels of access, managing that per person, per file, becomes its own ongoing job.

Put an identity provider (Keycloak or any OIDC-compliant provider) in front of the cluster instead. Access should follow an account and its group membership, not a certificate file. Configure it with a public OIDC client using PKCE, not a confidential client with a secret. Access changes become identity operations: add someone to a group, remove someone from a group. No file distribution required.

Architecture: three moving parts

The integration has three components that need to agree with each other:

  • kubectl, with the kubelogin exec plugin. Starts the login, gets a token from the identity provider, and attaches it to every API request.
  • The identity provider (Keycloak). Authenticates the user and issues an ID token carrying their username and group membership.
  • kube-apiserver, configured with --oidc-issuer-url, --oidc-client-id, and --oidc-groups-claim. Validates the token, extracts username and groups, and lets RBAC decide what that identity can do.
Figure 1: kubectl authenticates against the identity provider, then presents the resulting token to kube-apiserver, which validates it and hands it off to RBAC.

kubectl authenticates against the identity provider, then presents the resulting token to kube-apiserver, which validates it and hands it off to RBAC.

kubectl never talks to the API server first. A kubectl exec-credential plugin (kubelogin, also distributed as “kubectl oidc-login”) intercepts the request, drives the browser-based login against the IdP, and hands the resulting ID token back to kubectl as a bearer credential. The API server validates that token directly against the IdP’s public signing keys. It never needs network access to the IdP itself beyond fetching those keys once.

The client configuration decision that matters

Configure this client as public, not confidential. A confidential client issues a client secret, which then gets pasted into the kubelogin plugin config, and ships to every machine that needs cluster access.

A secret that has to be distributed to every client that uses it isn’t functioning as a secret. It’s a shared static credential with extra steps, and rotating it means a coordinated config push to every machine rather than disabling one compromised identity.

OAuth 2.1 already settles this for native and command-line applications. Make the client public. Issue no secret at all. Use PKCE (Proof Key for Code Exchange) instead, which stops anyone who intercepts the authorization code from redeeming it. PKCE works by having the client generate a random value locally, send a hash of it with the initial login request, then prove possession of the original value when exchanging the code for a token. An interceptor holding only the code can’t complete that proof.

The client, in Keycloak’s admin console, ends up configured as:

  • Client ID: kubernetes
  • Client authentication: Off (public client, no secret issued)
  • Standard flow: On
  • Direct access grants: Off
  • Require PKCE: On, method S256
  • Valid redirect URIs: http://127.0.0.1:* and http://localhost:* (loopback only, nothing external)
  • Web origins: http://127.0.0.1:* and http://localhost:*
  • Client scopes: openid, profile, email, groups
Figure 2: General settings: the Kubernetes client, registered as OpenID Connect

General settings: the Kubernetes client, registered as OpenID Connect

Figure 3: Access settings: redirect URIs and web origins locked to loopback only

Access settings: redirect URIs and web origins locked to loopback only

Figure 4: Capability config: Client authentication Off, Standard flow, Require PKCE On (S256)

Capability config: Client authentication Off, Standard flow, Require PKCE On (S256)

Deployment walkthrough

1. Add the groups claim mapper

Kubernetes has no concept of “users” as a first-class object. RBAC binds to usernames and groups asserted by the token, so the IdP needs to actually put group membership into the ID token. “In Keycloak”is a protocol mapper on the client scope, of type Group Membership, mapped to the claim name groups.

Figure 5: Group Membership mapper on the realm-level 'groups' client scope: Token Claim Name 'groups', Add to ID token On

Group Membership mapper on the realm-level ‘groups’ client scope: Token Claim Name ‘groups’, Add to ID token On

2. Point kube-apiserver at the issuer

--oidc-issuer-url=https://<your-keycloak-host>/realms/<realm>
 --oidc-client-id=kubernetes
 --oidc-username-claim=preferred_username
 --oidc-groups-claim=groups

If Keycloak’s certificate isn’t signed by a publicly trusted CA (the common case for a self-hosted) on-prem identity provider, add one more flag pointing at that CA’s certificate:

--oidc-ca-file=/etc/kubernetes/pki/oidc-ca.crt

The API server needs to trust this connection to fetch the issuer’s signing keys. Without it, OIDC authentication fails with a TLS verification error that has nothing to do with the login flow itself, which makes it a confusing one to debug the first time you hit it.

3. Configure the kubectl side

kubeconfig gets an exec-credential entry instead of embedded certs or a static token:

users:
 - name: oidc
   user:
 	exec:
   	apiVersion: client.authentication.k8s.io/v1
   	command: kubectl
   	args:
     	- oidc-login
     	- get-token
     	- --oidc-issuer-url=https://<your-keycloak-host>/realms/<realm>
     	- --oidc-client-id=kubernetes

No secret field. There’s nothing to put there.

4. Bind groups to RBAC

apiVersion: rbac.authorization.k8s.io/v1
 kind: ClusterRoleBinding
 metadata:
   name: platform-viewers
 subjects:
   - kind: Group
 	name: platform-viewer
 	apiGroup: rbac.authorization.k8s.io
 roleRef:
   kind: ClusterRole
   name: view
   apiGroup: rbac.authorization.k8s.io

Access changes now happen entirely in the IdP. Add someone to the platform-viewer group, and the next token they mint carries that group. The binding above applies immediately. No cluster-side change. No new kubeconfig to distribute.

Try it out

Using kubelogin as the exec plugin, a first login looks like this from the terminal:

$ kubectl get pods
 Opening in existing browser session.
 NAME                    	READY   STATUS    RESTARTS   AGE
 web-7f9c9c4d8-2xk9p     	1/1 	Running   0      	3d

The first call opens a browser window against the IdP; every call after that reuses the cached token until it expires, at which point kubelogin silently uses the refresh token to get a new one without another browser round-trip.

To confirm what identity and groups actually landed in the token:

$ kubectl auth whoami
ATTRIBUTE   VALUE
 Username	jane.doe@example.com
 Groups  	[platform-viewer system:authenticated]

The bottom line

None of this requires reworking how the cluster runs. It is a public OIDC client, one group membership mapper, a handful of RBAC bindings, and a kubectl plugin many engineers already have installed for other clusters. The setup cost is a single afternoon, not a platform migration.

What changes is what the cluster gets in return. Access follows group membership in the identity provider instead of a certificate file, so granting or revoking a level of access becomes a group change, not a search for every copy of a file across every laptop. There is a deeper benefit too. Kubernetes can log every request that hits its API server, but that audit trail is only as useful as the identity attached to each entry. A shared kubeconfig authenticating everyone as the same generic identity, often literally ‘cluster-admin’ means every audit log entry says the same thing no matter who actually ran the command. Federate identity through an OIDC provider instead, and every request the API server logs carries the person who actually made it. The audit trail stops being a list of anonymous actions and becomes an actual record of who did what, and it costs far less to set up than most teams assume.

  •  

Platform engineering maturity: From toolchain to self-service

Most platform engineering conversations tend to split into two rooms pretty quickly.

The first room is full of teams who don’t have a platform yet. Scattered scripts, tribal knowledge, and every team is doing the same task in a different way. The teams know something needs to change, but building a platform feels like a six-month project nobody has budgeted for.

The second room has already shipped a platform. There is a golden path, developer portal, a CLI and even AI agent in some cases. Adoption looks reasonable from the outset, but the platform team is still handling requests by hand, which becomes the bottleneck for anything outside the paved road, still wondering why “self-service” hasn’t actually reduced their workload.

These two might look like opposite problems, but they’re not. Both rooms are describing the same thing: they don’t know what the next stage of their platform interface looks like.

The first room thinks the answer is to build a platform. The second room believes the answer is to add more capabilities to the platform. Neither of them is wrong.

The right question isn’t “Do we have a platform?” It’s “How do developers actually interact with what we’ve built?” That gap between capabilities that exist and capabilities that are genuinely self-serviceable is an interface maturity problem.

In this post, we will look at the CNCF Platform Engineering Maturity Model that defines four stages of that journey. We will break down each stage but look specifically at interfaces and understand why most teams plateau at Stage 2 without realizing it, and what the path forward actually looks like.

Understanding The CNCF Platform Engineering Maturity Model

The CNCF Platform Engineering Maturity Model defines five aspects of platform engineering maturity: Investment, Adoption, Interfaces, Operations, and Measurement. Each is scored independently. An organization does not move through the model as a whole – it moves through each aspect on its own timeline, at its own pace.

Each aspect has four levels: Provisional, Operational, Scalable, and Optimizing. The model is a diagnostic framework that tells you where you are, but it does not tell you how to get to the next stage.

We’ll focus on one aspect of this maturity model – Interfaces. How developers actually interact with platform capabilities – the forms, the CLIs, the portals, the APIs – and why most teams stall at Level 2 without realizing it.

The Four Stages of Interfaces Maturity

The Interfaces aspect of the CNCF platform maturity model describes how developers interact with and consume platform capabilities. It has four levels where each one reflects how much the platform team still needs to be in the loop for things to happen.

Level 1: Custom Processes

Level 1 is custom processes, which consists of a collection of varying processes with no consistency of interface. Capabilities are provisioned through manual requests; knowledge is shared from person to person, and deep support from the capability provider is usually required to get anything done.

In practice, this is where most teams without a formal platform already live, whether they recognize it or not. The scripts, the runbooks, the “ask Jessica, she knows how to set up the database” culture. All of these constitute a Level 1 interface. The absence of a named platform does not mean the absence of a stage.

Level 2: Standard Tooling

The CNCF model describes Standard Tooling as consistent, standard interfaces for provisioning and observing capabilities. Golden paths and paved roads exist in some form. There are documentation and templates, so users can identify what is available and request it.

This is where most teams who have “built a platform” actually are. And it looks like success because adoption numbers improve, onboarding gets faster, and the metrics move in the right direction. However, everything outside the paved path still requires a human from the platform team to implement it. The interface is standardized but it is not self-sufficient.

Level 3: Self-Service Solutions

Level 3 is for self-service solutions, where there is genuine autonomy for users, requiring little support from maintainers. One-click provisioning for most of the asks where the platform team is not in the loop. Most of the routine tasks are great entry points to start the self-service journey.

The signal here is behavioral, not metric-based. Teams stop filing tickets for routine provisioning and start checking the internal platform first. Newly hired engineers ship their first meaningful change within days rather than weeks. The clearest confirmation comes from the backlog – organizations that reach Level 3 report exception requests drop by 40-60% after adding self-service configuration options. The platform team’s work shifts from actioning individual requests to improving the framework that handles them.

Level 4: Integrated Services

At level 4, integrated services, the platform capabilities are transparently integrated into the tools and processes teams already use. Some capabilities are provisioned automatically. The interface becomes invisible until you need to go deeper.

The sign that you are at level 4 of platform maturity is an absence of conversation. Developers stop thinking about infrastructure entirely because everything is bolted on the platform. When a new service is created, monitoring, logging, and security are integrated automatically – a developer doesn’t need to explicitly wrestle with these configurations. The security team defines policies that the platform enforces without a developer negotiating them. The observability team builds the capabilities which integrate automatically. The platform team’s success is measured by how rarely anyone mentions the platform.

Where Most Teams get Stuck

Self-service means a developer can get what they need without the platform team in the loop. Standard tooling means a developer can get what the platform team anticipated they would need, with the platform team standing by for everything else.

They might look and feel identical at level 2, but at scale, they diverge completely.

Based on interactions with organizations across industries, here’s why teams are stuck.

Queue problem

Golden paths cover the common cases that teams face. They do not cover the edge cases – and in any organization of meaningful size, edge cases are not edge cases. They are 30% of the work. Every request that falls outside the golden path lands on the platform team’s desk, increasing the backlog. The team that was supposed to reduce toil becomes the source of it.

At a discussion during a platform engineering round table, I spoke to a group that worked with a retail organization, they built a golden path for Kubernetes deployments using Helm charts and ArgoCD. Within six months, 85% of teams were using it. But the platform team’s backlog had grown from zero to 40 pending exception requests, and they were spending 60% of their time handling configurations that fell outside the golden path.

Expertise problem

Platform teams build capabilities for domains they generalize across organizations. A streaming pipeline built by a platform team without streaming expertise will work. It will not work as well as one built by the team that runs streaming workloads daily. The gap compounds over time. Specialized teams stop trusting the platform for specialized needs and building their own.

Another team that I spoke to had observed this with a financial services organization that had built a comprehensive internal developer platform with self-service infrastructure provisioning – documentation, office hours, extensive guides. Teams could follow the platform but they could not extend it. The interdependencies between Terraform modules, CI/CD pipelines, monitoring integrations, and service mesh configuration lived entirely in the platform team’s heads. Application teams had no mental model of how the components interacted. The gap compounded over time and teams started building their own capabilities.

Maintenance trap

Shipping capabilities is the easiest part. Maintaining them is the job nobody accounted for. Thirty capabilities shipped over two years means thirty capabilities to patch when a CVE drops, thirty things to test when Kubernetes upgrades, thirty surfaces where things can quietly break. The platform team that was hiring to build starts hiring to keep up with the patches and updates.

Working with an e-commerce organization, they had created shared Helm charts that abstracted Kubernetes complexity and accelerated deployments significantly. Eighteen months later, those charts had accumulated deprecated APIs, unused parameters, dependencies on specific cloud provider features, and hardcoded networking assumptions. The platform team was afraid to update them because every change required coordinated testing across dozens of applications. Application teams were afraid to customize because they would own the consequences. The golden path had become legacy code that everyone used and nobody wanted to touch.

Rigidity issue

Every golden path is built on assumptions about how work gets done. Those assumptions were accurate when the path was designed. With time, teams change, technologies change, processes change, and requirements shift. The golden path that removed friction at launch starts generating it when the organization outgrows the assumptions baked into it. Workarounds keep growing, and shadow infrastructure quietly appears until it breaks loudly.

At one of the KubeCon + CloudNativeCons, I spoke to a platform lead who had worked with a healthcare organization that standardized their Kubernetes, Istio, Prometheus, and ArgoCD deployments. The golden path assumed that stack entirely. When a team needed to deploy a legacy application that could not run in containers, or required a different database, or needed an alternative deployment pattern, the platform team built a custom exception. Then maintained it. Then built another one. The platform team became an exception factory, spending their time on one-off solutions rather than improving the core platform.

What connects these four scenarios is the same moment when the platform team realized the tools had scaled the requests without scaling the ability to handle them.

Moving past that requires a different kind of decision than the ones that got them here, and we’ll look into that transition in the next section.

The Transition: Moving Between Stages

The CNCF model is a diagnostic tool, not a prescription. It will tell you which level you are at. It will not tell you how to get to the next one. The path forward depends on what you have already built, how your organization is structured, and where the friction actually lives.

Here’s how you can climb the interface maturity ladder.

Level 1 to Level 2: Name it before you build it

The Level 1 team’s first mistake is usually building a developer portal before understanding what developers actually need. Portals are Level 3 infrastructure. At Level 1, priority is recognition before construction.

Name what already exists

Every Level 1 team has a platform – it is just unmanaged. The tribal knowledge bottleneck where one person’s absence stops work. Three teams are doing the same deployment in three different ways. The recurring Slack message is received by the same infrastructure engineer every time a database needs provisioning.

These are not gaps. They are your current interface. Mapping them – which requests are most common, which consumes the most time, which follow the same steps every time – tells you what your first golden path should be. Not what seems strategically important. The highest-volume, most repeatable, most painful manual process.

Build one thing and make it genuinely better

The principle that matters most here: a golden path is a documented, supported, opinionated way of doing one thing well. Start with one and make it genuinely better than the alternative to earn the trust before building the catalog. For a deeper exploration of golden path design and adoption patterns,read this guide to golden path implementation patterns that covers real implementation examples.

Level 2 to Level 3: Stop being the human in the loop

At Level 2, the team builds and operates capabilities. At Level 3, the team owns the interface through which capabilities are consumed while other people build and operate the capabilities within it.

That shift requires three concrete moves.

Decoupled Golden Paths

A golden path that can only be followed one way will always generate exceptions. The idea here is to parameterize it and give it escape hatches for legitimate edge cases.

We worked with a logistics organization that had a path hardcoding cloud regions, fixed resource limits, and assuming a specific service mesh. Any deviation meant waiting on the platform team. The fix was parameterization – validated options instead of hardcoded values, policy-based constraints instead of fixed limits. Teams went from waiting on every configuration change to self-servicing 80% of their needs within guardrails.

Instrument before you automate

The platform team’s first instinct is to automate what they understand best. The better signal is what actually arrives most often. We observed this with a media organization that spent three months logging every incoming request before building anything new. They found that 20% of request types drove 80% of the volume. They built self-service for those patterns first. Their backlog dropped 60% in six months. Build for the patterns that exist, not the ones you assume exist.

Treat Interface as Product

Treat the interface as the product. Not the capabilities behind it. How discoverable is it? How much does a developer need to know before they can use it successfully? How does it behave when a request falls outside what was anticipated?

We worked with a manufacturing organization where the platform team was building every deployment instance, each Helm chart, each environment config, each integration. They were falling behind. The change was not technical. They stopped building deployments and started owning the deployment interface, the schemas, the validation rules, the contracts. Application teams handled their own instances within those contracts. The platform team’s impact scaled because they were enabling rather than doing.

Level 3 to Level 4: Make the interface disappear

Moving from Level 3 to Level 4 requires a fundamental shift in thinking. At Level 3, developers are still conscious of the platform – they interact with it, configure it, and think about it. At Level 4, the platform becomes ambient infrastructure. The interface doesn’t disappear; it becomes so deeply integrated into the development workflow that developers rarely need to think about it explicitly.

This transition has three primary drivers.

Automate the common paths entirely

At Level 3, “self-service” still means a developer makes a choice and triggers a workflow. At Level 4, you anticipate those choices through convention and sensible defaults. When a new service is pushed to a repository, the platform detects it, validates it against policy, provisions monitoring, logging, and tracing automatically – a developer doesn’t fill out a form. The decision tree collapses into automation.

The shift is not about removing choice; it’s about encoding smart defaults with policy-enforced escape hatches. We worked with a fintech organization that reached Level 4 when they moved from “developers request a database” to “databases are provisioned alongside services using Terraform providers baked into the development workflow.” The interface was still there – developers could override defaults – but the common case required zero interaction.

Build platform capabilities into developer tools

The most effective Level 4 platforms are not separate experiences. They are integrated into the tools developers already use daily: Git, their IDE, their CI/CD system. When a pull request is opened, the platform automatically runs security checks, validates infrastructure assumptions, and flags potential production issues before human review. When a developer saves a Kubernetes manifest, their IDE knows about platform policies and prevents invalid configurations from being committed.

This requires the platform team to deeply understand the developer workflow and meet them where they work, not ask them to come to the platform.

Distribute capability ownership

At Level 3, the platform team still owns the framework. At Level 4, the framework is stable enough that capability ownership can be distributed. Observability teams own the observability capabilities within the platform interface. Security teams own the security policies. Database teams own database provisioning. The platform team shifts from operating every capability to operating the contracts and integration points between them.

This distribution works because the platform team has already established the interfaces, validation rules, and escape hatches. Domain experts can extend the platform within those boundaries without destabilizing it. We observed a SaaS organization structured this way see delivery times drop 40% once specialized teams could ship capabilities directly into production without platform team approval – because the platform contracts guaranteed safety.

The cultural shift

Level 4 requires organizational change as much as technical change. The platform team’s success metric inverts. Instead of “how many platforms did we build,” it becomes “how rarely do developers think about infrastructure.” Hiring shifts from generalists who build everything to specialists who own specific capability domains. Documentation shifts from “how to use the platform” to “what are the platform’s constraints and guarantees.”

The clearest sign you’ve reached Level 4 is when a new engineer onboards and ships production code within 48 hours without a single question about infrastructure. The platform is working so well that it’s invisible.

Conclusion

Most platform interfaces were designed with one actor in mind: a developer filling out a form, running a CLI command, or clicking through a portal. That assumption is starting to break.

AI agents are beginning to request platform capabilities the same way developers do – provisioning environments, spinning up pipelines, and requesting secrets. But they do not fill forms. They do not read documentation. They call APIs, and they call them at a frequency and pattern no human workflow was designed for. The self-service interface you built for your developers is not the same thing as a machine-consumable interface for agents. That gap is the next maturity conversation, and it is arriving faster than most platform teams have planned for.

If you want to understand where your platform’s Interfaces maturity sits today, the CNCF community self-assessment tool is a useful starting point.

Where on this line does your platform sit? Join the CNCF Platform Engineering Technical Community Group to discuss these maturity transitions with practitioners across the ecosystem and contribute to the ongoing development of platform engineering guidance.

  •  

Building an AI factory on Kubernetes

An AI factory is not just a model or a cluster. It is a pool of GPUs that many teams draw from at once: one team fine-tuning, another serving inference, a third running evaluations, all on the same accelerators. NVIDIA frames it as “infrastructure for the full AI lifecycle, from data preparation through training, fine-tuning, and high-volume inference”. In an enterprise that means one fleet, many teams, and different quotas, policies, and trust boundaries layered on top. The hard question is no longer how to train a model. It is how to give every team safe, isolated access to the same expensive hardware without anyone stepping on anyone else.

Two years ago every platform team was building a developer platform. Kubernetes already had mature primitives for containers, RBAC, autoscaling, and policy. What it did not have was a clean answer for accelerators, or for keeping tenants apart on the same nodes. That is the gap an AI factory has to close, and the cloud native ecosystem now supplies most of the parts to close it.

The bottleneck is utilization, not model serving

Accelerators are the dominant capital expense in the building, and the metric that decides whether that spend pays off is utilization, not a peak tokens-per-second number from a single run. The market grades GPU clouds the same way. SemiAnalysis’s ClusterMAX scores providers on security, networking, storage, reliability, and support rather than raw throughput, and its security criteria reward hard per-tenant isolation, down to per-tenant Kubernetes clusters and DPU-based isolation, while flagging weak boundaries like putting many tenants on one cluster. The wrapper around the GPUs is what gets judged.

Two things keep utilization low. First, the resource model: in the traditional device-plugin model a pod asks for nvidia.com/gpu: 1 and pins a whole accelerator even at ten percent use. Dynamic Resource Allocation (DRA), GA in Kubernetes 1.34, lets the scheduler treat accelerators as rich devices with attributes, memory, and topology, though it does not by itself carve a GPU into fractions; density comes from the device layer underneath. Second, the isolation model: to keep teams apart, platforms default to a dedicated cluster or a dedicated set of GPUs per team, which is the safe choice when trust is strict and wastes most of the hardware.

The same pattern recurs in the field: operators managing tenants with a bare metal provisioner and manual workarounds, or handing each customer a dedicated block of GPUs and turning away demand they cannot isolate cleanly. The fix is not a new model server. It is a stack that allocates accelerators so capacity is neither stranded nor unsafe, and isolates tenants so packing them together holds up.

The stack, layer by layer

An AI factory is an assembly problem. Most layers are Kubernetes native or CNCF projects, with a few OSS tools such as NVIDIA’s MIG, vCluster and Dynamo. The diagram below shows the shape, and the table lists the job each layer does.

LayerJobBuilding blocks
Hardware lifecycleProvision and validate bare metalMetal3 / Ironic, Tinkerbell, Redfish, NetBox
Cluster lifecycleCreate and version clusters, GitOpsCluster API, Argo CD or Flux, NVIDIA AICR
Node inventoryLabel GPUs, NICs, MIG, topologyNode Feature Discovery, GPU & Network Operator (NVIDIA, AMD)
Tenant isolationKeep teams apart on the same hardwareTenant clusters (vCluster), sandboxed runtimes
GPU allocationAllocate and schedule acceleratorsDRA, MIG, HAMi, time-slicing; KAI Scheduler (Topology Aware Scheduling), Volcano, Kueue
Inference and servingRun models behind an APIvLLM, KServe, llm-d
Batch and HPCRun SLURM workloadsSlinky (SLURM on Kubernetes)
VMsRun virtual machines for tenantsKubeVirt
Gateway and autoscalingRoute and scale endpointsGateway API, Envoy, LiteLLM, KEDA, HPA
NetworkingMove data between GPUs, isolate tenantsCilium, Multus, SR-IOV,  RDMA based networking solution 
Storage and dataPersist datasets, checkpoints, modelsCSI, Rook / Ceph, parallel filesystem CSI, object storage
ObservabilitySee utilization, logs, tracesPrometheus, OpenTelemetry, DCGM exporter
Identity and policyAuthn, authz, quotas, guardrailsKeycloak (OIDC), ResourceQuota / Kueue, Kyverno or OPA
Secrets and securitySecrets, runtime, supply chainOpenBao + External Secrets, Falco, Trivy
Reliability and remediationDetect and recover from node failuresDCGM health checks, Node Problem Detector, drain / cordon
Self-service and billingProvision and charge tenantsAPI / OpenTofu / GitOps, OpenCost, DCGM GPU-seconds

From bare metal to validated capacity

Everything starts at the rack.

Take a typical modern AI supercomputing platform as an example. Before any GPU can run a workload, something has to turn raw servers into a usable pool. That is the provisioning layer, often a proprietary hardware manager that ships with the system.

It works in steps. First it discovers each node, taking inventory: which GPUs and how many, whether memory is healthy (ECC state, meaning error-correction is on and not logging faults), and the identities of the network cards (InfiniBand GUIDs and NIC MACs, the permanent hardware IDs used to wire up and boot the node). Next it network boots the node and installs an OS image with the GPU driver and the CUDA and NCCL libraries baked in, so it can compute the moment it comes up. It then applies BIOS settings that match the node’s goal: baseline, performance, or confidential-compute.

Before a node joins the pool, it is tested. This burn-in runs the node under load to catch early failures, and an NCCL test confirms the GPUs actually talk to each other at full bandwidth. The result is written to a source of truth like NetBox, which also tracks IP address assignments (IPAM). Retiring a node runs the flow in reverse: wipe the disks, reset the remote-management login (eg the BMC), and return the clean node to the pool.

That proprietary manager is the vendor’s all-in-one take on this layer, tightly coupled to its own systems. The other path is to assemble the same loop from open building blocks: Metal3 driving Ironic, or a solution like vMetal, giving you the same discover, image, validate, and reclaim cycle on your own terms instead of adopting the vendor’s stack wholesale. That is the build-versus-assemble choice, and it recurs at every layer above. I come back to it at the end.

GPU allocation: the layer that makes the economics work

Figure 1. Whole-GPU allocation versus a partitioned GPU.

Figure 1. Whole-GPU allocation versus a partitioned GPU.

DRA gives a richer device-claim model, but fractional density comes from the device implementation. HAMi, a CNCF Incubating project, enforces per-pod memory and compute limits in software so several pods run on one card with guardrails between them, and it spans multiple accelerator vendors

Operators heading toward confidential computing do not place untrusted tenants on the same physical GPU; they give each tenant a whole GPU and reserve partitioning for workloads inside a single trust domain. MIG does isolate memory and faults in hardware, but its use as a boundary between hostile tenants is contested, so the conservative default is whole-GPU per tenant. The layer has two jobs: whole-GPU allocation for tenant isolation, and partitioning for density within a tenant. Scheduling is separate: KAI Scheduler and Volcano handle gang and topology-aware placement, and Kueue handles queueing, admission, and quota.

The workload layers: serving, Slurm, and VMs

Above allocation sit the things teams actually run. For inference, vLLM is a common engine and KServe, a CNCF incubating project, wraps it with autoscaling and standard endpoints, while NVIDIA Dynamo and llm-d push disaggregated inference for larger deployments. In front, Gateway API handles routing and LiteLLM adds an OpenAI-compatible gateway so dozens of specialized models speak one API.

Training customers usually live in Slurm, and the pattern has converged on running it on Kubernetes through SchedMD’s Slinky, which represents the Slurm daemons as CRDs and integrates with the GPU Operator and DRA for topology-aware scheduling, with pyxis and enroot, GPUDirect RDMA at full NCCL bandwidth, and prolog and epilog health checks. And some tenants want plain virtual machines rather than pods; KubeVirt runs VMs as Kubernetes workloads, so one platform hands out both containers and VMs from the same pooled fleet under the same RBAC and quotas.

Networking, storage, and observability

Training and disaggregated inference are bandwidth-bound, so the network is part of the design. Cilium handles the primary CNI and network policy; for the fast path, Multus and SR-IOV expose the NIC directly and RDMA over RoCEv2 or InfiniBand carries inter-node GPU traffic, with the isolation layer kept off that data path. 

A real cloud also gives tenants the cloud-edge services they expect: elastic IPs, NAT, and L4 load balancing from a gateway in front of the fabric. Storage needs per-tenant persistence, usually CSI with Rook and Ceph or a parallel filesystem, governed by per-tenant StorageClasses and quotas. 

For observability, OpenTelemetry is the neutral collection layer that keeps backends swappable, with Prometheus for metrics and VictoriaLogs for logs; the DCGM exporter publishes GPU telemetry that becomes per-tenant only with labels and a cost pipeline, and OpenCost turns GPU-seconds into chargeback.

Reliability and security

At fleet scale GPUs fail constantly: ECC errors, cards that fall off the bus, NVLink and thermal faults. The operator’s job is to catch these before a tenant does, which makes health a first-class layer rather than a dashboard afterthought. 

Active and passive checks on DCGM watch for degradation, Node Problem Detector turns hardware signals into node conditions, and a remediation loop cordons and drains a suspect node before new work lands on it. This is one of the categories the rating systems weigh most, because reliability, not peak throughput, is what a customer feels first. Identity and policy round it out: Keycloak over OIDC, OpenBao with the External Secrets Operator, Kyverno or OPA for guardrails, and Falco and Trivy for runtime and supply chain, with audit logs exported to the observability stack and traffic encrypted in transit.

Isolating tenants

Every layer above assumes one thing: that you can safely run more than one team on the same hardware. That is the tenant-isolation problem, and it has two halves worth keeping separate.

The first is the control plane. The tenant-cluster pattern gives each team a virtual control plane: a full Kubernetes API server with its own CRDs, admission webhooks, versions, and RBAC, running as a workload on a single underlying cluster, with no view into another tenant. Several CNCF and open source projects implement this pattern, like vCluster. Because each tenant cluster is conformant Kubernetes, plain kubectl, Helm, and Argo CD with no proprietary extensions, the model gives tenants a clean exit path rather than lock-in.

Figure 2. Tenant clusters on one underlying cluster, drawing from a pooled GPU fleet.

Figure 2. Tenant clusters on one underlying cluster, drawing from a pooled GPU fleet.

In practice operators run two tiers. High-trust or enterprise tenants get a dedicated cluster, sometimes dedicated hardware, where the boundary is physical; smaller or cost-sensitive tenants get a tenant cluster on pooled capacity. The same control plane drives both. Reliability follows from the same design: because a tenant control plane runs as pods, Kubernetes reschedules it on failure, and the open question is blast radius, so operators cap how many tenants ride one underlying cluster.

The second half is the data plane, which a tenant cluster does not solve on its own. You still need network isolation, storage isolation, quotas, Pod Security, and a runtime boundary. Network isolation usually comes from the fabric rather than from Kubernetes: a control plane carves per-tenant VPCs with VXLAN and EVPN on the Ethernet side and partition keys on InfiniBand. Increasingly that enforcement is pushed into hardware, where DPUs (Data Processing Units) such as NVIDIA BlueField or AMD Pensando move isolation and encryption off the host CPU, which is also how operators reach a confidential computing posture.

For the runtime boundary on shared nodes, the options range from dedicated nodes to sandboxed runtimes such as vNode. The bar for a real cloud is hardware-enforced isolation, not namespaces and good intentions.

What makes it a cloud, not just infrastructure

The line between a pile of GPUs and a cloud is that a customer can provision it themselves and get a bill that makes sense. Both are cloud native problems. Self-service means API-first with no UI-only paths: a tenant creates and deletes clusters through an API, a Terraform provider, or GitOps, with resources expressed as declarative CRDs reconciled by Flux or Argo CD, and access scoped by RBAC through OIDC. 

The bill comes from the metering layer: DCGM-driven GPU-seconds and OpenCost allocation, exported per tenant. None of this is glamorous, and it is usually the widest gap between a lab and a product. It is also, more than raw performance, what customers experience day to day.

From demo to production

Put the stack together and the demo is simple: two teams, two tenant clusters, two model endpoints, one physical GPU partitioned by MIG or software limits, each with its own RBAC, network policy, metrics, and cost line, neither aware of the other. This has been shown live on stage at KubeCon + CloudNativeCon with a single modern GPU serving two models at once. 

Two things turn it into production. The first is conformance: tooling like NVIDIA’s AI Cluster Runtime validates cluster configurations against the hardware you actually have and emits reproducible Helm or GitOps artifacts, and the Kubernetes AI Conformance program, introduced in the 1.35 release, pushes the same idea at the platform level. 

The second is scale: the design has to hold at hundreds of GPU nodes and several data centers, not the handful you prove it on, which is the real reason the foundation is GitOps, declarative tenants, and a single source of truth. There is a strategic choice here too, because the hardware vendor is moving into this layer with an integrated suite, NVIDIA’s DSX OS, so an operator decides layer by layer whether to adopt it, assemble the equivalent from cloud native projects, or compose the two.

The Takeaway

An AI factory is not another AI platform or model serving product. It is an operating model for running GPU infrastructure at scale on Kubernetes. Just as Kubernetes became the operating system for cloud native applications, it is becoming the foundation for AI infrastructure, making GPUs schedulable resources, providing isolated environments for tenants, and enabling on demand compute. The challenge is not deploying technologies like MIG, DRA, HAMi, or vLLM, but combining them into a platform that balances utilization, isolation, and cost while allowing multiple teams to safely share expensive GPU infrastructure without compromising performance or security.

Software is only half of it. The hardware layer is just as hard, often harder. Topology decides performance: which GPUs share an NVLink or NVSwitch domain, how each node attaches to a rail-optimized InfiniBand or RoCE fabric, whether the GPU, NIC, and CPU sit on the same NUMA node, and whether GPUDirect RDMA has a clean path. Schedule work without accounting for any of it and collective operations stall on the slowest hop, no matter how healthy the platform looks on paper. The stack has to be topology-aware, not just resource-aware.

The hard part is not naming the tools. It is making density, isolation, and chargeback work together, with hardware-enforced boundaries where the trust model demands them, without hiding the GPU data path behind an abstraction.

  •  

Cloud Native platform sovereignty through multi-plane architecture

When people talk about cloud sovereignty, the conversation often starts with regions: where a workload runs and where its data is stored. But choosing a region is only part of the story. The architecture of the platform matters just as much, particularly how it separates control, runtime, build, and observability responsibilities across clusters.

A recent CNCF community post, “From data residency to digital sovereignty: architectural patterns for cloud native platforms,” made the case well. Under regimes like the EU Data Act, NIS-2, DORA, and the UK Data Use and Access Act, platform teams now have to show more than where workloads run. They have to show how the platform is operated, secured, and governed, all the way down to the control plane.

That article laid out the requirements and introduced the tenant-cluster pattern as one way to draw isolation boundaries. In this article, we look at the same requirements from a different but complementary angle: what happens when you treat sovereignty as a property of a platform’s plane topology. We will use OpenChoreo, an open source internal developer platform and a CNCF Sandbox project, as an inspectable example. The architectural ideas apply broadly, though.

What auditors ask platform teams

The earlier post narrows down the regulatory and procurement noise into four practical properties. Rather than repeat them, we will restate them as the questions an auditor or a procurement team actually asks a platform team:

  1. For every component that can touch tenant data, including the control plane and the logs and metadata around it, can you name the legal jurisdiction it runs under?
  2. If your vendor’s hosted service disappeared tomorrow, could your team continue operating the workload, rebuild it, and move it elsewhere?
  3. Can anyone outside the boundary reach your keys, your cluster state, or an admin credential?
  4. If the provider, hardware, or country changes, does the workload move, or does it need to be rewritten?

Look at what these questions have in common. Almost none of them are about location. They are about where control and state live, and who can reach them.

A single shared Kubernetes cluster makes those questions difficult to answer cleanly. One API server, one etcd, one set of controllers and admission webhooks serve every tenant. That makes it hard to point to a clear architectural boundary during an audit.. The tenant-cluster pattern fixes this by giving each boundary its own control plane. A multi-plane platform fixes the same problem one layer up. The two patterns also work well together, as we will see later.

A multi-plane topology

OpenChoreo splits the platform into planes. Each plane is its own cluster, with its own lifecycle, its own scaling behavior, and its own security boundary:

  • A control plane holds desired state through declarative APIs and runs the reconciliation controllers. It orchestrates. It does not run tenant workloads.
  • One or more data planes are conformant Kubernetes clusters that actually run workloads. Each data plane has its own API server and its own state.
  • One or more observability planes collect and serve logs, metrics, and traces.
  • One or more workflow planes execute CI and GitOps workflows.
  • An experience plane provides the developer portal, CLI, and API/MCP surfaces.
Figure 1: The multi-plane topology. Arrow direction shows who initiates the connection.

Figure 1: The multi-plane topology. Arrow direction shows who initiates the connection.

The connection model is the detail that matters most for sovereignty, and it is why the arrows in Figure 1 point the way they do. The data, observability, and workflow planes each open an outbound, mutually authenticated (mTLS) connection to the control plane’s gateway. The control plane never dials into them.

Two things follow from this. First, because the connection is outbound only, the API servers of the clusters holding regulated workloads are never exposed to the internet. Second, the control plane holds the desired state, not the tenant runtime state. It translates higher-level resources into Kubernetes resources and hands reconciliation to each data plane’s own API server. So a data plane keeps serving traffic even if it loses its link to the control plane.

That separation allows a single orchestrating control plane to work with strong regional boundaries without becoming the place where all runtime states have to live.

How the topology holds up

Question 1: Can you name the jurisdiction? With a “one jurisdiction, one data plane” model, the answer can be represented directly in the architecture and defended during an audit. Figure 2 shows what this looks like with two jurisdictions. The observability design keeps the answer clean. Each data plane reports to a regional observability plane, and the portal queries that plane directly instead of routing telemetry back through the control plane. OpenChoreo’s documentation flags exactly this pattern for multi-regional deployments under regional data-privacy rules. A tenant’s runtime state lives in its regional data plane, and its logs never leave the region either.

Graphic: Regional boundaries in practice. Workloads and telemetry stay inside the dashed lines; only the outbound mTLS link crosses them.

Figure 2: regional boundaries in practice. Workloads and telemetry stay inside the dashed lines; only the outbound mTLS link crosses them.

Question 2: can you operate without the vendor? Each data plane is a full, conformant Kubernetes cluster, not an opaque managed endpoint, and it runs fine while disconnected from the control plane. The stack underneath is open source and largely built around CNCF and cloud -native projects, including Argo Workflows, Cloud Native Buildpacks, OpenSearch, Prometheus, OpenTelemetry, Flux, cert-manager, and Cilium. The design is modular, so a team can swap a component or put an existing observability stack behind the same query interface. No single hosted service sits on the critical path.

Question 3: can outsiders reach keys, state, or credentials? The outbound-only mTLS model keeps the sensitive clusters unreachable from the internet in the first place. Secrets and keys live in whichever External Secrets Operator-compatible store or vault the team chooses, which makes key ownership a decision rather than a default. Authorization is fine grained, down to individual namespaces, projects, and components, with groups mapped from any OAuth2/OIDC identity provider. The same authorization model applies whether the caller is a developer, the CLI, or an AI agent.

Question 4: does a change of provider mean a rewrite? Workloads are represented as standard Kubernetes resources and run on conformant Kubernetes clusters, whether those clusters are in a public cloud, on premises, or on bare metal. Promotion is a first class concept. A pipeline can move a component from development on one data plane to production on another, in a different geography or provider, applying environment-specific configuration on the way. Swapping the infrastructure under a jurisdiction becomes a topology change, not a migration project.

How the platform layer meets the infrastructure layer

The earlier post built its isolation story on the tenant-cluster pattern giving each tenant a virtual cluster of its own inside a shared host cluster.It is useful to look closely at how a platform layer can sit on top of that pattern, because the two solve different parts of the sovereignty problem.

Let’s start with what a virtual cluster gives you at the infrastructure layer. Each tenant gets a virtual control plane: its own API server and its own datastore, running as pods inside a shared host cluster. Tenant A cannot see tenant B’s resources, cannot be taken down by tenant B’s misbehaving CRDs or webhooks, and cannot touch tenant B’s cluster state. That is real control plane isolation, and it costs a fraction of a dedicated cluster because one node pool serves everyone.

Now look at what the tenant-cluster pattern, by design, does not decide. It does not decide which jurisdiction a tenant lands in. It does not decide where that tenant’s logs and traces are shipped. It does not define who may promote a workload from staging in one region to production in another. It does not give developers a paved road that keeps them from hand crafting kubeconfigs against raw clusters. These are not gaps in the pattern. They are platform layer concerns, and they are exactly the concerns the four sovereignty questions keep circling back to: where state lives, where telemetry flows, who can act across a boundary, and whether any of it is provable.

This is where the layering pays off. To the OpenChoreo control plane, a virtual cluster is just another conformant Kubernetes API. Register it as a data plane, and every platform layer control in this article now applies to it:

  • The DataPlane resource carries a jurisdiction label, so tenant placement becomes declarative and reviewable, not tribal knowledge.
  • The observabilityPlaneRef pins telemetry to the regional observability plane, so a tenant’s logs inherit the same residency guarantee as its workloads.
  • Promotion pipelines defined at the platform layer decide which environments a component may move between, so a workload cannot drift into the wrong jurisdiction through an ad-hoc deployment.
  • The same fine-grained authorization applies to every tenant, mapped from the same identity provider, whether the caller is a developer, the CLI, or an AI agent.
  • Developers get golden paths and a portal instead of raw cluster access, which shrinks the number of humans who ever hold credentials to the sensitive clusters.

Figure 3 shows the composed topology: one physical host cluster per jurisdiction, virtual clusters inside it as per-tenant data planes, one control plane orchestrating all of them over the same outbound mTLS link. In principle nothing about the registration changes because the data plane happens to be virtual; if you try this and hit an edge, that is a contribution waiting to happen.

Figure 3: the two layers composed. Virtual clusters draw the tenant boundaries inside the jurisdiction; the platform layer decides what may cross any boundary, and records why.

Figure 3: the two layers composed. Virtual clusters draw the tenant boundaries inside the jurisdiction; the platform layer decides what may cross any boundary, and records why.

Read the figure as a division of labor. The infrastructure layer answers “who is isolated from whom.” The platform layer answers “what is allowed to go where, and can we prove it.” Neither layer can answer the other’s question. Virtual clusters alone leave placement, telemetry routing, and promotion as manual policy enforced by hope. A platform alone, running tenants as namespaces on shared clusters, leaves every tenant one admission-webhook misconfiguration away from its neighbors. Together they cover all four sovereignty questions at a cost that scales with jurisdictions, not tenants.

One caveat belongs in the open, and Figure 3 makes it visible: tenants on the same host still share nodes and a kernel. If the threat model demands hardware isolation per tenant, a virtual cluster is not enough, and that tenant needs a physical data plane of its own. The point of a composable topology is that this, too, is just a registration decision, not a redesign.

Sovereignty as declarative configuration

The earlier post ends with an idea worth carrying forward: sovereignty should be something with a name, a template, and a commit history, rather than only a clause in a contract.

A plane topology expressed as declarative Kubernetes resources fits that idea well. Data planes, environments, and deployment pipelines are custom resources. The entire topology can live in Git: which region has which data plane, which observability sink it uses, and which promotion paths are allowed.

The manifest below is a simplified sketch with abbreviated field names. It is meant to show the shape, not an exact schema:

# Illustrative only. See the project docs for the real API surface.
kind: DataPlane
metadata:
 name: eu-west
 labels:
   jurisdiction: eu
spec:
 observabilityPlaneRef: eu-observability  # telemetry stays in-region
 # registry, gateway, and network settings scoped to the EU boundary
---
kind: Environment
metadata:
 name: production-eu
spec:
 dataPlaneRef: eu-west

Adding a new jurisdiction then becomes a reviewed pull request. And when someone asks “why is this tenant’s data in this jurisdiction?”, the answer is a commit history, not a screenshot of a console.

What the topology does and does not give you

First, the multi-plane topology does not change the legal jurisdiction of the organization operating the infrastructure. If a cluster operator is subject to a particular legal regime, that exposure still exists. Where the threat model requires sovereign hardware or a sovereign operator, that decision must be made at the infrastructure layer. The topology can partition exposure and reduce its scope, but it cannot remove the legal context of the operator.

Second, the topology draws boundaries; it does not enforce what happens inside them. Policy enforcement, supply-chain attestation and SBOMs, audit logging, and workload identity through something like SPIFFE/SPIRE are separate concerns, and a sovereign deployment may need all of them. The distinction here is between the architectural pattern and the platform implementation. The pattern alone does not provide these controls, but the platform layer is a natural place to integrate them. OpenChoreo’s modular architecture is intended to support this type of integration, in the same way it already orchestrates components such as Cilium, Flux, and cert-manager. Boundary comes first, enforcement is layered on top.

Third, more planes mean more to run. Every cluster is something to monitor, upgrade, and back up. The pattern earns its cost when the boundary you are drawing carries real legal or risk weight. It is overkill when it does not. This is also where the virtual cluster composition above pays off, by keeping the number of physical clusters tied to the number of jurisdictions rather than the number of tenants.

Takeaways

The reframing in the original post holds up. Sovereignty is less about a region on a dropdown and more about how control, state, keys, and audit trails are distributed. 

Seen through that lens, the interesting design question becomes how to split responsibilities: tenant clusters as the isolation primitive, a plane-separated topology as the boundary map, and, most powerfully, both together. Expressing those boundaries as declarative, version-controlled objects is what turns “sovereign” from a procurement promise into something a platform team can actually operate and audit.

OpenChoreo is an open -source CNCF Sandbox project. If you would like to explore or contribute, the code and community links are at openchoreo.dev.

  •  

Federating clusters for zero-downtime Kubernetes

Every multi-region setup eventually meets the same awkward moment: a whole cluster goes away, and the identical copy of your service running two regions over might as well not exist, because nothing is wired to treat them as one thing. Failover becomes a runbook: restore, repoint DNS, and wait for an outage that, on paper, you’d already paid to survive.

Linkerd’s multicluster extension closes that gap by letting several clusters present a service as a single, load-balanced endpoint. The part that the official tasks gloss over is that a real platform almost never picks one multicluster mode. Some services want federation (same service everywhere, one endpoint, automatic failover). While others want mirroring (reach a specific remote service by name). And you frequently want both patterns living on the same set of links. The docs walk through each mode on its own. This post wires all three together across three GKE clusters, with a full-mesh link topology, a chaos test that takes out an entire cluster, and scripts you can clone and run on a fresh GCP project.

Companion repo: Every script referenced here lives in this repository. Feel free to clone it, set your project ID, and run it.

Linkerd multicluster modes: Gateway, flat, and federated 

Linkerd’s multicluster extension supports three modes. The nice thing is they’re not mutually exclusive: on the same set of linked clusters, the mode is chosen per service via a label.

ModeLabelWhat happensNetwork Requirement
Hierarchical (gateway)mirror.linkerd.io/exported=trueService mirrored as <svc>-<cluster>, traffic routed through a gatewayGateway IP reachable
Flat (pod-to-pod)mirror.linkerd.io/exported=remote-discoveryService mirrored as <svc>-<cluster>, traffic goes directly to remote podsFlat network (pod IPs routable)
Federatedmirror.linkerd.io/federated=memberAll same-name services unioned into <svc>-federated, load balanced across all clustersFlat network (pod IPs routable)

The distinction that matters operationally is that hierarchical mirroring works on any network. Only the gateway IP needs to be reachable, while flat and federated modes need real pod-to-pod connectivity. On GCP, VPC-native GKE clusters on peered VPCs give you that flat network for free. So, you can run federated services for your core workloads over a flat network and still mirror a specialized service through a gateway from a cluster that isn’t on that network. Most platform teams I’ve seen end up with exactly this kind of mix.

Multi-region architecture: GKE cluster setup 

We have three GKE clusters across three regions, fully linked to each other (six directional links total). Three demo services, each using a different multicluster mode:

A chart showcasing the three demo services of the GCO Project.

frontend is federated and runs in all three clusters. A single federated frontend service in each cluster load-balances across all nine pods (3 replicas × 3 clusters). When a cluster goes down, the remaining six pods absorb the traffic with no application changes.

api is flat-mirrored and runs in `west` and `east`. The `north` cluster consumes it as `api-west` and `api-east`, which are explicit remote service names with traffic sent straight to the remote pods. This is what you reach for when the client needs to decide which backend it talks to, for example, to keep a request in-region for data locality.

analytics is gateway-mirrored and runs only in `east`. Exported through the Linkerd gateway so `west` and `north` reach it as `analytics-east-gw` without needing flat-network connectivity to `east`’s pods. It’s here mainly to prove that gateway mode coexists with flat and federated modes on the same links.

Deployment prerequisites: GKE, Linkerd, and CLI tools

  • A GCP account (free-tier credits cover this. Use three standard clusters with small node pools)
  • `gcloud` CLI, authenticated (`gcloud auth login`)
  • `kubectl` v1.28+
  • `step` CLI, `brew install step` (for certificate generation)
  • `helm` v3
  • ~30 minutes for the full setup

The infra script enables the `compute` and `container` APIs for you, so a brand-new project works out of the box.

Step 0: Configure

Clone the repo, create a local .env file from the example file, and customize it for your GCP project. The defaults are enough for the rest of the demo, so in most cases you only need to change the project ID.

```bash
git clone <your-repo-url>
cd blog-linkerd-federation
cp env.example .env
```

Open `.env` and set at least your project ID. The file ships with sensible defaults for everything else:

```bash
export GCP_PROJECT="your-project-id"

export REGION_WEST="us-central1"
export REGION_EAST="us-east1"
export REGION_NORTH="europe-west1"

# One zone per region. We pin node-locations to a single zone so num-nodes is
# the TOTAL node count — see the cost note below for why this matters.
export ZONE_WEST="us-central1-a"
export ZONE_EAST="us-east1-b"
export ZONE_NORTH="europe-west1-b"

export CLUSTER_MACHINE_TYPE="e2-medium"
export CLUSTER_NODE_COUNT="1"
export FRONTEND_REPLICAS="3"
```

At minimum, set GCP_PROJECT. Everything else ships with sensible defaults: three regions, one zone per region, and small node pools to keep the cost down. If you run cat .env, you should see the full set of variables populated.

Load the variables into your current shell so the scripts can read them:

```bash
source .env
```

Every script below reads from this file, and they all run with `set -euo pipefail`, so a missing variable fails loudly rather than silently. That’s why `env.example` carries the full set, the VPC and cluster names included, instead of just the project ID.

Step 1: Provision three GKE clusters with VPC peering

Run the infrastructure script to create the networks and clusters. This takes about 10–15 minutes, so it’s a good point to grab a coffee.

```bash
./scripts/01-infra.sh
```

This script does the following:

  1. Enables the `compute` and `container` APIs (no-op if they’re already on).
  2. Creates three VPCs with non-overlapping pod and service CIDRs, a hard requirement for VPC peering.
  3. Sets up full-mesh VPC peering (west↔east, east↔north, north↔west) with `–export-custom-routes` and `–import-custom-routes` so pod CIDRs are actually advertised. This is what gives us the flat network.
  4. Creates three GKE Standard clusters, one per VPC/region, each pinned to a single zone.
  5. Renames the kubectl contexts to `west`, `east`, `north`.

Here’s the address plan the script uses. The ranges are intentionally non-overlapping so VPC peering can route pod traffic correctly:

ClusterVPC SubnetPod CIDRService CIDR
west10.10.0.0/2010.100.0.0/1410.104.0.0/20
east10.20.0.0/2010.108.0.0/1410.112.0.0/20
north10.30.0.0/2010.116.0.0/1410.120.0.0/20

Non-overlapping ranges are non-negotiable. If pod CIDRs overlap across peered VPCs, routing breaks silently. Pods get responses from the wrong cluster, or connections time out with nothing useful in the logs. Ask me how I know.

One zone, not three. A GKE regional cluster places `–num-nodes` nodes in each of three zones by default. With `–num-nodes 1` that’s 3 nodes per cluster, 9 total, and triple the bill. The script pins `–node-locations` to a single zone so `CLUSTER_NODE_COUNT=1` really means one node per cluster.

Cost note: Three Standard clusters with one `e2-medium` node each run roughly $10–15/day total for this demo (management fee + nodes + a small gateway load balancer on `east`). The teardown script removes everything.

Step 2: Install Linkerd with a shared trust anchor

Install Linkerd into all three clusters using a shared trust anchor. The script generates the certificates, installs the control plane, and configures each cluster to trust the others for cross-cluster mTLS.

```bash
./scripts/02-linkerd-install.sh
```

This generates a root CA and per-cluster issuer certificates, then installs Linkerd on all three clusters:

```
root.crt (shared trust anchor)
├── issuer-west.crt + issuer-west.key
├── issuer-east.crt + issuer-east.key
└── issuer-north.crt + issuer-north.key


Per-cluster issuer certs are a production habit worth keeping: if one cluster’s issuer is compromised you rotate it in isolation, without touching the others. The shared root is what lets cross-cluster mTLS work at all. Every proxy can verify every other proxy’s identity back to the same anchor.

To keep resource usage (and cost) down, this installs the control plane only with no Viz add-on.

Step 3: Install multicluster and create a full-mesh link topology

Set up the multicluster components and create a full-mesh topology between the clusters. After this step, every cluster can consume services from every other cluster.

```bash
./scripts/03-multicluster-setup.sh
```

This is the step with the most going on. We create six directional links, every cluster linked to every other cluster, so every cluster gets a `<svc>-federated` service for federated workloads, and every cluster can consume mirrored services from any other.

The wrinkle is the gateway. Only `east` needs one (it’s the only cluster exporting `analytics` hierarchically), so we enable the gateway in east’s install and leave everyone else gatewayless. One install per cluster, all flags at once, no re-running install a second time to bolt a gateway on afterward:

```bash
# west: gatewayless, with one controller per cluster it consumes from
linkerd --context west multicluster install --gateway=false \
  --set controllers[0].link.ref.name=east \
  --set controllers[1].link.ref.name=north \
  --set controllers[2].link.ref.name=east-gw \
  | kubectl --context west apply -f -

# east: gateway enabled here, controllers for the clusters it consumes
linkerd --context east multicluster install --gateway=true \
  --set controllers[0].link.ref.name=west \
  --set controllers[1].link.ref.name=north \
  | kubectl --context east apply -f -

# north: gatewayless, controllers for west, east, and east's gateway link
linkerd --context north multicluster install --gateway=false \
  --set controllers[0].link.ref.name=west \
  --set controllers[1].link.ref.name=east \
  --set controllers[2].link.ref.name=east-gw \
  | kubectl --context north apply -f -


Note the controller count. The service-mirror controller runs on the consuming side, one per link. `west` and `north` each consume `analytics` via the gateway, so they get a third controller for the `east-gw` link; `east` doesn’t consume its own analytics, so it only needs two.

Then we generate the links. Flat/federated links use `–gateway=false`; the gateway-aware link to `east` (for the analytics export) is a separate link named `east-gw`:

```bash
# Flat links (no gateway) — for federated + flat-mirrored services
linkerd --context east multicluster link-gen --cluster-name=east --gateway=false \
  | kubectl --context west apply -f -
linkerd --context west multicluster link-gen --cluster-name=west --gateway=false \
  | kubectl --context east apply -f -
# ... (all six directions)

# Gateway-aware link from east (for the analytics hierarchical export)
linkerd --context east multicluster link-gen --cluster-name=east-gw \
  | kubectl --context west apply -f -
linkerd --context east multicluster link-gen --cluster-name=east-gw \
  | kubectl --context north apply -f -
```


After this, `linkerd multicluster check` on any cluster should report every linked cluster healthy.

Step 4: Deploy the demo services

Deploy the demo workloads. The next sections label them for federation, flat mirroring, and gateway mirroring and show what each mode creates.

```bash
./scripts/04-deploy-app.sh
```

Three services, three modes, and deliberately the same `buoyantio/bb` image for all of them, a tiny HTTP server that echoes a fixed string. The application isn’t the point. The point is that one `kubectl label` changes how Linkerd treats the service across clusters, with everything else held constant.

frontend (federated)

Deploy to all three clusters with a per-cluster response string, then labeled for federation:

```bash
for ctx in west east north; do
  kubectl --context $ctx -n mc-demo label svc/frontend mirror.linkerd.io/federated=member
done
```

Within a few seconds, `frontend-federated` shows up in all three clusters:

```bash
$ kubectl --context west -n mc-demo get svc
NAME                 TYPE        CLUSTER-IP     PORT(S)    AGE
frontend             ClusterIP   10.104.1.50    8080/TCP   45s
frontend-federated   ClusterIP   10.104.2.100   8080/TCP   10s
```

api (flat-mirrored)

Label the api service in `west` and `east` for flat export:

```bash
kubectl --context west -n mc-demo label svc/api mirror.linkerd.io/exported=remote-discovery
kubectl --context east -n mc-demo label svc/api mirror.linkerd.io/exported=remote-discovery
```

Now `north` can see `api-west` and `api-east` as separate services:

```bash
$ kubectl --context north -n mc-demo get svc
NAME                 TYPE        CLUSTER-IP      PORT(S)    AGE
frontend             ClusterIP   10.120.1.50     8080/TCP   45s
frontend-federated   ClusterIP   10.120.2.100    8080/TCP   10s
api-west             ClusterIP   10.120.3.20     8080/TCP   5s
api-east             ClusterIP   10.120.3.21     8080/TCP   5s
```

The client in `north` picks `api-west` or `api-east` explicitly. Traffic will go straight to the remote pods with no gateway in the path.

analytics (gateway-mirrored)

Next, deploy only to `east`, labeled for hierarchical (gateway) export:

```bash
kubectl --context east -n mc-demo label svc/analytics mirror.linkerd.io/exported=true
```

This creates `analytics-east-gw` in `west` and `north`, routed through east’s Linkerd gateway:

```bash
$ kubectl --context west -n mc-demo get svc analytics-east-gw
NAME               TYPE        CLUSTER-IP     PORT(S)    AGE
analytics-east-gw  ClusterIP   10.104.5.10    8080/TCP   5s
```

The endpoints for this service point at east’s gateway IP, not the analytics pods directly. That’s the right trade when you can’t guarantee flat-network connectivity, or when you specifically want the gateway handling load balancing and mTLS termination.

Step 5: Verify all three modes

Generate traffic against all three service patterns and verify that each resolves the way you expect.

```bash
./scripts/05-verify.sh
```

This deploys a traffic generator in `north` that hits all three service patterns in a loop and tails the logs. The response strings come straight from the deployments, so you’ll see which cluster served each request:

```
[federated]  frontend from east
[federated]  frontend from west
[federated]  frontend from north
[flat-west]  api from west
[flat-east]  api from east
[gateway]    analytics from east
```

You can also inspect endpoints to see how differently each mode resolves:

```bash
# Federated: endpoints span all three clusters
$ linkerd --context west diagnostics endpoints frontend-federated.mc-demo.svc.cluster.local:8080
NAMESPACE   IP            PORT   POD                        SERVICE
mc-demo     10.100.1.15   8080   frontend-xxx-west          frontend.mc-demo
mc-demo     10.108.0.42   8080   frontend-xxx-east          frontend.mc-demo
mc-demo     10.116.0.33   8080   frontend-xxx-north         frontend.mc-demo

# Flat mirror: endpoints are the remote pod IPs
$ linkerd --context north diagnostics endpoints api-west.mc-demo.svc.cluster.local:8080
NAMESPACE   IP            PORT   POD                        SERVICE
mc-demo     10.100.2.10   8080   api-xxx-west               api.mc-demo

# Gateway mirror: the endpoint is east's gateway IP on port 4143
$ kubectl --context west -n mc-demo get endpoints analytics-east-gw
NAME               ENDPOINTS             AGE
analytics-east-gw  35.186.xxx.xxx:4143   30s
```

Three modes, one mesh, one set of clusters, and the only difference between them is a label.

Step 6: The chaos test, kill a cluster

This is where federation earns its keep. We simulate a full cluster failure and watch how each service type reacts.

```bash
./scripts/06-chaos-test.sh
```

The script scales every deployment in `east` to zero replicas (standing in for a cluster outage), then samples traffic from `north` across all three patterns.

Federated service (`frontend-federated`):

```
Before:  west=33% east=33% north=33%
After:   west=50% north=50%              ← automatic rebalance, zero errors
```


Traffic redistributes immediately. No errors, no config changes. As east’s pods drop out of the endpoint list, Linkerd’s load balancer simply spreads requests across what’s left.

Flat-mirrored service (`api-east`):

```
Before:  api-east responds normally
After:   api-east returns 503s           ← expected: the remote pods are gone
```

This is the correct behavior. The client explicitly asked for `api-east`, and east is down. Handling that is the client’s job: fail over to `api-west`, retry, or front the two with a TrafficSplit. Mirroring hands you control; federation hands you automation.

Gateway-mirrored service (`analytics-east-gw`):

```
Before:  analytics-east-gw responds normally
After:   analytics-east-gw returns 502s  ← the gateway is down too
```


Same story here, the client asked for a specific remote, and that remote is gone.

Bring east back:

```bash
kubectl --context east -n mc-demo scale deploy --all --replicas=1
kubectl --context east -n mc-demo scale deploy/frontend --replicas=3
```

(The script restores `frontend` to its full `FRONTEND_REPLICAS` count rather than leaving it at 1, otherwise east would rejoin the federation under-weighted, landing around 14% instead of an even third.) Within 15–30 seconds all three patterns recover: the federated service rebalances back to 33/33/33, and the mirrored services start answering again.

The lesson worth carrying out of this: federation is the right default for anything that should simply be available everywhere. Mirroring, flat or gateway, is the right call when the client genuinely needs to know which cluster it’s talking to.

Step 7: Teardown

When you’re finished with the demo, run the teardown script to remove all the infrastructure and avoid ongoing GCP charges.

```bash
./scripts/99-teardown.sh
```

This removes all three clusters, the VPC peerings, subnets, firewall rules, and VPCs created by the earlier steps. Run it when you’re done so the meter stops.

Selecting your Linkerd multicluster architecture strategy

After running all three side by side, here’s the decision framework I’d hand a teammate:

Question→ Mode
Should the client be cluster-agnostic?Federated
Does the client need to pick a specific cluster?Flat mirror
Is there no flat network between clusters?Gateway mirror
Do you need automatic failover with no app changes?Federated
Do you need traffic splitting with explicit weights?Flat mirror + TrafficSplit
Is the service a singleton (only in one cluster)?Mirror (flat or gateway)

And you can mix them freely in the same mesh. The label on each service decides its behavior independently of the others.

Linkerd multicluster gotchas and configuration lessons

The gotchas that cost us time and don’t jump out of the docs:

VPC peering route exchange. Creating the peering isn’t enough. You have to pass `–export-custom-routes` and `–import-custom-routes` on both sides, or the pod CIDRs never get advertised. The symptom is brutal to diagnose: DNS resolves fine, then connections just hang. Maddening to debug.

Regional clusters multiply your nodes. A regional cluster with `–num-nodes 1` quietly gives you three nodes (one per zone). We pin `–node-locations` to a single zone to keep it at one. Easy to miss until the bill arrives.

Overlapping CIDRs. GKE auto-allocates large ranges out of the `10.0.0.0/8` space by default, and three clusters built with defaults will overlap, at which point peering fails silently. Always set explicit, non-overlapping `–cluster-ipv4-cidr` and `–services-ipv4-cidr`.

Controller count matters. Each cluster needs one service-mirror controller per link it consumes. Miss one and the Link CR is created, but nothing mirrors, and `linkerd multicluster check` still looks green, so you’ll stare at it for a while before the penny drops.

Federated service naming is fixed. The federated service is always `<svc>-federated`; you can’t change the suffix. Clients have to target `frontend-federated`, not `frontend`. Plan your naming around it, or use a TrafficSplit to point `frontend` at `frontend-federated`.

Gateway and flat can’t share one link. A single Link CR is either gateway or flat, not both. To get both behaviors to the same cluster you create two links with different names. That’s why our setup uses `east` (flat) and `east-gw` (gateway) as separate links, with a matching controller for each on the consuming clusters.

Production checklist

  • uncheckedBidirectional links between all clusters (full mesh) so every cluster has the federated service
  • uncheckedcert-manager with a shared CA instead of hand-rolled `step` certificates
  • uncheckedSeparate issuer certs per cluster (don’t skip it!)
  • uncheckedNetworkPolicies restricting cross-cluster traffic to only the services that need it
  • uncheckedLinkerd authorization policies for fine-grained access control
  • uncheckedMonitoring: pipe Linkerd-Viz metrics into your Prometheus/Grafana stack, and alert on a federated service’s endpoint count dropping
  • uncheckedGitOps: keep Link CRs and multicluster config in version control
  • uncheckedTest failover regularly: scale a cluster to zero in staging and confirm traffic redistributes

Key takeaways: Mastering multi-region Linkerd deployments 

The docs show each multicluster mode in isolation; real platforms need all three at once. Federation covers the common case: The same service everywhere, automatic failover, and nothing to change in the app. Flat mirrors give you explicit, cluster-aware routing when data locality matters. Gateway mirrors get you cross-cluster reach when a flat network isn’t on the table.

What surprised me most about building this is how little of it is genuinely complex. It’s mostly wiring. Once the trust anchor is shared and the links are up, adding a service to the federation is a single `kubectl label`, and removing a cluster is as simple as letting it go down. The mesh adjusts on its own.

For teams running across regions, that’s a real chunk of operational toil gone: your services run everywhere, traffic finds the healthy copies, and you pick the multicluster mode per service based on what that service actually needs.

References:

– [Linkerd Federated Services Task Guide](https://linkerd.io/2-edge/tasks/federated-services/): official walkthrough for federation

– [Linkerd Multicluster Reference](https://linkerd.io/2-edge/reference/multicluster/): architecture and mode details

– [Linkerd Multicluster Communication Guide](https://linkerd.io/2-edge/tasks/multicluster/): hierarchical mirroring walkthrough

– [Installing Multicluster Components](https://linkerd.io/2-edge/tasks/installing-multicluster/): installation reference

– [Linkerd 2.17 Announcement](https://linkerd.io/2024/12/05/announcing-linkerd-2.17/): federated services introduction

– [GKE VPC-native Clusters](https://docs.cloud.google.com/kubernetes-engine/docs/concepts/alias-ips): flat networking on GCP

  •  

Launch of the AI Infra SIG under the CNCF Japan chapter: First meetup and call for speakers

Japanese article follows English one.

As we all know, AI is advancing from generative AI to agents, driving growing demand for scalable, efficient infrastructure. Kubernetes and the broader Cloud Native ecosystem are becoming increasingly critical foundations for modern AI workloads.

To bring together Japan-local engineers, researchers, and platform builders to share best practices, operational experiences, and emerging technologies for Cloud Native AI infrastructure, Cloud Native Community Japan (CNCJ) is launching AI Infra SIG under the CNCF Japan chapter.

In this post, we introduce the motivation behind the AI Infra SIG, its mission, and the FIRST meetup—along with an open call for speakers.

Why We Are Launching the AI Infra SIG

The emergence of infrastructure challenges for AI workloads.

The Cloud Native ecosystem has historically evolved around web applications and microservices. These workloads are typically stateless, CPU- and memory-oriented, and exhibit relatively predictable traffic and scaling patterns.

Modern AI workloads, however—particularly those powered by LLMs and AI agents—present a new set of infrastructure challenges with heavy dependence on specialized hardware, complex distributed processing and orchestration, and unique traffic and scaling characteristics.

To address these challenges, the Kubernetes community has been actively advancing its vision of AI Readiness, introducing new capabilities, projects, and standards for AI-native infrastructure.

Japan is also home to a growing community of practitioners, contributors, and organizations actively working in this space.

Many engineers from Japan are already involved in open source projects and community discussions across the Cloud Native AI ecosystem.

At the same time, innovation is extending well beyond Kubernetes and CNCF. Across the broader Linux Foundation ecosystem, AI and agent technologies are rapidly emerging as major areas of community-driven development.

The PyTorch Foundation has become a hub for open-source AI innovation, supporting projects such as PyTorch and vLLM, while the Agentic AI Foundation has brought together communities working on open standards and interoperability across the rapidly evolving AI agent ecosystem.

These developments are not limited to global communities. Japan has an active and growing ecosystem of engineers, researchers, and open-source contributors participating across CNCF, the PyTorch Foundation, the Agentic AI Foundation, and other Linux Foundation projects. Events such as AGNTCon + MCPCon Japan 2026 further highlight the growing momentum around Cloud Native AI and Agentic AI technologies within Japan.

As innovation, we believe there is a unique opportunity to strengthen collaboration within Japan while connecting local practitioners to the broader global open-source ecosystem. This is one of the key motivations behind launching the CNCJ AI Infra SIG.

Mission and Activities

The CNCJ AI Infra SIG aims to bring together engineers, researchers, and platform builders who are shaping the future of AI infrastructure on Cloud Native technologies.

Our mission is to:

  • Share optimization techniques, operational best practices and real-world experiences for building and operating AI infrastructure and workloads.
  • Strengthen Japan’s engagement with upstream open-source projects and contribute back to the broader AI-related Cloud Native and Linux Foundation ecosystems.

To achieve these goals, the AI Infra SIG will organize regular community activities to facilitate knowledge sharing and collaboration around Cloud Native AI infrastructure, including:

  • Meetups covering practical experiences, emerging technologies, and updates from upstream communities.
  • Joint initiatives with other CNCJ communities, such as Cloud Native Security Japan, Cloud Native Platform Engineering Japan, and CoHDI Japan.
  • Collaborative events with companies, research institutions, and broader open source communities across the Linux Foundation ecosystem, including the PyTorch Foundation and Agentic AI Foundation.

The AI Infra SIG welcomes discussions and contributions across the Cloud Native AI infrastructure ecosystem, including but not limited to:

  • Scheduling: Dynamic Resource Allocation (DRA), Workload-Aware Scheduling, Kueue
  • Orchestration: JobSet, LeaderWorkerSet, KubeRay
  • AI Deployment Platforms: KServe, llm-d, N, AIBrix 
  • Networking: Gateway API Inference Extension, kgateway, Envoy AI Gateway
  • Agent Infrastructure: agentgateway (AAIF) and Agent Sandbox 
  • Standards and Conformance: AI Conformance and related standardization efforts

The AI Infra SIG organizers are (in alphabetical order):

  • Masaya Aoyama, Senior Software Engineer / Product Owner, CyberAgent, Inc.
  • Sunyanan Choochotkaew (Pang), Senior Research Scientist, IBM Research – Tokyo
  • Toru Komatsu (Toru), Engineer, Preferred Networks, Inc
  • Shingo Omura, Principal Architect of AI Infrastructure, LY Corporation
  • Kenta Tada, Member, CNCF End User Technical Advisory Board & eBPF Foundation Governing Board
  • Kenji Tagashira, Senior Manager, Fsas Technologies Inc

Join Our First Meetup & Call For Speaker!

To celebrate the launch of the SIG, we are delighted to invite you to our first meetup, themed “Let’s Get Started with the CNCJ AI Infra SIG — Where Are We Today?”

Together, we’ll explore the current landscape of Cloud Native AI infrastructure and discuss the latest developments across Kubernetes, CNCF, and the broader open-source ecosystem.

Call for Speakers: We’d also love to hear from practitioners, researchers, and contributors working on AI infrastructure, platforms, and open source projects. If you have experiences, insights, or emerging ideas to share, please consider submitting a session proposal or lightning talk.

Get Involved

The CNCJ AI Infra SIG is open to anyone interested in Cloud Native AI infrastructure—from engineers and platform builders to researchers, operators, and open-source contributors.

There are many ways to participate:

Whether you’re running AI workloads in production, contributing to upstream projects, exploring new technologies, or just getting started, we’d love to have you join us.

We look forward to seeing your proposals and welcoming you to the CNCJ AI Infra SIG community!

CNCJ AI Infra SIG発足のお知らせ & 第1回 Meetup 開催・スピーカー募集開始!

本記事は, Shingo Omura, Principal Architect of AI Infrastructure, LY Corporation, Sunyanan Choochotkaew (Pang), CNCF Ambassador, Senior Research Scientist, IBM Research – Tokyoによって執筆されました。

近年、生成AIや自律型エージェントをはじめとするAI/機械学習技術の進化は目覚ましく、それらを支えるインフラとしてのKubernetesやCloud Nativeエコシステムの重要性がかつてないほど高まっています。

このような大きな技術変革のうねりを受け、このたびCloud Native Community Japan (CNCJ)に、AI/機械学習ワークロードを動かすためのベストプラクティスや最適化、現場の実践知の共有にフォーカスした新しいコミュニティAI Infra SIG (Special Interest Group)を立ち上げます!

この記事では、新SIG設立の背景やミッション、そして記念すべき第1回 Meetupの開催と講演(プロポーザル)募集についてご案内します。

AI Infra SIG 設立の背景

AIワークロードによる新たなインフラ課題の台頭

これまで、Kubernetesを中心とするCloud Nativeエコシステムは、マイクロサービスに代表されるようなWebサービスを主なワークロードとして発展してきました。Webサービスはステートレスで、CPUとメモリを主眼に置き、比較的予測可能なリクエスト特性・スケーリング特性を持っていました。

しかし、LLM(大規模言語モデル)に代表される現代のAIワークロードでは、この前提が大きく変化しており、GPUなどの専用ハードウェアへの強い依存、複雑な分散処理、これまでとは異なるリクエスト特性やスケール特性といった、新たなインフラ課題を生んでいます。

こうしたギャップを埋めるため、Kubernetesのアップストリームコミュニティは近年「AI Readiness」を掲げ、驚くべきスピードで機能開発と標準化が進んでいます。

  • スケジューリングとリソース最適化: デバイス管理を柔軟にする Dynamic Resource Allocation (DRA) や、マルチテナントなジョブキュー制御を担うKueue、複数のPodをスケジューリングPrimitiveとして扱うWorkload Aware Scheduling
  • 分散ジョブのオーケストレーション: JobSet や LeaderWorkerSet といった、これまで以上に複雑な分散ワークロードのオーケストレーション。
  • ネットワークとエージェント: Gateway API Inference Extension によるAI推論向けのトラフィック制御や、安全なエージェント実行環境を目指す Agent Sandbox 。
  • AI Readiness要件の標準化: Kubernetes SIG Architectureを中心に議論が進む AI Conformance。「AI Ready」なKubernetesの機能要件の定義や、環境間での互換性・適合性を検証するための共通規格の策定。

このインフラの進化は、Kubernetes/CNCFだけにはとどまりません。Linux Foundation全体を見渡しても、AI/エージェントのエコシステムは急速に発展しています。2022年に設立されたPyTorch Foundationでは、PyTorchだけでなく、vLLM や Ray、AIBrix といったLLM/Agent向けの技術が、CNCFのプロジェクトと親和性を持ちながら活発に開発されています。さらに昨年末には、AIエージェント領域における技術の断片化を抑え、開発者と企業が信頼して使えるAI標準を確立すべくAgentic AI Foundationが設立され、2026年09月にはAGNTCon + MCPCon Japanも開催されます。

このように、AI/エージェントの潮流は、Cloud Nativeという枠を超え、インフラ、プラットフォーム、フレームワーク、Agentic AIの実行レイヤーに至るまで、コミュニティやレイヤーを超えて技術的なイノベーションがまさに現在進行形で起きています。

なぜ今、日本でAI Infra SIGを立ち上げるのか?

こうした世界規模で技術革新の波が押し寄せる中、日本国内においてもAI/MLインフラの構築・運用に取り組む企業や開発者が急増している一方で、急速に進化するCloud Native AIエコシステムのベストプラクティスを体系化し、実用的な知見として共有する場はまだ十分に存在していないのが現状ではないでしょうか。

こうした背景から、私たちは、国内のエンジニア、研究者、プラットフォーム開発者、企業、組織がCloud Native × AIのフロンティアを切り拓いていくためのコミュニティとしてAI Infra SIGを立ち上げます。

SIGの概要: ミッションと活動内容

CNCJ AI Infra SIGは、次の2つをコアミッションとして活動します。

  1. AI/機械学習ワークロードを実行するインフラとしてのベストプラクティス/最適化/実践知見の共有
  2. AI/機械学習ワークロード関連のCNCFアップストリームへの日本からの貢献拡充

主な活動予定

AI Infra SIGの主な関連トピック

  • Scheduling: Dynamic Resource Allocation, Workload Aware Scheduling, Kueue
  • Orchestration: JobSet, LeaderWorkerSet, KubeRay
  • Deployment Stack: KServe, llm-d, NVIDIA Dynamo, AIBrix
  • Networking: Gateway API Inference Extension, kgateway, Envoy AI Gateway
  • Agent Infra: agentgateway (AAIF), Agent Sandbox
  • Standardization: AI Conformance
  • その他AI Infraに関連する内容

AI Infra SIG Organizers

(Alphabetical Order)

  • Masaya Aoyama, Senior Software Engineer / Product Owner, CyberAgent, Inc.
  • Sunyanan Choochotkaew (Pang), Senior Research Scientist, IBM Research – Tokyo
  • Toru Komatsu (Toru), Engineer, Preferred Networks, Inc
  • Shingo Omura, Principal Architect of AI Infrastructure, LY Corporation
  • Kenta Tada, Member, CNCF End User Technical Advisory Board & eBPF Foundation Governing Board
  • Kenji Tagashira, Senior Manager, Fsas Technologies Inc

【第1回】CNCJ AI Infra SIG Meetup 開催決定!

SIGの設立を記念し、記念すべき第1回目のミートアップを開催します! テーマは「CNCFのAI Infraインフラとしての現在地を把握しよう」です!最先端のCloud Native AIインフラに興味があるエンジニアの皆様、ぜひご参加・ご登壇ください。

講演スピーカー(プロポーザル)を募集します!

第1回ミートアップを一緒に盛り上げてくださるスピーカーを募集します。上記の通り初回開催として、「CNCFのAI Infraインフラとしての現在地を把握しよう」をテーマ、幅広いトピックをカバーするセッションを募集します。 上記に挙げた技術トピックに関する解説・検証実績、自社でのAI/ML基盤の運用事例、アップストリームへのコントリビューション経験など、大小問わず大歓迎です!

「こんなテーマで話してみたい」「まだ検証段階だけど知見を共有したい」という方も、ぜひお気軽にご応募ください。

  • 募集セッション枠:
    • 一般セッション(20〜30分程度)
    • ライトニングトーク(LT)(5〜10分程度)
  • 応募方法: 以下のGoogle Formより必要事項をご記入の上、ご応募ください。応募多数の場合はご講演いただけない場合があります。
  • 応募URL: https://ocgroups.dev/cncf/group/bqd97by/event/7fu3gfg 
  • 応募締め切り: 2026/8/28

おわりに

AIや自律型エージェントがもたらす新しいパラダイムは、アプリケーションの作り方だけでなく、それを支えるインフラやプラットフォーム全体の設計、そして運用のあり方を根本から再定義しつつあります。これまでWebシステムの運用で私たちが培ってきたCloud Nativeの強みやアプローチを、この新しいAIというフロンティアにおいてどう進化させ、スケールさせていくのか。まさに今、その非常に面白い転換点にいるのではないでしょうか。

このAI Infra SIGは、最先端を追いかけるだけの場ではありません。国内のエンジニアやプラクティショナーが直面している課題や実践知を共有し体系化していく場になっていけば良いなと思っています。そしてCNCFやLinux Foundationのアップストリームへと還元を促進していけたら良いなとも思っています。

この刺激的な技術変革をコミュニティ一丸となって、ぜひ一緒に盛り上げ、創り上げていきましょう。

皆様からのプロポーザル、そしてミートアップへのご参加を、心よりお待ちしております!

  •  

Operating OpenTelemetry at scale with OpAMP

As more organizations move to use OpenTelemetry in production at scale, with multiple Collectors across heterogeneous environments, a new challenge arises: how to remotely manage, configure, and update this agent fleet in a consistent and secure way?

This is where Open Agent Management Protocol (OpAMP) comes into the picture: it provides a standardized protocol that lets a central backend automatically configure agents, push updates, monitor their health, and collect status information.

In a recent episode of OpenObservability Talks, I sat down with Andy Keller, OpAMP maintainer and Principal Engineer at BindPlane, to hear what OpAMP is and how it makes large-scale observability deployments much easier to operate and control. We also covered project status and roadmap, including a hot KubeCon update you don’t want to miss.

OpenObservability Talks: Operating OpenTelemetry at Scale with OpAMP

Why OpAMP: The management challenge at scale

As OpenTelemetry adoption has exploded, organizations are finding themselves managing increasingly complex collector deployments. Before OpAMP, the landscape was fragmented and challenging. Andy shared their journey: “We probably developed in-house three, four, maybe five different agent management protocols. Some were HTTP-based, long polling. We used WebSockets. We used protobufs. We used JSON.”

The problem becomes acute when you consider the scale and variety of deployments. We’re not just talking about a handful of collectors — organizations are deploying collectors everywhere from massive gateways to embedded devices. Each deployment model brings its own management challenges, and the teams responsible for deploying collectors are often different from the observability teams who need to configure them. This disconnect creates operational friction that can undermine your entire observability strategy.

Scale and variety of OTel Collector deployments

I found the sub-story about the diversity of OpenTelemetry collector deployments staggering. “We see anything from a couple massive OpenTelemetry gateways where really what you’re doing is managing the configuration of the gateway and doing all the processing there,” said Andy, “but then we also even see people deploying collectors to embedded devices. We have collectors in point of sale machines. We have collectors on laptops collecting Windows events for security tracking.”

The scale ranges from dozens to millions of collectors. When you factor in IoT and embedded use cases, the numbers become truly massive. As Andy noted, when you get into the embedded space, it gets to millions of collectors that you need to start reasoning about.

What is the Open Agent Management Protocol (OpAMP)

OpAMP (Open Agent Management Protocol) is a standardized protocol that provides remote management capabilities for observability agents, primarily the OpenTelemetry Collector (which is why it resides under the OpenTelemetry project). It enables central backends to automatically configure agents, push updates, monitor their health, and collect status information — all in real-time over WebSocket or HTTP connections.

Managing OpenTelemetry Collectors with OpAMP. Source: opentelemetry.io

Managing OpenTelemetry Collectors with OpAMP. Source: opentelemetry.io

What’s particularly interesting is how OpAMP has evolved beyond simple configuration management. As Andy explained: “It started to really focus on configuration management and with agent health and component health and things like that, really moving into this observability for your observability realm. Because observability is something that is so critical to operations that you need to know is your observability actually working?”

To me, this evolution reflects a crucial insight: your observability infrastructure is too important to be a black box. You need observability for your observability. OpAMP addresses this by providing real-time visibility into collector health, configuration drift, and operational status. There was a great talk at last year’s KubeCon North America 2025, in which Nike’s observability platform engineers shared how they built an enterprise-grade implementation of OpAMP for their scale and use case.

OpAMP protocol and components

OpAMP, as the name suggests, is first and foremost a network protocol specification, used to remotely manage large fleets of data collection Agents. The protocol is elegantly simple: just two messages — server-to-agent and agent-to-server — defined using Protocol Buffers. The specification lives in the opamp-spec repository under OpenTelemetry, while opamp-go provides the reference implementation in Go.

The architecture includes several key components. The OpAMP extension is a read-only component that reports current configuration and health status. The OpAMP supervisor sits as a separate process alongside the collector, implementing both read and write capabilities. As Andy described it: “It kind of sits between the management platform and the collector. It speaks to the collector on behalf of the management platform, and it can accept changes.”

The supervisor’s approach is particularly clever — it writes new configurations to disk, shuts down the collector, and restarts it with the new configuration. Critically, it includes safety mechanisms: “If it doesn’t start, it will revert the config and run with the last known good config so that we’re not breaking your telemetry pipelines remotely.”

Supervisor-based management with OpAMP. Source: gihub.com/open-telemetry

Supervisor-based management with OpAMP. Source: gihub.com/open-telemetry

Beyond OTel Collector: OpAMP for Kubernetes, SDKs and more

What makes OpAMP powerful is its protocol-level flexibility. The configuration payload is intentionally generic — just a map of name-value pairs. This allows OpAMP to manage not just OpenTelemetry collectors, but any type of agent. In fact, it’s already used to manage SDKs and Kubernetes deployments.

To manage Kubernetes deployments, OpAMP utilizes the OpAMP Bridge, which acts as an intermediary between OpAMP-speaking management platforms and Kubernetes-native deployment mechanisms. Andy explained the architecture: “rather than communicating with [OTel] Collectors, you’re communicating with this OpAMP Bridge. The OpAMP Bridge is communicating within the cluster with the OpenTelemetry Operator, and that Operator reads CRDs and deploys Collectors.”

Andy mentioned that the OpenTelemetry Java SDK can also speak OpAMP and receive remote configuration. The use cases are compelling: imagine remotely adjusting sampling rates across your entire microservices architecture to investigate an issue, or enabling debug logging for specific services without redeploying. Reconfiguring SDKs, however, requires a different operational model, as we can’t shut down applications to reconfigure the SDK. SDKs require hot-reloading capabilities rather than the restart-based approach used for collectors.

Another interesting use case is managing a fleet of Fluent Bit agents. This brand-new project was recently open-sourced these days by Phil Wilkins, with Fluentd support planned next. Check out the GitHub repo and Phil’s blog post for more details.

The protocol’s agent-agnostic design is intentional. The remote config message is just a map of name-value pairs, where that value can be anything. This flexibility means OpAMP can manage security agents, custom telemetry collectors, or any other agent-based software. The protocol defines the communication contract, but doesn’t dictate what agents do with the configuration they receive.

Hot off the press: OpAMP Gateway Extension

One of the most exciting developments is the upcoming OpAMP Gateway Extension launching these days around KubeCon Europe 2026. This addresses a critical scaling challenge: WebSocket connection limits.

Andy described the problem and solution: “Let’s say I’ve got 100,000 collectors deployed across my many different clusters in my organization, instead of all the 100,000 [collectors] connecting to the management platform, I can deploy 100 OpAMP gateways, have 1,000 collectors connect to each one, and then those 100 connect to the management platform.”

OpAMP Gateway — Connection Fan-In. Source: bindplane.com

OpAMP Gateway — Connection Fan-In. Source: bindplane.com

The OpAMP Gateway is an OpenTelemetry Collector extension. It runs inside a collector and acts as a multiplexer, aggregating OpAMP messages from thousands of edge collectors and relaying them through a smaller number of upstream connections. You can think about it as similar to how OpenTelemetry gateways work for telemetry data, but for the control plane.

The benefits are substantial: reduced connection overhead on management platforms (i.e., OpAMP Server), support for network-segmented environments where edge collectors can’t directly reach external management systems, and more efficient use of network resources. For organizations operating at IoT scale — millions of collectors — the gateway becomes essential infrastructure. The OpAMP Gateway Extension is launched in Alpha, check out the launch blog for more details.

OpAMP roadmap

OpAMP is currently in beta, with different components at varying maturity levels, and the community is actively working toward stability.

Configuration diff support is another priority — sending only configuration changes rather than complete configurations becomes critical when dealing with large, complex collector configurations. There’s also significant interest in true hot-reloading capabilities that wouldn’t require collector restarts.

Perhaps most intriguing is the telemetry policy OTEP (OpenTelemetry Enhancement Proposal) currently in draft. This would introduce policy as a concept distinct from configuration — communicating intent (like “filter out these log messages” or “add this attribute”) rather than specific implementation details. The SDK and collector could then implement the same policy differently based on their capabilities.

Andy expects additional SDKs to support OpAMP, expanding remote management capabilities to more languages and platforms.

Want to learn more? Check out the OpenObservability Talks episode: Operating OpenTelemetry at Scale with OpAMP.

 

Ambassador post originally published on Medium by Dotan Horovits.

  •  

Building a Cluster-Aware AI Agent with Kubernetes, Argo CD, and GitOps

A practical walkthrough of running a self-hosted, read-only AI agent inside a Kubernetes cluster, with the full CI/CD chain handled by GitHub Actions and Argo CD Image Updater. No data leaves the cluster, no cloud AI provider involved.


Why a Cluster-Aware Agent Is an Interesting Pattern

Most “AI for Kubernetes” tooling today is a hosted SaaS that consumes cluster data and returns advice. The model lives elsewhere. The data leaves the network.

This article walks through the opposite design: an agent that runs inside the cluster, observes live state through the Kubernetes API, and reasons with a local LLM. Every layer is visible, every credential is scoped, and the only network egress is a model pull at startup.

The interesting properties of this pattern for platform engineers:

PropertyWhat it provides
Cluster-awareThe agent reads live pods, events, and logs and reasons about real state rather than generic Kubernetes facts.
Read-only by designA dedicated ServiceAccount + ClusterRole with get/list verbs only. The agent can observe but cannot mutate the cluster, regardless of what the model produces.
Just another K8s workloadThe agent is a Deployment + Service + PersistentVolumeClaim. No special runtime, no operator, no custom scheduler.
Full GitOpsPrompts, model selection, and RBAC live in Git. Argo CD reconciles them. The agent’s behavior is auditable through git log.

Source code: github.com/MaryamTavakkoli/local-k8s-ai-agent


LLM vs. AI Agent: The Distinction That Matters

A Large Language Model answers from training data alone. It has no awareness of the environment it’s deployed into. An agent, in the sense used here, performs an extra step before reasoning: it observes the real world and incorporates that observation into the prompt.

Graphic: LLM Alone vs. AI Agent

The contrast in output is concrete. A generic LLM call returns “CrashLoopBackOff usually means the container is failing health checks or exiting unexpectedly…” An agent call returns “Pod api-7b8d has restarted 14 times in the last hour with ImagePullBackOff against registry.local. Run kubectl describe pod api-7b8d to confirm.”

The second answer is grounded. The first answer is correct but not actionable for this cluster.

This project demonstrates both modes through two REST endpoints:

  • POST /ask — LLM alone, useful for general questions like “What is a StatefulSet?”
  • POST /diagnose — the agent: reads live cluster state, then reasons over it

Architecture

The system has two halves: a CI/CD chain on the top and a Kubernetes runtime on the bottom.

Runtime side:

  • An Ollama pod serves a local Mistral 7B model on port 11434
  • A FastAPI pod exposes the agent’s HTTP API and chat UI on port 8000
  • A PersistentVolumeClaim holds the model weights so pulls aren’t repeated
  • A dedicated ServiceAccount mounted in the FastAPI pod has a ClusterRole permitting only read operations on pods, events, logs, services, and deployments

Delivery side:

  • A push to the application source in Git triggers GitHub Actions to build a multi-architecture image (linux/amd64 + linux/arm64) tagged with the 7-character commit SHA
  • Argo CD Image Updater (from argoproj-labs) polls Docker Hub on a 2-minute interval, detects new tags matching the configured regex, and commits the new tag back into the repository’s kustomization.yaml
  • Argo CD detects the manifest change and reconciles the cluster
Graphic: Architecture: Local AI Agent on Kubernetes

The two halves are decoupled. Argo CD has no awareness of the registry. GitHub Actions has no awareness of the cluster. Image Updater is the small operator that bridges them, and it does so by writing to Git, which preserves a single source of truth.


The AI Concepts You’ll Actually Touch

Here are a few concepts that every AI engineer uses every day, explained in plain language.

1. LLM (Large Language Model)

A statistical model trained on enormous amounts of text. It doesn’t “know” facts; it predicts the most likely next word given everything that came before. That’s it. The magic is that this simple task, done at scale, produces something that feels like reasoning.

This project uses Mistral 7B, a 7-billion-parameter open-source model. “Parameters” are the numbers the model learned during training, similar to the strengths of connections in a brain.

2. Local LLM

Most commercial AI services send your text to a remote cloud provider. The trade-off is capability: a local 7B model isn’t as expansive as a massive foundational model running on cloud infrastructure.. But for experimenting, it’s more than enough. And nothing leaves your network.

3. Ollama (The Model Serving Runtime)

Ollama is not an AI model. It’s a server that runs AI models. Think of it like a web server for LLMs: it downloads the model files, loads them into memory, and exposes a REST API on port 11434 so anything (including our FastAPI app) can send prompts and get responses.

Without Ollama, you’d be wrestling with PyTorch, CUDA, and tokenizer libraries. With it, running an LLM is ollama pull mistral followed by an HTTP POST.

4. System Prompt (The Personality)

This is the single most important AI concept for application developers, and you can master it in about ten minutes.

A system prompt is the instructions you give the model before the user’s question. The model reads it first and uses it to shape every response.

In our project, the system prompt for /ask is:

“You are a DevOps assistant specializing in Kubernetes.

When given an error or question, you:

1. Explain what it means clearly

2. Provide the exact kubectl commands to diagnose or fix it

3. Explain why the fix works

Be concise and practical.”

Without that prompt, Mistral is a general assistant. With it, Mistral is a Kubernetes specialist who always returns structured answers. No retraining was needed. This is called prompt engineering, and it’s how almost every AI product you use was built.

5. RAG (Retrieval-Augmented Generation)

The fancy term for what /diagnose does. RAG means: before asking the model, retrieve real-world data and augment the prompt with it.

RAG is why contemporary AI assistants work. A code assistant reads your local workspace repository; our agent reads your live cluster state.. Same pattern, different data source.


The Two Modes: Where the Agent Becomes Real

Here’s where the “agent” idea earns its name.

Mode 1: Ask (LLM alone)

You type a question, FastAPI prepends the system prompt, sends it to Ollama. The model answers from its training data. Useful for general K8s questions like “What is a StatefulSet?”

Mode 2: Diagnose Cluster (true agent)

You type a question and a namespace. FastAPI does something new: it calls the Kubernetes API and reads:

All pods in that namespace (phase, restart count, waiting reason)

The last 10 events

The last 20 lines of logs from any non-Running pod

That entire context is injected into the prompt. Then Mistral reasons, but now it’s reasoning about your actual cluster, not generic Kubernetes knowledge.

Graphic: Local K8s AI Agent

The chat UI even shows you the exact context the agent read, in a collapsible panel under each answer.


Read-Only by Design

The agent runs with a ServiceAccount bound to a ClusterRole that exposes only read verbs:

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: ai-devops-api-reader
rules:
  - apiGroups: [""]
    resources: ["pods", "pods/log", "events", "services", "configmaps", "namespaces"]
    verbs: ["get", "list"]
  - apiGroups: ["apps"]
    resources: ["deployments", "replicasets", "statefulsets", "daemonsets"]
    verbs: ["get", "list"]

This is the most important design decision in the project, and it generalizes beyond AI workloads. An agent that can delete pod based on its own reasoning is a production incident waiting to happen. Hallucinations multiplied by write access is a poor combination.

Read-only RBAC inverts the trust model. The agent is allowed to be wrong because being wrong has no consequences. The Kubernetes API server enforces the boundary; the LLM’s output cannot bypass it. Iteration on prompts and models becomes cheap because the worst-case behavior is bounded.

The same pattern scales: start every agent read-only, then earn each additional capability one verb at a time, each with its own RBAC rule and review.


The CI/CD Chain in Detail

The delivery half of the architecture uses three independent components, each with one responsibility.

GitHub Actions builds and pushes the image. The workflow uses docker buildx with QEMU emulation to produce a manifest list covering both linux/amd64 (GitHub-hosted runners) and linux/arm64 (Apple Silicon developer machines). The tag is the 7-character commit SHA, an immutable reference.

Argo CD Image Updater polls the registry on a 2-minute interval. Configuration lives in an ImageUpdater Custom Resource that names the target Argo CD Application, the image to track, an allowTags regex (^[0-9a-f]{7}$), and the update strategy (newest-build). When a new matching tag is found, the operator rewrites the newTag field in k8s/kustomization.yaml and commits the change to the main branch.

apiVersion: argocd-image-updater.argoproj.io/v1alpha1
kind: ImageUpdater
metadata:
  name: local-k8s-ai-agent
  namespace: argocd
spec:
  writeBackConfig:
    method: git
    gitConfig:
      branch: main
      writeBackTarget: "kustomization:."
  applicationRefs:
    - namePattern: "local-k8s-ai-agent"
      images:
        - alias: api
          imageName: marytvk/local-k8s-ai-agent
          commonUpdateSettings:
            updateStrategy: newest-build
            allowTags: "regexp:^[0-9a-f]{7}$"

Argo CD watches the repository and reconciles the cluster on each commit. Because the source manifests are managed by Kustomize, Argo CD applies the rendered output, which now includes the updated image tag.


Try It Yourself: It’s a Starting Point, Not a Destination

The repo is here: github.com/MaryamTavakkoli/local-k8s-ai-agent

Step-by-step setup with exact commands is in the README. Total time from git clone to working chat UI is about 30 minutes (most of that is the Mistral download).

To bring this article full circle: if you’re a DevOps or Platform Engineer who’s been hearing “AI agents are coming” and wondering what that actually means in practice, this is meant to be your starting point, not your finish line. Once you’ve seen the agent loop running, you’ll be in a much better position to continue.

The point of starting local isn’t that local is always the right answer. It’s to understand how the full circle works behind the scenes.

  •