❌

Vue lecture

Google is a Leader in the 2026 Gartner Magic Quadrant for Container Management

We’re excited and proud to share that Gartner has recognized Google as a Leader for the fourth year in a row in the 2026 Gartner® Magic Quadrant™ for Container Management, based on its Completeness of Vision and Ability to Execute. Google was positioned highest in Ability to Execute of all vendors evaluated and we believe this validates the success of our mission to deliver a container platform that’s highly optimized for both performance and efficiency. We help global customers to build and run their most demanding and complex workloads at scale, including the next generation of AI and agentic applications. 

In the accompanying 2026 Gartner Critical Capabilities for Container Management report, Google Cloud was ranked first in every use case: New Cloud Native Applications, Containerized Existing Applications, AI Training, AI Inference, Edge Applications, and Hybrid Applications.

Gartner predicts1 that “By 2028, 95% of new AI deployments will use Kubernetes, up from less than 30% in 2025.” Containers power today’s most innovative apps and businesses — and deliver the infrastructure customers demand as they transform their businesses in the agentic era.

2026 Gartner Magic Quadrant for Container Management

Google Cloud spearheaded the industry-wide cloud-native revolution when we introduced Kubernetes in 2014 and launched Google Kubernetes Engine (GKE), the world’s first managed Kubernetes service, in 2015. Our commitment to container platforms and the vibrant, innovative Kubernetes ecosystem has only grown stronger and deeper since. Alongside GKE, our serverless container platforms GKE Autopilot and Cloud Run dramatically lower operational costs and help developers deliver amazing containerized apps faster than ever before. 

The massive acceleration in enterprise AI has inspired us to redefine infrastructure management for the AI era. In 2026 so far we’ve introduced a wide range of foundational improvements to shift GKE and Cloud Run into agent-native, high-performance platforms designed for autonomous AI systems, massive inference workloads, and secure runtime isolation. Whether you’re training AI at the frontier, launching an AI startup, or leading your enterprise AI transformation, we have the container platform you need. Important highlights include:

Delivering leading performance and efficiency for AI infrastructure

  • GKE predictive latency boost: Built into the GKE Inference Gateway, this ML-driven capability uses capacity-aware routing rather than static configurations to reduce Time-to-First-Token (TTFT) by up to 70%.

  • GKE automatic KV Cache storage tiering: Automatically shifts KV cache data across RAM, Local SSD, and Cloud Storage. This reduces memory bottlenecks, improving TTFT by 40% via RAM offloading and increasing throughput by 70% via Local SSDs for large prompt contexts. [1]

  • GKE accelerated container and model startups: GKE node spin-up times are up to 4x faster, and pod startup speeds have improved by up to 80%. Additionally, native run:AI Model Streamer integration pulls heavy models from Cloud Storage 5x faster.

  • Cloud Run on-demand serverless GPU scale-to-zero: Cloud Run supports NVIDIA RTX PRO 6000 Blackwell GPUs, allowing teams to serve 70B+ parameter models on-demand. Your services can go from zero to a fully provisioned GPU — with all drivers pre-installed — in under 5 seconds. Once active inference or fine-tuning runs complete, Cloud Run automatically scales instances back to zero, eliminating idle infrastructure costs.

Evolving Kubernetes for agentic infrastructure security and scale

  • GKE Agent Substrate: As an open-source, secure-by-default agent execution runtime, Agent Substrate is engineered to run millions of sandboxes with 10x higher density than standard container runtimes. Purpose-built for the era of autonomous agents, Substrate delivers sub-500ms resume operations at over 500 suspend/resume activations per second with a native zero-trust kernel and network isolation. Agent Substrate is available as an open-source solution that runs on any Kubernetes infrastructure and is optimized for GKE.

  • GKE Agent Sandbox: Built on gVisor kernel-isolation technology, Agent Sandbox isolates the host environment from untrusted, multi-agent AI code execution. It provides secure execution at scale, processing up to 300 sandboxes per second with sub-second latency and delivering up to 30% better price-performance when running on Axion processors than comparable hyperscaler cloud providers. 

  • GKE Dataplane V2 scalability limits: Architectural capacity bounds for GKE clusters implementing active NetworkPolicies doubled from 7,500 nodes to 15,000 nodes per cluster, supporting the massive infrastructure needs of large enterprise and AI customers.

  • GKE intent-based autoscaling: GKE can now natively autoscale horizontally using application intent and custom metrics beyond basic hardware metrics. This reduces resource allocation reaction times from 25 seconds down to just 5 seconds.

  • Filestore agent volumes: a new offering that attaches and detaches NFS mounts in milliseconds, allowing agents to start/resume near-instantaneously, along with native Read-Write-Many (RWX) access and POSIX-compliant file locking to enable safe multi-agent collaboration without write collisions. 

Next-gen developer experience with serverless containers

Whether you’re hosting a standard web API, running a heavy batch data job, processing an asynchronous message queue, or deploying a complex AI agent, Cloud Run handles it all under a single, unified serverless model that delivers an unmatched developer experience and maximum engineering velocity. 

  • One-click prototyping in Google AI Studio: You can build and deploy full-stack applications directly within Google AI Studio, making it an exceptional environment for rapid prototyping and experimentation. With a single click, you can instantly package and publish your vibe-coded applications to Cloud Run.

  • Cloud Run instances: This new primitive manages individual, addressable, long-running singleton resources with integrated Cloud Storage volume mounts, allowing persistent background agents like OpenClaw to be deployed cost-effectively. With baseline shared-CPU configurations starting at a highly predictable flat rate of ~$5.70 per month (for 1 vCPU and 1 GiB of RAM), Cloud Run instances delivers an always-on, VM-like experience while bypassing the idle-cost penalties and operational overhead of traditional VMs.

  • Cloud Run sandboxes: Hard-isolated environments spin up in under 500 milliseconds to safely execute untrusted, model-generated code, protecting the host system from unauthorized access.

Take the next steps

As we reach for new heights of performance, security, and scale for our container platforms, we continue to build the future in the open. We invite you to explore Agent Sandbox and Agent Substrate today. We can’t wait to shape the future of agent infrastructure together with our customers and partners. Check out these resources to continue your learning journey:


1. Gartner report: Critical Capabilities for Container Management, 8 September 2026

Gartner, Magic Quadrant for Container Management, Dennis Smith, et al, 2 September 2026
Gartner, Critical Capabilities for Container Management, By Tony Iams, Wataru Katsurashima, Lucas Albuquerque, Dennis Smith, Bhuvie Chhabra, 8 September 2026. 
Gartner and Magic Quadrant are trademarks of Gartner, Inc. and/or its affiliates.
Disclaimer: Gartner does not endorse any company, vendor, product or service depicted in its publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner publications consist of the opinions of Gartner’s business and technology insights organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this publication, including any warranties of merchantability or fitness for a particular purpose.

  •  

Introducing GKE agentic migration for AI-assisted EKS-to-GKE migrations with built-in governance

Enterprises are increasingly standardizing on Google Kubernetes Engine (GKE) to run their most critical and AI-driven workloads. From Cloud Storage FUSE for high-throughput data access to custom compute classes (CCC) and advanced GPU slicing, GKE provides the scale and efficiency required for modern applications.

However, migrating complex Kubernetes environments from AWS EKS to GKE has traditionally been a daunting, high-friction engineering endeavor. Your platform teams must manually dissect sprawling infrastructure-as-code (IaC), navigate cloud-specific architectural differences, and build custom translation scripts.

While your engineering teams often experiment with general-purpose LLMs to draft conversions, ad-hoc prompting quickly can become an operational trap. Raw models hallucinate non-existent resource properties, drop critical network or identity configurations, and lose context across interdependent files. The time platform engineers spend auditing, untangling, and debugging model errors ends up cannibalizing any upfront speed gains, creating manual toil and unpredictability. 

Today, we are excited to announce the open-source release of GKE agentic migration, a purpose-built agent plugin that replaces brittle, ad-hoc prompting with an AI-assisted migration pipeline protected by deterministic guardrails. 

“For large enterprise clients, the biggest barrier to cloud modernization is execution risk and unpredictability. Unlike raw chat prompts that lose context and hallucinate configurations, Google’s GKE agentic migration pairs the speed of generative AI with the deterministic guardrails enterprises need: structured state persistence, multi-persona boundaries between platform and app teams, and non-negotiable human approval gates. It gives our global engineering practice a provable, compiler-grade migration factory that slashes delivery risk.- Rahul Shrivastava, EVP, Persistent

The challenges of infrastructure migrations

When talking to customers about their infrastructure migration journeys, we consistently hear about several governance challenges:

  • The automation trust gap: Refactoring Kubernetes configurations manually can be agonizingly slow. Yet, using generic AI coding assistants introduces unacceptable risk. Standard LLMs can hallucinate infrastructure code, use deprecated API fields, or omit critical security rules. Generating code that is "almost right" simply shifts the bottleneck from writing code to debugging it.

  • The danger of live cluster mutability (ClickOps): Legacy migration tools often connect directly to live clusters and deploy via API calls. This bypasses the organization's Git repository (the true source of truth), breaks CI/CD pipelines, and makes rollbacks incredibly difficult.

  • The siloed handoff bottleneck: Migrations are often long-running, multi-week operations. Platform engineers build the landing zone and your application developers migrate the workloads. Standard AI tools lose context across the handoff.

  • The fragmented toolchain: Backup tools like Velero are excellent for disaster recovery but capture exact AWS-specific configurations (like ALBs) without translating them for Google Cloud. Reverse-engineering tools, meanwhile, generate flat configurations that strip away the developer's original logical intent.

Introducing the GKE agentic migration

The GKE agentic migration addresses these challenges by combining the reasoning capabilities of LLMs with strict, deterministic tooling. Designed as a compilation of agent skills and a local Model Context Protocol (MCP) server, it uses AI to translate complex AWS EKS IaC and Kubernetes manifests directly into GKE landing zones via automated Pull Requests.

Here are the key capabilities that set the GKE agentic migration apart:

1. Hybrid verification — LLM-generated, deterministically validated. To combat dangerous IaC hallucinations, LLM workers handle the complex authoring of Terraform and Kubernetes YAML, while the server runs deterministic transforms for exact mappings such as Workload Identity annotations and image registries. Crucially, these AI-generated translations are then submitted to strict deterministic validations (e.g., terraform validate, Kubernetes manifest contracts) before they are presented to the user. This approach helps maintain safety against hallucinations while gating everything behind human-in-the-loop (HITL) approval.

2. GitOps-native PR workflows: The plugin never applies changes directly to a live cluster. Instead, it reads your source of truth, generates the target state, and opens a Pull Request. This helps route all changes through your standard human-in-the-loop (HITL) CI/CD review process. No "ClickOps."

3. Protected separation of translation vs. transport: The plugin automates the tedious logic of architectural translation, but it intentionally does not transport stateful data. To protect your most sensitive assets, the plugin generates contextual runbooks that guide your team in using purpose-built, SLA-backed tools (like Google Cloud's Database Migration Service or Storage Transfer Service).

4. Multi-persona state management: Migrations are team efforts. The plugin persists the long-running migration state.  This enables protected, asynchronous handoffs: Platform engineers establish the baseline landing zone, while app developers independently join the workspace from their own machines to translate individual workloads within permission-isolated folders.

How it works: The migration lifecycle

Under the hood, the GKE agentic migration utilizes a migration state graph of executable functions, systematically passing context down the chain. Packaged as an open-source agent plugin, there are no custom CLI binaries to install and no central control planes to manage — your team collaborates through your existing development harness, delivering validated pull requests and actionable runbooks directly into your source repositories. This provides:

  • Deep EKS repository discovery: The plugin clones the source Git repository or performs a live scan of your EKS cluster, programmatically indexes the source manifests, maps dependencies, and builds an inventory
  • Assessment & blocker governance: It generates a readiness report identifying architectural incompatibilities. Before design can unlock, every blocker must have an assigned owner and target resolution date. The Platform Engineer signs off on the migration boundaries before translation begins.
  • Landing zone design: The plugin scaffolds the foundational Google Cloud Terraform modules (VPC, subnets, GKE cluster, org policies) based on explicit platform decisions (such as GKE Autopilot vs. GKE Standard).
  • AI-assisted cloud translation: The plugin handles proprietary shifts, including translating AWS IRSA to Workload Identity, mapping ALB ingress to the Gateway API, and converting Karpenter node claims to GKE Node Auto Provisioning (NAP) or Custom Compute Classes (CCC).
  • Offline validation: Generated modules and manifests are compiled and verified offline (terraform validate, manifest structure checks, and output contracts). 
  • Deployment via Pull Request: The finalized configuration is verified locally and opens a PR for review. 

Getting started

The GKE agentic migration transforms cloud migrations from disjointed refactoring exercises into predictable, AI-assisted, and reviewable GitOps workflows. Ready to accelerate your journey to GKE?

  •  

Scale your own way, using HPA with built-in support for PromQL metrics queries in GKE

Earlier this year, we announced native support for Google Kubernetes Engine (GKE) custom metrics. This milestone allowed you to scrap external adapters and instead collect autoscaling metrics directly from your pods. By routing these metrics straight to the Horizontal Pod Autoscaler (HPA), we cut metrics reading latency down to 5 seconds.

Today, we are excited to introduce built-in support for processing Prometheus metrics, allowing you to use expressive PromQL queries to customize autoscaling triggers. With this update, HPA can now directly process autoscaling metrics present in Cloud Monitoring using Google Managed Service for Prometheus. Reading metrics from these backends will not require third-party adapters, leveraging the AutoscalingMetric integration used to support pod-level metrics. After the preview, we plan to support self-hosted Prometheus servers as we move to general availability. 

The challenge: Setting up Cloud Monitoring metrics

Support for custom pod-level metrics made autoscaling more straightforward, but production workloads often need to scale on multiple, complex infrastructure metrics. Common examples include scaling:

  • a worker pool based on the number of unacknowledged messages in a Pub/Sub topic

  • an inference service based on query-per-second (QPS) metrics stored in Cloud Monitoring / Prometheus

  • a webserver farm based on the 95th percentile of their measured response time

To achieve this, you used to need to deploy an external adapter like the Stackdriver Custom Metrics Adapter or the Prometheus adapter to retrieve the metrics from an external logging environment. While this sounds straightforward at first, these adapters introduce a lot of operational friction:

  • Management overhead: Platform teams have to install, configure, patch, and monitor these third-party components.

  • Reliability and inefficiency: Intermediate adapter pods reading from external systems introduce failure points in critical autoscaling loops. 

  • IAM complexity: Enabling secure cross-component communication requires setting up Kubernetes service account mappings to Cloud service accounts including their permissions.

And while setting up this system and maintaining it not impossible, it’s complex and features a complicated architecture:

1

How processing Prometheus Metrics in GKE can help

Extending the AutoscalingMetric object drastically simplifies this setup. Now you can read metrics from monitoring directly via PromQL and provide them to HPA via a high-performance, low-latency autoscaling pipeline, resulting in a simplified environment.

2

To prevent inefficiencies, we built this feature with minimal resource consumption in mind. The controller runs on the GKE control plane. It monitors your AutoscalingMetric custom resources and only deploys the system pod on your user nodes when a PromQL metric is actively requested. If no Prometheus metrics are configured, the controller is shut down, so there’s no resource overhead.

Configuring built-in Prometheus metrics

Configuring GKE to use PromQLl metrics is easy; here’s a sample configuration file providing PubSubs message queue depth as scaling metric:

code_block
<ListValue: [StructValue([('code', 'apiVersion: autoscaling.gke.io/v1beta1\r\nkind: AutoscalingMetric\r\nmetadata:\r\n name: gmp-metric\r\nspec:\r\n metrics:\r\n - promql:\r\n name: pubsub-queue-depth\r\n query: |\r\n {\r\n "pubsub.googleapis.com/subscription/num_undelivered_messages",\r\n subscription_id="my-subscription"\r\n }'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f522d455250>)])]>

Linking Prometheus metrics to your HPA

Once defined in your AutoscalingMetric resource, you can reference the metric in your standard HorizontalPodAutoscaler using the same intuitive format as raw custom metrics: autoscaling.gke.io|<custom-resource-name>|<metric-name>.

Scaling globally (Prometheus metric)

For global metrics like a queue size that returns a single aggregate value:

code_block
<ListValue: [StructValue([('code', 'apiVersion: autoscaling/v2\r\nkind: HorizontalPodAutoscaler\r\nmetadata:\r\n name: worker-hpa\r\nspec:\r\n scaleTargetRef:\r\n apiVersion: apps/v1\r\n kind: Deployment\r\n name: worker-deployment\r\n maxReplicas: 10\r\n metrics:\r\n - type: External\r\n external:\r\n metric:\r\n name: autoscaling.gke.io|gmp-metric|pubsub-queue-depth\r\n target:\r\n type: AverageValue\r\n averageValue: 100 # maintain queue size at ~100 per pod'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f522d485390>)])]>

Scaling on Cloud Monitoring per-Pod metrics

GKE natively supports scale based on the most recent gauge metric values, but PromQL offers greater flexibility, allowing you to scale across time windows and calculate rates or histogram percentiles.

To use this capability, configure your PromQL metric to include a label for the pod name, then assign type: Pods within your AutoscalingMetric manifest. Below is an example that calculates a Pod's average memory usage over a five-minute rolling window.

code_block
<ListValue: [StructValue([('code', 'apiVersion: autoscaling.gke.io/v1beta1\r\nkind: AutoscalingMetric\r\nmetadata:\r\n name: per-pod-stored-metric\r\nspec:\r\n metrics:\r\n - promql:\r\n name: container-memory-metric\r\n query: |\r\n sum by ("pod")\r\n (avg_over_time({"container_memory_working_set_bytes"}[5m]))\r\n type: Pods # The promql query returns per-pod metrics'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f522d486c10>)])]>

Key benefits

  • No adapter maintenance: No pods to install, configure, or upgrade. The entire lifecycle is fully managed within GKE.

  • Streamlined security: Out of the box, the Kubernetes Default Node Service Agent has read permissions to Cloud Monitoring and Google Managed Prometheus in the same project. No extra IAM service accounts, keys, or federation parameters are required.

  • Low latency and fast scalability: The new Autoscaling Metric system polls the backend every 15 seconds, helping ensure fast scaling reactions.

  • Rich query capabilities: Leverage the full power of PromQL (including rate calculations, averages, and percentiles) to translate high-level business and user-experience objectives directly into scaling.

  • Support for the new HPA scale-to-zero capability: Utilize it for scaling workloads to zero replicas when demand hits zero (e.g., Pub/Sub queue size) and, more crucially, back up from zero replicas quickly using CapacityBuffers API.

Try it today 

By natively supporting both custom container metrics and Prometheus metrics, GKE now  offers a more robust, performant, and low-friction autoscaling experience. Built-in support for Prometheus Metrics is in preview now. To learn more about setting up your first AutoscalingMetric resource, check out the latest GKE autoscaling documentation.

  •  

GKE becomes more elastic: Scale to zero, save costs, and keep workloads responsive

True elasticity has long been the holy grail of cloud-native engineering. And while Kubernetes has revolutionized resource management, workloads that run sporadically (e.g., batch processors, event-driven workers, and development environments) still consume compute resources while they wait for work, driving up costs.

We’re addressing this head-on in Google Kubernetes Engine (GKE) 1.37 with a native way to scale to and from zero. A new collection of features allows you to scale down your workloads completely to zero replicas so that they stop consuming resources. At the same time, you can quickly and easily restart these workloads on GKE capacity buffers when demand returns, so you waste less infrastructure. This isn't just about saving money, but about decoupling the cost of always-on infrastructure from workload readiness.

The evolution: HPA-based scale-to-zero vs. KEDA

For years, Kubernetes Event-Driven Autoscaling (KEDA), an optional Kubernetes component, was the go-to solution for scaling to zero. While powerful, KEDA adds complexity to an environment. 

Feature

GKE scale-to-zero

KEDA-based setups

Operational toil

Managed service; no extra components.

Requires management of ScaledObject CRDs & operators.

Configuration

Native HPA & CRDs (minimal YAML).

Can exceed 10,000 lines of YAML for large fleets.

Latency

Internalized signal path reduces reaction time.

Polling intervals and hop-counts increase cold-start delays.

By baking scale-to-zero directly into the GKE control plane, we eliminate the need for add-on operators and thousands of lines of configuration. The logic moves from "sidecar management" to a native attribute of the workload.

Under the hood: HPA with AutoscalingMetric and KEP-2021

The magic behind scaling to zero within GKE lies in the integration of two critical components:

  1. HPA with AutoscalingMetric: This is the managed metrics signal pipeline that now supports direct reading of external signals from Google Cloud Managed Service for Prometheus. HorizontalPodAutoscaler (HPA) with AutoscalingMetric provides a unified, high-performance path for metrics from Pub/Sub, Cloud Monitoring, or Load Balancer signals to reach the autoscaler, without the complexity of an adapter.

  2. KEP-2021: Built on the Kubernetes Enhancement Proposal that enables minReplicas: 0 in the HPA, this mechanism allows the HPA to stop all pods when metrics fall below a threshold. It also ensures the HPA can "wake up" the deployment as soon as the metric indicates pending work.

Configuring your first scale-to-zero workload

To implement native scale-to-zero, you need two primary objects: a metric definition and an HPA. In the following example, we scale a worker based on the number of undelivered messages in a Pub/Sub subscription.

Define the metric source

Use the AutoscalingMetric CRD to map an external Cloud Monitoring metric to your cluster.

code_block
<ListValue: [StructValue([('code', 'apiVersion: autoscaling.gke.io/v1beta1\r\nkind: AutoscalingMetric\r\nmetadata:\r\n name: my-autoscalingmetric\r\nspec:\r\n metrics:\r\n - promql:\r\n name: pubsub-undelivered\r\n query: >\r\n {\r\n "pubsub.googleapis.com/subscription/num_undelivered_messages",\r\n subscription_id="my-subscription"\r\n }'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f522d485350>)])]>

Configure the HPA with minReplicas: 0

Reference the metric in your HPA and explicitly set the minimum replicas to zero.

code_block
<ListValue: [StructValue([('code', 'apiVersion: autoscaling/v2\r\nkind: HorizontalPodAutoscaler\r\nmetadata:\r\n name: worker-hpa\r\nspec:\r\n scaleTargetRef:\r\n apiVersion: apps/v1\r\n kind: Deployment\r\n name: worker-deployment\r\n minReplicas: 0\r\n maxReplicas: 50\r\n metrics:\r\n - type: External\r\n pods:\r\n metric:\r\n name: autoscaling.gke.io|my-autoscalingmetric|pubsub-undelivered\r\n target:\r\n type: AverageValue\r\n averageValue: 10'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f522d4c3a10>)])]>

There you go — you’ve allowed your workload to scale to and from zero based on an external metric.

Scale-to-zero capabilities are made possible by support in GKE for external metrics from Cloud Monitoring. By extending the AutoscalingMetric custom resource, you can now query metrics from Google Managed Service for Prometheus, without complex, third-party adapters. This reduces latency, simplifies security, and serves as a key foundation for configuring native scale-to-zero workloads. To learn more about this integration, read our companion blog post on native support for external metrics in GKE.

Managing startup latency with capacity buffers

The biggest challenge with scaling from zero is the so-called cold start — the time it takes for GKE to provision a node and for the container to pull it and start it. This is where GKE capacity buffers come in.

Capacity buffers act as pooled warm capacity. By maintaining a small amount of warm compute resources that can be shared by multiple workloads that can all scale to zero, GKE ensures that when your HPA jumps from 0 to 1, the pod has resources that it can claim immediately. This eliminates the 60-90 second wait for a new GKE node to spin up, reducing startup latency from minutes to an instant, all while maintaining zero cost for the workload. 

Capacity buffers come in two flavors: active and standby. A small active buffer can serve hundreds of workloads that are scaled to zero; instead of each of the workloads maintaining a replica, the active buffer acts as wildcard capacity that serves the whole cluster. A larger standby buffer, which costs a fraction of an active buffer, quickly refills the active buffer for any sustained load encountered by the cluster. By using them together, you get both instant scaling and can maintain low costs. 

What’s ahead

We continue to expand our roadmap for GKE elasticity. For example, imagine you want your development environments to scale to zero at 8:00 PM and scale back up at 7:00 AM. Be on the lookout for methods to exert finer-grained control over recurring scaling, so you can proactively define your scale-to-zero windows. 

Get started with scaling-to-zero today

The days of paying for idle resources are numbered. By enabling GKE's native scale-to-zero capabilities for event-driven and sporadic workloads, you can slash costs without sacrificing startup performance. To get started with scale-to-zero, follow these steps:

  1. Identify a workload with fluctuating demand that has periods of idleness.

  2. Configure your AutoscalingMetric, and set your minReplicas to zero. 

  3. Add capacity buffers to your cluster or workload to keep response times snappy.

For more, check out the documentation on Scaling GKE workloads to and from zero using HPA.

  •  

Scale your AI workloads faster and more efficiently with GKE Pod snapshots

When running modern AI workloads, there’s often a conflict between performance and cost. Workloads like large language models (LLMs) load massive files, and may serve thousands of AI agents that need to execute code instantly. If each component is starting “cold” with a full data-load process, all this provisioning takes time, often forcing organizations to overprovision their infrastructure just to meet scaling requirements.

To solve this, we introduced Google Kubernetes Engine (GKE) Pod snapshots, a new feature that lets you save the running state of your workload, including CPU and GPU memory, and restore it on demand.

GKE Pod snapshots reduce AI inference start-up by as much as 89%, loading 70B parameter models in just 37 seconds and 8B parameters models in just 15 seconds. This speed allows your infrastructure to scale as fast as your demand, significantly reducing the need for overprovisioning.

1

The high cost of cold starts — resuming instead of restarting

The cold start problem isn't unique to AI; it’s a challenge for any application that requires significant initialization time — from game servers to complex Java monoliths. However, the cold start problem is particularly acute in AI workloads. Inference servers must initialize, then download and load gigabytes of model weights into GPU memory — a process that can take several minutes. Further, many agentic AI workloads, including code execution and computer use tools, require isolated sandboxes for each request, and they need to be started quickly and suspended when idle.

In both scenarios, startup latency degrades the user experience and prevents rapid auto-scaling during traffic spikes. Consequently, engineers often resort to overprovisioning expensive infrastructure, or building sophisticated, custom systems to quickly restore state at the application level.

Scaling AI inference without the wait

For generative AI, GKE Pod snapshots solves the linear scaling penalty of model loading. Typically, every new replica you add to a cluster must independently download model weights and load them into accelerator memory. For models with tens of billions of parameters, this step alone often accounts for the majority of the startup time.

With Pod snapshots, you perform this initialization once to create the initial snapshot. GKE captures the fully loaded state including the CPU and GPU memory and persists it in high-throughput Cloud Storage. When the workload needs to scale up, new replicas restore directly from this state, bypassing the initialization phase entirely. In our benchmarks this approach reduced startup latency by as much as 89% for large models like llama3-70b. This speed allows platform teams to shift from expensive overprovisioning strategies to on-demand autoscaling, to help you meet service level objectives while significantly reducing idle GPU costs.

2

Optimizing agentic workflows and sandboxes

GKE Pod snapshots also provide distinct advantages for agentic workflows where agents delegate code execution and computer use to isolated sandboxes. Isolating untrusted, LLM-generated code and commands means one sandbox per user or discrete workflow. In these scenarios, both startup latency and idle sandboxes can result in significant overprovisioning and underutilization. 

Pod snapshots addresses both of these challenges:

  1. To improve startup latency, a snapshot can be captured once of the initial agent sandbox environment, and later used to quickly initialize new sandboxes.

  2. To reduce idle sandboxes, a sandbox can be suspended when idle, capturing its entire compute resources. Later it can be resumed nearly instantly when the environment is needed.

This approach is showing significant success by our customers. For instance, Retake, an AI-powered photo editing platform built by Codeway, faced a significant performance bottleneck with its GPU workloads. By adopting Pod snapshots, they were able to replace a complex custom caching layer and reduce startup time to seconds.

"At Retake, serving personalized models to millions of users requires a massive, unified pipeline for both fine-tuning training and real-time inference on A3 H100 GPUs. We initially engineered a complex custom caching layer for compiled artifacts, which reduced startup time to 1 minute. However, this solution added significant maintenance overhead and still limited our ability to autoscale aggressively. We resolved this by replacing that complexity with GKE Pod snapshots, slashing startup latency to just 8 seconds. By eliminating the initialization penalty, we can now dynamically spin up H100s for specific fine-tuning or inference jobs instantly and shut them down immediately after, drastically reducing idle GPU costs and simplifying our codebase." - Ahmet Furkan Çomak, Lead DevOps Engineer, Codeway

Flexible configuration for any workload

We designed Pod snapshots to improve startup performance and fit naturally into existing Kubernetes workflows. Adopting Pod snapshots to your workload is easy: just define a new declarative policy using Pod snapshot CRDs. The policy allows you to define which Pods to snapshot and where to store the data, and handles the end-to-end storage lifecycle and management. 

You can take snapshots at any stage of the workload — either at workload startup using a workload signal, or during the lifecycle of the Pod using an on-demand trigger. You can further control storage and  restore behavior, setting snapshots retention for cost optimization, choosing between the default behaviour of restoring from the last taken snapshot, or specifying an explicit snapshot during a new Pod deployment.

While the primary use cases for GKE Pod snapshots are AI inference and agent sandboxes, this feature is workload-agnostic. You can use it to speed up any application with a long initialization phase, such as complex Java applications, game servers, or legacy monoliths.

Get started

You can begin optimizing your startup latency today with GKE Pod snapshots. Check out the documentation to learn how to get started and we look forward to your feedback.

  •  

Global AI routing with <1% overhead on multi-cluster GKE Inference Gateway

Demand for AI infrastructure is at an all-time high. Global accelerator shortages mean engineering teams can rarely get all the compute they need from just one data center — capacity comes a cluster here, a cluster there, often an ocean apart. At the same time, workloads are getting hungrier: Today’s long-running agentic workloads often have context windows of 100k to 800k+ tokens, which consume accelerator memory faster than any previous generation of AI traffic.

In this environment, the goal is to maximize "intelligence per dollar." Fragmented, poorly balanced infrastructure is rarely up to the task though, allowing expensive accelerators to sit idle, while requests queue up somewhere else.

To close that gap, we built a layered routing architecture that makes globally scattered capacity behave like a single pool behind a single entry point. At the edge, the multi-cluster GKE Inference Gateway focuses on global, multi-region traffic distribution and high availability. Beneath that, the LLM-d router handles the complex, memory-aware scheduling algorithms that keep utilization high. 

This architecture is deliberately runtime-, model-, and accelerator-agnostic — it works across serving frameworks, model families, and GPU or TPU hardware. To make the results concrete rather than abstract, we recently benchmarked managing production-level global request routing at scale across a multi-region GKE deployment of 17,000 compute nodes spread across the US and Europe. The deployment served a leading Mixture of Experts (MoE) foundation model using SGLang. 

The results: Scaling to three clusters achieved a near-linear throughput boost while maintaining a 99.9% success rate under heavy multi-client concurrency. Additionally, routing traffic through the multi-cluster GKE Inference Gateway added less than 1% overhead, delivering 99.5% of the throughput of a direct, local cluster call.

Read on to learn how it works, more on the benchmark results, and what it means for your own distributed inference deployment.

Three regions, one endpoint

The deployment spanned three GKE clusters in three geographic regions: us-east5 (the config cluster), us-west8, and europe-west4. However, from the client’s perspective, none of that geography exists. Requests hit a single global virtual IP, and the gateway decides — in real time — which cluster should serve each one. 

What makes that decision smart rather than blind is telemetry. Instead of traditional round-robin routing at the network layer, the multi-cluster load balancer is configured to route traffic based on live application signals. Specifically, the Endpoint Picker Proxy (EPP) reads the KV-cache token utilization natively exposed by the underlying inference engines and emits it as a metric for the load balancer. When the load balancer sees a region running hot based on this emitted metric, it spills traffic to the next healthy region.

1

Multi-cluster GKE Inference Gateway topology. The config cluster holds routing configuration but sits outside the request path; each target cluster runs its own EPP and reports KV-cache utilization back to the load balancer.

Distributed LLM engines also operate differently than standard web apps. In a typical inference engine's distributed mode (such as tensor parallelism across multiple nodes), only the master (rank-0) pod serves the API. GKE already handles local routing using standard Service selectors and LeaderWorkerSet (LWS) to direct traffic exclusively to leader pods. The multi-cluster Inference Gateway also integrates with this foundation: It routes global traffic to the correct regional services, helping your cross-region load balancing respects your underlying multi-node topologies out of the box.

The net effect: Three isolated regional data centers start behaving like one cohesive global accelerator fleet, with failover and load balancing driven by what the models are actually doing. 

Measuring the routing overhead 

The first question every team asks about a global routing tier is almost always, ‘How much throughput am I giving up for cross-region capability?’ 

The benchmarks answer this directly: Deploying the multi-cluster GKE Inference Gateway to maximize your accelerator fleet doesn't have to come at the cost of throughput. Routing traffic through the Gateway added less than 1% overhead, delivering 99.5% of the throughput of a direct, local cluster call.

2

That’s the whole trade-off. All the benefits of global load balancing, essentially for free.

Linear scaling across regions 

A bigger test is scale. In our test, growing the fleet from one cluster to three, spanning the US and Europe, while every client request originated from a single region (us-east5), put real pressure on the Gateway: If it couldn’t distribute load efficiently across those distances, throughput would flatten as hardware was added. 

Instead, throughput multiplied almost exactly in line with capacity: 

Fleet topology 

Request 

throughput

Token 

throughput

Success 

rate

1 cluster (us-east5-a) 

0.72 req/s 

2,898 tok/s 

99.87%

2 clusters (+ us-west8-a) 

1.40 req/s 

6,380 tok/s 

99.95%

3 clusters (+ europe-west4- b) 

2.10 req/s 

8,457 tok/s 

99.90%

3

Scaling to three clusters achieved a near-linear throughput boost while maintaining a 99.9% success rate under heavy multi-client concurrency.

Memory-aware routing in action 

Round-robin load balancing is inadequate for serving LLMs because it treats every request as equal. They aren’t. Heavy prompts saturate GPU compute cores, long generations stress memory bandwidth, and long-context conversations quietly eat VRAM until the engine can’t schedule anything new. 

Here, the pressure on memory bandwidth came from the routing signal chosen for this deployment. By mapping Inference Engine's native token-usage metric onto the Gateway’s KV-cache signal, the routing plane gained a real-time view of memory pressure across the entire 17,000 fleet. (Depending on the workload, the Gateway can route on other signals too, like queue depth or running concurrency.) 

Under live production loads, as the primary region climbed toward its high-bandwidth memory (HBM) limits, the Gateway detected the saturation the moment the cluster crossed its 40% KV-cache utilization threshold; it then automatically began routing the overflow to the next healthy region. No operator intervention was needed. The complexity of running in multiple regions simply never reached the user. 

The payoff

By routing traffic based on live KV-cache utilization, this GKE Inference Gateway setup effectively pools globally scattered compute capacity into one unified engine. For this deployment, the result was a near-linear throughput boost across three global regions, with virtually zero routing overhead.

This translates directly into maximizing 'intelligence per dollar,' extracting near-perfect proportional performance out of every accelerator you add to your fleet, rather than letting capital go to waste.

What this means for your team 

If you’re planning your own distributed inference deployment, five lessons from this work stand out:

  • Smarter load balancing pays for itself. Round-robin routing wastes expensive GPU capacity because it can’t see memory or compute pressure. Routing on real-time application signals turns fragmented regional clusters into one efficient fleet — the difference between stranded hardware and 90%+ utilization of scarce compute.

  • Agentic workloads change the bottleneck. Long-running agents with extreme context windows exhaust memory long before there’s no more compute. If your routing layer can’t see memory pressure, your compute will strand compute behind full VRAM. Make KV-cache utilization a first-class routing signal.

  • AI traffic breaks web-era assumptions. Traditional load balancers are tuned for sub-second transactions; LLM requests can run for minutes. Plan connection limits and timeouts for AI-scale latency early, or expect aborted connections in production.

  • Your routing layer must integrate with native serving patterns. Distributed LLM engines have master-worker topologies where only certain pods can serve traffic. By pairing your Gateway with native Kubernetes constructs like LeaderWorkerSet (LWS), your global routing respects local pod topologies out of the box, saving your team from building custom proxy infrastructure.

  • For large foundation model builders, bet on an open, portable stack. Teams operating at frontier scale face the most acute capacity fragmentation, forcing them to hunt for compute resources across whichever regions have availability capacity. An open, portable inference stack such as LLM-d on GKE lets you absorb that capacity wherever it lands, rather than hard-wiring your serving architecture to any single cluster, region, or bespoke infrastructure.

Next steps 

Ready to maximize your distributed accelerator efficiency and set up global cross-region load balancing with multi-cluster GKE Inference Gateway? 

  1. Deploy it yourself: Set up the multi-cluster GKE Inference Gateway. 

  2. Understand the architecture: About multi-cluster GKE Inference Gateway. 

  3. Learn about the cross-region spillover behavior featured in this post: About elastic cross-region high availability and Configure elastic cross-region high availability.

  •  

For SeaVerse, GKE Agent Sandbox reduces infrastructure costs by 60%

Editor’s note: Today we hear from SeaVerse, a gaming startup from SeaArt that is building a platform for playable AI experiences, where users can open lightweight games, character chats, and interactive apps, or create their own experiences from a prompt. To support that creative loop, SeaVerse needed infrastructure that could run dynamic, multi-tenant sandbox workloads with strong isolation, low latency, better observability, and more flexible costs. Google Kubernetes Engine (GKE) and GKE Agent Sandbox gave SeaVerse the managed foundation from which to execute these AI workloads, helping the team reduce their infrastructure costs by up to 60%, while giving creators a faster path from idea to playable experiences. Read on to learn more.


What if AI were a playground? Welcome to SeaVerse, a creation-first platform for playable AI experiences. Here, an AI creation can be as peaceful as drawing a path for a snake to follow, or as chaotic as a music-backed stickman simulation. Some people come to play lightweight games. Others come to chat with AI characters, try interactive apps, create visual patterns, share what they made, or remix an idea into something new. 

We built SeaVerse around a simple promise: Every experience should feel immediate and easy to share. A creator should be able to describe an idea in plain language, refine the result, and publish it in moments, without a traditional coding workflow. 

Delivering that simplicity requires serious infrastructure. Every creation that users make moves through the same chain: generate, run, preview, debug, publish, remix. If any part of that chain is slow, unstable, or poorly isolated, users feel it immediately. That’s why we turned to GKE and GKE Agent Sandbox. 

The infrastructure challenge of instant interaction 

What looks effortless to a user is anything but on our end. Every creation on SeaVerse runs as a distinct workload and is expected to behave reliably from the first interaction. 

Because each workload runs in its own environment, we needed clear security boundaries between users, creations, and sandboxes. But overly strict isolation could slow the very creative loop we were trying to protect, and when something went wrong, diagnosing it was costly. Our engineers had to trace problems across multiple parts of the execution chain with little visibility into what was happening inside the environment. 

We explored existing sandbox approaches, but needed deeper kernel-level isolation and native observability at scale to support fast diagnosis across multi-tenant environments. Something had to change.

Building on GKE and GKE Agent Sandbox 

We chose GKE because we needed a reliable, secure way to operate Kubernetes without turning our engineering team into a cluster maintenance team. GKE brought together the proven ecosystem and operational tooling we needed, freeing us to focus on building the platform rather than managing the infrastructure beneath it. 

As a Kubernetes primitive designed for agent code execution and computer use, GKE Agent Sandbox addressed our requirement for strong isolation, enforcing strong security boundaries without slowing down the creation experience. By utilizing GKE Agent Sandbox with Kata Containers+Cloudhypervisor (microVM), we’ve achieved the perfect balance of multi-cloud flexibility and robust security, option to switch isolation runtime between microVM and gVisor, running our AI sandboxes safely. GKE empowers us to scale toward our long-term vision of supporting over a million sandboxes. Built on gVisor, it provides kernel-level isolation for dynamic sandbox workloads while preserving the Kubernetes orchestration model, so that they can be managed through the same scheduling, monitoring, and operations as the rest of the cluster.  

With SeaVerse, users can generate interactive experiences from a single prompt. After an experience is generated, GKE Agent Sandbox supports the run, test, integration, and verification steps needed to make it ready to preview, refine, and publish. At general availability, it supports allocating up to 300 sandboxes per second, per cluster, with 90% of allocations completing in 200 milliseconds. Together, GKE and GKE Agent Sandbox gave us a reliable foundation for AI-generated interactive workloads that helped keep our team focused on the product experience. 

From black box to glass box 

Before GKE Agent Sandbox, a failed sandbox workload could feel like flying blind. We could often see that something had gone wrong, but didn’t have enough runtime status, metrics, or failure signals to understand why. 

Now, Google Cloud’s native logging and monitoring reach directly into those sandboxed environments, giving us a clearer view of workload behavior, faster issue resolution, and a stronger foundation for managing multi-tenant workloads. 

That visibility matters to developers, but it also matters to the platform’s users: A creator never sees the logs, the cluster, or the orchestration layer. They see whether an experience opens quickly, whether it responds when they draw, click, chat, or share, and whether they can keep building without friction. 

Flexibility that translates to savings

GKE Agent Sandbox also changed how we think about cost. Previously, running secure sandboxed environments meant stronger dependencies on specific server types, which limited how precisely we could match resources to each workload. With GKE Agent Sandbox, we can run secure, isolated workloads on appropriately sized cloud VMs. This gives us greater flexibility in resource allocation and helped us cut our infrastructure costs by up to 60%. 

That same flexibility extended to storage. Not all SeaVerse creations are built in a single session. Some evolve over time as creators return to refine them, build on earlier ideas, or invite others to remix what they’ve made. Our previous architecture didn’t support the persistent file-system capabilities those more complex use cases demanded, but that gap is gone now. We can attach persistent storage where workloads require it while maintaining the isolation boundaries that multi-tenant AI experiences need. For creators, that means experiences that are fast to open and easier to refine, revisit, and build on over time. 

The next remix 

Supporting creations that can evolve and deepen is central to what we’re building. It’s still early in what playable AI can become. As the platform grows, we need to keep strengthening what matters most: stability, observability, elastic scaling, and cost efficiency, all in service of a creator experience that stays fast, reliable, and expressive. 

We’re also exploring additional Google Cloud tools to support smarter analytics and creation assistance. Gemini and agent models could help operators and creators better understand how experiences perform. BigQuery AI and ML capabilities can support use cases such as churn prediction, LTV and ROI prediction, and user segmentation. Multimodal tools such as Imagen and Veo on Gemini Enterprise Agent Platform open up new possibilities for material analysis, creative generation, and AI interactive content production. 

Our goal is to make AI experiences feel immediate, expressive, and connected. With GKE and GKE Agent Sandbox, we have a stronger foundation for the next generation of playable AI.

  •  

Bringing gVisor sandboxes to distributed Ray clusters

The reinforcement learning (RL) ecosystem is rapidly adopting Ray as the unified compute runtime for complex post-training workflows. Across Google Cloud, we see customers using Ray for workloads ranging from multimodal data pipelines to frontier RL. But as agentic and reasoning models evolve, a critical bottleneck has emerged: orchestrating secure, isolated sandboxes at scale to safely execute dynamic rollouts, code generation, and multi-turn tool interactions. Today, in partnership with Anyscale, we are excited to introduce an experimental library for Ray that leverages agentic AI technologies being developed at Google to bring native, high-performance sandboxing directly into distributed Ray clusters.

Sandboxes as Ray Primitives

Ray has become a common runtime for orchestrating post-training workloads. Frameworks including veRL, NeMo-RL, SLIME, MILES, and SkyRL already use Ray to coordinate distributed trainers, inference engines, rollout workers, and other components.

When we designed Ray Sandboxing, an important goal was to make it fit naturally into the existing Ray programming model rather than introduce a separate abstraction for isolated execution. A sandbox has many of the same properties as other resources managed by Ray: it needs to be placed on a machine, assigned resources, created and destroyed, recovered from failures, and scaled with the surrounding workload. This led us to represent each high-level sandbox through a Ray Actor:

image1

The Ray scheduler decides which node should run a sandbox and reserves the corresponding CPU and memory resources. The sandbox Actor manages its lifecycle, while gVisor provides the isolated execution environment on that node.

Starting in Ray 2.58, framework authors and researchers can manage sandboxed environments using the same Ray APIs and patterns they already use for the rest of their workload. For example:

code_block
<ListValue: [StructValue([('code', 'import ray\r\nfrom ray.experimental import sandbox\r\n\r\nray.init()\r\n# Create a gVisor sandbox environment and return an actor handle for a proxy actor\r\nsb = sandbox.create(\r\n cpu=1.0,\r\n memory="512Mi",\r\n image="python:3.12-slim"\r\n)\r\n# Execute code inside the sandbox\r\nresult = ray.get(sb.exec.remote("python -c \'import sys; print(sys.version)\'"))\r\nprint(result.stdout)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f094035b130>)])]>

This creates a gVisor sandbox from an OCI-compatible image and returns a Ray Actor handle. Calls to exec are normal Ray Actor calls, so the sandbox can live anywhere in the cluster. The created actor is a proxy that will forward the operations to gVisor.

The sandbox API covers the basic lifecycle needed by agentic workloads:

  • Create environments from OCI container images

  • Set CPU and memory limits

  • Configure environment variables, working directories, and networking

  • Execute commands

  • Read, write, upload, and download files

  • Inspect sandbox state

  • Terminate or delete environments.

For lower-level use cases, SandboxRuntime provides direct access to local gVisor sandboxes and lets users modify the OCI specification before it is handed to gVisor. Here is an example how this API can be used to build a pool of local sandboxes inside of an actor:

code_block
<ListValue: [StructValue([('code', 'import ray\r\nfrom ray.experimental.sandbox.runtime import SandboxRuntime\r\n\r\n@ray.remote\r\nclass SandboxPool:\r\n def __init__(self, size: int = 3, image: str = "python:3.10-slim"):\r\n self.runtime = SandboxRuntime()\r\n self.sandboxes = [\r\n self.runtime.create(image=image, memory="512Mi")\r\n for _ in range(size)\r\n ]\r\n\r\n def run_command(self, index: int, command: str):\r\n return self.runtime.exec(self.sandboxes[index], command)\r\n\r\n def close(self):\r\n for sb_id in self.sandboxes:\r\n self.runtime.delete(sb_id)\r\n\r\n# Deploy an actor managing a pool of local sandboxes\r\npool = SandboxPool.remote(size=3)\r\nresult = ray.get(pool.run_command.remote(0, "python3 -c \'print(\\"Hello from pool!\\")\'"))\r\nprint(result.stdout)\r\nray.get(pool.close.remote())'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f094035b1c0>)])]>

Why gVisor?

Running model-generated code means treating the code inside the environment as untrusted. Ray Sandboxing uses gVisor, Google's open-source application kernel, as its initial sandbox runtime. gVisor implements a substantial portion of the Linux system-call interface in userspace, putting an additional isolation boundary between workloads and the host kernel. It is OCI-compatible, works with standard container images, and does not require exposing a Docker daemon or host Docker socket to the sandbox.

This combination is particularly useful for agentic workloads: environments remain lightweight enough to create dynamically while providing stronger isolation than executing generated code directly in ordinary containers. gVisor also provides sub-second sandbox startup and low per-sandbox memory overhead, making it possible to use sandboxes as relatively fine-grained distributed resources.

In future versions of Ray, we plan to extend support to other sandboxing runtimes such as Agent Substrate or Kata Containers.

Try Ray sandboxing on GKE

Check out the Ray documentation to learn more about Ray Sandboxes. To try out these sandboxing capabilities on GKE, head over to the Ray sandboxing User Guide. Have feedback or ideas? Join the discussion on the GitHub issue to collaborate on the future of Ray for reinforcement learning.

  •  

Google is a Leader in the 2026 Gartner® Magic Quadrant™ for Cloud-Native Application Platforms

We are thrilled to announce that Google has been recognized as a Leader for the third year in a row in the 2026 Gartner® Magic Quadrant™ for Cloud-Native Application Platforms (CNAP). We believe this placement in the Leaders quadrant validates our commitment to providing an accessible, developer-centric platform that accelerates onboarding and supports rapid prototyping across modern workloads.

Figure_1_Magic_Quadrant_for_CloudNative_Application_Platforms

Our vision for an application-centric cloud focuses on enabling developers to prioritize writing code and building agents or traditional apps by removing infrastructure complexity. Google Cloud provides a unified execution environment supporting serverless, containerized, and agentic deployment options. We believe our placement highlights Google's unique readiness to power both standard enterprise microservices and the next generation of autonomous AI applications.  

Some key features and capabilities of our platform are highlighted below. 

From idea to implementation

Generative AI has ushered in a wave of vibe coding, allowing anyone to go from an idea to a deployed application in a fraction of the time it used to take. To make this even smoother, Google Cloud integrates its serverless infrastructure with AI vibe-coding and prototyping tools. We also simplify access to Google Cloud resources with tools like managed MCP servers and agentic skills — packaged sets of instructions, scripts, and resources to teach an AI how to complete specialized, multi-step workflow. 

  • One-click prototyping in Google AI Studio: Developers can build and deploy full-stack applications directly within Google AI Studio, making it a great environment for prototyping and experimentation. With a single click, you can instantly package and publish your vibe-coded applications to Cloud Run. 

  • Google-managed MCP servers: To enable AI agents to interact with cloud resources, we support official, fully managed remote MCP servers. An example is the Cloud Run MCP server (via run.googleapis.com/mcp), which allows developers to easily launch endpoints and deploy server-side logic. The MCP tools are deployed with a simple config, skipping cloud builds to launch code in seconds, saving valuable developer time. These fully managed servers are integrated with IAM and VPC Service Controls, and they leverage Model Armor for content security.

  • Google's official Skills Repository: Level up your agents with additional, condensed expertise on various Google Cloud technologies. Published and available in Agent Registry, the repository includes skills for Cloud Run, the Well-Architected Pillar (security, reliability, and cost optimization), and more.

Ready to try vibe coding yourself? Get hands on with this codelab to build a vibe-coded app and deploy it to Cloud Run.

From implementation to enterprise-ready

Translating prototypes into production-grade, secure, and cost-effective enterprise software is where Google Cloud excels, with a full suite of developer, architect, and platform engineering tools. From designing your application to optimizing day 2 operations, we offer the services and tools to help you build, operate, and deploy applications across their entire lifecycle, and you have the freedom to build with any language, any library, and any framework.

Build

  • Build with Google Antigravity: At Google, we’re simplifying and expanding our development ecosystem behind the Antigravity harness, collapsing developer silos into a unified orchestration layer. By integrating multi-step AI reasoning directly into the developer workflow, Antigravity natively brings local codebase development to our cloud-native application platforms (e.g., Cloud Run).

  • Design and deploy with Application Design Center (ADC): Now, you can bridge the gap between developer velocity and enterprise control, using Application Design Center to eliminate manual Terraform and YAML configuration. This platform engineering component helps teams design, standardize, and deploy template-driven applications on Google Cloud. It is also integrated as part of Gemini Cloud Assist design agent and published as an MCP server. With ADC, you can visually design your architecture using Cloud Run services, databases, and event brokers backed by automated Gemini Cloud Assist security templates. Beyond human-guided design, ADC enables programmatic orchestration at the time of no HITL (Human-in-the-Loop), allowing automated pipelines to provision policy-governed Terraform configurations directly and autonomously.

Operate 

  • Intelligent investigations: Integrating Gemini Cloud Assist with native telemetry creates an AI-driven framework for Day-2 incidents. When alerts fire, operators engage Gemini Cloud Assist to instantly synthesize logs and metrics, pinpoint root causes, and generate remediations — context that can be handed off to accelerate support escalations. Crucially, IAM permissions strictly govern all AI recommendations, and help to ensure explicit human-in-the-loop approval are required before any infrastructure changes occur.

  • Cost analysis and optimizations: Machine learning algorithms learn natural seasonal traffic cycles to detect cost anomalies within minutes, triggering notifications to protect your bottom line without risking destructive infrastructure shutdown.  

Deploy

  • Reliability and high availability: Cloud Run is a regional service by default, but you can deploy an app to multiple regions via a single gcloud command. Integrated with service health, Cloud Run automates cross-region failover and failback. If a service in one region becomes unhealthy, traffic is automatically routed to the next-closest healthy region, failing back once the issue is resolved.

  • An open platform: As a long-time and top contributor to the Cloud Native Computing Foundation (CNCF), we operate with an open-source-first strategy. By integrating foundational, community-driven technologies, we help enable application portability for enterprise customers who are increasingly demanding multi-cloud flexibility

Ready to start deploying your apps to Google Cloud? Get hands on with these codelabs:  

From enterprise-ready to autonomous

AI agents are software’s next frontier. They offer more than just increased productivity and efficiency; they can unlock exponential growth. To provide enterprises with robust agentic deployment options, Google Cloud provides a dedicated infrastructure stack tailored specifically to host, govern, and secure autonomous agent fleets. This stack seamlessly integrates with our Agent Development Kit (ADK) as well as other leading agentic frameworks to give developers maximum flexibility.

Gemini Enterprise Agent Runtime
At the core of this stack is Gemini Enterprise Agent Platform and its dedicated Agent Runtime, which delivers the serverless and containerized deployment options you need for enterprise-scale agent development, including the following capabilities:

  • Native personalization: Built-in sessions and memory banks manage context and long-term state, preventing costs from ballooning.

  • Agent observability and tracing: Built on OpenTelemetry (OTel) standards and agentic schemas, turnkey dashboards feature agent topology graphs and interactive trace logs that detail sessions, tool calls, and reasoning paths.

  • Agent evaluation and simulation: Automated simulation tools allow developers to test agents against golden sets with side-by-side comparisons and simulate thousands of interactions to test edge cases.

Hosting agents on Cloud Run
For customers requiring additional flexibility, granular control, or specific regulatory compliance, Cloud Run serves as an excellent serverless alternative to host your agents. Some of its latest features include:

  • Cloud Run instances (coming soon): This primitive manages individual, addressable, long-running singleton resources with integrated Cloud Storage volume mounts, allowing persistent background agents to be deployed cost-effectively.

  • Cloud Run sandboxes: Hard-isolated environments spin up in under 500 milliseconds to safely execute untrusted, model-generated code, protecting the host system from unauthorized access.

Agent security, governance, and auditability
To securely deploy AI agents and prevent unmanaged shadow AI, enterprises need an ironclad governance framework. Google Cloud delivers this through Agent Identity (non-human IAM with cryptographic IDs) to provide an auditable trail of all actions and reasoning; a centralized Agent Registry to manage approved agents, skills, tools and application artifacts,  and prevent unauthorized tool integrations; and an Agent Gateway to proxy traffic, enforce Model Armor policies, and actively block destructive actions. These features are available on Agent Runtime today and will be available soon on Cloud Run and Google Kubernetes Engine (GKE).

Ready to start deploying agents? Check out various codelabs featuring Gemini Enterprise Agent Platform here.

Build the future of cloud-native applications

Whether you’re a vibe coder deploying your first full-stack application, a software architect standardizing production microservices, or an enterprise team scaling a fleet of secure AI agents, Google Cloud delivers the simplicity, elasticity, and security you need. Read the full report: Download your complimentary copy of the 2026 Gartner® Magic Quadrant™ for Cloud-Native Application Platforms (CNAP).


Magic Quadrant for Cloud-Native Application Platforms, By Mukul Saha, Alex Coqueiro, Prasanna Lakshmi Narasimha, Richard Watson, 3 August 2026

Gartner and Magic Quadrant are trademarks of Gartner, Inc. and/or its affiliates. This graphic was published by Gartner, Inc. as part of a larger research document and should be evaluated in the context of the entire document. The Gartner document is available upon request from Google. 

Gartner does not endorse any company, vendor, product or service depicted in its publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner publications consist of the opinions of Gartner’s business and technology insights organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this publication, including any warranties of merchantability or fitness for a particular purpose.

  •  

ClusterNetworkPolicy in GKE: Balancing control and autonomy for your microservices

Managing network security in a multi-tenant Kubernetes environment typically requires balancing two distinct needs: developers need their microservices to communicate effectively, while platform and security teams must maintain compliance, prevent lateral movement, and establish cluster-wide guardrails.

Historically, the standard Kubernetes NetworkPolicy has been the primary tool for this. While effective for single-namespace isolation, standard NetworkPolicy is scoped strictly to individual namespaces and designed around developer self-service. When cluster administrators attempt to use it for global security enforcement, it can lead to policy conflicts and operational challenges.

To address this, we introduced ClusterNetworkPolicy (CNP), an open-source standard developed by the Kubernetes SIG-Policy Working Group (WG), to Google Kubernetes Engine (GKE). Designed for scale, CNP is a cluster-wide resource that allows administrators to manage network security centrally, providing a mechanism for those responsible for global security to implement consistent, non-bypassable policies.

Read on for technical details about CNP, some common use cases, an example policy, and how to get started. 

Structuring policies with tiers

A core capability of ClusterNetworkPolicy is its hierarchical tier system. Rather than attempting to reconcile flat, conflicting peer rules simultaneously, CNP establishes a deterministic, top-to-bottom evaluation hierarchy:

  1. The admin tier: The highest precedence level. Rules here are enforced before any other policies.

  2. The network policy tier: The standard namespace level, where developers manage their specific application policies.

  3. The baseline tier: The lowest precedence, establishing the cluster’s default behavior when no other policies apply. This can be overridden using namespace scoped policies.

tiers

This tiered structure helps align network security with organizational roles. Using standard role-based access control (RBAC), you can manage the admin tier to enforce compliance mandates, while platform teams can use the baseline tier to set a default "deny-all" zero-trust posture across the cluster. At the same time, developers can write standard network policies for their applications without overriding core security mandates.

This deterministic, top-to-bottom evaluation method resolves conflicts between different teams' policies. The admin tier introduces an explicit Pass action. This allows security teams to inspect traffic against global rules and then delegate the final Accept or Deny decision down to the developer's namespace policy, facilitating both central oversight and distributed management.

Common network security scenarios

This tiered architecture translates complex security requirements into centralized rules. Here are common scenarios where ClusterNetworkPolicy provides a practical solution:

  • Isolating sensitive workloads: You can apply an admin-tier global deny rule to isolate specific namespaces — such as those used for payment processing or compliance data — from the rest of the cluster. This action overrides any permissive developer policies that might otherwise expose these environments.

  • Protecting core services: To prevent configurations that might disrupt internal operations, administrators can create an admin-tier global allow rule for critical services like kube-dns. This allows these services to remain accessible regardless of any misconfigured namespace policies.

  • Managing external egress: By utilizing IP address range matching, egress traffic can be controlled at the cluster level. This functionality allows you to explicitly restrict or permit access to corporate intranets or external IP ranges, serving as a safeguard against unauthorized data exfiltration.

Example scenario

Consider a common enterprise requirement: Application workloads across all namespaces must be permitted to reach central platform infrastructure (such as shared authentication and telemetry services), while access to sensitive environments — like a restricted vault namespace — is strictly prohibited. Meanwhile, routine microservice traffic is delegated to developer-managed, namespace-scoped policies.

ClusterNetworkPolicy makes this straightforward. A platform administrator simply defines an admin-tier guardrail centrally:

code_block
<ListValue: [StructValue([('code', 'apiVersion: policy.networking.k8s.io/v1alpha2\r\nkind: ClusterNetworkPolicy\r\nmetadata:\r\n name: platform-isolation-guardrail\r\nspec:\r\n tier: Admin\r\n priority: 10\r\n subject:\r\n # Target all application tenant namespaces, excluding system and core infrastructure\r\n namespaces:\r\n matchExpressions:\r\n - key: kubernetes.io/metadata.name\r\n operator: NotIn\r\n values: ["kube-system", "shared-services", "restricted-vault"]\r\n egress:\r\n # 1. Mandate access to central shared platform services\r\n - name: allow-shared-services\r\n action: Accept\r\n to:\r\n - namespaces:\r\n matchLabels:\r\n kubernetes.io/metadata.name: shared-services\r\n\r\n # 2. Enforce strict block on accessing the restricted vault namespace\r\n - name: block-restricted-vault\r\n action: Deny\r\n to:\r\n - namespaces:\r\n matchLabels:\r\n kubernetes.io/metadata.name: restricted-vault\r\n\r\n # 3. Explicitly delegate all remaining traffic to developer namespace policies\r\n - name: delegate-remaining-egress\r\n action: Pass\r\n to:\r\n - namespaces: {}\r\n - networks:\r\n - 0.0.0.0/0\r\n - ::/0'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f2488c1ef50>)])]>

Extending open-source foundations

Instead of building this functionality as proprietary extensions, we worked with the Kubernetes community to design the ClusterNetworkPolicy API (policy.networking.k8s.io), distinguishing it from the namespace-scoped NetworkPolicy API (networking.k8s.io). Furthermore, we collaborated closely with the Cilium community to build its implementation of the API.

Because it is built on open-source standards, GKE helps ensure that security configurations remain portable across different environments. The ClusterNetworkPolicy API natively supports tier selection, enabling clear and deterministic policy evaluation. This approach lets administrators enforce robust security guardrails while maintaining the operational flexibility that development teams depend on.

ClusterNetworkPolicy on GKE elevates workload network security — shifting operations from namespace-scoped rules to unified, cluster-wide governance. It is currently in preview in version 1.36 and later. To learn more and get started, check out:

  •  

Do more with less: How GKE can reduce your cost per agent by 75%

In today’s agentic era, modern cloud applications are evolving from a set of passive tools to fleets of autonomous digital workers that reason, plan, and take action across a wide range of tasks. 

For platform engineering teams designing these environments, the simplest approach is often to deploy an agent on to an open-source framework like OpenClaw and Hermes running on  a virtual machine (VM). But as those workloads move into production and scale to support additional users or use cases, teams quickly hit a critical challenge: AI agents tend to operate in bursts; for a while they actively process requests or execute code, followed by long periods of inactivity while awaiting user input or external triggers. If you rely on static compute allocations, idle agents are still consuming valuable CPU and memory. 

The question becomes: how do you safely pack more agents onto a fixed compute footprint without sacrificing reliability, scalability, or efficiency?

The answer is to incorporate orchestration upfront as a holistic part of your architecture. Orchestration helps you unlock dramatically improved unit economics and scalability, ease of use, and reliability from day one. Google Kubernetes Engine (GKE) offers sophisticated orchestration capabilities. To help you make the most of your compute capacity, we tested the maximum number of AI agents that can be packed onto a single GKE node running on a fixed Google Compute Engine VM instance (n2-standard-48) — without performance degradation, or repeated failures. Using an OpenClaw profile, we applied progressive optimizations to demonstrate the meaningful role that orchestration can play in running agentic workloads at scale — read on to learn more.

agent_density_blog_img_1

Baseline: Running OpenClaw on microVMs 

Running untrusted, multi-agent workloads securely requires strong isolation. A common approach is to run each agent inside a dedicated microVM (such as Kata containers) on a Kubernetes deployment, which provides strong hardware-level isolation. 

While this provides the necessary security boundary, it hits a scaling wall almost immediately. Every microVM requires its own guest operating system that consumes memory and CPU resources, limiting the actual resources available for your actual agents. In this baseline scenario, we hit a scaling wall at 61 OpenClaw agents on a standard GKE node before reliability dropped and workload health checks began to fail regularly.

Optimization 1: Pushing density with GKE Agent Sandbox

To address this, we migrated the same agent workload from microVMs to GKE Agent Sandbox, a Kubernetes primitive that’s designed specifically for the security and performance requirements of running agents.

Instead of relying on heavy guest operating systems, GKE Agent Sandbox leverages the open-source secure container sandbox, gVisor. gVisor uses a user-space kernel (the Sentry) to intercept and filter system calls. This provides secure, production-grade isolation for untrusted code execution while maintaining the lightweight footprint of standard Kubernetes containers.

This reduced overhead improves the efficiency of the sandbox itself, resulting in being able to deploy 88 OpenClaw agents inside the same VM before failure — a 44% increase in the number of agents you can run on the same fixed capacity while maintaining a highly reliable security perimeter.

It’s no surprise then, that when GKE Agent Sandbox reached General Availability in May, its usage grew more than 7x in under four weeks.

Key takeaway: In our tests, migrating OpenClaw-type agents to GKE Agent Sandbox enabled us to run more than 40% more agents per vCPU, and reduced the cost per agent by more than 30%, all while maintaining a similar performance profile.

Optimization 2: The value of orchestration

While the GKE Agent Sandbox optimizes active workloads, solving the problem of idle AI agents requires making workload orchestration a central part of your agent architecture.

Rather than keeping idle agents running in the background, you can use GKE Pod snapshots to checkpoint (freeze) them to persistent storage, which releases their physical CPU and memory resources back to the cluster. When a new task trigger arrives, a lightweight Kubernetes controller or event gateway intercepts the request and signals GKE to resume the agent from the snapshot. This happens in milliseconds. 

This pattern lets you reliably oversubscribe physical compute resources based on workload behavior, so you can fit more agents on the same node. However, oversubscription isn’t a one-size-fits-all approach and comes with a set of tradeoffs: Different AI agents have different latency requirements and execution models. If you treat all agents the same, you will either degrade your user experience with latency, or bankrupt your project with over-provisioned VMs.

With GKE, you can run an agent platform that supports tailored deployments for different types of agents and use cases, each fine-tuned to their unique performance and cost requirements.

image1

GKE supports a spectrum of agent workload behaviors, balancing latency sensitivity against resource density.

Here are some examples of agentic workloads with very different performance requirements:

  • Real-time coding assistant (latency-sensitive): Direct developer-facing agents need sub-second startup times (<1s) and have zero tolerance for queueing. By pairing GKE Pod snapshots with Agent Sandbox Warm Pools, GKE maintains pre-warmed, isolated sandboxes that can be executed nearly instantaneously.

  • Autonomous teammate (balanced): Interactive background agents can tolerate average startup times (a few seconds). GKE suspend and resume functionality restores these agents on demand, so they don’t consume compute resources while they are idle.

  • Headless background agent (latency-tolerant): Scheduled daily research or analysis cron jobs can tolerate queueing delays; you’re not going to compromise business outcomes by waiting to execute these jobs for an hour while cluster capacity becomes available. To save on costs for these kinds of agents, go ahead and use maximum resource oversubscription.

In other words, rather than forcing you into a single cluster-wide strategy, GKE supports different behaviors simultaneously across node pools and workload configurations.

Consider the "thundering herd" problem, where a surge of agents all wake and demand compute simultaneously. GKE offers a tunable dial with features like Agent Sandbox warm pools and suspend and resume to balance potential cost savings against guaranteed performance based on your specific requirements — performance- or cost-optimized:

  • Performance-optimized: If your use case requires guaranteed, sub-second performance during massive, sudden traffic spikes, you can provision buffers using Agent Sandbox warm pools. In this configuration, we were able to run 133 OpenClaw agents on the same node.

  • Cost-optimized: For workloads that are latency-tolerant or that can be staggered, higher oversubscription ratios significantly increase node density. In this configuration, we ran 274 agents on the same node (>3x more agents than the baseline) while keeping startup times under five seconds.

Key takeaway: By combining GKE Agent Sandbox with GKE’s suspend and resume capabilities, you can freeze idle agents to oversubscribe fixed compute capacity. For agents with intermittent activity, this can enable up to 3.5x greater agent density and cost reductions of up to 75% per agent.

Scale your agents, not your budget

Scaling your agents shouldn’t mean linearly scaling your infrastructure budget. As our examples show, adopting the right platform features and considering orchestration from the get-go can dramatically alter the value you get from your compute capacity. GKE allows you to easily align your infrastructure with your business goals — whether that means prioritizing aggressive cost savings or optimizing for performance.

And this is just the beginning. At Google Cloud, we’re continuously innovating new ways to help you manage the demands of the agentic era. Ready to get more out of your compute capacity? Check out the GKE Agent Sandbox documentation and learn how GKE is helping teams innovate faster for less.

  •