❌

Vue lecture

Google is a Leader in the 2026 Gartner Magic Quadrant for Container Management

We’re excited and proud to share that Gartner has recognized Google as a Leader for the fourth year in a row in the 2026 Gartner® Magic Quadrant™ for Container Management, based on its Completeness of Vision and Ability to Execute. Google was positioned highest in Ability to Execute of all vendors evaluated and we believe this validates the success of our mission to deliver a container platform that’s highly optimized for both performance and efficiency. We help global customers to build and run their most demanding and complex workloads at scale, including the next generation of AI and agentic applications. 

In the accompanying 2026 Gartner Critical Capabilities for Container Management report, Google Cloud was ranked first in every use case: New Cloud Native Applications, Containerized Existing Applications, AI Training, AI Inference, Edge Applications, and Hybrid Applications.

Gartner predicts1 that “By 2028, 95% of new AI deployments will use Kubernetes, up from less than 30% in 2025.” Containers power today’s most innovative apps and businesses — and deliver the infrastructure customers demand as they transform their businesses in the agentic era.

2026 Gartner Magic Quadrant for Container Management

Google Cloud spearheaded the industry-wide cloud-native revolution when we introduced Kubernetes in 2014 and launched Google Kubernetes Engine (GKE), the world’s first managed Kubernetes service, in 2015. Our commitment to container platforms and the vibrant, innovative Kubernetes ecosystem has only grown stronger and deeper since. Alongside GKE, our serverless container platforms GKE Autopilot and Cloud Run dramatically lower operational costs and help developers deliver amazing containerized apps faster than ever before. 

The massive acceleration in enterprise AI has inspired us to redefine infrastructure management for the AI era. In 2026 so far we’ve introduced a wide range of foundational improvements to shift GKE and Cloud Run into agent-native, high-performance platforms designed for autonomous AI systems, massive inference workloads, and secure runtime isolation. Whether you’re training AI at the frontier, launching an AI startup, or leading your enterprise AI transformation, we have the container platform you need. Important highlights include:

Delivering leading performance and efficiency for AI infrastructure

  • GKE predictive latency boost: Built into the GKE Inference Gateway, this ML-driven capability uses capacity-aware routing rather than static configurations to reduce Time-to-First-Token (TTFT) by up to 70%.

  • GKE automatic KV Cache storage tiering: Automatically shifts KV cache data across RAM, Local SSD, and Cloud Storage. This reduces memory bottlenecks, improving TTFT by 40% via RAM offloading and increasing throughput by 70% via Local SSDs for large prompt contexts. [1]

  • GKE accelerated container and model startups: GKE node spin-up times are up to 4x faster, and pod startup speeds have improved by up to 80%. Additionally, native run:AI Model Streamer integration pulls heavy models from Cloud Storage 5x faster.

  • Cloud Run on-demand serverless GPU scale-to-zero: Cloud Run supports NVIDIA RTX PRO 6000 Blackwell GPUs, allowing teams to serve 70B+ parameter models on-demand. Your services can go from zero to a fully provisioned GPU — with all drivers pre-installed — in under 5 seconds. Once active inference or fine-tuning runs complete, Cloud Run automatically scales instances back to zero, eliminating idle infrastructure costs.

Evolving Kubernetes for agentic infrastructure security and scale

  • GKE Agent Substrate: As an open-source, secure-by-default agent execution runtime, Agent Substrate is engineered to run millions of sandboxes with 10x higher density than standard container runtimes. Purpose-built for the era of autonomous agents, Substrate delivers sub-500ms resume operations at over 500 suspend/resume activations per second with a native zero-trust kernel and network isolation. Agent Substrate is available as an open-source solution that runs on any Kubernetes infrastructure and is optimized for GKE.

  • GKE Agent Sandbox: Built on gVisor kernel-isolation technology, Agent Sandbox isolates the host environment from untrusted, multi-agent AI code execution. It provides secure execution at scale, processing up to 300 sandboxes per second with sub-second latency and delivering up to 30% better price-performance when running on Axion processors than comparable hyperscaler cloud providers. 

  • GKE Dataplane V2 scalability limits: Architectural capacity bounds for GKE clusters implementing active NetworkPolicies doubled from 7,500 nodes to 15,000 nodes per cluster, supporting the massive infrastructure needs of large enterprise and AI customers.

  • GKE intent-based autoscaling: GKE can now natively autoscale horizontally using application intent and custom metrics beyond basic hardware metrics. This reduces resource allocation reaction times from 25 seconds down to just 5 seconds.

  • Filestore agent volumes: a new offering that attaches and detaches NFS mounts in milliseconds, allowing agents to start/resume near-instantaneously, along with native Read-Write-Many (RWX) access and POSIX-compliant file locking to enable safe multi-agent collaboration without write collisions. 

Next-gen developer experience with serverless containers

Whether you’re hosting a standard web API, running a heavy batch data job, processing an asynchronous message queue, or deploying a complex AI agent, Cloud Run handles it all under a single, unified serverless model that delivers an unmatched developer experience and maximum engineering velocity. 

  • One-click prototyping in Google AI Studio: You can build and deploy full-stack applications directly within Google AI Studio, making it an exceptional environment for rapid prototyping and experimentation. With a single click, you can instantly package and publish your vibe-coded applications to Cloud Run.

  • Cloud Run instances: This new primitive manages individual, addressable, long-running singleton resources with integrated Cloud Storage volume mounts, allowing persistent background agents like OpenClaw to be deployed cost-effectively. With baseline shared-CPU configurations starting at a highly predictable flat rate of ~$5.70 per month (for 1 vCPU and 1 GiB of RAM), Cloud Run instances delivers an always-on, VM-like experience while bypassing the idle-cost penalties and operational overhead of traditional VMs.

  • Cloud Run sandboxes: Hard-isolated environments spin up in under 500 milliseconds to safely execute untrusted, model-generated code, protecting the host system from unauthorized access.

Take the next steps

As we reach for new heights of performance, security, and scale for our container platforms, we continue to build the future in the open. We invite you to explore Agent Sandbox and Agent Substrate today. We can’t wait to shape the future of agent infrastructure together with our customers and partners. Check out these resources to continue your learning journey:


1. Gartner report: Critical Capabilities for Container Management, 8 September 2026

Gartner, Magic Quadrant for Container Management, Dennis Smith, et al, 2 September 2026
Gartner, Critical Capabilities for Container Management, By Tony Iams, Wataru Katsurashima, Lucas Albuquerque, Dennis Smith, Bhuvie Chhabra, 8 September 2026. 
Gartner and Magic Quadrant are trademarks of Gartner, Inc. and/or its affiliates.
Disclaimer: Gartner does not endorse any company, vendor, product or service depicted in its publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner publications consist of the opinions of Gartner’s business and technology insights organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this publication, including any warranties of merchantability or fitness for a particular purpose.

  •  

Deploy personal AI agents with Cloud Run instances

Need a low-cost, high-performance way to run long-lived, stateful workloads such as AI agents? Today, we introduced Cloud Run instances, which let you do just that.  

Consider AI agents such as OpenClaw or Hermes, which are intended for individual developers or personal use. Because these agents often work continuously and tend to serve only one user at a time, their infrastructure requirements look quite different from stateless, high-throughput web services that typically run on Cloud Run services.

Cloud Run services scale to zero when requests stop, so they aren’t ideal for a long-lived agent that expects exactly one copy to be running continuously. On the other hand, the alternative — running a dedicated VM — means paying for full compute 24/7, managing operating system updates, opening firewall ports, and provisioning your own HTTPS endpoints.

Cloud Run instances provide dedicated, singleton compute runtimes on Cloud Run. They have the following attributes:

  • Runs just one instance with no autoscaling

  • Up to 7-day continuous runtime, with automatic restart policy configured by default

  • Every instance gets a HTTPS URL that remains unchanged across updates and restarts.

  • You can stop each instance when you aren't using it and resume it whenever you need it

The cost to run a Cloud Run instance with 1 vCPU and 1 GiB of memory continuously for 30 days is $5.70. Cloud Run instances use shared vCPU with vCPU burst budgets to run continuously for a low, predictable price. This model is also ideal for long-lived agents that aren’t doing compute-intensive work all the time, and only spike in usage when asked to perform a task.

Example: Deploy OpenClaw on a Cloud Run instance

OpenClaw is an open-source personal AI agent that can perform various tasks on your behalf, and become a better assistant over time. Many OpenClaw users start out running it on their own laptops, until they realize they need somewhere to run it where it won’t shut down every time their laptop goes to sleep.

Deploying OpenClaw to a Cloud Run instance is easy. Once you’ve uploaded OpenClaw’s configuration files to a Cloud Storage bucket, you can deploy your OpenClaw agent to a Cloud Run instance with just one command:

code_block
<ListValue: [StructValue([('code', 'gcloud beta run instances create openclaw-instance \\\r\n --image ghcr.io/openclaw/openclaw:latest \\\r\n --port 18789 \\\r\n --public \\\r\n --add-volume mount-path=/home/node/.openclaw,type=cloud-storage,mount-options="uid=1000;gid=1000;file-mode=0700;dir-mode=0700",bucket=${BUCKET} \\\r\n --set-env-vars "OPENCLAW_GATEWAY_PASSWORD=${PASSWORD},GEMINI_API_KEY=${GEMINI_API_KEY}"'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7febd4c1b4f0>)])]>

Once deployed, you can keep this OpenClaw instance running for as long as you want. You can interact with it over Telegram, WhatsApp, or the social media platform of your choice, and connect it to any tools you want it to use, as well.

For the full instructions on how to deploy OpenClaw, refer to this codelab.

Coming soon, we’re also launching SSH access for both Cloud Run instances and Cloud Run services. Sign up for private access here.

What users are saying

Cloud Run instances are helping Google Cloud users achieve their goals for running AI agents and other long-lived workloads at low cost and high performance.

OffDeal, an AI-powered investment bank for small businesses, is running long-lived agents on Cloud Run instances:

“We're currently using Cloud Run instances as our primary infrastructure for our long-running agent. It reduced cold starts by 88%. Everything was very straightforward to implement, and it has been very reliable.” - Luis Ruiz Morel, Member of Technical Staff @ OffDeal

Learn more

Currently in preview, Cloud Run instances are a cost-effective way to run a new kind of workload, without sacrificing performance. For more information about Cloud Run instances, check out the following resources:

  •  

Making highly available, multi-region Cloud Run services just got easier

Application downtime for mission-critical services can directly impact your reputation and bottom line. To avoid that, you need to be able to deploy regionally resilient workloads that detect and automatically recover from failures. But setting up multi-region, highly available deployments often involves complex configurations, and responding to incidents or outages is usually a manual process. 

Multi-region services on Cloud Run provide a one-command approach to deploying the same service configuration across multiple regions. When deployed with a global external application load balancer, you can serve traffic from different regions.

Now, we’ve made it easier to detect regional service disruptions and automatically fail over to a healthy region within seconds with new capabilities:

  1. Readiness probes provide instance-level health checks for your Cloud Run service to determine exactly when your containers are ready to serve traffic. You can also use these probes to monitor how many healthy or unhealthy instances exist for your service in each region.

  2. Service health aggregates instance-level health checks from readiness probes to calculate the health of your service in each region. This aggregate health is exposed via serverless network endpoint groups (NEGs) in each region. When your service is connected to a global application load balancer, traffic automatically fails away from regions with unhealthy services. Service health can be used with both single and multi-region services.

Let’s take a closer look at some scenarios where these new capabilities can come in handy.

Use cases

To make your Cloud Run applications highly available, it is essential to minimize the downtime for each incident. In high availability scenarios, readiness probes can help you detect regional service failures and automatically fail over, minimizing service degradation or disruptions. 

To achieve automated failover, one key thing to consider is whether you plan to support application traffic from the public internet or from within your private network (VPC).

  • Public internet applications: When you have a public-facing website or API, configure Cloud Run with a global external application load balancer for automatic detection and failover capabilities. 

  • Private network applications: When you have private applications with internal traffic, configure Cloud Run with a cross-regional internal application load balancer for automatic detection and failover capabilities. 

Design Considerations

Cloud Run’s new service health excels at quickly detecting and recovering outages in active-active configurations, where two or more regions are actively configured to serve traffic. Some things to consider when designing your multi-region setup:

  • Single points of failure: As you design your application, ensure that each layer of your application, including your database layer, has regional redundancies to avoid any single points of failure. For three-tiered applications on Cloud Run, consider setting up your web tier and application tier with distinct multi-region architectures to handle public internet and private networking respectively.

  • Data replication: When replicating data across regions and evaluating your recovery point objective (RPO), consider whether you require zero data loss. Cloud Run service health works best with read- and write-heavy applications that actively synchronize data across regions. 

  • Data residency: Google Cloud offers several multi-region database configurations with managed multi-region solutions including Firestore, Spanner, Cloud Storage, and Cloud SQL. These all work great for multi-region architectures on Cloud Run that have strict data sovereignty requirements.

Get started

Cloud Run’s enhanced multi-region high availability services are currently available in all Cloud Run regions at no additional cost. You only pay for the standard CPU and memory required to run the readiness probes. To learn more, check out our documentation.

  •