❌

Vue normale

Reçu avant avant-hierCloud Blog

Google is a Leader in the 2026 Gartner Magic Quadrant for Container Management

24 septembre 2026 à 21:00

We’re excited and proud to share that Gartner has recognized Google as a Leader for the fourth year in a row in the 2026 Gartner® Magic Quadrant™ for Container Management, based on its Completeness of Vision and Ability to Execute. Google was positioned highest in Ability to Execute of all vendors evaluated and we believe this validates the success of our mission to deliver a container platform that’s highly optimized for both performance and efficiency. We help global customers to build and run their most demanding and complex workloads at scale, including the next generation of AI and agentic applications. 

In the accompanying 2026 Gartner Critical Capabilities for Container Management report, Google Cloud was ranked first in every use case: New Cloud Native Applications, Containerized Existing Applications, AI Training, AI Inference, Edge Applications, and Hybrid Applications.

Gartner predicts1 that “By 2028, 95% of new AI deployments will use Kubernetes, up from less than 30% in 2025.” Containers power today’s most innovative apps and businesses — and deliver the infrastructure customers demand as they transform their businesses in the agentic era.

2026 Gartner Magic Quadrant for Container Management

Google Cloud spearheaded the industry-wide cloud-native revolution when we introduced Kubernetes in 2014 and launched Google Kubernetes Engine (GKE), the world’s first managed Kubernetes service, in 2015. Our commitment to container platforms and the vibrant, innovative Kubernetes ecosystem has only grown stronger and deeper since. Alongside GKE, our serverless container platforms GKE Autopilot and Cloud Run dramatically lower operational costs and help developers deliver amazing containerized apps faster than ever before. 

The massive acceleration in enterprise AI has inspired us to redefine infrastructure management for the AI era. In 2026 so far we’ve introduced a wide range of foundational improvements to shift GKE and Cloud Run into agent-native, high-performance platforms designed for autonomous AI systems, massive inference workloads, and secure runtime isolation. Whether you’re training AI at the frontier, launching an AI startup, or leading your enterprise AI transformation, we have the container platform you need. Important highlights include:

Delivering leading performance and efficiency for AI infrastructure

  • GKE predictive latency boost: Built into the GKE Inference Gateway, this ML-driven capability uses capacity-aware routing rather than static configurations to reduce Time-to-First-Token (TTFT) by up to 70%.

  • GKE automatic KV Cache storage tiering: Automatically shifts KV cache data across RAM, Local SSD, and Cloud Storage. This reduces memory bottlenecks, improving TTFT by 40% via RAM offloading and increasing throughput by 70% via Local SSDs for large prompt contexts. [1]

  • GKE accelerated container and model startups: GKE node spin-up times are up to 4x faster, and pod startup speeds have improved by up to 80%. Additionally, native run:AI Model Streamer integration pulls heavy models from Cloud Storage 5x faster.

  • Cloud Run on-demand serverless GPU scale-to-zero: Cloud Run supports NVIDIA RTX PRO 6000 Blackwell GPUs, allowing teams to serve 70B+ parameter models on-demand. Your services can go from zero to a fully provisioned GPU — with all drivers pre-installed — in under 5 seconds. Once active inference or fine-tuning runs complete, Cloud Run automatically scales instances back to zero, eliminating idle infrastructure costs.

Evolving Kubernetes for agentic infrastructure security and scale

  • GKE Agent Substrate: As an open-source, secure-by-default agent execution runtime, Agent Substrate is engineered to run millions of sandboxes with 10x higher density than standard container runtimes. Purpose-built for the era of autonomous agents, Substrate delivers sub-500ms resume operations at over 500 suspend/resume activations per second with a native zero-trust kernel and network isolation. Agent Substrate is available as an open-source solution that runs on any Kubernetes infrastructure and is optimized for GKE.

  • GKE Agent Sandbox: Built on gVisor kernel-isolation technology, Agent Sandbox isolates the host environment from untrusted, multi-agent AI code execution. It provides secure execution at scale, processing up to 300 sandboxes per second with sub-second latency and delivering up to 30% better price-performance when running on Axion processors than comparable hyperscaler cloud providers. 

  • GKE Dataplane V2 scalability limits: Architectural capacity bounds for GKE clusters implementing active NetworkPolicies doubled from 7,500 nodes to 15,000 nodes per cluster, supporting the massive infrastructure needs of large enterprise and AI customers.

  • GKE intent-based autoscaling: GKE can now natively autoscale horizontally using application intent and custom metrics beyond basic hardware metrics. This reduces resource allocation reaction times from 25 seconds down to just 5 seconds.

  • Filestore agent volumes: a new offering that attaches and detaches NFS mounts in milliseconds, allowing agents to start/resume near-instantaneously, along with native Read-Write-Many (RWX) access and POSIX-compliant file locking to enable safe multi-agent collaboration without write collisions. 

Next-gen developer experience with serverless containers

Whether you’re hosting a standard web API, running a heavy batch data job, processing an asynchronous message queue, or deploying a complex AI agent, Cloud Run handles it all under a single, unified serverless model that delivers an unmatched developer experience and maximum engineering velocity. 

  • One-click prototyping in Google AI Studio: You can build and deploy full-stack applications directly within Google AI Studio, making it an exceptional environment for rapid prototyping and experimentation. With a single click, you can instantly package and publish your vibe-coded applications to Cloud Run.

  • Cloud Run instances: This new primitive manages individual, addressable, long-running singleton resources with integrated Cloud Storage volume mounts, allowing persistent background agents like OpenClaw to be deployed cost-effectively. With baseline shared-CPU configurations starting at a highly predictable flat rate of ~$5.70 per month (for 1 vCPU and 1 GiB of RAM), Cloud Run instances delivers an always-on, VM-like experience while bypassing the idle-cost penalties and operational overhead of traditional VMs.

  • Cloud Run sandboxes: Hard-isolated environments spin up in under 500 milliseconds to safely execute untrusted, model-generated code, protecting the host system from unauthorized access.

Take the next steps

As we reach for new heights of performance, security, and scale for our container platforms, we continue to build the future in the open. We invite you to explore Agent Sandbox and Agent Substrate today. We can’t wait to shape the future of agent infrastructure together with our customers and partners. Check out these resources to continue your learning journey:


1. Gartner report: Critical Capabilities for Container Management, 8 September 2026

Gartner, Magic Quadrant for Container Management, Dennis Smith, et al, 2 September 2026
Gartner, Critical Capabilities for Container Management, By Tony Iams, Wataru Katsurashima, Lucas Albuquerque, Dennis Smith, Bhuvie Chhabra, 8 September 2026. 
Gartner and Magic Quadrant are trademarks of Gartner, Inc. and/or its affiliates.
Disclaimer: Gartner does not endorse any company, vendor, product or service depicted in its publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner publications consist of the opinions of Gartner’s business and technology insights organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this publication, including any warranties of merchantability or fitness for a particular purpose.

What’s new in cloud-native apps?

15 avril 2022 à 22:00

Developers and IT operations pros of all stripes come to Google Cloud to build modern, cloud-first and cloud-native applications. Here’s the latest from Google Cloud on everything app dev, containers, Kubernetes, DevOps, serverless and open source, all in one place.

Week of Apr 11 - Apr 15, 2022

Listen to a Prodcast
Google’s SRE team has launched a “Prodcast” focusing on concepts from its SRE book. Available from wherever you get your podcasts. 

Run Apache Spark on a modern container base
Dataproc, our managed version of Apache Spark, is now generally available on Google Kubernetes Engine (GKE), allowing you to create a Dataproc cluster and submit Spark jobs on a self-managed GKE cluster. Read all about it. 

Loads of new runtimes in App Engine and Cloud Functions
Java, Ruby, Python and PHP developers, rejoice! You can now update or develop new App Engine apps and Cloud Functions using Java 17, Ruby 3, Python 3.10 and PHP 8.1.

BeReal shows you how modern app development is done
Social media company BeReal discusses how it uses Google Cloud services including Firebase, Cloud Functions and GKE to build its app.  

Build fast without breaking things
In this three-part series, learn about the Supply-chain Levels for Software Artifacts (SLSA) framework designed to improve the integrity of your software packages and infrastructure. Start with, How to SLSA Part 1 - The Basics, then move on to part 2 and part 3.

Week of Apr 4 - Apr 8, 2022

How to migrate a container from a VM to Cloud Run
With Cloud Run, you can migrate a legacy VM to a container and save money – even if you don’t know Kubernetes. This video shows you how. 

Receive Error Reporting notifications through Slack and Webhooks
Error Reporting can analyze, aggregate, and notify DevOps teams about crashes that happened in their cloud services, right to their preferred channels. Learn more in this blog. 

Cloud-native architecture is in the cards at NCR
Earlier this year, NCR Authentic Cards talked about how it built a transaction processing platform on Google Cloud. NCR and its consulting partner Opus Systems are back for part two of the migration story, taking a detailed look at all the components that went into the cloud-based architecture. 

How to easily share a service with Cloud Run 
Have you ever written a script that you wanted to make available to others? Cloud Run makes it easy to deploy a processing service quickly and easily. In this blog post, Developer Advocate Laurent Picard creates an image processing service that generates coloring pages, then makes it available to others — all in under 200 lines of Python and JavaScript. Follow along in this tutorial.

Week of Mar 28 - Apr 1, 2022

Another cool thing you can do with Cloud Functions
Got data you want to ingest from Cloud Storage to BigQuery? Cloud Functions can help with that. This tutorial shows you how.  

Add custom severity levels to Cloud Monitoring alert policies
Not all alerts are created equal. In this blog post, learn how to add static and dynamic severity levels to a Cloud Monitoring alert policy, with enhanced notification channels including email, webhooks, Cloud Pub/Sub and PagerDuty. 

Learn how to use CPU allocation controls in Cloud Run
Last fall, we added “always-on CPU” capabilities to Cloud Run, making it a better fit for running background- and other asynchronous-processing tasks. In this post, Developer Advocate Wesley Chun uses a weather alerting app to demonstrate how to use the feature, and along the way, reduces the app’s average user response latency by over 80%.

Week of Mar 21 - Mar 25, 2022

Get Going with latest Go 1.18 release
With the release of version 1.18, the Go programming language now includes support for generic code using parameterized types, integrated fuzz testing, and a new Go workspace mode that makes it simple to work with multiple modules. Learn more here.

Week of Mar 14 - Mar 18, 2022

Create EventArc triggers with Terraform
In addition to the Google Cloud Console or gcloud, you can also use a Terraform resource to create an Eventarc trigger. Mete Atamel shows you how. 

Scaling to new markets with Cloud Run
French publisher Les Echos Le Parisien Annonces switched from dedicated on-prem infrastructure to Cloud Run to supplement its main news site with regional variations. Les Echos shares its website architecture here. 

The serverless way to celebrate Pi Day
In honor of Pi Day, Google Cloud Developer Advocate Emma Haruka Iwao shows you how to use the new Cloud Functions (2nd gen) to calculate π — serverlessly.

Week of Mar 07 - Mar 11, 2022

Rhode Island moves to Google Cloud-based job board
When the pandemic hit, the State of Rhode Island moved its workforce development operations entirely online on a foundation of Google Workspace and Google Cloud resources, including Firestore, Cloud Functions, and Kubernetes, among others. Check out how they did it. 

Containerized microservices at Lowe’s
Lowe’s already told us how they use SRE. They’re at it again, describing how they built an e-commerce website using a containerized microservices architecture and Kubernetes, with Istio for service mesh and Cloud Operations for good measure.

Cruise AVs hit the road with Google Cloud services
Autonomous Vehicle (AV) startup Cruise detailed how it’s using data analytics and machine learning on a foundation of Google Kubernetes Engine (GKE) and other services to develop and test its self-driving cars. Read the guest post. 

L’Oréal’s data analytics gets a makeover with serverless
We’re hurtling toward a programmable cloud — a world where developers use cloud-native serverless tools like Cloud Functions to quickly prototype and build powerful, data-driven business insights. L’Oréal is a great example.  

Better telemetry for your Anthos clusters
Anthos Service Mesh Dashboard is now available (public preview) on the Anthos clusters on Bare Metal and Anthos clusters on VMware. Now, you can get out-of-the-box telemetry dashboards to see a services-first view of your application on the Cloud Console.

Instrument your Java apps
With the new version of the Google Cloud Logging Java library, you can wire your application logs with more information — without adding a single line of code.

Visualize metrics from Cloud Spanner
Building an app on top of Cloud Spanner but can’t assess how well it’s operating? The new OpenTelemetery receiver for Cloud Spanner provides an easy way for you to process and visualize metrics from Cloud Spanner System tables, and export these to the APM tool of your choice. Read more here.

Week of Feb 28 - Mar 4, 2022

Introducing Cloud SDK
The rebranded Cloud SDK is a collection of all the libraries and tools (including Google Cloud CLI) you need to interact with Google Cloud products and services. Learn more here. 

Cloud CLI, meet Terraform
Google Cloud CLI’s new Declarative Export for Terraform allows you to export the current state of your Google Cloud infrastructure into a descriptive file compatible with Terraform (HCL) or Google’s KRM declarative tooling, and is now available in preview. 

Knative graduates to incubating project 
Congratulations to Knative, which has been accepted by the Cloud Native Computing Foundation, or CNCF, as an incubating project, enabling the next phase of serverless architecture. 

We manage Prometheus so you don’t have to
Google Cloud Managed Service for Prometheus is now generally available! Get all the benefits of open source-compatible monitoring with the ease of use of Google-scale managed services. Learn more here.

Deploy personal AI agents with Cloud Run instances

27 août 2026 à 18:00

Need a low-cost, high-performance way to run long-lived, stateful workloads such as AI agents? Today, we introduced Cloud Run instances, which let you do just that.  

Consider AI agents such as OpenClaw or Hermes, which are intended for individual developers or personal use. Because these agents often work continuously and tend to serve only one user at a time, their infrastructure requirements look quite different from stateless, high-throughput web services that typically run on Cloud Run services.

Cloud Run services scale to zero when requests stop, so they aren’t ideal for a long-lived agent that expects exactly one copy to be running continuously. On the other hand, the alternative — running a dedicated VM — means paying for full compute 24/7, managing operating system updates, opening firewall ports, and provisioning your own HTTPS endpoints.

Cloud Run instances provide dedicated, singleton compute runtimes on Cloud Run. They have the following attributes:

  • Runs just one instance with no autoscaling

  • Up to 7-day continuous runtime, with automatic restart policy configured by default

  • Every instance gets a HTTPS URL that remains unchanged across updates and restarts.

  • You can stop each instance when you aren't using it and resume it whenever you need it

The cost to run a Cloud Run instance with 1 vCPU and 1 GiB of memory continuously for 30 days is $5.70. Cloud Run instances use shared vCPU with vCPU burst budgets to run continuously for a low, predictable price. This model is also ideal for long-lived agents that aren’t doing compute-intensive work all the time, and only spike in usage when asked to perform a task.

Example: Deploy OpenClaw on a Cloud Run instance

OpenClaw is an open-source personal AI agent that can perform various tasks on your behalf, and become a better assistant over time. Many OpenClaw users start out running it on their own laptops, until they realize they need somewhere to run it where it won’t shut down every time their laptop goes to sleep.

Deploying OpenClaw to a Cloud Run instance is easy. Once you’ve uploaded OpenClaw’s configuration files to a Cloud Storage bucket, you can deploy your OpenClaw agent to a Cloud Run instance with just one command:

code_block
<ListValue: [StructValue([('code', 'gcloud beta run instances create openclaw-instance \\\r\n --image ghcr.io/openclaw/openclaw:latest \\\r\n --port 18789 \\\r\n --public \\\r\n --add-volume mount-path=/home/node/.openclaw,type=cloud-storage,mount-options="uid=1000;gid=1000;file-mode=0700;dir-mode=0700",bucket=${BUCKET} \\\r\n --set-env-vars "OPENCLAW_GATEWAY_PASSWORD=${PASSWORD},GEMINI_API_KEY=${GEMINI_API_KEY}"'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7febd4c1b4f0>)])]>

Once deployed, you can keep this OpenClaw instance running for as long as you want. You can interact with it over Telegram, WhatsApp, or the social media platform of your choice, and connect it to any tools you want it to use, as well.

For the full instructions on how to deploy OpenClaw, refer to this codelab.

Coming soon, we’re also launching SSH access for both Cloud Run instances and Cloud Run services. Sign up for private access here.

What users are saying

Cloud Run instances are helping Google Cloud users achieve their goals for running AI agents and other long-lived workloads at low cost and high performance.

OffDeal, an AI-powered investment bank for small businesses, is running long-lived agents on Cloud Run instances:

“We're currently using Cloud Run instances as our primary infrastructure for our long-running agent. It reduced cold starts by 88%. Everything was very straightforward to implement, and it has been very reliable.” - Luis Ruiz Morel, Member of Technical Staff @ OffDeal

Learn more

Currently in preview, Cloud Run instances are a cost-effective way to run a new kind of workload, without sacrificing performance. For more information about Cloud Run instances, check out the following resources:

Google is a Leader in the 2026 Gartner® Magic Quadrant™ for Cloud-Native Application Platforms

20 août 2026 à 18:00

We are thrilled to announce that Google has been recognized as a Leader for the third year in a row in the 2026 Gartner® Magic Quadrant™ for Cloud-Native Application Platforms (CNAP). We believe this placement in the Leaders quadrant validates our commitment to providing an accessible, developer-centric platform that accelerates onboarding and supports rapid prototyping across modern workloads.

Figure_1_Magic_Quadrant_for_CloudNative_Application_Platforms

Our vision for an application-centric cloud focuses on enabling developers to prioritize writing code and building agents or traditional apps by removing infrastructure complexity. Google Cloud provides a unified execution environment supporting serverless, containerized, and agentic deployment options. We believe our placement highlights Google's unique readiness to power both standard enterprise microservices and the next generation of autonomous AI applications.  

Some key features and capabilities of our platform are highlighted below. 

From idea to implementation

Generative AI has ushered in a wave of vibe coding, allowing anyone to go from an idea to a deployed application in a fraction of the time it used to take. To make this even smoother, Google Cloud integrates its serverless infrastructure with AI vibe-coding and prototyping tools. We also simplify access to Google Cloud resources with tools like managed MCP servers and agentic skills — packaged sets of instructions, scripts, and resources to teach an AI how to complete specialized, multi-step workflow. 

  • One-click prototyping in Google AI Studio: Developers can build and deploy full-stack applications directly within Google AI Studio, making it a great environment for prototyping and experimentation. With a single click, you can instantly package and publish your vibe-coded applications to Cloud Run. 

  • Google-managed MCP servers: To enable AI agents to interact with cloud resources, we support official, fully managed remote MCP servers. An example is the Cloud Run MCP server (via run.googleapis.com/mcp), which allows developers to easily launch endpoints and deploy server-side logic. The MCP tools are deployed with a simple config, skipping cloud builds to launch code in seconds, saving valuable developer time. These fully managed servers are integrated with IAM and VPC Service Controls, and they leverage Model Armor for content security.

  • Google's official Skills Repository: Level up your agents with additional, condensed expertise on various Google Cloud technologies. Published and available in Agent Registry, the repository includes skills for Cloud Run, the Well-Architected Pillar (security, reliability, and cost optimization), and more.

Ready to try vibe coding yourself? Get hands on with this codelab to build a vibe-coded app and deploy it to Cloud Run.

From implementation to enterprise-ready

Translating prototypes into production-grade, secure, and cost-effective enterprise software is where Google Cloud excels, with a full suite of developer, architect, and platform engineering tools. From designing your application to optimizing day 2 operations, we offer the services and tools to help you build, operate, and deploy applications across their entire lifecycle, and you have the freedom to build with any language, any library, and any framework.

Build

  • Build with Google Antigravity: At Google, we’re simplifying and expanding our development ecosystem behind the Antigravity harness, collapsing developer silos into a unified orchestration layer. By integrating multi-step AI reasoning directly into the developer workflow, Antigravity natively brings local codebase development to our cloud-native application platforms (e.g., Cloud Run).

  • Design and deploy with Application Design Center (ADC): Now, you can bridge the gap between developer velocity and enterprise control, using Application Design Center to eliminate manual Terraform and YAML configuration. This platform engineering component helps teams design, standardize, and deploy template-driven applications on Google Cloud. It is also integrated as part of Gemini Cloud Assist design agent and published as an MCP server. With ADC, you can visually design your architecture using Cloud Run services, databases, and event brokers backed by automated Gemini Cloud Assist security templates. Beyond human-guided design, ADC enables programmatic orchestration at the time of no HITL (Human-in-the-Loop), allowing automated pipelines to provision policy-governed Terraform configurations directly and autonomously.

Operate 

  • Intelligent investigations: Integrating Gemini Cloud Assist with native telemetry creates an AI-driven framework for Day-2 incidents. When alerts fire, operators engage Gemini Cloud Assist to instantly synthesize logs and metrics, pinpoint root causes, and generate remediations — context that can be handed off to accelerate support escalations. Crucially, IAM permissions strictly govern all AI recommendations, and help to ensure explicit human-in-the-loop approval are required before any infrastructure changes occur.

  • Cost analysis and optimizations: Machine learning algorithms learn natural seasonal traffic cycles to detect cost anomalies within minutes, triggering notifications to protect your bottom line without risking destructive infrastructure shutdown.  

Deploy

  • Reliability and high availability: Cloud Run is a regional service by default, but you can deploy an app to multiple regions via a single gcloud command. Integrated with service health, Cloud Run automates cross-region failover and failback. If a service in one region becomes unhealthy, traffic is automatically routed to the next-closest healthy region, failing back once the issue is resolved.

  • An open platform: As a long-time and top contributor to the Cloud Native Computing Foundation (CNCF), we operate with an open-source-first strategy. By integrating foundational, community-driven technologies, we help enable application portability for enterprise customers who are increasingly demanding multi-cloud flexibility

Ready to start deploying your apps to Google Cloud? Get hands on with these codelabs:  

From enterprise-ready to autonomous

AI agents are software’s next frontier. They offer more than just increased productivity and efficiency; they can unlock exponential growth. To provide enterprises with robust agentic deployment options, Google Cloud provides a dedicated infrastructure stack tailored specifically to host, govern, and secure autonomous agent fleets. This stack seamlessly integrates with our Agent Development Kit (ADK) as well as other leading agentic frameworks to give developers maximum flexibility.

Gemini Enterprise Agent Runtime
At the core of this stack is Gemini Enterprise Agent Platform and its dedicated Agent Runtime, which delivers the serverless and containerized deployment options you need for enterprise-scale agent development, including the following capabilities:

  • Native personalization: Built-in sessions and memory banks manage context and long-term state, preventing costs from ballooning.

  • Agent observability and tracing: Built on OpenTelemetry (OTel) standards and agentic schemas, turnkey dashboards feature agent topology graphs and interactive trace logs that detail sessions, tool calls, and reasoning paths.

  • Agent evaluation and simulation: Automated simulation tools allow developers to test agents against golden sets with side-by-side comparisons and simulate thousands of interactions to test edge cases.

Hosting agents on Cloud Run
For customers requiring additional flexibility, granular control, or specific regulatory compliance, Cloud Run serves as an excellent serverless alternative to host your agents. Some of its latest features include:

  • Cloud Run instances (coming soon): This primitive manages individual, addressable, long-running singleton resources with integrated Cloud Storage volume mounts, allowing persistent background agents to be deployed cost-effectively.

  • Cloud Run sandboxes: Hard-isolated environments spin up in under 500 milliseconds to safely execute untrusted, model-generated code, protecting the host system from unauthorized access.

Agent security, governance, and auditability
To securely deploy AI agents and prevent unmanaged shadow AI, enterprises need an ironclad governance framework. Google Cloud delivers this through Agent Identity (non-human IAM with cryptographic IDs) to provide an auditable trail of all actions and reasoning; a centralized Agent Registry to manage approved agents, skills, tools and application artifacts,  and prevent unauthorized tool integrations; and an Agent Gateway to proxy traffic, enforce Model Armor policies, and actively block destructive actions. These features are available on Agent Runtime today and will be available soon on Cloud Run and Google Kubernetes Engine (GKE).

Ready to start deploying agents? Check out various codelabs featuring Gemini Enterprise Agent Platform here.

Build the future of cloud-native applications

Whether you’re a vibe coder deploying your first full-stack application, a software architect standardizing production microservices, or an enterprise team scaling a fleet of secure AI agents, Google Cloud delivers the simplicity, elasticity, and security you need. Read the full report: Download your complimentary copy of the 2026 Gartner® Magic Quadrant™ for Cloud-Native Application Platforms (CNAP).


Magic Quadrant for Cloud-Native Application Platforms, By Mukul Saha, Alex Coqueiro, Prasanna Lakshmi Narasimha, Richard Watson, 3 August 2026

Gartner and Magic Quadrant are trademarks of Gartner, Inc. and/or its affiliates. This graphic was published by Gartner, Inc. as part of a larger research document and should be evaluated in the context of the entire document. The Gartner document is available upon request from Google. 

Gartner does not endorse any company, vendor, product or service depicted in its publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner publications consist of the opinions of Gartner’s business and technology insights organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this publication, including any warranties of merchantability or fitness for a particular purpose.

Making highly available, multi-region Cloud Run services just got easier

20 juillet 2026 à 18:00

Application downtime for mission-critical services can directly impact your reputation and bottom line. To avoid that, you need to be able to deploy regionally resilient workloads that detect and automatically recover from failures. But setting up multi-region, highly available deployments often involves complex configurations, and responding to incidents or outages is usually a manual process. 

Multi-region services on Cloud Run provide a one-command approach to deploying the same service configuration across multiple regions. When deployed with a global external application load balancer, you can serve traffic from different regions.

Now, we’ve made it easier to detect regional service disruptions and automatically fail over to a healthy region within seconds with new capabilities:

  1. Readiness probes provide instance-level health checks for your Cloud Run service to determine exactly when your containers are ready to serve traffic. You can also use these probes to monitor how many healthy or unhealthy instances exist for your service in each region.

  2. Service health aggregates instance-level health checks from readiness probes to calculate the health of your service in each region. This aggregate health is exposed via serverless network endpoint groups (NEGs) in each region. When your service is connected to a global application load balancer, traffic automatically fails away from regions with unhealthy services. Service health can be used with both single and multi-region services.

Let’s take a closer look at some scenarios where these new capabilities can come in handy.

Use cases

To make your Cloud Run applications highly available, it is essential to minimize the downtime for each incident. In high availability scenarios, readiness probes can help you detect regional service failures and automatically fail over, minimizing service degradation or disruptions. 

To achieve automated failover, one key thing to consider is whether you plan to support application traffic from the public internet or from within your private network (VPC).

  • Public internet applications: When you have a public-facing website or API, configure Cloud Run with a global external application load balancer for automatic detection and failover capabilities. 

  • Private network applications: When you have private applications with internal traffic, configure Cloud Run with a cross-regional internal application load balancer for automatic detection and failover capabilities. 

Design Considerations

Cloud Run’s new service health excels at quickly detecting and recovering outages in active-active configurations, where two or more regions are actively configured to serve traffic. Some things to consider when designing your multi-region setup:

  • Single points of failure: As you design your application, ensure that each layer of your application, including your database layer, has regional redundancies to avoid any single points of failure. For three-tiered applications on Cloud Run, consider setting up your web tier and application tier with distinct multi-region architectures to handle public internet and private networking respectively.

  • Data replication: When replicating data across regions and evaluating your recovery point objective (RPO), consider whether you require zero data loss. Cloud Run service health works best with read- and write-heavy applications that actively synchronize data across regions. 

  • Data residency: Google Cloud offers several multi-region database configurations with managed multi-region solutions including Firestore, Spanner, Cloud Storage, and Cloud SQL. These all work great for multi-region architectures on Cloud Run that have strict data sovereignty requirements.

Get started

Cloud Run’s enhanced multi-region high availability services are currently available in all Cloud Run regions at no additional cost. You only pay for the standard CPU and memory required to run the readiness probes. To learn more, check out our documentation.

❌