❌

Vue normale

Reçu avant avant-hiercloud

Google is a Leader in the 2026 Gartner Magic Quadrant for Container Management

24 septembre 2026 à 21:00

We’re excited and proud to share that Gartner has recognized Google as a Leader for the fourth year in a row in the 2026 Gartner® Magic Quadrant™ for Container Management, based on its Completeness of Vision and Ability to Execute. Google was positioned highest in Ability to Execute of all vendors evaluated and we believe this validates the success of our mission to deliver a container platform that’s highly optimized for both performance and efficiency. We help global customers to build and run their most demanding and complex workloads at scale, including the next generation of AI and agentic applications. 

In the accompanying 2026 Gartner Critical Capabilities for Container Management report, Google Cloud was ranked first in every use case: New Cloud Native Applications, Containerized Existing Applications, AI Training, AI Inference, Edge Applications, and Hybrid Applications.

Gartner predicts1 that “By 2028, 95% of new AI deployments will use Kubernetes, up from less than 30% in 2025.” Containers power today’s most innovative apps and businesses — and deliver the infrastructure customers demand as they transform their businesses in the agentic era.

2026 Gartner Magic Quadrant for Container Management

Google Cloud spearheaded the industry-wide cloud-native revolution when we introduced Kubernetes in 2014 and launched Google Kubernetes Engine (GKE), the world’s first managed Kubernetes service, in 2015. Our commitment to container platforms and the vibrant, innovative Kubernetes ecosystem has only grown stronger and deeper since. Alongside GKE, our serverless container platforms GKE Autopilot and Cloud Run dramatically lower operational costs and help developers deliver amazing containerized apps faster than ever before. 

The massive acceleration in enterprise AI has inspired us to redefine infrastructure management for the AI era. In 2026 so far we’ve introduced a wide range of foundational improvements to shift GKE and Cloud Run into agent-native, high-performance platforms designed for autonomous AI systems, massive inference workloads, and secure runtime isolation. Whether you’re training AI at the frontier, launching an AI startup, or leading your enterprise AI transformation, we have the container platform you need. Important highlights include:

Delivering leading performance and efficiency for AI infrastructure

  • GKE predictive latency boost: Built into the GKE Inference Gateway, this ML-driven capability uses capacity-aware routing rather than static configurations to reduce Time-to-First-Token (TTFT) by up to 70%.

  • GKE automatic KV Cache storage tiering: Automatically shifts KV cache data across RAM, Local SSD, and Cloud Storage. This reduces memory bottlenecks, improving TTFT by 40% via RAM offloading and increasing throughput by 70% via Local SSDs for large prompt contexts. [1]

  • GKE accelerated container and model startups: GKE node spin-up times are up to 4x faster, and pod startup speeds have improved by up to 80%. Additionally, native run:AI Model Streamer integration pulls heavy models from Cloud Storage 5x faster.

  • Cloud Run on-demand serverless GPU scale-to-zero: Cloud Run supports NVIDIA RTX PRO 6000 Blackwell GPUs, allowing teams to serve 70B+ parameter models on-demand. Your services can go from zero to a fully provisioned GPU — with all drivers pre-installed — in under 5 seconds. Once active inference or fine-tuning runs complete, Cloud Run automatically scales instances back to zero, eliminating idle infrastructure costs.

Evolving Kubernetes for agentic infrastructure security and scale

  • GKE Agent Substrate: As an open-source, secure-by-default agent execution runtime, Agent Substrate is engineered to run millions of sandboxes with 10x higher density than standard container runtimes. Purpose-built for the era of autonomous agents, Substrate delivers sub-500ms resume operations at over 500 suspend/resume activations per second with a native zero-trust kernel and network isolation. Agent Substrate is available as an open-source solution that runs on any Kubernetes infrastructure and is optimized for GKE.

  • GKE Agent Sandbox: Built on gVisor kernel-isolation technology, Agent Sandbox isolates the host environment from untrusted, multi-agent AI code execution. It provides secure execution at scale, processing up to 300 sandboxes per second with sub-second latency and delivering up to 30% better price-performance when running on Axion processors than comparable hyperscaler cloud providers. 

  • GKE Dataplane V2 scalability limits: Architectural capacity bounds for GKE clusters implementing active NetworkPolicies doubled from 7,500 nodes to 15,000 nodes per cluster, supporting the massive infrastructure needs of large enterprise and AI customers.

  • GKE intent-based autoscaling: GKE can now natively autoscale horizontally using application intent and custom metrics beyond basic hardware metrics. This reduces resource allocation reaction times from 25 seconds down to just 5 seconds.

  • Filestore agent volumes: a new offering that attaches and detaches NFS mounts in milliseconds, allowing agents to start/resume near-instantaneously, along with native Read-Write-Many (RWX) access and POSIX-compliant file locking to enable safe multi-agent collaboration without write collisions. 

Next-gen developer experience with serverless containers

Whether you’re hosting a standard web API, running a heavy batch data job, processing an asynchronous message queue, or deploying a complex AI agent, Cloud Run handles it all under a single, unified serverless model that delivers an unmatched developer experience and maximum engineering velocity. 

  • One-click prototyping in Google AI Studio: You can build and deploy full-stack applications directly within Google AI Studio, making it an exceptional environment for rapid prototyping and experimentation. With a single click, you can instantly package and publish your vibe-coded applications to Cloud Run.

  • Cloud Run instances: This new primitive manages individual, addressable, long-running singleton resources with integrated Cloud Storage volume mounts, allowing persistent background agents like OpenClaw to be deployed cost-effectively. With baseline shared-CPU configurations starting at a highly predictable flat rate of ~$5.70 per month (for 1 vCPU and 1 GiB of RAM), Cloud Run instances delivers an always-on, VM-like experience while bypassing the idle-cost penalties and operational overhead of traditional VMs.

  • Cloud Run sandboxes: Hard-isolated environments spin up in under 500 milliseconds to safely execute untrusted, model-generated code, protecting the host system from unauthorized access.

Take the next steps

As we reach for new heights of performance, security, and scale for our container platforms, we continue to build the future in the open. We invite you to explore Agent Sandbox and Agent Substrate today. We can’t wait to shape the future of agent infrastructure together with our customers and partners. Check out these resources to continue your learning journey:


1. Gartner report: Critical Capabilities for Container Management, 8 September 2026

Gartner, Magic Quadrant for Container Management, Dennis Smith, et al, 2 September 2026
Gartner, Critical Capabilities for Container Management, By Tony Iams, Wataru Katsurashima, Lucas Albuquerque, Dennis Smith, Bhuvie Chhabra, 8 September 2026. 
Gartner and Magic Quadrant are trademarks of Gartner, Inc. and/or its affiliates.
Disclaimer: Gartner does not endorse any company, vendor, product or service depicted in its publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner publications consist of the opinions of Gartner’s business and technology insights organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this publication, including any warranties of merchantability or fitness for a particular purpose.

Global AI routing with <1% overhead on multi-cluster GKE Inference Gateway

21 septembre 2026 à 18:00

Demand for AI infrastructure is at an all-time high. Global accelerator shortages mean engineering teams can rarely get all the compute they need from just one data center — capacity comes a cluster here, a cluster there, often an ocean apart. At the same time, workloads are getting hungrier: Today’s long-running agentic workloads often have context windows of 100k to 800k+ tokens, which consume accelerator memory faster than any previous generation of AI traffic.

In this environment, the goal is to maximize "intelligence per dollar." Fragmented, poorly balanced infrastructure is rarely up to the task though, allowing expensive accelerators to sit idle, while requests queue up somewhere else.

To close that gap, we built a layered routing architecture that makes globally scattered capacity behave like a single pool behind a single entry point. At the edge, the multi-cluster GKE Inference Gateway focuses on global, multi-region traffic distribution and high availability. Beneath that, the LLM-d router handles the complex, memory-aware scheduling algorithms that keep utilization high. 

This architecture is deliberately runtime-, model-, and accelerator-agnostic — it works across serving frameworks, model families, and GPU or TPU hardware. To make the results concrete rather than abstract, we recently benchmarked managing production-level global request routing at scale across a multi-region GKE deployment of 17,000 compute nodes spread across the US and Europe. The deployment served a leading Mixture of Experts (MoE) foundation model using SGLang. 

The results: Scaling to three clusters achieved a near-linear throughput boost while maintaining a 99.9% success rate under heavy multi-client concurrency. Additionally, routing traffic through the multi-cluster GKE Inference Gateway added less than 1% overhead, delivering 99.5% of the throughput of a direct, local cluster call.

Read on to learn how it works, more on the benchmark results, and what it means for your own distributed inference deployment.

Three regions, one endpoint

The deployment spanned three GKE clusters in three geographic regions: us-east5 (the config cluster), us-west8, and europe-west4. However, from the client’s perspective, none of that geography exists. Requests hit a single global virtual IP, and the gateway decides — in real time — which cluster should serve each one. 

What makes that decision smart rather than blind is telemetry. Instead of traditional round-robin routing at the network layer, the multi-cluster load balancer is configured to route traffic based on live application signals. Specifically, the Endpoint Picker Proxy (EPP) reads the KV-cache token utilization natively exposed by the underlying inference engines and emits it as a metric for the load balancer. When the load balancer sees a region running hot based on this emitted metric, it spills traffic to the next healthy region.

1

Multi-cluster GKE Inference Gateway topology. The config cluster holds routing configuration but sits outside the request path; each target cluster runs its own EPP and reports KV-cache utilization back to the load balancer.

Distributed LLM engines also operate differently than standard web apps. In a typical inference engine's distributed mode (such as tensor parallelism across multiple nodes), only the master (rank-0) pod serves the API. GKE already handles local routing using standard Service selectors and LeaderWorkerSet (LWS) to direct traffic exclusively to leader pods. The multi-cluster Inference Gateway also integrates with this foundation: It routes global traffic to the correct regional services, helping your cross-region load balancing respects your underlying multi-node topologies out of the box.

The net effect: Three isolated regional data centers start behaving like one cohesive global accelerator fleet, with failover and load balancing driven by what the models are actually doing. 

Measuring the routing overhead 

The first question every team asks about a global routing tier is almost always, ‘How much throughput am I giving up for cross-region capability?’ 

The benchmarks answer this directly: Deploying the multi-cluster GKE Inference Gateway to maximize your accelerator fleet doesn't have to come at the cost of throughput. Routing traffic through the Gateway added less than 1% overhead, delivering 99.5% of the throughput of a direct, local cluster call.

2

That’s the whole trade-off. All the benefits of global load balancing, essentially for free.

Linear scaling across regions 

A bigger test is scale. In our test, growing the fleet from one cluster to three, spanning the US and Europe, while every client request originated from a single region (us-east5), put real pressure on the Gateway: If it couldn’t distribute load efficiently across those distances, throughput would flatten as hardware was added. 

Instead, throughput multiplied almost exactly in line with capacity: 

Fleet topology 

Request 

throughput

Token 

throughput

Success 

rate

1 cluster (us-east5-a) 

0.72 req/s 

2,898 tok/s 

99.87%

2 clusters (+ us-west8-a) 

1.40 req/s 

6,380 tok/s 

99.95%

3 clusters (+ europe-west4- b) 

2.10 req/s 

8,457 tok/s 

99.90%

3

Scaling to three clusters achieved a near-linear throughput boost while maintaining a 99.9% success rate under heavy multi-client concurrency.

Memory-aware routing in action 

Round-robin load balancing is inadequate for serving LLMs because it treats every request as equal. They aren’t. Heavy prompts saturate GPU compute cores, long generations stress memory bandwidth, and long-context conversations quietly eat VRAM until the engine can’t schedule anything new. 

Here, the pressure on memory bandwidth came from the routing signal chosen for this deployment. By mapping Inference Engine's native token-usage metric onto the Gateway’s KV-cache signal, the routing plane gained a real-time view of memory pressure across the entire 17,000 fleet. (Depending on the workload, the Gateway can route on other signals too, like queue depth or running concurrency.) 

Under live production loads, as the primary region climbed toward its high-bandwidth memory (HBM) limits, the Gateway detected the saturation the moment the cluster crossed its 40% KV-cache utilization threshold; it then automatically began routing the overflow to the next healthy region. No operator intervention was needed. The complexity of running in multiple regions simply never reached the user. 

The payoff

By routing traffic based on live KV-cache utilization, this GKE Inference Gateway setup effectively pools globally scattered compute capacity into one unified engine. For this deployment, the result was a near-linear throughput boost across three global regions, with virtually zero routing overhead.

This translates directly into maximizing 'intelligence per dollar,' extracting near-perfect proportional performance out of every accelerator you add to your fleet, rather than letting capital go to waste.

What this means for your team 

If you’re planning your own distributed inference deployment, five lessons from this work stand out:

  • Smarter load balancing pays for itself. Round-robin routing wastes expensive GPU capacity because it can’t see memory or compute pressure. Routing on real-time application signals turns fragmented regional clusters into one efficient fleet — the difference between stranded hardware and 90%+ utilization of scarce compute.

  • Agentic workloads change the bottleneck. Long-running agents with extreme context windows exhaust memory long before there’s no more compute. If your routing layer can’t see memory pressure, your compute will strand compute behind full VRAM. Make KV-cache utilization a first-class routing signal.

  • AI traffic breaks web-era assumptions. Traditional load balancers are tuned for sub-second transactions; LLM requests can run for minutes. Plan connection limits and timeouts for AI-scale latency early, or expect aborted connections in production.

  • Your routing layer must integrate with native serving patterns. Distributed LLM engines have master-worker topologies where only certain pods can serve traffic. By pairing your Gateway with native Kubernetes constructs like LeaderWorkerSet (LWS), your global routing respects local pod topologies out of the box, saving your team from building custom proxy infrastructure.

  • For large foundation model builders, bet on an open, portable stack. Teams operating at frontier scale face the most acute capacity fragmentation, forcing them to hunt for compute resources across whichever regions have availability capacity. An open, portable inference stack such as LLM-d on GKE lets you absorb that capacity wherever it lands, rather than hard-wiring your serving architecture to any single cluster, region, or bespoke infrastructure.

Next steps 

Ready to maximize your distributed accelerator efficiency and set up global cross-region load balancing with multi-cluster GKE Inference Gateway? 

  1. Deploy it yourself: Set up the multi-cluster GKE Inference Gateway. 

  2. Understand the architecture: About multi-cluster GKE Inference Gateway. 

  3. Learn about the cross-region spillover behavior featured in this post: About elastic cross-region high availability and Configure elastic cross-region high availability.

Changing the game: Using agentic AI to secure infrastructure code

18 septembre 2026 à 18:00

AI is accelerating software development at an unprecedented pace. But as code generation scales, so do the challenges of securing the code, especially emerging AI-based vulnerability exploitations. To meet these challenges, the Google AI and Infrastructure team is transforming how we approach security. In this article, we discuss new AI-native agentic methods that we’ve developed that systematically embed high-precision, pervasive vulnerability scanning and patching directly into Google’s software development lifecycle. By continuously scanning every code change across hundreds of millions of lines of code that we deploy onto our infrastructure, we are preventing hundreds of vulnerabilities per month from ever reaching our code base or production, defending our global network, AI infrastructure and our users. 

Solution architecture and implementation

image1

Pervasive pre-submit agentic scanning: security as part of ongoing software development

Traditionally, the technology industry relies on large one-off security scans that are slow and lack sufficient context. As a result, they often find vulnerabilities too late. Our approach instead focuses on pre-submit scanning, where we evaluate each code check-in (across every layer of the stack) in real-time using AI agents. By integrating the pre-submit scan into the tools developers already use, security becomes a continuous routine process, similar to rule checkers, readability reviews or other software development tools. Also, from an AI perspective, scanning each individual code change requires much less context than performing a large one-off scan, significantly improving the scan’s effectiveness. 

The importance of localized threat models

For this initiative, we evolved Mantis, our open-source multi-agent review harness, to increase the precision of our security agents by matching them with a cohort of robust localized threat models. Rather than relying on static decoupled documents, the threat models use live codebase metadata. The scanning agent improves its accuracy further using a dependence call graph across packages and libraries to expand and refine its threat model context. Making threat models part of our ongoing vulnerability scanning encourages developers to continuously update threats and dependencies, keeping the models up-to-date. Using localized and precise threat model data translates to dramatic accuracy improvements, bringing our false-positive rates down to 3% in some cases.

Specialized triage agents speed up development

Vulnerability scanning as part of code check-in requires it to respond quickly to the developer or agents generating the code, so as not to impede engineering productivity. To get responses with low latency, we run a two-step validation process. First, we run a quick lightweight scan that validates its findings against a specialized triage agent. This agent programmatically checks the actual structure of the code (using abstract syntax tree parsing, call-graph traversal, and pre-indexed domain safety rules) to prove that the vulnerable path is actually reachable by an attacker. This agent gets over 92% precision and completes its work in less than a minute. Then, a post-submit scan as part of nightly integration testing serves as a second layer of defense, using off-peak cycles to test for vulnerabilities that may have been introduced across multiple changes. 

Bug fix agents close the loop

Finding vulnerabilities is only half the battle. The last component of our solution is an automated bug-fix agent that uses the scan results and generated proofs (snippet of code that demonstrates how the vulnerability is exercised) to autonomously construct precise fixes that are consistent with our internal coding standards. The agent submits the fixes for human review as part of the original change request’s review, further reducing the time between detection and resolution. 

Learnings and call to action 

Embedding continuous scanning directly into the software development lifecycle has been a game changer at Google; its suggestions are widely adopted, and it’s prevented a multitude of vulnerabilities from being introduced into the codebase. But any organization wishing to improve security can adopt a similar AI-native approach, following these principles: 

  1. Keep systems separate: To prevent bias, keep the harnesses, rules, and context for each of your development, scanning, triage agents separate. Pair lightweight AI scans with deterministic, structural validation to drive down latency and improve accuracy. 

  2. Use context wisely: Feed your agents your existing threat models. Precise context is the answer to reducing false positives, and up-to-date threat models set a high floor on a team's security posture by improving the rate of true positives in presubmit scanning.

  3. Build a good harness: While the choice of the underlying model is important, using a multi-agent harness can have substantial impact, by helping compensate for variability in model choice. 

  4. Automate the fix: Use agents to also propose human-in-the-loop fixes, to further reduce time-to-resolution. 

If you want to get started on your own AI-native security transformation, Mantis is now available as open source for you to use and benefit from. You can also learn more about the fundamentals of cybersecurity and the other platforms that power this agentic pipeline: Google Cloud, Gemini Enterprise and Gemini models running on Trillium and Ironwood TPUs. And you can get inspiration from how agentic vulnerability scanning and remediation defends Google Cloud customers as an integral part of Google Cloud’s secure software development lifecycle (SDLC) effort.


With special recognition to critical team members who made this delivery possible: Stella Voutsina (Lead Program Manager), Yulong Zhang (Senior Staff Security Engineer, Mantis), and Nick Galloway (Staff Security Engineer, Mantis).

For SeaVerse, GKE Agent Sandbox reduces infrastructure costs by 60%

16 septembre 2026 à 18:00

Editor’s note: Today we hear from SeaVerse, a gaming startup from SeaArt that is building a platform for playable AI experiences, where users can open lightweight games, character chats, and interactive apps, or create their own experiences from a prompt. To support that creative loop, SeaVerse needed infrastructure that could run dynamic, multi-tenant sandbox workloads with strong isolation, low latency, better observability, and more flexible costs. Google Kubernetes Engine (GKE) and GKE Agent Sandbox gave SeaVerse the managed foundation from which to execute these AI workloads, helping the team reduce their infrastructure costs by up to 60%, while giving creators a faster path from idea to playable experiences. Read on to learn more.


What if AI were a playground? Welcome to SeaVerse, a creation-first platform for playable AI experiences. Here, an AI creation can be as peaceful as drawing a path for a snake to follow, or as chaotic as a music-backed stickman simulation. Some people come to play lightweight games. Others come to chat with AI characters, try interactive apps, create visual patterns, share what they made, or remix an idea into something new. 

We built SeaVerse around a simple promise: Every experience should feel immediate and easy to share. A creator should be able to describe an idea in plain language, refine the result, and publish it in moments, without a traditional coding workflow. 

Delivering that simplicity requires serious infrastructure. Every creation that users make moves through the same chain: generate, run, preview, debug, publish, remix. If any part of that chain is slow, unstable, or poorly isolated, users feel it immediately. That’s why we turned to GKE and GKE Agent Sandbox. 

The infrastructure challenge of instant interaction 

What looks effortless to a user is anything but on our end. Every creation on SeaVerse runs as a distinct workload and is expected to behave reliably from the first interaction. 

Because each workload runs in its own environment, we needed clear security boundaries between users, creations, and sandboxes. But overly strict isolation could slow the very creative loop we were trying to protect, and when something went wrong, diagnosing it was costly. Our engineers had to trace problems across multiple parts of the execution chain with little visibility into what was happening inside the environment. 

We explored existing sandbox approaches, but needed deeper kernel-level isolation and native observability at scale to support fast diagnosis across multi-tenant environments. Something had to change.

Building on GKE and GKE Agent Sandbox 

We chose GKE because we needed a reliable, secure way to operate Kubernetes without turning our engineering team into a cluster maintenance team. GKE brought together the proven ecosystem and operational tooling we needed, freeing us to focus on building the platform rather than managing the infrastructure beneath it. 

As a Kubernetes primitive designed for agent code execution and computer use, GKE Agent Sandbox addressed our requirement for strong isolation, enforcing strong security boundaries without slowing down the creation experience. By utilizing GKE Agent Sandbox with Kata Containers+Cloudhypervisor (microVM), we’ve achieved the perfect balance of multi-cloud flexibility and robust security, option to switch isolation runtime between microVM and gVisor, running our AI sandboxes safely. GKE empowers us to scale toward our long-term vision of supporting over a million sandboxes. Built on gVisor, it provides kernel-level isolation for dynamic sandbox workloads while preserving the Kubernetes orchestration model, so that they can be managed through the same scheduling, monitoring, and operations as the rest of the cluster.  

With SeaVerse, users can generate interactive experiences from a single prompt. After an experience is generated, GKE Agent Sandbox supports the run, test, integration, and verification steps needed to make it ready to preview, refine, and publish. At general availability, it supports allocating up to 300 sandboxes per second, per cluster, with 90% of allocations completing in 200 milliseconds. Together, GKE and GKE Agent Sandbox gave us a reliable foundation for AI-generated interactive workloads that helped keep our team focused on the product experience. 

From black box to glass box 

Before GKE Agent Sandbox, a failed sandbox workload could feel like flying blind. We could often see that something had gone wrong, but didn’t have enough runtime status, metrics, or failure signals to understand why. 

Now, Google Cloud’s native logging and monitoring reach directly into those sandboxed environments, giving us a clearer view of workload behavior, faster issue resolution, and a stronger foundation for managing multi-tenant workloads. 

That visibility matters to developers, but it also matters to the platform’s users: A creator never sees the logs, the cluster, or the orchestration layer. They see whether an experience opens quickly, whether it responds when they draw, click, chat, or share, and whether they can keep building without friction. 

Flexibility that translates to savings

GKE Agent Sandbox also changed how we think about cost. Previously, running secure sandboxed environments meant stronger dependencies on specific server types, which limited how precisely we could match resources to each workload. With GKE Agent Sandbox, we can run secure, isolated workloads on appropriately sized cloud VMs. This gives us greater flexibility in resource allocation and helped us cut our infrastructure costs by up to 60%. 

That same flexibility extended to storage. Not all SeaVerse creations are built in a single session. Some evolve over time as creators return to refine them, build on earlier ideas, or invite others to remix what they’ve made. Our previous architecture didn’t support the persistent file-system capabilities those more complex use cases demanded, but that gap is gone now. We can attach persistent storage where workloads require it while maintaining the isolation boundaries that multi-tenant AI experiences need. For creators, that means experiences that are fast to open and easier to refine, revisit, and build on over time. 

The next remix 

Supporting creations that can evolve and deepen is central to what we’re building. It’s still early in what playable AI can become. As the platform grows, we need to keep strengthening what matters most: stability, observability, elastic scaling, and cost efficiency, all in service of a creator experience that stays fast, reliable, and expressive. 

We’re also exploring additional Google Cloud tools to support smarter analytics and creation assistance. Gemini and agent models could help operators and creators better understand how experiences perform. BigQuery AI and ML capabilities can support use cases such as churn prediction, LTV and ROI prediction, and user segmentation. Multimodal tools such as Imagen and Veo on Gemini Enterprise Agent Platform open up new possibilities for material analysis, creative generation, and AI interactive content production. 

Our goal is to make AI experiences feel immediate, expressive, and connected. With GKE and GKE Agent Sandbox, we have a stronger foundation for the next generation of playable AI.

Introducing Filestore agent volumes: fully managed storage for agent workspaces

15 septembre 2026 à 18:00

From running build tools, to data analysis pipelines, to collaborative research, executing data-driven tasks is essential for any enterprise agent. 

Today, platform teams often stitch together custom workarounds to address agent storage requirements, which could include shuttling state back and forth between agent sandboxes and centralized storage or manually managing local disks and/or self-hosted file systems. However, as agent fleets scale, these approaches force difficult trade-offs between cold-start latency, operational complexity, and the cost of idle, pre-allocated storage.

As organizations scale agent sandboxes to thousands or even millions of concurrent sessions, storage must evolve to overcome these trade-offs and meet the needs of these dynamic workloads, which require strict workspace isolation, instant session resumption, elastic pay-per-use economics, and fluid multi-agent collaboration.

To meet these emerging demands, we’re expanding our AI storage portfolio and announcing availability of Filestore agent volumes, a new, fully managed capability purpose-built to deliver high-performance, elastic file storage for scaling agentic workloads on Google Cloud.

Purpose-built storage for AI agent workspaces

Autonomous agents require isolated runtime environments to safely execute dynamic code, install third-party packages, and run tools without putting host infrastructure or tenant data at risk. While Agent Substrate on GKE and GKE Agent Sandbox provide the dedicated compute environments needed to run high-density agent fleets, those sandboxes also need dedicated persistent workspaces to operate on.

Filestore agent volumes within Google Cloud Filestore, give you purpose-built agentic storage to complement your agentic compute via a dynamic provisioning architecture designed specifically for the scale and elasticity of AI agent fleets. Co-designed with Agent Substrate to support agentic fleets at scale, Filestore agent volumes provide GKE sandboxes with instantaneous access to isolated, persistent file storage. When configured to leverage Filestore, GKE storage management happens behind the scenes: Every time GKE launches a sandbox for a new agent task, Filestore automatically allocates and attaches a dedicated, isolated file workspace to that environment in milliseconds. Platform teams don't need to manually create, attach, or tear down storage volumes for individual agent runs; instead, the system handles the entire volume lifecycle automatically as your agent fleet scales up and down. The result is an efficient, end-to-end infrastructure solution for cost-effective agent management that provides: 

  • Granular isolation and enterprise guardrails: Agent platforms face security and data leakage risks when running untrusted, autonomous code. Filestore agent volumes enforce strict boundary controls and granular access permissions per workspace, ensuring agents operate exclusively within their designated directories and keeping dynamic toolchains strictly isolated across tenants.

  • Sub-second session resumption: Traditional storage provisioning approaches can introduce cold-start latency that stalls interactive agent sessions. Agent volumes attach and detach in milliseconds, making it possible for orchestrators to aggressively suspend idle sandboxes to save compute costs, and resume instantly when new tasks or user inputs arrive.

  • Smart lifecycle economics and pay-per-use pricing: Pre-allocating fixed-size, high-performance storage for thousands of short-lived or intermittent agent tasks can create massive storage waste. With agent volumes, platforms pay only for the storage capacity consumed and benefit from automatic lifecycle tiering. This means you get high performance without wasted spend: When your agents aren’t actively reading/modifying code or analyzing datasets, you can automatically shift idle workspace state to lower-cost storage.

  • Multi-agent collaboration: Coordinating multi-agent swarms can result in brittle data-passing pipelines and risk of file collisions. Built with native Read-Write-Many (RWX) support and POSIX file locking, agent volumes allow orchestrators to attach a single shared workspace across multiple agents. Collaborating agents can safely co-author, test, and review project files concurrently with file-level consistency and protection against write conflicts.

Powering next-generation agentic workloads

By providing an elastic, high-performance, and isolated file tier, Filestore agent volumes unlock a wide spectrum of agentic workloads and use cases in production:

  • Software engineering and coding sandboxes: Agentic coding platforms can spin up thousands of isolated workspaces where agents safely install libraries, write multi-file patches, run build tools, and execute unit tests, all leveraging standard POSIX file semantics with no need for storage-specific customization.

  • Collaborative multi-agent swarms: Complex workflows, such as a lead orchestrator delegating tasks to dedicated research, code generation, and validation sub-agents, can directly share a unified file tree. RWX support allows agents to co-author and review project files concurrently without write conflicts.

  • Interactive long-horizon workflows: For user-in-the-loop applications (such as agents that require asynchronous user approval or run multi-hour data analysis pipelines), platforms can suspend idle agent sandboxes to minimize compute waste, then resume execution on demand with sub-second responsiveness.

Get started today 

If you are building an Agent-as-a-Service platform, scaling coding assistants, or deploying enterprise agent fleets, your storage tier should accelerate your innovation — not hinder it.

Filestore agent volumes are now available to all Google Cloud customers for non-production workloads. GA support for production workloads is available via allowlist. This new offering features out-of-the-box integrations with Agent Substrate on GKE and GKE Agent Sandbox to help you build responsive, scalable, and cost-efficient agent platforms today.

To request access to Filestore agent volumes, submit this form and visit the Filestore documentation and GKE documentation to learn more.

Announcing: New Windows App client-side endpoints for Azure Virtual Desktop

14 septembre 2026 à 20:24
Beginning in early October 2026, Windows App will begin using three new wildcard fully qualified domain names (FQDNs) for client-side service traffic to Azure Virtual Desktop. These domains are already included in the cloud-side connectivity requirements.

Google is a leader in The Forrester Wave™: Public Cloud Platforms, Q3 2026

14 septembre 2026 à 18:00

We are excited to share that Google Cloud was named a Leader and received the highest score in the ‘current offering’ category in the Forrester Wave™: Public Cloud Platforms, Q3 2026 report, which examines the 10 most significant public cloud providers across 30 comprehensive criteria, Google also received the highest possible score in 23 out of 30 evaluation criteria, including, but not limited to vision, innovation, AI development services, database services, analytics services, containers and kubernetes services, modernization services, and security services. We believe Forrester’s recognition confirms our belief that to lead in the agentic era, you need a complete, integrated platform that’s engineered from the ground up, from silicon to systems to models.

Build on co-designed infrastructure proven in global enterprises

For over a decade, our infrastructure engineers, application developers, and AI researchers worked side by side to co-design infrastructure to power Gemini, Search, YouTube, Maps, and Gmail. We couldn't simply buy the platform and infrastructure we needed; we had to invent it. This led to the creation of everything from TPUs, the Transformer architecture, Kubernetes, Axion, and now Gemini.

In the agentic era, you need an integrated AI stack, where compute, orchestration software, modernization tools, and global networks operate together to give you more value from your investments — even if you’re not working at the frontiers of AI research. At Google Cloud, we’ve worked tirelessly to bring these breakthrough innovations to leading enterprises, startups, and frontier labs to help them achieve new levels of scale and efficiency, and we believe Forrester’s evaluation validates that strategy: 

“Google Cloud’s vision is to enable the ‘agentic enterprise,’ and AI already permeates its platform, positioning the company to push further up the tech stack toward business users who increasingly shape AI adoption in the enterprise. Google Cloud is a good fit for enterprises seeking rapid technology innovation and a broad AI-enabled cloud platform.” - The Forrester Wave™: Public Cloud Platforms, Q3 2026 report

Run agents quickly on a secure, flexible platform

Most traditional infrastructure can’t keep pace with agents, and enterprises need a scalable alternative. But you don’t want a new, greenfield platform just for AI agents. Kubernetes is the proven industry standard for modern enterprise applications — from microservices and transactional databases to real-time LLM inference. We are evolving Google Kubernetes Engine (GKE) and our operations tooling so organizations can scale autonomous agents alongside traditional workloads on a single, proven platform.

Forrester gave Google Cloud the highest scores possible in Container and Kubernetes services, Serverless/FaaS services, and Operations management services, noting:

“Operators will find strong offerings in operations management as well as containers and Kubernetes services. Our evaluation did not identify significant capability gaps.”

Over the past three months, we’ve enhanced our infrastructure portfolio to help teams scale agentic workloads with enterprise predictability. Recent updates let you:

  • Safely execute untrusted agent code alongside traditional workloads with default-deny security using GKE Agent Sandbox (GA) and Cloud Run Sandboxes (preview), which provision lightweight, gVisor-isolated boundaries for your agent in under a second (and up to 300 sandboxes/sec per cluster).

  • Eliminate up to 90% of idle compute costs by serializing your container RAM state directly to Google Cloud Storage with GKE Pod Snapshots, allowing you to suspend idle agent sessions in ~100ms and resume them in ~280ms.

  • Cut time-to-first-token (TTFT) up to 70% and double cache-hit rates with predictive routing in GKE Inference Gateway, which uses a continuously trained ML model to make routing decisions based on real-time traffic data.

Ground your agents with real-time enterprise data

Agents are only as effective as the context that grounds them. Traditional distributed data topologies separate operational databases from analytical systems through fragmented, multi-hop pipelines. In the agentic era, this divide introduces multi-hop latency, stale context, and governance friction.

Our Agentic Data Cloud evolves the enterprise data platform from a static repository into a dynamic reasoning engine. It unifies transaction processing and analytical intelligence into an active system of action, providing the real-time context and deterministic responsiveness that autonomous workflows require. Google received 5/5 scores across the Database services, Analytics services, Data integration services, and Data Governance services criteria:

“Google Cloud’s traditional strength in database services and analytics drives strong performance, including multicloud and hybrid capabilities, along with an Agentic Data Cloud that bridges analytics and transactional systems.”

Over the past three months, we’ve introduced key capabilities to the Agentic Data Cloud to help customers unify their data estates:

  • Enable agents to query live financial and supply chain records without costly data movement using SAP BDC Connect for BigQuery (GA), which provides bi-directional, zero-copy data sharing between your SAP systems and BigQuery.

  • Map and infer business meaning across your entire data estate with Knowledge Catalog. You can now aggregate native context across your Google and partner data platforms, semantic models, and third-party catalogs, unifying them into a single, governed source of truth.

  • Access live data from Iceberg and BigQuery from the PostgreSQL data plane with Lakehouse federation. Perform live joins between AlloyDB's transactional data and historical insights in BigQuery or Iceberg without any data movement. You can also replicate data continuously to BigQuery and, importantly, to Iceberg tables directly from AlloyDB with Datastream.

The benchmark is set: Build what’s next on Google Cloud

We are honored that Forrester has named Google Cloud a Leader in The Forrester Wave™: Public Cloud Platforms, Q3 2026. We believe this recognition validates decades of foundational research, disciplined full-stack co-design, and our commitment to building an open, reliable cloud.

The era of fragmented infrastructure has come to an end. Whether your organization is an AI research lab scaling models across one million accelerator chips, a global financial exchange settling trillions in clearing systems, or an enterprise empowering millions of users with autonomous workflows, Google Cloud delivers the performance, scale, security, and data foundation to build what’s next.

Take the next step in your cloud journey:

Google Cloud partners with CIQ to provide an enterprise-grade experience for Rocky Linux

6 avril 2022 à 18:00

At Google Cloud, we strive to offer a great customer experience for enterprises by building a robust and supported platform for running all Linux-based workloads.

This mission is why we were one of the first cloud providers to offer purpose-built Rocky Linux images when Rocky Linux debuted last year as a replacement option for CentOS. We were also one of the first hyperscalers to sponsor the Rocky Enterprise Software Foundation (RESF) to support the open source community behind this Linux distribution. With these efforts, we’re pleased that many customers are already running Rocky Linux in Google Cloud today.

Today, we’re excited to announce that we’re taking another step in furthering the support we provide for Rocky Linux. We’re partnering with CIQ—the company started by CentOS co-founder and Rocky Linux founder Gregory Kurtzer featuring core expertise across Linux, cloud, HPC, containers and security— so we can provide customers a new and improved experience for Rocky Linux on Google Cloud. 

Starting today, customers can leverage Google’s support offerings to file support cases for Rocky Linux. Google support teams and the Rocky Linux experts at CIQ are working together to address customer issues to help ensure they get enterprise-grade support. If you already have a paid support plan with Google, you will be able to open a case for an issue related to Rocky Linux. Google teams can expediently help resolve the issues, backed by CIQ expertise, giving you an integrated experience of using Rocky Linux on Google Cloud. 

"We asked ourselves, how do we bring the best value to everyone? Through this partnership, anytime you use our Rocky Linux on Google Cloud, both Google and CIQ jointly have your back! From the cloud platform itself, all the way through the enterprise operating system, every aspect of using Google Cloud is supported by a single call to Google, and together, we are your escalation team.”—Gregory Kurtzer, CEO of CIQ and Founder/Director of Rocky Linux and the RESF

In addition to CIQ-backed support for Rocky Linux, Google is also working with CIQ to provide a streamlined product experience - with plans to include performance-tuned Rocky Linux images, out-of-the-box support for specialized Google infrastructure, tools to help support easy migration, and more. We’re doing these updates in a community-friendly way. Together with CIQ, Google is helping to create a Rocky Linux Cloud SIG that aims to provide optimized, standardized, and simplified Rocky Linux experience. 

If you’re currently looking for alternatives to CentOS as it reaches end of life, Rocky Linux on Google Cloud can have you covered both from a product and support perspective. So, take Rocky for a spin if you haven’t already, and if you have questions or suggestions on how we can help you, please don’t hesitate to reach out to us. To learn more, please also join us for a webinar discussion on April 6th 2022 at 11.00am PT.

Introducing Topaz — the first subsea cable to connect Canada and Asia

6 avril 2022 à 15:00

There’s a new subsea cable in town: Topaz, the first-ever fiber cable to connect Canada and Asia. 

Once complete, Topaz will run from Vancouver to the small town of Port Alberni on the west coast of Vancouver Island in British Columbia, and across the Pacific Ocean to the prefectures of Mie and Ibaraki in Japan. We expect the cable to be ready for service in 2023, not only delivering low-latency access to Search, Gmail and YouTube, Google Cloud, and other Google services, but also increasing capacity to the region for a variety of network operators in both Japan and Canada. 

Google is spearheading construction of the project, joined by a number of local partners in Japan and Canada to deliver the full Topaz subsea cable system. Other networks and internet service providers will be able to benefit from the cable’s additional capacity, whether for their own use or to provide to third parties. And, similar to other cables we’ve built, with Topaz we will exchange fiber pairs with partners who have systems along similar routes. This is a longstanding practice in the industry that strengthens the intercontinental network lattice for network operators, for Google, and for users around the world.

topaz.jpg

Network infrastructure investments like Topaz bring significant economic activity to the regions where they land. For example, according to a recent Analysys Mason study, Google’s historical and future network infrastructure investments in Japan are forecasted to enable an additional $303 billion (USD) in GDP cumulatively between 2022 and 2026. 

The width of a garden hose, the Topaz cable will house 16 fiber pairs, for a total capacity of 240 Terabits per second (not to be confused with TSPs). It includes support for Wavelength Selective Switch (WSS), an efficient and software-defined way to carve up the spectrum on an optical fiber pair for flexibility in routing and advanced resilience. We’re proud to bring WSS to Topaz and to see the technology is being implemented widely across the submarine cable industry.  

While Topaz is the first trans-Pacific fiber cable to land on the West Coast of Canada, it’s not the first communication cable to connect to Vancouver Island. In the 1960s, the Commonwealth Pacific Cable System (COMPAC) was a copper undersea cable linking Vancouver with Honolulu (United States), Sydney (Australia), and Auckland (New Zealand), expanding high-quality international phone connectivity. Today, COMPAC is no longer in service but its legacy lives on. The original cable landing station in Vancouver — the facility where COMPAC made landfall on Canadian soil — has been upgraded to fit the needs of modern fiber optics and will house the eastern end of the Topaz cable.

Traditional and treaty rights, and local communities, are deeply important to our infrastructure projects. The Topaz cable is built alongside the traditional territories of the Hupacasath, Maa-nulth, and Tseshaht, and we have consulted with and partnered with these First Nations every step of the way. 

"Tseshaht is very proud of this collaboration and our partnership with Google, who has been very respectful and thoughtful in its engagement with our Nation. That’s how we carry ourselves and that's how we want business to carry themselves in our territory.“ — Tseshaht First Nation - Elected Chief Councillor-Ken Watts 

“The five First Nations of the Maa-nulth Treaty Society are pleased that we have concluded an agreement with Google Canada and have consented to the installation of a new, high-speed fiber optic cable through our traditional territories. This agreement, in which both Google Canada and our Nations benefit, is based on respect for our constitutionally protected treaty and aboriginal rights and enhances the process of reconciliation. We would also like to acknowledge the sensitivity that Google Canada expressed during our talks in regard to the pain and trauma experienced by our people as a result of residential school experience. We look forward to a long and mutually beneficial relationship with Google Canada.” —Chief Charlie Cootes, President of the Maa-nulth Treaty Society

“Google's respect towards our Nation is appreciated and has good energy behind it.” —Hupacasath First Nation - Elected Chief Councilor - Brandy Lauder 

With the addition of Topaz today, we have announced investments in 20 subsea cable projects. This includes Curie, Dunant, Equiano, Firmina and Grace Hopper, and consortium cables like Blue, Echo, Havfrue and Raman — all connecting 29 cloud regions, 88 zones, 146 network edge locations across more than 200 countries and territories. Learn about Google Cloud’s network and infrastructure on our website and in the below video.

Dynamic capacity management for AI infrastructure

26 août 2026 à 15:30

The internet connected billions of people and mobile devices, putting computers in every hand. Now, we’re in the middle of the next big technology shift, deploying millions of autonomous AI agents to work alongside employees and end users. Today, we announced new FinOps controls for Gemini Enterprise to help organizations manage project-level AI spend and eliminate token shock. But the sheer scale of the agentic era is placing new constraints at every layer of the stack, including infrastructure. AI workloads are notoriously difficult to architect, resource-intensive, and bursty, which can also lead to scaling bottlenecks and large pools of underutilized — or misutilized — compute resources. 

Organizations need insights to help them extract more value from their infrastructure investments. In this blog, we outline best practices for dynamic capacity management — scheduling and utilization strategies to help you run enterprise and AI applications on a single, flexible foundation with predictable cost and performance. These capabilities are designed to augment our on-demand, Spot and committed use discount (CUD) consumption models, which provide flexible pricing and discounting for your workloads. Let’s jump in.

Here's a quick summary

Three ways you can implement dynamic capacity management:

  1. Schedule capacity for planned events. Schedule mission-critical resources (GPUs, TPUs and select VM families) ahead of planned events using calendar mode, or optimize costs for batch jobs with flexible start times using flex-start mode in Dynamic Workload Scheduler. Once you obtain the capacity, those resources are guaranteed for the specified duration.

  2. Maintain service continuity by creating a fallback plan for every application. Define automated, prioritized hardware fallback lists using managed instance groups (MIGs) so your apps automatically pivot to the next approved compute option when your preferred option isn’t available.

  3. Automate your entire capacity management lifecycle on a single, adaptive control plane. Google Kubernetes Engine (GKE) provides an agent-native environment to orchestrate the entire process — from fallback lists using Custom ComputeClasses, to granular hardware slicing with dynamic resource allocation, so agents can rapidly spin up in secure sandboxes and containers while it dynamically reallocating resources on the fly.

Why architectural flexibility matters

Ninety percent of enterprises want to deploy agents within the next three years, but only 17% of IT leaders feel confident their current IT setup can handle the load. Because these workloads have unique performance needs, organizations are racing to adopt specialized infrastructure, including accelerators (GPUs, TPUs) and CPUs with customized compute, memory, and storage ratios. However, agents also require access to enterprise applications and databases — often at a volume and scale that vastly exceeds typical human usage. Handling the intense demands of both agents and the applications they interact with requires a dynamic infrastructure. Infrastructure teams can leverage custom-designed processors like Google’s Axion to meet these needs, but hardware isn’t a complete solution. They also need ways to use that infrastructure wisely, solving execution inefficiencies to enable more flexibility across the stack.

How to overcome infrastructure constraints

Achieving this kind of flexibility requires a two-pronged approach: securing resources for the demand you can predict, and building automation to respond to the demand you can't. Combining the two, you can preschedule capacity for planned events and your infrastructure can adapt to unexpected changes without manual intervention.

1. Schedule capacity for planned events

You can secure mission-critical capacity ahead of scheduled milestones, offline training, or anticipated demand surges using Dynamic Workload Scheduler. By scheduling the resources you need up front, you optimize your spend and ensure you get access to the compute resources you need. Dynamic Workload Scheduler supports hardware accelerators (TPUs and GPUs) and select CPUs with two distinct modes:

  • Flex-start mode: Use this for latency-tolerant workloads like batch processing, model training, or offline fine-tuning. Instead of requiring resources immediately, you submit a defined duration request and the system intelligently queues your job, provisioning the resources as soon as capacity becomes available. This maximizes cost-efficiency and drastically improves your ability to obtain high-demand accelerators.

  • Calendar mode: Use this for mission-critical, time-bound events like a major product launch, a scheduled migration, or a seasonal traffic surge. By specifying the exact start and end dates of your event, you create a future reservation. This guarantees the requested capacity will be available when the event begins.

1

2. Maintain service continuity by creating a fallback plan for every application

Not every spike in traffic is predictable. You also need to plan for unexpected traffic from, say, a breaking news cycle or a sudden market shift that drives a surge in user activity. To help your services get the resources they need without interruption, you need a fallback plan — an automated, prioritized sequence of acceptable hardware configurations. This strategy:

  • Decouples your workloads from a single VM shape, size, or configuration. This allows them to run without manual intervention if your preferred option is unavailable

  • Allows you to execute a progressive tech refresh by adopting the newest VM generations as your primary choice while keeping older generations as an automatic fallback option.

If you run non-containerized workloads on Google Compute Engine, you can dynamically manage capacity with instance flexibility in managed instance groups (MIGs) and bulk VM creation. Instance flexibility lets you specify multiple machine types for your VM instances rather than being limited to a single machine type.

How it works: If your preferred machine type is temporarily unavailable, the MIG automatically provisions a compatible alternative from your list based on real-time capacity. When combined with location flexibility — by specifying multiple zones your MIGs can search within a region — you can drastically improve your provisioning success rate. If your MIGs use Spot VMs, Compute Engine automatically integrates with Spot capacity signals to prioritize machine types that offer longer estimated uptimes and lower risk of pre-emption.

2

You can also extend instance flexibility to your block storage layer by setting baseline disk defaults and configuring disk overrides so your storage adapts when a VM falls back to a different machine type. 

How it works: Most of the time you can simply rely on our default options, omitting ‘disk type’ from the instance template entirely. However, for data disks that will outlive their associated VMs, it’s possible to enable a fast, durable Hyperdisk across multiple VM generations.

While Compute Engine provides instance flexibility for organizations working with virtual machines, GKE goes a step further and automates the entire capacity lifecycle from a single control plane. With GKE custom ComputeClasses, platform teams can design multi-dimensional fallback lists, automatically combine different VM machine families, sizes, and ratios, scale across multiple zones, and shift between on-demand and Spot VMs. By using Dynamic Workload Scheduler as a capacity target, and custom ComputeClasses to define the policy and priority, you can fully automate the capacity management lifecycle.

How it works: Once you’ve set up ComputeClasses, GKE automatically detects when a preferred node configuration is unavailable and falls back to your pre-approved alternative options in order of priority. When active migration is enabled, GKE gracefully migrates workloads back to higher-priority node configurations as capacity becomes available. For short-lived disks such as boot disks, GKE dynamically picks the right defaults based on the instance family. However, for long-term disks that will outlive the VM, you can use Hyperdisk.

3

Another GKE feature, dynamic resource allocation, helps eliminate wasteful, all-or-nothing hardware assignments by letting developers define advanced rules that dictate how resources are consumed.

How it works: Instead of claiming an entire GPU or TPU, your application specifies its exact parameters — such as total memory or number of cores — and the system allocates the perfect slice of hardware, helping to maximize utilization and reduce costs. 

 

4

Take the next step toward dynamic infrastructure

Scaling AI shouldn’t mean linearly scaling your infrastructure budget or accumulating more tech debt. As these examples show, the right tools can help you overcome constraints and dramatically alter the value you get from your compute investments. Here are three steps to get started:

  1. Audit your workloads for immediate cost-savings: Identify any applications currently tightly coupled to a single VM family, machine type, or availability zone, and map out viable alternative hardware shapes. Look beyond your existing configurations to evaluate new compute options that might better serve or act as alternatives based on your workload-level objectives. Then use Compute Engine MIGs, bulk VM creation or GKE Custom ComputeClasses to adopt them automatically, integrating them into your fallback lists.

  2. Commit to a minimum spend for deeply discounted prices: Receive automatic discounts for sustained use, or up to 63% off when you sign up for Compute flexible committed use discounts, where your discount is tied to the resources you use regardless of the specific machine type or location.

  3. Engage your account team: Reach out to your Google Cloud account team to craft a tailored capacity management strategy and configure your automated fallback lists.

Bringing gVisor sandboxes to distributed Ray clusters

25 août 2026 à 18:00

The reinforcement learning (RL) ecosystem is rapidly adopting Ray as the unified compute runtime for complex post-training workflows. Across Google Cloud, we see customers using Ray for workloads ranging from multimodal data pipelines to frontier RL. But as agentic and reasoning models evolve, a critical bottleneck has emerged: orchestrating secure, isolated sandboxes at scale to safely execute dynamic rollouts, code generation, and multi-turn tool interactions. Today, in partnership with Anyscale, we are excited to introduce an experimental library for Ray that leverages agentic AI technologies being developed at Google to bring native, high-performance sandboxing directly into distributed Ray clusters.

Sandboxes as Ray Primitives

Ray has become a common runtime for orchestrating post-training workloads. Frameworks including veRL, NeMo-RL, SLIME, MILES, and SkyRL already use Ray to coordinate distributed trainers, inference engines, rollout workers, and other components.

When we designed Ray Sandboxing, an important goal was to make it fit naturally into the existing Ray programming model rather than introduce a separate abstraction for isolated execution. A sandbox has many of the same properties as other resources managed by Ray: it needs to be placed on a machine, assigned resources, created and destroyed, recovered from failures, and scaled with the surrounding workload. This led us to represent each high-level sandbox through a Ray Actor:

image1

The Ray scheduler decides which node should run a sandbox and reserves the corresponding CPU and memory resources. The sandbox Actor manages its lifecycle, while gVisor provides the isolated execution environment on that node.

Starting in Ray 2.58, framework authors and researchers can manage sandboxed environments using the same Ray APIs and patterns they already use for the rest of their workload. For example:

code_block
<ListValue: [StructValue([('code', 'import ray\r\nfrom ray.experimental import sandbox\r\n\r\nray.init()\r\n# Create a gVisor sandbox environment and return an actor handle for a proxy actor\r\nsb = sandbox.create(\r\n cpu=1.0,\r\n memory="512Mi",\r\n image="python:3.12-slim"\r\n)\r\n# Execute code inside the sandbox\r\nresult = ray.get(sb.exec.remote("python -c \'import sys; print(sys.version)\'"))\r\nprint(result.stdout)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f094035b130>)])]>

This creates a gVisor sandbox from an OCI-compatible image and returns a Ray Actor handle. Calls to exec are normal Ray Actor calls, so the sandbox can live anywhere in the cluster. The created actor is a proxy that will forward the operations to gVisor.

The sandbox API covers the basic lifecycle needed by agentic workloads:

  • Create environments from OCI container images

  • Set CPU and memory limits

  • Configure environment variables, working directories, and networking

  • Execute commands

  • Read, write, upload, and download files

  • Inspect sandbox state

  • Terminate or delete environments.

For lower-level use cases, SandboxRuntime provides direct access to local gVisor sandboxes and lets users modify the OCI specification before it is handed to gVisor. Here is an example how this API can be used to build a pool of local sandboxes inside of an actor:

code_block
<ListValue: [StructValue([('code', 'import ray\r\nfrom ray.experimental.sandbox.runtime import SandboxRuntime\r\n\r\n@ray.remote\r\nclass SandboxPool:\r\n def __init__(self, size: int = 3, image: str = "python:3.10-slim"):\r\n self.runtime = SandboxRuntime()\r\n self.sandboxes = [\r\n self.runtime.create(image=image, memory="512Mi")\r\n for _ in range(size)\r\n ]\r\n\r\n def run_command(self, index: int, command: str):\r\n return self.runtime.exec(self.sandboxes[index], command)\r\n\r\n def close(self):\r\n for sb_id in self.sandboxes:\r\n self.runtime.delete(sb_id)\r\n\r\n# Deploy an actor managing a pool of local sandboxes\r\npool = SandboxPool.remote(size=3)\r\nresult = ray.get(pool.run_command.remote(0, "python3 -c \'print(\\"Hello from pool!\\")\'"))\r\nprint(result.stdout)\r\nray.get(pool.close.remote())'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f094035b1c0>)])]>

Why gVisor?

Running model-generated code means treating the code inside the environment as untrusted. Ray Sandboxing uses gVisor, Google's open-source application kernel, as its initial sandbox runtime. gVisor implements a substantial portion of the Linux system-call interface in userspace, putting an additional isolation boundary between workloads and the host kernel. It is OCI-compatible, works with standard container images, and does not require exposing a Docker daemon or host Docker socket to the sandbox.

This combination is particularly useful for agentic workloads: environments remain lightweight enough to create dynamically while providing stronger isolation than executing generated code directly in ordinary containers. gVisor also provides sub-second sandbox startup and low per-sandbox memory overhead, making it possible to use sandboxes as relatively fine-grained distributed resources.

In future versions of Ray, we plan to extend support to other sandboxing runtimes such as Agent Substrate or Kata Containers.

Try Ray sandboxing on GKE

Check out the Ray documentation to learn more about Ray Sandboxes. To try out these sandboxing capabilities on GKE, head over to the Ray sandboxing User Guide. Have feedback or ideas? Join the discussion on the GitHub issue to collaborate on the future of Ray for reinforcement learning.

Empowering autonomous agents with advanced security governance

24 août 2026 à 18:00

AI agents are the ultimate insiders. We grant them permission to read emails, query databases, and trigger API calls. They don’t just retrieve information, they take action. 

Agents offer incredible potential for increased productivity and better customer experiences, but they also come with new security concerns. In our new State of AI infrastructure report, 79% of tech leaders cite security, governance, or operations as their most significant challenge to scaling inference.

While there’s still a crucial role for traditional security tools, the threat model has fundamentally changed. Autonomous workflows have redefined enterprise risk, so it's crucial that we give agents the access they need without compromising security.

Blog 4_Infographic 1
Blog 4_Infographic 2

The agentic paradox

The path to success starts with viewing governance as a driver for innovation. To be useful and secure, an agent needs access — and also guardrails. Yet 35% of senior IT decision makers cite insufficient security for multi-system access as a primary issue preventing agentic deployment.

Agents expand the surface area that defenders need to protect, and can introduce new threats, including tool poisoning and indirect prompt injection, where an attacker can hijack an agent’s logic through the data it processes. Managing the dynamic permissions that agents need to succeed at their tasks can also be a significant challenge, particularly as legacy security wasn’t designed for today’s automated threats.

Blog 4_Infographic 3

Securing the chain of thought

Along with securing more identity and access issues, it’s important for defenders to secure both the network layer and the model.

Security leaders are increasingly shifting their focus from preventing breaches to verifying provenance to guard against misuse, including indirect prompt injection.

From an infrastructure perspective, what are your top security concerns related to AI?

From blocking to managing

We’ve looked at the new security challenges posed by agentic AI. You can’t solve them by simply locking down the system, as that defeats the purpose of autonomous agents.

Many organizations are turning to integrated, full-stack cloud platforms to give them greater oversight. 69% of surveyed executives now rate a full-stack platform as a critical requirement, and 80% say data compliance is the primary factor dictating that choice.

By adopting frameworks like the Secure AI Framework (SAIF) and moving to a central control plane, purpose-built platforms such as Gemini Enterprise Agent Platform, organizations can manage risk in three main areas:

  • Secure-by-default design: Embedding security directly into the AI development process to proactively guard against threats including prompt injection.

  • Agent governance and oversight: Adopting purpose-built permission and identity management for agents — giving greater control over agent interactions, exposing blind spots and limiting risks tools.

  • Human-in-the-loop control: Enforcing clear rules that automatically flag when an agent requires human approval before moving forward with a critical action.

Governance will guide you to success

The true value of a modern security foundation is its ability to encourage innovation. By embedding robust governance directly into a unified foundation, organizations can deploy agents with confidence across their most sensitive, business-critical workloads. 

The leaders of the agentic era are re-architecting their stack to use security as a launchpad — empowering them to innovate securely and scale faster than their competition.

Find out more about how enterprise leaders are rethinking security for the agentic era in the State of AI infrastructure report.

Expanding connectivity in the Americas: Introducing Alisios, Canoa, and OlaLuz subsea cables

11 août 2026 à 12:00

The global digital economy relies on robust, resilient, and highly secure infrastructure to support everything from daily internet use to telehealth and scientific research breakthroughs. Today, Google is announcing a major expansion of our global network infrastructure in the Americas with three new subsea cable systems — Alisios, Canoa, and OlaLuz. These systems, alongside a new branch for the Firmina cable and previous investments in Curie, Nuvem, and Sol, form Americas Connect.

The Alisios subsea cable system will connect the Dominican Republic, Panama, and Chile. The cable is named after the vientos alisios — the Spanish term for the trade winds that blow across the tropics and the Caribbean, historically used by sailors to navigate between continents. The name "Alisios" nods to the flows of data that will navigate these new subsea routes to connect communities.

Furthering regional connectivity, Canoa is a new subsea cable system connecting the Dominican Republic to Bermuda. Canoa comes from the Taíno word for canoe, harkening to the longstanding maritime culture in the Caribbean and North Atlantic. When combined with the Nuvem and Sol cable systems, the Canoa cable will enable a highly reliable and resilient connection from Latin America and the Caribbean directly to Europe and North America. 

The OlaLuz subsea cable system will connect the Dominican Republic directly to Florida. The name "OlaLuz" combines the Spanish words for "wave" (ola) and "light" (luz), referencing the waves of light that carry data along the cable on the ocean floor. This new route significantly strengthens network capacity between the Caribbean and the U.S. East Coast, providing a pathway that enhances resilience for the region.

To further strengthen network diversity, Google is also introducing a new branch of the Firmina subsea cable that will land directly in the Dominican Republic. Originally built to connect North and South America, this new branch extends Firmina’s high-capacity reach and brings additional network connectivity to the Caribbean.

Reach, reliability and resilience

With its diverse subsea cable routes to geographically separated cloud regions, Americas Connect improves the reach, reliability and resilience of connectivity infrastructure. 

Pacific Coast: In the Pacific, a subsea ring will create redundant paths connecting Chile to Panama with a third link from Panama to the U.S. West Coast, providing connectivity to Google Cloud regions in Chile, Los Angeles, and Las Vegas.

1 Americas

Caribbean Sea: In the Caribbean, the Alisios cable from Panama to the Dominican Republic will interlink with a ring from the Dominican Republic to the U.S. East Coast and Google Cloud regions in South Carolina and Virginia.

2 Americas

Atlantic Ocean: The Canoa subsea cable from the Dominican Republic to Bermuda will interlink with the Nuvem and Sol cables, while cable rings to both the United States and Europe provide connectivity to the Google Cloud regions on the U.S. East Coast (South Carolina and Virginia) and Europe (Madrid).

3 Americas

“As we welcome the announcement of the Alisios, OlaLuz, and Canoa subsea cable systems, the Dominican Republic is pleased to continue partnering with Google to bolster digital connectivity and innovation as part of Americas Connect.  Our participation in the Caribbean and Latin America telecommunications network allows us to join our partners in strengthening connectivity across the Americas, serving as an integral node in this expanding digital ecosystem. Our shared vision supports a leap forward in the DR government’s quest to bridge the digital divide, foster local talent, and drive tech-based economic opportunity for our people. This milestone is yet another piece of our collaboration with Google propelling the Dominican Republic into a new era of global digital integration.” - Luis Rodolfo Abinader Corona, President of the Dominican Republic

“We are delighted to welcome this new Google project and take pride in the fact that Panama is considered a strategic location for this subsea infrastructure. The announcement of the new subsea cables marks a transformative milestone and aligns with Panama’s Digital Hub Initiative and the digital infrastructure pillar of the National Digital Strategy. We recognize that building a more prosperous future requires this kind of transformative infrastructure. It fosters opportunities for both our country and the region, propelling us into a new era of secure and competitive global integration. By strengthening regional connectivity, we accelerate our efforts to bridge the digital divide, cultivate local tech talent, and drive high-tech economic growth. Building on this momentum, we look forward to continuing our collaboration with Google on other strategic initiatives to ensure that digital transformation benefits the citizens of Panama.” - José Raúl Mulino, President of Panama

“We are pleased to welcome Canoa as the first subsea cable system to connect Bermuda directly to the Dominican Republic. It follows the Nuvem and Sol landings, and the expansion of the Bermuda Digital Exchange Port announced in June 2026. The name carries a maritime heritage that moved people and goods across these waters long before any cable did. This connection creates new digital pathways and partnerships between Bermuda and the Caribbean Community, while strengthening resilience for Bermudian households and businesses. We are encouraged that global infrastructure leaders such as Google continue to choose Bermuda. This investment reflects growing confidence in our island as a trusted digital gateway and strategic landing point in the Atlantic.” - The Hon. Alexa N.H. Lightbourne, JP, MP, Minister of Home Affairs, Government of Bermuda

“The new Alisios submarine cable, which will connect Chile, Panama, and the Dominican Republic, will be a strategic milestone, as it will strengthen our country’s position as a digital leader in Latin America. Beyond this technological achievement, this infrastructure will translate into direct benefits for our citizens. Providing faster, more resilient, and more secure connectivity would undoubtedly boost local businesses, foster economic growth, and bring the opportunities of the global digital economy directly into the homes of millions of Chileans.” - Louis de Grange, Minister of Transportation and Telecommunications and Minister of Public Works, Government of Chile

With these new investments, Google continues to pave the way for a more connected and open global cloud network centered on increasing reach, reliability, and resilience for all. New systems can further facilitate future connectivity through strategically placed branching units, mirroring the approach of our Pacific Connect initiative to ensure the network and our partners can grow alongside the needs and opportunities of the region.

Digital sovereignty in the age of AI: You don’t have to choose between control and innovation

6 août 2026 à 18:00

For enterprises and governments with strict compliance and sovereignty requirements, keeping sensitive data on-premises often means missing out on the latest AI. These organizations are managing three major risks:

  1. Jurisdictional risk: Shifting local regulations, the need to protect intellectual property and the potential of foreign data access requests make local data handling essential.

  2. Economic independence: Reliance on foreign infrastructure providers could leave critical services vulnerable.

  3. Geopolitical risk: A need to safeguard critical local services against unpredictable global disruptions.

In a recent survey of over 1,400 senior IT leaders for our State of AI Infrastructure report, 48% of leaders stated they are prioritizing infrastructure with data residency, controls, supporting compliance, with local data security laws.

image1

However, staying on-premises no longer means being cut off from the latest innovation. Organizations are increasingly deploying hybrid (on-premises and multicloud solutions) to bridge this gap. Our research shows that 52% of organizations now have a hybrid cloud approach to AI.

This approach allows enterprises to balance the massive raw power of the public cloud with the sovereignty and compliance benefits of local environments — allowing them to control where their data resides and who has access to it. In the past, organizations with such strict data rules couldn't easily access advanced AI. Building their own AI systems was also too slow and costly.

That is why we introduced Google Distributed Cloud (GDC). GDC brings Google Cloud to wherever you need it — in your own data center or at the edge. It is offered in two deployment models to meet your AI workload sovereignty requirements:

  • Air-gapped: A fully disconnected solution that does not require connectivity to Google Cloud or the public internet. It cannot be remotely shut down by Google.
  • Connected: An integrated, Google-managed software lifecycle that runs directly on your existing hardware.

GDC offers a complete, on-premises AI solution with infrastructure optimized for AI workloads, a choice of Gemini or open models, and cost-effective inference services. This foundation empowers you to build and run secure AI agents while maintaining total control over your data.

Meet your sovereign AI needs on-premises

You no longer have to choose between data control and AI innovation. With Google Distributed Cloud, we bring the world's leading AI directly into your environment — keeping your data entirely yours.

Explore the hybrid strategies of leading enterprises in the State of AI infrastructure report.

Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applications

6 août 2026 à 15:00

Nearly every major AI lab uses Google Cloud infrastructure, including for training of models, inference for agents, and new frontier research. Google Cloud also continues to be the platform of choice for new, high-growth AI startups who are driving much of the industry’s research and innovation.

Today, we’re announcing that Mirendil, an exciting frontier AI lab focused on accelerating AI development, will also utilize Google Cloud’s AI Hypercomputer. This includes using a mix of Google’s TPU AI accelerators and full-stack NVIDIA AI infrastructure running on Google Cloud; this purpose-built AI infrastructure will support model pre-training and post-training applications for Mirendil. 

The Mirendil team is building new AI systems that can help accelerate and democratize AI research and development. This means managing complex, end-to-end training workflows from initial model pre-training through post-training, and powering reinforcement learning on a massive scale. The ability to choose a mix of both TPU and NVIDIA’s full-stack accelerated computing platform through Google Cloud meant that Mirendil could access critical compute very quickly, and continue to match its workloads to the architecture best-suited to it over time.

We closely partnered with Mirendil on end-to-end design and deployment of combined TPU and NVIDIA AI infrastructure across compute, storage, networking, and control planes. We also collaborated on a system that uses managed training clusters running in Gemini Enterprise Agent Platform, which effectively streamlines the provisioning and management of both TPU and GPU environments for Mirendil. Mirendil is already live with a cluster of TPU v5P chips, with NVIDIA AI accelerated computing systems coming online soon.

"Progress in AI has been bounded by how fast humans can run the research loop - designing experiments, evaluating results, and iterating," said Behnam Neyshabur, cofounder and CEO of Mirendil. "We're building AI systems that can accelerate and improve that loop itself. Expanding on Google Cloud gives us the scale and flexibility to push those systems further and put frontier AI research capabilities in the hands of many more scientists and engineers to run that loop faster and at a greater scale."

You can read more about our partnership on Mirendil’s blog.

What’s new in AI infrastructure and orchestration in August

31 août 2026 à 18:00

Welcome back to What’s new in AI infrastructure and orchestration this month, a collection of product updates, how-tos, customer stories, research and other resources about all the AI compute, networks, storage, frameworks, and orchestration software that you can find at Google Cloud. To be honest, we thought August would be a slow month, but nothing could be further from the truth. Read on and you’ll see what we mean.

August 2026

Product, technology, and tools updates

  • Product update: Filestore, Google Cloud’s first-party, secure, scalable NFS file service, has emerged as a popular storage platform for AI and agentic workflows, and now, it’s even better suited to the task, with a new backend storage layer built directly on Colossus, Google’s foundational distributed storage system. This new backend lets you provision IOPS independently from storage capacity, and is deeply integrated with GKE. In AI environments, this can help you service so-called agentic swarms — large groups of agents that need to read and write to a common dataset — without a drop off in performance. For more, check out the blog post. 

  • New feature: gVisor sandboxes are now available in distributed Ray clusters on GKE. In partnership with Anyscale, we introduced an experimental library for Ray that brings gVisor, Google’s open-source application kernel, directly into distributed Ray clusters. gVisor provides lightweight environments with stronger isolation than ordinary containers, plus fast startup times and low memory overhead. To try out these sandboxing capabilities on GKE, head over to the Ray sandboxing User Guide.

  • Product update: Looking for high-performance, easy-to-use infrastructure on which to run a personal AI agent, but don’t want to spend a lot of money? New Cloud Run instances are dedicated, singleton compute runtimes on Cloud Run that won’t shut down when the agent is idle. Better yet, the cost to run a Cloud Run instance with 1 vCPU and 1 GiB of memory continuously for 30 days is just $5.70.  

Practitioner guides, documentation and how-tos

  • How-to guide: Big news in Model Context Protocol (MCP) land: As of the 2026-07-28 specification, the protocol core is “completely stateless. The handshake is gone. The initialize / initialized handshake (SEP-2575) and the logical Mcp-Session-Id header (SEP-2567) have been removed entirely. Instead, every request is now self-describing and independent.” Whoa. Learn more about the changes that the latest MCP specification brings, and more importantly, how to implement them, in this Google Developers blog.  
  • Guide: Real-time AI systems make a mess of traditional network load balancing techniques. “Instead of handling isolated requests, the backend has to manage a continuous, live bidirectional stream. You’re dealing with a constant stream of audio chunks, transcripts, model outputs, and synthesized speech flowing back and forth simultaneously.” Things only get worse when the user gets involved. “The server has to immediately halt its current speech generation, pivot to update the context, maybe trigger a new tool, and start drafting a different response; this must be done without dropping the connection.” For a new approach to managing load in the AI era, read Scaling real-time AI agents with session-aware load balancing.
  • How-to: Learn how to build an elastic, scalable LLM inference platform on GKE, even with a mix of different GPU accelerators. The proposed architecture combines Capacity Advisor and Compute Advisor, plus high-performance storage like RunAI:model streamer or GCPFuse with parallel downloads. Get all the details here.
  • Documentation: The thing about hosts with GPUs or TPUs is that you can’t use live migration to update them, setting up a maintenance challenge. In this new docs page, learn how to update accelerator-equipped hosts according to your tolerance for downtime for your training and inference workloads.    
  • Documentation: Advanced Compute Images, or ACIs, are standardized image stacks for AI/ML and HPC infrastructure, so you don’t need to manually build your own custom images. In this new docs page, learn how to create an ACI image using the Google Cloud CLI, console, or SchedMD's Slurm workload manager. 
  • Guide: AI workloads are notoriously difficult to architect, resource-intensive, and bursty, which can also lead to scaling bottlenecks and large pools of underutilized — or misutilized — compute resources. A new blog outlines the three main ways to achieve dynamic capacity management in Google Cloud: 1) scheduling capacity for planned downtime; 2) maintaining automated fallback capacity for unplanned downtime; and 3) relying on GKE’s core orchestration capabilities to automate resource allocation. 

Customer and partner updates

  • Business orchestration software provider UiPath was dealing with spiky workloads, and wanted more predictable costs. To get there, it re-architected its infrastructure, moving from isolated clusters to a shared Google Cloud GPU fleet that included both A3 VM instances (NVIDIA H100 GPUs) for training with G4 VM instances (NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs) for inference. You can read more about their architecture here. 

  • Mirendil, an frontier AI lab focused on accelerating AI development, announced that it is using AI Hypercomputer with both TPUs and NVIDIA GPUs to support its model pre-training and post-training applications. 

  • Replenit, a retail CRM provider, built its AI decision engine in Google Cloud, using BigQuery, Gemini Enterprise Agent Platform, and open-source Gemma models that it runs on Cloud TPUs. This latter combination provided Replenit with 90% lower pipeline costs than their previous cloud provider, the company reports. Read the full case study for more. 

  • Malachyte architected its AI-powered e-commerce recommendation platform on top of Bigtable, Managed Service for Apache Kafka, Pub/Sub, Compute Engine, and last but not least, GKE. See how it all comes together in this blog.


July 2026

Product, technology, and tools updates

  • Product update: Google Cloud Managed Lustre is now GA, and available in four distinct performance tiers that deliver throughput ranging from 125 MB/s, 250 MB/s, 500 MB/s, to 1000 MB/s per TiB of capacity — with the ability to scale up to 8 PB of storage capacity. The Managed Lustre solution is powered by DDN’s EXAScaler, combining DDN's decades of leadership in high-performance storage with Google Cloud's expertise in cloud infrastructure.

  • Product update: C4N network and storage optimized VMs are now GA. C4N is our first network- and block-storage-optimized VM series built to eliminate data-transfer bottlenecks. Powered by 5th Gen Intel Xeon Scalable processors and built on Google's Titanium offloading hardware, it achieves 400 Gbps network bandwidth, 95 million packets per second (MPPS), and up to 25 GiB/s of block storage throughput when paired with Hyperdisk Extreme.

  • New feature: GKE Dataplane V2 up to 15K Nodes with Network Policies (GA). This capability enables standard GKE clusters to scale up to 15,000 nodes while maintaining full active Network Policy enforcement, supporting the massive infrastructure needs of large enterprise and AI/ML customers.

  • New feature: Co-operative time-slicing in llm-d. If you’re running reinforcement learning (RL) workloads, you can now interleave independent RL jobs onto shared physical hardware, increasing aggregate accelerator duty cycles from a ~40% baseline up to 70% without impacting model convergence or accuracy. 

  • New AI security tool: Looking to secure your AI supply chain on GKE, deploy AI workloads safely, and cut down on shadow AI? We open-sourced k8s-aibom, a lightweight, unprivileged Kubernetes controller that continuously monitors container clusters to automatically detect running AI runtimes (like vLLM and Triton) and generate standard CycloneDX Machine Learning Bill of Materials (ML-BOMs). Check out the k8s-aibom project and get involved.

Practitioner guides and how-tos

  • How-to guide: On July 27, Google announced Day 0 support for Moonshot AI’s Kimi K3 2.8-trillion-parameter open-weight model, the day weights were released. Whichever your preferred deployment path — via Model Garden, custom orchestration, or GKE with llm-d recipes — this guide offers detailed step-by-step instructions to help you evaluate and pilot Kimi K3 in Google Cloud. 

  • How-to guide: Google Kubernetes Engine (GKE) managed DRANET supports both GPUs and TPUs. There are several configurations to use this implementation, including standard cluster (where you have full control) and autopilot cluster (where Google does the heavy configs for you). Take a deeper dive in the hands-on lab, GKE Autopilot clusters with TPUs, GKE managed DRANET and Gemma 4.

  • How-to guide: Learn to run Ray on TPUs, not GPUs. In Part 1 of this two-part series, we discuss TPU slices (hint: Ray thinks of them as just another accelerator on which to schedule), then walk through Ray’s various AI libraries (Part 2).

  • How-to guide: Evaluate TPUs for sample workloads using a new microbenchmark suite that helps you accurately assess whether a device is achieving its theoretical performance specifications, and to identify specific performance gaps or architecture-specific bottlenecks. Dive in here. 

  • How-to guide: Scale your agents without killing your budget. Learn how GKE orchestration can help you safely pack more agents onto a fixed compute footprint with GKE Agent Sandbox and Pod snapshots. Whether your goal is performance or cost optimization, we teach you how to turn the right dials for optimal agent efficiency. 

  • Technical blueprint: Inside the optimization of Mistral 3 large inference on Ironwood. This blog outlines how one Google team optimized Mistral 3 large MoE model inference on Google’s Ironwood (TPU v7x), achieving a 1.5x performance gain. They did so with hybrid sharding, replacing linear VPU summations with tree reductions, optimizing GMM/MLA kernels, and adopting asynchronous scheduling. As a result, they boosted throughput by up to 48% while maintaining benchmark accuracy neutrality. Read the full blog here.

Research, reports and deep-dives

  • Report: Google was named a Leader in the inaugural GartnerⓇ Magic Quadrant™ for AI Infrastructure, positioned highest for ‘Ability to Execute’ and furthest for ‘Completeness of Vision’. Gartner called out Google’s proprietary scalable compute, integrated AI Hypercomputer architecture, and the scale of our AI compute capacity as key strengths. Download a copy here.

  • Report: We recently surveyed more than 1,400 senior IT leaders for our State of AI Infrastructure report, and a resounding pattern emerged: The gap between AI ambition and infrastructure reality is widening. In fact, 83% of organizations say they require infrastructure upgrades to support production-grade agentic AI. Read the accompanying blog to understand how adapting your infrastructure to meet the demands that agentic applications place on your systems will help you move from pilot to production.


June 2026

Product, technology and tool updates

Practitioner guides and how-tos

  • How-to guide: Learn how to build high availability into an AI inference workload running on GKE Inference Gateway with TPUs, Cloud Storage FUSE and Dynamic Resource Allocation (DRA). This blog provides an overview, or you can get all the technical details in the hands-on codelab.

  • How-to guide: Did you know you can connect your AI agents to unstructured data in Cloud Storage via Model Context Protocol (MCP)? In this blog, learn about why would want to do that from three customer examples, then how to do it, choosing either a fully managed service, or a self-managed local server for more customization and control. 

Research, reports and deep-dives

  • Report: According to an independent benchmark report, GKE Inference Gateway outperforms the next leading managed Kubernetes service with 15.7% higher throughput, 92.8% shorter wait times, and 62.6% lower inter-token latency. This performance can be attributed to its use of prefix caching, which optimizes LLM performance by storing the KV cache (activation states) of long, repetitive prompt prefixes. Learn more in the blog. 

  • Architecture deep dive: A closer look at the cold start problem, this time for TPUs and GKE, and how the Run:ai Model Streamer can help change the dynamic. 

Customer and partner updates


May 2026

Product, technology and tool updates

  • Product update: GKE Agent Sandbox is now generally available.

  • New open-source project: Agent Substrate is a new open-source project aimed at continuing to push the limits of agentic infrastructure density

  • New feature: Google AI Edge Portal, a solution for testing and benchmarking on-device machine learning (ML) at scale, now supports benchmarking and debugging on-device LLMs. Read more here. 

  • Product deep dive: We went into depth about Cloud Storage Rapid, a new family of high-performance storage offerings for AI workloads. At launch, offerings include Rapid Bucket (formerly Rapid Storage), a high-performance zonal object storage offering, and Rapid Cache (formerly Anywhere Cache), which accelerates reads on-demand and colocates compute and data for workloads in existing buckets. 

Research, reports and deep dives

Customer and partner updates

Minimize idle accelerators: Native RL job interleaving with co-operative time-slicing in llm-d

23 juillet 2026 à 19:00

The math behind reinforcement learning (RL) post-training for large language models (LLMs) is notoriously unforgiving. As frontier AI labs push the boundaries of reasoning and coding models using RL post-training algorithms like Group Relative Policy Optimization (GRPO), they routinely hit hard architectural and infrastructure constraints. While much of the industry's focus remains on acquiring raw accelerator capacity, infrastructure efficiency is equally critical for achieving the high velocity needed to run multiple RL jobs and drive models to higher levels of intelligence. At scale, distributed RL suffers from severe resource bottlenecks because synchronous sampling and training run as strictly sequential phases, causing trainer and sampler resources to alternate sitting idle. Meanwhile, asynchronous architectures attempt to overlap these phases, but trainers still experience frequent idle gaps while waiting for specific trajectory batches to finish before starting the next cycle. 

Today, we are introducing a solution to this structural waste: co-operative time-slicing through the llm-d project. By treating discrete RL steps — such as sampling rollouts and gradient training — as dynamic, schedulable entities, we can interleave independent RL jobs onto shared physical hardware. Our initial benchmarks show that this platform-level multiplexing increases aggregate accelerator duty cycles from a ~40% baseline up to 70% without impacting model convergence or accuracy. This improves price-performance and lowers TCO significantly by eliminating wasted compute accrued over time.

For synchronous setups, the platform interleaves both samplers and trainers to minimize alternating idle windows, while asynchronous workloads leverage time-slicing to dynamically reclaim and utilize the fragmented idle gaps between RL-trainer iterations.

image_1

Throughout this blog, we will describe the time-slicing solution, detailing the technical flows, current release and future roadmap. 

llm-d for RL infrastructure efficiency (the bigger picture)

From the get-go, we anticipated the severe infrastructure bottlenecks of large-scale RL post-training and invested in addressing infrastructure inefficiency for RL workloads. 

We have built llm-d into a highly composable infrastructure stack for inference, agentic and RL workloads focused on eliminating accelerator idle time. The llm-d stack for RL features:

  1. Throughput-driven inference (llm-d-router): A mature, production-tested engine deployed across RL workloads and focused on maximizing rollout generation throughput to continuously saturate the pipeline.

  2. High-velocity Agent Sandbox (recipe): Tested for scale and density, and helping deliver secure, sub-second tool-use and isolated code execution during rollout generation and evals. Agent Sandbox serves as the high-speed intake manifold for reward signal generation, helping ensure the Sandbox never becomes the latency bottleneck that starves your time-sliced NVIDIA GPUs.

  3. Core pipeline primitives: To combat reliability and speed in weight transfer, we are building Weight Propagation Interface (WPI), as well as focusing on improving overall observability and reliability for RL. 

The efficiency problem with RL loops

Distributed RL post-training operates as a fragmented, continuous cycle alternating between generation (sampling rollouts) and optimization (gradient updates). Because traditional cloud infrastructure is designed for continuous, steady-state workloads, standard Kubernetes clusters can’t adapt to this alternating cadence.

image_2

At scale, this structural cadence introduces two massive systemic inefficiencies:

  • Idle accelerators: Because these phases occur sequentially, GPU clusters sit completely idle (0% utilization) for 40% to 60% of their lifecycle. Trainers sit idle waiting for sampling rollouts to finish; samplers sit idle during gradient updates and weights distribution. This could represent millions in wasted capital annually.

  • Locked-in context: RL training and samplers hold their accelerator allocations for the entirety of their runtime even during idle phases because the NVIDIA CUDA context and all device memory needs to remain resident. Standard schedulers treat these pods as static, siloed allocations rather than aligning them to the alternating, phase-level states of the live RL loop, leaving valuable hardware locked up even during inactive phases.

Importantly, this is not just a synchronous RL problem. Asynchronous variants overlap generation and training, but they do not fully mitigate idle time. Generation remains the inherent bottleneck of the RL loop, meaning trainer accelerators still starve while waiting for rollout data to accumulate. The closer an asynchronous job runs to on-policy, the larger those idle windows become — bounded staleness limits how far generation and training can drift apart, stalling the pipeline whenever fresh rollouts are not ready. 

How co-operative time-slicing (RL job interleaving) helps

To eliminate idle accelerators during RL jobs, co-operative time-slicing under the llm-d project allows the infrastructure to dynamically interleave independent RL jobs onto shared hardware blocks rather than forcing hardware to wait on upstream phases. This helps drive aggregate accelerator utilization up without altering the underlying model convergence or accuracy.

When Job A goes idle at a phase boundary in synchronous RL (or stalls on fresh rollout data in asynchronous RL), the infrastructure time-slices the physical accelerators, swapping in the active sampling or training phase of Job B. Under the hood, a swap is a checkpoint/restore: Job A's entire device state is checkpointed out of accelerator memory into host DRAM, and Job B's previously saved state is restored in its place. Because only one job's state ever occupies the accelerator at a time, steps alternate safely without framework-level interference or out-of-memory (OOM) faults.

image_3

Time-slicing: High-level architecture 

The time-slicing system architecture is organized into three layers: workload-scoped (application logic), cluster-scoped (coordination), and node-scoped (hardware management).

Workload-scoped layer (application runtime)
This is where the user's code runs — training loops, inference servers, and RL frameworks. The new addition is the time-slice client library, which exposes two gRPC APIs on the time-slice orchestrator: acquire() to request exclusive accelerator access, and yield() to release it. The user wraps any accelerator-touching phase with these calls to signal phase boundaries to the orchestrator. Everything else — the ML framework (PyTorch FSDP, vLLM, etc.), the CUDA context, the accelerator memory allocations — runs unmodified.

Cluster-scoped layer (control and orchestration plane)
This layer decides which job gets accelerator access, and when. Jobs that share the same physical accelerators — for example, two RL jobs interleaving on the same set of GPU nodes — are placed into a group. For each group, the time-slice orchestrator maintains a lock queue — an ordered list of jobs waiting for exclusive access to that group's accelerators. Only the job at the head of the queue holds the lock and runs on the hardware; all the other jobs wait, blocked on their acquire() call. When the running job calls yield(), the orchestrator passes the lock to the next job in the queue and triggers a coordinated context switch across every node in the group. In the future, a workload placement optimizer will be able to profile workload phase patterns and automatically pair jobs with complementary idle phases, removing the need for the user to explicitly indicate job groupings.

Node-scoped layer (hardware and data plane isolation)
This layer performs the checkpoint/restore swap on each accelerator node. The snapshot agent, a privileged DaemonSet, receives directives from the orchestrator and translates them into hardware-level operations — pausing accelerator processes, serializing device state to host DRAM, and restoring it when the job regains access. The agent is built around a pluggable backend interface, with cuda-checkpoint as the first implementation (more to come). Future backends will introduce faster snapshot mechanisms and more selective approaches, such as offloading specific memory addresses like LoRA adapters instead of full device state. The agent itself is designed to run standalone outside Kubernetes for bare metal and Slurm environments.

The flow: How it all comes together

image_4

When a workload finishes its current accelerator phase, its time-slice client library calls yield() to the time-slice orchestrator to release access. The orchestrator initiates the context switch by sending directives to the snapshot agent on each node in the group. The agent freezes the yielding workload's processes and moves its device state from accelerator memory into host DRAM.

With the accelerators vacated, the orchestrator grants the group lock to the next workload waiting in the queue. It directs the Snapshot Agents on those nodes to restore that workload's previously saved state from host DRAM back into accelerator memory, then unblocks the workload's pending acquire() call. The workload resumes execution exactly where it left off — no container restart, no framework reinitialization, no model reload from storage.

The yielding workload remains warm in host DRAM. When the orchestrator grants it the lock again, the Snapshot Agents perform the same swap in reverse.

Developer experience (client-side)
Researchers want to focus on core modeling logic rather than wrestling with low-level CUDA context switching or custom scheduling loops. If you use Ray or a similar platform to orchestrate your RL job, using time-slicing will have a minimal impact on the client side. In fact, there may not be any impact on the client side at all if you are queuing the training and sampling jobs separately at the platform level.

code_block
<ListValue: [StructValue([('code', 'from timeslice import TimeSliceOrchestratorClient\r\n\r\norchestrator = TimeSliceOrchestratorClient(target="orchestrator:50051")\r\n\r\n@orchestrator.on_accelerators(group_id="trainer-group")\r\ndef train_phase(model, trajectories):\r\n return model.update(trajectories)\r\n\r\n@orchestrator.on_accelerators(group_id="sampler-group")\r\ndef generate_phase(model, prompts):\r\n return model.generate(prompts)\r\n\r\n# Standard sequential loop — interleaved with other jobs under the hood\r\nfor epoch in range(EPOCHS):\r\n trajectories = generate_phase(policy, dataset)\r\n rewards = compute_rewards(trajectories)\r\n train_phase(policy, rewards)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fa062f02cd0>)])]>

Current release and future outlook

Today we are releasing the full time-slicing stack: the Snapshot Agent, the Accelerator Orchestrator, and the Python client libraries, each with a user guide for integrating time-slicing into your RL workloads. 

Key roadmap highlights include:

  • Latency and state optimization: Expanding the Snapshot Agent with faster checkpoint/restore backends to minimize context-switch overhead, alongside application-aware backends for selective memory region snapshotting (e.g., swapping LoRA adapters instead of full model weights).

  • Automated scheduling and onboarding: Introducing an automated scheduler to profile running processes, identify time-sliceable structures, and handle job placement dynamically. 

  • Cross-hardware compatibility: Extending data plane support beyond GPUs to TPUs and custom accelerator architectures.

Get started 

Building robust, highly optimized RL infrastructure requires tight collaboration with the engineers and researchers running these workloads at scale.

If you are currently wrestling with low GPU utilization, synchronization stalls, or complex scheduling logic in your post-training pipelines, time-slicing can help. To get started, check out the following resources, and don’t forget to leave us your feedback!

  • Start using time-slicing during your RL run immediately with these user guides.

  • Try llm-d-router (kubernetes native) or the RL Scheduler (python library) user-guide for improved sampling throughput during the RL generation phase.

  • Explore the Weight Propagation Interface repo. 

  • Join the discussion in the #sig-rl channel in the llm-d Slack.

  • Contribute by sharing your reference implementations, benchmarks, and edge cases to help us refine this path.


Thank you to Dolev Ish Am and Bogdan Berce for their contributions to this blog post.

Your AI agents are ready. Is your data?

23 juillet 2026 à 18:00

What’s one of the biggest bottlenecks stopping organizations from scaling their AI initiatives? It isn’t the capabilities of today’s models — it’s their access to business context and semantic meaning. 

In the agentic era, enterprises need to go beyond simply storing data to activating it with trusted context, moving from passive systems of record to proactive systems of action.

But AI agents operate with nonlinear speed; for example, a single prompt can trigger the agent to independently browse, query, and execute across multiple systems, placing stress on the underlying infrastructure. If the compute, networking, and storage layers aren't optimized for agentic AI, the data platform sitting on top of them will buckle.

It’s no wonder that, according to our State of infrastructure report, 83% of organizations believe they require infrastructure upgrades to support production-grade agentic AI systems.

2

To solve this problem, we introduced the Agentic Data Cloud at Google Cloud Next 2026; unifying your data, AI models, and operational databases into a single System of Action. To make an Agentic Data Cloud work, it must be AI-native from the chip to the model. The underlying infrastructure must be able to accommodate agentic load.

3

Google’s Agentic Data Cloud

4

Let’s explore how the right infrastructure foundation empowers an Agentic Data Cloud to solve the biggest data challenges organizations face today.

Overcoming a lack of context

To be effective, agentic systems require access to context that is often found in fragmented data systems and legacy architectures. This can make it hard for agents to get this context, leading to incomplete, inaccurate results. In fact, our report found that 43% of IT leaders cite “difficulty integrating with legacy APIs and data sources” as their biggest agentic AI infrastructure gap.

But organizations cannot simply move massive datasets and connect them to AI without increasing complexity and cost. 

Our Agentic Data Cloud solves this by leveraging a borderless Lakehouse running on open, flexible infrastructure. By accessing powerful native engines like BigQuery and Spanner over open standards (Apache Spark, Apache Iceberg), agents can read, reason over, and activate data across environments as if it were local, bypassing the latency and costs of traditional setups.

Escaping unnecessary manual work 

Scaling agents on a patchwork of disconnected systems can create significant bottlenecks. In our research, 81% of leaders called out operational complexity and engineering overhead as top unforeseen expenses when scaling AI, citing the time engineers spend doing manual work to patch together AI agents across disparate systems.

To move from thinking to doing, agents must be able to connect real-time data across both analytical and operational sources. This requires vertical integration. When an Agentic Data Cloud is built on an AI-native infrastructure where the models, data systems, and underlying accelerators are co-designed, there are fewer network hops and tooling is better integrated. This unified system allows an agent to reach an insight and trigger secure transactions without the typical engineering overhead.

Bringing trust and knowledge to the data

It’s not enough for agents to just discover and query data. To take safe, accurate actions, agents also need rich context and business logic. Yet, 36% of leaders cite a lack of specialized, high-throughput vector databases used for AI model grounding, as a key infrastructure gap, hindering their ability to give agents context.

In order to work to their full potential, agents need a foundation which is built to read and write data systems in real-time, including legacy ERPs and third-party CRMs. It also gives them the long-term memory to recall a user’s preference from, say, three weeks ago, while executing a complex task today. And without this real-time automation, agents have to re-process data for every single query.

To provide context for AI, organizations are using Knowledge Catalog to aggregate and enrich data in their data lakes, and enable agentic searches. By extracting meaning from unstructured data and automatically generating semantics, the catalog acts as an active reasoning layer. That catalog in turn, must be backed by high-throughput infrastructure, so that agents can retrieve the right context.

The path forward

To turn AI into a true competitive advantage, it’s time to build a connected, active data ecosystem. Giving your agents seamless access to all of your data is a must to move from pilots to production, and this must be supported by an infrastructure that can handle the demands of the agentic era. The winners in 2026 and beyond won’t necessarily be the ones with the smartest agents. They’ll be the ones who can feed those agents the right knowledge — securely, cost-effectively, and at scale. Is your data ready for the agentic era? 

See how leaders are taking an AI-optimized approach to architecture in the State of infrastructure in the agentic AI era report.

IDC: Why the right networking approach is foundational to agentic AI

15 juillet 2026 à 18:00

Editor’s note: Today we hear from IDC on the results of its 2026 AI in Networking Special Report Survey exploring the enterprises' concerns about networking infrastructure to support the rise of agentic AI in their organizations. The survey was sponsored by Google Cloud.


Enterprises are moving quickly on AI pilots, but the move from pilot to production remains uneven. While AI models remain important, IDC research indicates that the pilot-to-production bottleneck is primarily infrastructure-centric, with core networking concerns emerging as one of the leading drivers of AI project delays and abandonment. In IDC's 2026 AI in Networking Special Report Survey:

  • 32.6% of respondents cite security concerns: As AI workflows become more distributed and autonomous, enforcing consistent security and governance becomes more difficult.

  • 26.8% of respondents cite challenges in automation: Manual operations and fragmented controls can slow deployment and make AI environments harder to scale.

  • 24.7% of respondents cite staff time and talent restrictions: Limited skills and operational bandwidth can constrain an organization's ability to move AI initiatives into production. 

Agentic AI specifically heightens these concerns by introducing more distributed and dynamic interactions across applications, services, APIs, tools, and data sources. In production environments, these interactions often span different agent frameworks, model providers, clouds, open-source tools, SaaS APIs, and internal applications, expanding both the operational scope and the security and governance surface area. 

Networking for operational control, security, and governance at scale

Networking is the primary enabler of agentic interactions and plays a foundational role for intracloud and intercloud network- and services-layer connectivity, end-to-end security, and consistent governance. In agentic systems, networking increasingly extends into tighter service-centric controls that govern how distributed services identify one another, communicate, and exchange data securely. While AI workloads in general are increasing east-west traffic demands, agentic AI adds an additional layer of complexity by creating dynamic interactions that require tighter policy, visibility, and control closer to the application workflow.

From an infrastructure perspective, networking is much more than just a connectivity function. It is part of the infrastructure platform control plane that applies policy-based controls, supports observability, and helps maintain consistent security and governance across an AI agent's activity. This is significant because framework-level controls alone become insufficient in environments where agents and services span different runtimes, clouds, deployment models, and operating domains.

That is why an infrastructure-level approach becomes key. It does not replace application frameworks or orchestration environments, but it provides broader and more consistent policy implementation across a complex architectural landscape. As agentic AI becomes more autonomous and distributed, organizations need these controls built in as part of the infrastructure to reduce fragmented observability, inconsistent policy application, and unmanaged shadow agent activities. From a cloud infrastructure standpoint, this is where cloud network services become strategically important.

Balancing act: A platform vs. best-of-breed approach

Agentic AI systems are inherently fragmented because of underlying distributed workflows. Enterprises are already navigating a rapidly evolving landscape of business requirements, open-source components, emerging protocol standards, and new architecture patterns. In this context, choices between best-of-breed point solutions and platform-based approaches should be strategic rather than ideological.

Best-of-breed capabilities may be necessary to address specific technical requirements. But it is also true that point solutions introduced across a distributed agentic AI landscape can create inconsistent policies, operational complexity, and governance gaps. IDC research reflects this tension. In IDC’s 2026 AI in Networking Special Report Survey, organizations remained divided between platform and best-of-breed preferences for AI workloads; among respondents who favored platforms, the main reasons cited were stronger security (32.9%), reduced complexity (27.7%), and faster deployment (24.2%).

In IDC's view, a balance is important. Platforms can provide a consistent operational and policy foundation for AI deployments, but at the same time, they need to be modular and extensible to allow the inclusion of best-of-breed functionality as part of the platform toolset. The right platform for agentic AI should be open, flexible, and able to evolve. It should support integration with third-party and open-source tools, allow insertions of needed security and observability functions, and adapt without complete architectural rework.

This is a period of technology disruption. Businesses must meet their AI objectives while carefully managing dynamic agentic AI systems. In this environment, networking not only remains a connectivity piece of the AI infrastructure but becomes foundational to how organizations establish operational control, apply policy consistently, and maintain end-to-end trust across agentic workflows. 

As agentic AI systems continue to evolve, the demands they place are unlikely to be addressed through best-of-breed point solutions alone. Operationalizing agentic AI at scale will require organizations to leverage the right networking approach, supported by infrastructure platforms that are open, flexible, and extensible, enabling a cohesive and adaptable security and governance framework.

Message from the sponsor
The autonomous and non-deterministic communications of agentic applications pose challenges for which the infrastructure and governance models of the cloud-native era are not prepared. In the agent-native era, an infrastructure-led approach is required to enable agentic applications at scale in production with effective governance and observability. An extensible platform based on open standards is critical in enabling the agentic journey today and through its maturity. Learn about the infrastructure imperatives and open standards that make a viable agentic infrastructure here.

Claude at scale on Google Cloud: Frontier AI, built for enterprise production

14 juillet 2026 à 18:00

Running frontier AI in production is demanding — accelerators to manage, latency to hold steady across continents, regulated data to keep in-region, and long-context requests to serve reliably. Claude on Google Cloud is built for exactly this. 

Like Monet and water lilies, frontier models and the enterprise platforms are often better together. In our case, Claude brings the reasoning, and Google Cloud brings the managed infrastructure, global reach, and compliance posture that enterprises already run on. Calling Claude becomes operationally identical to calling any other Google Cloud service — same Identity and Access Management (IAM), same VPC Service controls, same observability — so teams are able to spend their time building features instead of running inference infrastructure.

This post walks through what Claude on Google Cloud delivers in production across four areas: 

  1. Managed infrastructure that gives engineers their time back 

  2. Global endpoints that hold latency low, and uptime high for a worldwide user base 

  3. Security and data-sovereignty controls inherited straight from Google Cloud

  4. Serving-layer features that keep cost and performance optimized at scale.

Managed infrastructure that frees engineering time

Claude on Google Cloud runs on fully managed infrastructure, so enterprise teams ship features instead of building clusters. Compute provisioning, auto-scaling logic, load balancing, and failover at frontier-model scale are handled by the platform — work that would otherwise occupy multiple teams full-time. 

Claude is available through Agent Platform's Model Garden as a Model-as-a-Service offering, ready to use over standard REST / JSON over HTTP/1.1 or HTTP/2 endpoints. Invoking Claude is operationally identical to invoking any other Google Cloud service: the same IAM policies, the same VPC controls, and the same observability stack via Cloud Logging and Cloud Monitoring. 

Serving Claude takes a few lines of Python using the AnthropicVertex client:

code_block
<ListValue: [StructValue([('code', 'from anthropic import AnthropicVertex\r\n\r\nclient = AnthropicVertex(\r\n project_id="your-project-id",\r\n region="us"\r\n)\r\n\r\nmessage = client.messages.create(\r\n model="claude-opus-4-8",\r\n max_tokens=1024,\r\n messages=[{"role": "user", "content": "Analyze this system architecture."}]\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fd94a6a1d30>)])]>

The same AnthropicVertex client handles prompt caching, tool use, structured outputs, streaming, and adaptive thinking; for batch inference, use Vertex AI Batch Prediction. Authentication uses Application Default Credentials; requests automatically inherit your project's IAM and VPC configuration. 

Global reach with consistent latency and built-in failover

Serving a worldwide user base from a single endpoint produces high tail latency and a single point of failure. Most enterprises can't replicate inference infrastructure across continents while keeping performance consistent.

Agent Platform exposes three endpoint types for Claude, each solving a different production requirement:

  1. Global endpoints route requests to a region with available AI compute capacity. For example, if us-central1 is capacity-constrained, traffic redirects to europe-west1 or another region with available capacity. That’s automatic failover and geographic load balancing without application-side routing logic. Global endpoints are ideal for maximum availability and lowest cost.

  2. Regional endpoints like us-east5 or europe-west1 keep prompts, completions, and intermediate state inside a specific geographical boundary, making it ideal for low latency and data-residency requirements.

  3. Multi-region endpoints give U.S. or EU data residency without single-region dependency. They dynamically route across regional endpoints  providing built-in resilience against regional outages and capacity constraints.

The diagram below shows how applications reach Claude through these endpoint types, and how the Agent Platform serving layer routes traffic to the Compute AI clusters across regions:

1 _GC_BlogGraphics_Anthropic

Serving Claude Models From Regional & Global Endpoints

2_GC_BlogGraphics_Anthropic

Serving Claude Models From Multi-Region Endpoints

3_GC_BlogGraphics_Anthropic

Serving Claude Models From Regional Endpoints

Enterprise security and data sovereignty built in

Regulated workloads — financial services, healthcare, and government — get enterprise-grade security and data sovereignty without trading compliance for convenience, and without re-engineering the hardest layer to control: inference, where prompts, completions, and intermediate state all flow through the serving stack.

Claude on Agent Platform inherits Google Cloud's full security posture. FedRAMP High and HIPAA compliance enable deployment in government, healthcare, and financial services environments. VPC Service Controls let organizations define a perimeter around Agent Platform resources, preventing data exfiltration. IAM-native access control governs Claude endpoints with the same roles and policies that protect every other Google Cloud resource — no separate API keys to manage or rotate. Cloud Logging and Cloud Monitoring provide near real-time visibility into token usage, error rates, latency, and quota consumption.

Combined with the regional and multi-region endpoints above, this gives regulated customers a path to running frontier AI in production without re-auditing their compliance posture.

Optimized for cost and performance at scale

In production, cost and performance drive every architectural decision. Getting both right requires capabilities from two layers: Claude's native model features, and Google Cloud's serving infrastructure. Agent Platform supports both, so teams can optimize across the stack without managing them separately.

Claude-native capabilities, fully supported on Agent Platform

These features are built into Claude and available on Agent Platform without any additional configuration:

  • Prompt caching stores and reuses shared prefixes — long system prompts, legal documents, codebases — reducing request latency by up to 80% and cost by up to 90%.

  • Streaming responses over server-sent events deliver tokens as they are generated, critical for chat interfaces and coding assistants where perceived latency matters.

  • Extended and adaptive thinking lets Claude dynamically determine when and how much to reason through complex, multi-step problems — and allows users to dial the thinking effort directly, for example to control cost. Optimized for use cases like advanced code generation, mathematical reasoning, and multi-document analysis.

  • Extended context windows up to 1M tokens (for Claude Opus 4.6,Sonnet 4.6 and newer models) enable long-document analysis, large codebase reasoning, and multi-turn conversations at depth.

Google Cloud serving infrastructure

Agent Platform adds its own serving-layer capabilities on top of Claude's native features:

  • Batch prediction handles large-scale offline workloads — document classification, content moderation, bulk summarization — asynchronously at lower priority and reduced cost.

  • Provisioned throughput reserves dedicated inference capacity for mission-critical workloads, isolating them from public traffic and ensuring predictable performance during peak demand.

  • Memory management and scheduling for long-context requests is handled at the infrastructure layer,.

Together, these two layers give teams the full range of optimization levers — from model-level efficiency to infrastructure-level capacity control — on a single, unified platform.

From inference to agents

The same infrastructure that serves Claude inference powers the agent layer of Agent Platform on Google Cloud. The build-and-register flow has three steps:

  1. Build with Claude. Claude is well-suited as an orchestration backbone — its extended context window, native tool use, and adaptive thinking make it effective at planning multi-step tasks and delegating to sub-agents. Pick Claude Opus, Sonnet, or Haiku from the Model Garden, then build with the Agent Development Kit (ADK) — code-first in Python, Go, Java, or TypeScript — deploy to Agent Runtime, Cloud Run or Google Kubernetes Engine.

  2. Deploy the Agent to a Runtime. Depending on your use case, select Agent Runtime, Google Kubernetes Engine or GKE Agent Sandbox to run your deployed agents.

  3. Interoperate over A2A. The Agent2Agent protocol runs at 150+ organizations, letting a registered Claude-powered agent delegate tasks to agents from SaaS and other service providers.

The result: a planning agent built on Claude can orchestrate sub-tasks across the broader agent ecosystem, under unified IAM, fully auditable, on the same infrastructure that serves the underlying inference.

Start building

Open the Agent Platform console, enable Claude in the Model Garden, and make your first API call with the AnthropicVertex SDK. Add prompt caching, provisioned throughput, and other features as your workload demands. When you're ready to go agentic, learn more about Claude on Agent Platform.

Reach out to your Google Cloud sales representative to discuss bringing Claude into your production environment at scale.

❌