❌

Vue lecture

Storage-optimized Z4D machine family, now GA, is designed for IO-intensive workloads

Today, we’re excited to announce the general availability of our next-generation Storage-optimized Z4D machine series in Google Compute Engine with both Virtual Machine (VM) and bare-metal instances. 

We built Z4D for IO-intensive and business-critical workloads that require large local storage capacity and high storage performance, including SQL, NoSQL, KVrocks and vector databases, data analytics and data search. Powered by 5th Gen AMD EPYC processors (Turin) and paired with the latest enhancements in Titanium, Z4D provides up to 84,000 GiB of Local SSD (LSSD) storage. It improves the performance of these demanding workloads by up to 40% compared to the prior-generation Z3 instances, so you can increase your applications throughput while right-sizing your cloud investment. Z4D’s large LSSD capacity also makes it a strong storage solution for AI/ML training and inference workloads and running distributed parallel file systems at scale.

Z4D VMs and bare-metal instances

The Z4D VM portfolio lets you rightsize your infrastructure and scale your clusters to meet workloads requirements by providing large total local SSD capacity and high local SSD capacity per vCPU. Z4D offers two different VM types: the Z4D-highmem-standardlssd VM type, which includes seven VM shapes and offers 219 GiB of LSSD per vCPU. These VMs are optimized for data analytics (OLAP), and SQL databases like MySQL and Postgres. The Z4D-highmem-highlssd VM type includes seven different VM shapes, with 438 GiB of LSSD per vCPU and is optimized for distributed databases, data streaming, large parallel file systems and data search. In addition, you can easily scale your existing Z3-based workloads by expanding into Z4D clusters.

Z4D bare metal instances give you direct access to the physical hardware without a virtualization layer. This reduces latency for latency-sensitive workloads, custom hypervisors and workloads with specific licensing needs. Further, Z4D bare metal will be the very first AMD-based instance to support Nutanix Cloud Clusters (NC2), a hybrid multi-cloud platform that works across several cloud providers. Z4D bare-metal instances deliver the LSSD capacity and low latency that agentic AI architectures using microVMs require, allowing developers to run thousands of isolated sandboxes per host with native performance and efficiency. 

In addition to 84,000 GiB of local SSD storage, Z4D VMs and bare metal instances offer up to 384 vCPUs and up to 3,072 GiB of memory. Z4D instances are based on Titanium SSDs, which offload local storage processing from CPU resources to deliver real-time data processing, low-latency, high-throughput storage performance and enhanced storage security. Z4D delivers up to 15,600K random read IOPS and up to 75,600 MiB/s sequential read throughput, improving the LSSD storage performance by up to 70% compared to Z3. Z4D also reduces write latency by up to 25% and improves mixed read-write IOPS by up to 30% without increasing the IO latency vs Z3. At the same time, Z4D instances provide the connectivity and storage performance that enterprise and AI/ML workloads need by doubling the networking throughput compared to Z3 and offering up to 400 Gbps of standard networking bandwidth.

What customers and partners are saying

elastic

“Elastic is committed to delivering best-in-class performance with the Elasticsearch Platform that powers observability, security, search, and AI solutions. Performance and cost efficiency are critical for teams running these workloads at scale and our initial testing of the new Z4D virtual machines shows up to 50% better indexing throughput compared to previous generation Z3 VMs. We look forward to bringing these benefits to Elasticsearch users deploying on Google Cloud.” - Yuvraj Gupta, Principal Product Manager, Elastic

immunai

"Migrating to the Z4D instance reduced our processing runtime by ~70% while reducing overall costs. This improvement enables Immunai to process large-scale immune data significantly faster and turn it into biological insights that support pharma companies in making better-informed decisions throughout drug discovery and development." - Guy Yachdav, Senior Director of Software Engineering, immumeai

nutanix

"We are thrilled to expand our technical collaboration with Google Cloud and to deepen our strategic partnership with AMD to bring Nutanix Cloud Clusters (NC2) to the new Z4D bare metal instances. This marks a significant milestone for our customers, as NC2 on Z4D will be a first-of-its-kind offering — the very first AMD metal instance on which NC2 is supported. By uniting AMD's cutting-edge compute performance with Google Cloud's robust infrastructure and the Nutanix hybrid cloud platform, we are delivering unprecedented flexibility, scale, and choice to empower enterprise workloads." - Saveen Pakala, Vice President of Product Management, Nutanix

redis

"Redis powers real-time data infrastructure and AI workloads for thousands of organizations worldwide. Flex extends that to larger datasets without all-RAM economics — local SSD handles the scale while memory provides the speed. We compared Google Cloud's new Z4D storage-optimized VMs to our current generation C3D instances using our Flex benchmark suite on the same flash-heavy workloads. The results were impressive: up to 4.3 times higher throughput and up to 77% lower latency on our smaller shapes. Z4D also sustained nearly 1.8 million operations per second at full RAM hit ratio. The more data we served from flash, the more that advantage grew — which is exactly the profile Flex is built for. Z4D gives us a clear path to deliver the same performance tier on a smaller footprint, and we're looking forward to expanding our testing as Z4D moves toward GA.” - Benjamin Renaud, CTO, Redis

shopify

“Shopify looks at what's best for our fleet, and Z4D gives us high CPU density alongside fast locally attached storage and fast networking. Shopify's user experience depends on how fast data can be served, and AI shopping agents query storefronts far more aggressively than people do. These shapes give us the throughput to keep up. We saw roughly 20% better throughput on Z4D than on Z3.” - Brad Dietrich, Distinguished Engineer, Shopify

silk

"Google's Z4D VMs are a significant leap forward. In our testing we observed up to 60% better performance than previous Gen 2 VMs, and up to 36% better performance than Z3 VMs. With high local SSD density and strong cost-efficiency together with improved system reliability, Z4D gives Silk customers faster, more predictable performance for their most demanding workloads." - Adik Sokolovski, Chief R&D Officer, Silk

turbopuffer

"Connecting AI with petabytes of fresh data means we're searching more than ever before. We're excited about the new Z4D instance types. Compared to Z3, they gave us a 40% throughput improvement on real query and indexing workloads — which directly translates to faster and cheaper web-scale search for our customers" - Ben Linsay, Engineer, turbopuffer

Enhanced maintenance experience

Both Z4D VMs and bare-metal instances make it easier for you to plan ahead and schedule maintenance operations at a time of your choosing by providing notice from the system several days in advance of a required maintenance. Z4D VMs further enhance the maintenance experience by allowing you to live-migrate an instance during maintenance events for VMs with 42,000 GiB or less of local SSD storage. Z4D VMs with 84,000 GiB of local SSD and Z4D bare metal instances are terminated and restarted while preserving your data through the planned maintenance events.

Support for Hyperdisk

Z4D VMs and bare metal support Hyperdisk, Google Cloud’s workload-optimized block storage that lets you optimize the performance for each workload by independently tuning the storage performance and capacity for each instance.

Specifically, they are compatible with Hyperdisk Balanced, Hyperdisk Throughput, and Extreme Hyperdisk storage for scalable, high-performance network-attached storage, supporting up to 512 TiB of capacity per instance. For general-purpose workloads, Hyperdisk Balanced, with up to 320K IOPS per instance, offers a mix of performance and cost-efficiency. Hyperdisk Extreme delivers ultra-low latency and supports up to 500K IOPS and 12,500 MiB/s throughput per Z4D VM and bare metal instance, making it well-suited for demanding database workloads. 

Get started with Z4D today

Z4D VMs are available today in select regions worldwide and  Z4D bare metal instances are in preview - reach out to your account team for additional information and access. To start using Z4D instances, select Z4D under the Storage-Optimized machine family when creating a new VM or GKE node pool in the Google Cloud console. Learn more at the Z4D machine series page. Contact your Google Cloud sales representative for more information on regional availability.

  •  

M4N VM family, now GA: Highest per-core IOPS and throughput for I/O and memory-bound workloads

As enterprise organizations scale mission-critical applications, storage I/O and memory access can become severe operational bottlenecks. Whether its Oracle databases, in-memory databases like SAP HANA, or high-throughput SQL Server clusters, EHR systems, and real-time big data analytics, memory-bound databases often force enterprises to over-provision compute cores (vCPUs) to get the RAM capacity and storage bandwidth they need, driving up costly third-party software licensing fees.

Today, we are thrilled to announce the general availability of the M4N machine series in Google Compute Engine, purpose-built for I/O intensive, high-memory workloads, the second offering in our network- and block-storage optimized VM family. Compared to similar offerings from other hyperscalers M4N provides the highest per-core IOPS and throughput for high-memory instances, and over 20% TCO reduction for Oracle databases.

image4

M4N is also the industry’s first instance of network and block storage optimized with higher memory ratios (up to 26:1) and size (6TB). Powered by 5th Gen Intel® Xeon® Scalable processors and built on Google Cloud's custom Titanium offload architecture, M4N instances deliver up to 25,000 MiB/s (25 GiB/s) of aggregate host storage performance and up to 1 million IOPS when paired with Hyperdisk Extreme — doubling the block storage performance of current M4 instances.

M4N targets workloads that demand both extreme high-density RAM and uncompromising I/O performance, complementing our existing memory-optimized families (such as M1, M2, M3, M4, and X4) by solving specific storage and network bottlenecks for high-throughput enterprise applications.

Built for demanding workloads

Workload Category

Typical Applications

Why M4N Wins

Mission-critical enterprise DBs

Oracle, SAP HANA, SQL Server, IBM DB2, MySQL, PostgreSQL

Memory-to-core ratios (up to 26.57 GB/vCPU) paired with 25 GiB/s storage for rapid data ingestion, transaction logging, and zero-stall backup cycles.

Generative AI and RAG data layers

Milvus, Pinecone, Qdrant, Vespa, Redis, In-Memory Context Caching

Sub-millisecond similarity search across massive vector indexes in RAM, combined with 400 Gbps network bandwidth for distributed model retrieval.

Enterprise healthcare and ERP

Epic Systems (Operational Database), SAP ECC, SAP S/4HANA

Sustained I/O headroom that prevents query latency spikes during peak clinical/transactional hours.

Real-time analytics and EDA

Electronic Design Automation, Genomic Modeling, In-Memory OLAP

High memory capacity to load massive datasets entirely in RAM with maximum storage bandwidth for checkpoint dumps.

Optimizing Oracle licensing costs

Enterprise IT departments struggle with the rising cost of core-based software licensing. For workloads like Oracle database, licensing fees are typically calculated based on the number of vCPUs or physical cores assigned to the instance. Historically, this has forced a difficult trade-off: paying for more compute cores than necessary just to obtain the required amount of RAM and storage performance.

M4N changes this paradigm with its industry-leading high memory-to-vCPU ratio. By providing the highest per-core IOPS and throughput for high-memory instances of all the leading hyperscalers, M4N allows database administrators to:

  • Reduce TCO and licensing overhead: Stop over-provisioning of cores while meeting Oracle database performance density requirements, resulting in over 20% TCO reduction compared to similar offerings from leading hyperscalers.

  • Right-size infrastructure: Allocate the exact amount of compute power needed for the workload while still accessing massive memory pools.

  • Improve cache-hit ratios: With more memory available per core, larger portions of the database can reside in the system global area (SGA), reducing expensive I/O operations and further boosting efficiency.

What customers are saying

Early experiences with M4N show that workload-optimized infrastructure is the engine for transformation. 

“Before M4N, meeting our demanding I/O requirements on Google Cloud often required over-provisioning our compute to achieve the necessary performance density. The new M4N instances solve this by delivering high throughput across the smaller to larger shapes.” - Sherri Trojan, Sr Principal Solution Architect, Sabre

sabre

"We are delighted to see Google Cloud introduce this next-generation high-performance infrastructure for mission-critical database workloads. The new compute platform demonstrates tremendous potential for enterprise Oracle deployments requiring scalability, resiliency, and performance. We are excited about what this innovation means for customers running Oracle workloads on Google Cloud.” - Bala Kuchibhotla, Co-Founder and CEO, Tessell

tessel

"With M4N, Google Cloud continues to push the boundaries of platform co-design. By combining 5th Gen Intel Xeon Scalable processors with Google's custom Titanium offload architecture, M4N delivers the extreme memory capacity, high memory bandwidth, and uncompromising I/O throughput required for the world’s most demanding mission-critical data environments." -  Intel

intel

What’s new: Scaling extreme data layers with M4N

M4N bridges two previously separate paradigms in cloud infrastructure: large memory footprints and extreme I/O performance. Engineered with custom Titanium offloads, M4N minimizes I/O bottlenecks without requiring infrastructure add-ons or compromises on memory density. Let’s take a look at how M4N fits into these environments. 

1. Enabling high bandwidth data transfer

For workloads with large memory footprints, M4N provides: 

  • Superior VM-to-VM bandwidth: Delivers up to 400 Gbps aggregate VM-to-VM network bandwidth and up to 50 Gbps single-flow bandwidth within the same VPC, unlocking non-blocking data exchange for distributed database clusters and real-time streaming data layers.

  • Enhanced internet and egress throughput: Enjoy up to 200 Gbps internet egress bandwidth and up to 48 MPPS packet processing performance.

  • High bandwidth out-of-the-box: Achieve full performance without needing to purchase or configure premium Tier_1 networking add-ons.

2. Dynamic storage performance with Hyperdisk

Paired with Google Cloud's next-generation storage portfolio, M4N with Hyperdisk lets you independently tune IOPS, throughput, and capacity:

  • Hyperdisk Extreme (HdX): Delivers up to 25 GiB/s aggregate block storage throughput and 1,000,000 IOPS—double the storage performance of standard M4. This is great for rapid database recovery, transactional checkpointing, and instant in-memory index reloads.

  • Hyperdisk Balanced (HdB): Scales up to 20 GiB/s throughput and 640,000 IOPS for cost-effective enterprise storage at scale.

M4N machine types and specifications

M4N instances are offered across three distinct memory-to-vCPU ratio tiers, scaling from 16 to 224 vCPUs and up to 5,952 GB of DDR5 RAM. M4N also offers predefined VM shapes across three distinct memory-to-vCPU ratios to match specific workload requirements, with support for Resource-based Committed Use Discounts (CUDs).  Details here.

Get started today

The M4N instances are now available in select regions around the globe. To learn more about how the M4N family can enhance your memory- and I/O-bound applications and reduce your licensing costs, contact your account representative or explore the documentation.

  •  

Google is a leader in The Forrester Wave™: Public Cloud Platforms, Q3 2026

We are excited to share that Google Cloud was named a Leader and received the highest score in the ‘current offering’ category in the Forrester Wave™: Public Cloud Platforms, Q3 2026 report, which examines the 10 most significant public cloud providers across 30 comprehensive criteria, Google also received the highest possible score in 23 out of 30 evaluation criteria, including, but not limited to vision, innovation, AI development services, database services, analytics services, containers and kubernetes services, modernization services, and security services. We believe Forrester’s recognition confirms our belief that to lead in the agentic era, you need a complete, integrated platform that’s engineered from the ground up, from silicon to systems to models.

Build on co-designed infrastructure proven in global enterprises

For over a decade, our infrastructure engineers, application developers, and AI researchers worked side by side to co-design infrastructure to power Gemini, Search, YouTube, Maps, and Gmail. We couldn't simply buy the platform and infrastructure we needed; we had to invent it. This led to the creation of everything from TPUs, the Transformer architecture, Kubernetes, Axion, and now Gemini.

In the agentic era, you need an integrated AI stack, where compute, orchestration software, modernization tools, and global networks operate together to give you more value from your investments — even if you’re not working at the frontiers of AI research. At Google Cloud, we’ve worked tirelessly to bring these breakthrough innovations to leading enterprises, startups, and frontier labs to help them achieve new levels of scale and efficiency, and we believe Forrester’s evaluation validates that strategy: 

“Google Cloud’s vision is to enable the ‘agentic enterprise,’ and AI already permeates its platform, positioning the company to push further up the tech stack toward business users who increasingly shape AI adoption in the enterprise. Google Cloud is a good fit for enterprises seeking rapid technology innovation and a broad AI-enabled cloud platform.” - The Forrester Wave™: Public Cloud Platforms, Q3 2026 report

Run agents quickly on a secure, flexible platform

Most traditional infrastructure can’t keep pace with agents, and enterprises need a scalable alternative. But you don’t want a new, greenfield platform just for AI agents. Kubernetes is the proven industry standard for modern enterprise applications — from microservices and transactional databases to real-time LLM inference. We are evolving Google Kubernetes Engine (GKE) and our operations tooling so organizations can scale autonomous agents alongside traditional workloads on a single, proven platform.

Forrester gave Google Cloud the highest scores possible in Container and Kubernetes services, Serverless/FaaS services, and Operations management services, noting:

“Operators will find strong offerings in operations management as well as containers and Kubernetes services. Our evaluation did not identify significant capability gaps.”

Over the past three months, we’ve enhanced our infrastructure portfolio to help teams scale agentic workloads with enterprise predictability. Recent updates let you:

  • Safely execute untrusted agent code alongside traditional workloads with default-deny security using GKE Agent Sandbox (GA) and Cloud Run Sandboxes (preview), which provision lightweight, gVisor-isolated boundaries for your agent in under a second (and up to 300 sandboxes/sec per cluster).

  • Eliminate up to 90% of idle compute costs by serializing your container RAM state directly to Google Cloud Storage with GKE Pod Snapshots, allowing you to suspend idle agent sessions in ~100ms and resume them in ~280ms.

  • Cut time-to-first-token (TTFT) up to 70% and double cache-hit rates with predictive routing in GKE Inference Gateway, which uses a continuously trained ML model to make routing decisions based on real-time traffic data.

Ground your agents with real-time enterprise data

Agents are only as effective as the context that grounds them. Traditional distributed data topologies separate operational databases from analytical systems through fragmented, multi-hop pipelines. In the agentic era, this divide introduces multi-hop latency, stale context, and governance friction.

Our Agentic Data Cloud evolves the enterprise data platform from a static repository into a dynamic reasoning engine. It unifies transaction processing and analytical intelligence into an active system of action, providing the real-time context and deterministic responsiveness that autonomous workflows require. Google received 5/5 scores across the Database services, Analytics services, Data integration services, and Data Governance services criteria:

“Google Cloud’s traditional strength in database services and analytics drives strong performance, including multicloud and hybrid capabilities, along with an Agentic Data Cloud that bridges analytics and transactional systems.”

Over the past three months, we’ve introduced key capabilities to the Agentic Data Cloud to help customers unify their data estates:

  • Enable agents to query live financial and supply chain records without costly data movement using SAP BDC Connect for BigQuery (GA), which provides bi-directional, zero-copy data sharing between your SAP systems and BigQuery.

  • Map and infer business meaning across your entire data estate with Knowledge Catalog. You can now aggregate native context across your Google and partner data platforms, semantic models, and third-party catalogs, unifying them into a single, governed source of truth.

  • Access live data from Iceberg and BigQuery from the PostgreSQL data plane with Lakehouse federation. Perform live joins between AlloyDB's transactional data and historical insights in BigQuery or Iceberg without any data movement. You can also replicate data continuously to BigQuery and, importantly, to Iceberg tables directly from AlloyDB with Datastream.

The benchmark is set: Build what’s next on Google Cloud

We are honored that Forrester has named Google Cloud a Leader in The Forrester Wave™: Public Cloud Platforms, Q3 2026. We believe this recognition validates decades of foundational research, disciplined full-stack co-design, and our commitment to building an open, reliable cloud.

The era of fragmented infrastructure has come to an end. Whether your organization is an AI research lab scaling models across one million accelerator chips, a global financial exchange settling trillions in clearing systems, or an enterprise empowering millions of users with autonomous workflows, Google Cloud delivers the performance, scale, security, and data foundation to build what’s next.

Take the next step in your cloud journey:

  •  

Google Cloud partners with CIQ to provide an enterprise-grade experience for Rocky Linux

At Google Cloud, we strive to offer a great customer experience for enterprises by building a robust and supported platform for running all Linux-based workloads.

This mission is why we were one of the first cloud providers to offer purpose-built Rocky Linux images when Rocky Linux debuted last year as a replacement option for CentOS. We were also one of the first hyperscalers to sponsor the Rocky Enterprise Software Foundation (RESF) to support the open source community behind this Linux distribution. With these efforts, we’re pleased that many customers are already running Rocky Linux in Google Cloud today.

Today, we’re excited to announce that we’re taking another step in furthering the support we provide for Rocky Linux. We’re partnering with CIQ—the company started by CentOS co-founder and Rocky Linux founder Gregory Kurtzer featuring core expertise across Linux, cloud, HPC, containers and security— so we can provide customers a new and improved experience for Rocky Linux on Google Cloud. 

Starting today, customers can leverage Google’s support offerings to file support cases for Rocky Linux. Google support teams and the Rocky Linux experts at CIQ are working together to address customer issues to help ensure they get enterprise-grade support. If you already have a paid support plan with Google, you will be able to open a case for an issue related to Rocky Linux. Google teams can expediently help resolve the issues, backed by CIQ expertise, giving you an integrated experience of using Rocky Linux on Google Cloud. 

"We asked ourselves, how do we bring the best value to everyone? Through this partnership, anytime you use our Rocky Linux on Google Cloud, both Google and CIQ jointly have your back! From the cloud platform itself, all the way through the enterprise operating system, every aspect of using Google Cloud is supported by a single call to Google, and together, we are your escalation team.”—Gregory Kurtzer, CEO of CIQ and Founder/Director of Rocky Linux and the RESF

In addition to CIQ-backed support for Rocky Linux, Google is also working with CIQ to provide a streamlined product experience - with plans to include performance-tuned Rocky Linux images, out-of-the-box support for specialized Google infrastructure, tools to help support easy migration, and more. We’re doing these updates in a community-friendly way. Together with CIQ, Google is helping to create a Rocky Linux Cloud SIG that aims to provide optimized, standardized, and simplified Rocky Linux experience. 

If you’re currently looking for alternatives to CentOS as it reaches end of life, Rocky Linux on Google Cloud can have you covered both from a product and support perspective. So, take Rocky for a spin if you haven’t already, and if you have questions or suggestions on how we can help you, please don’t hesitate to reach out to us. To learn more, please also join us for a webinar discussion on April 6th 2022 at 11.00am PT.

  •  

Google named a Leader in 2026 Gartner® Magic Quadrant™ for Strategic Cloud Platform Services

For the ninth consecutive year, Gartner® has named Google a Leader in the Gartner Magic Quadrant™ for Strategic Cloud Platform Services, positioned furthest for Completeness of Vision.

scps-26

We believe this recognition reflects our longstanding dedication to helping customers build and scale their most demanding workloads reliably and securely on Google Cloud. As we enter the agentic era, we're accelerating their journeys with a dynamic infrastructure, and connecting enterprise apps, data and agents on a single, flexible platform for predictable cost and performance.

What’s driving this momentum? There are three major advantages that we feel set Google Cloud apart:

  1. A co-designed, unified technology stack across custom silicon and hardware systems, open software and orchestration, frontier models and agentic applications.

  2. A dynamic infrastructure that helps you securely connect and scale your users, data, apps, and agents everywhere.

  3. Digital sovereignty with genuine choice, giving organizations total control over their data without sacrificing essential cloud functionality.

We are committed to helping our customers innovate and deliver at scale while giving them the flexibility, performance, and control they need. Let’s dive into three design principles that Google Cloud lives by as we continue to build and enhance our infrastructure:

1. Accelerate AI with a co-designed stack without lock-in

Google Cloud is the only provider to deliver a complete, first-party AI stack that is deeply co-designed from silicon to agentic applications. Our infrastructure team works with Google DeepMind researchers to co-design and optimize every layer of our technology stack. From custom silicon, like Google TPUs and Arm-based Google Axion processors, to Google Kubernetes Engine (GKE) and Gemini models, our system delivers exceptional performance and predictable costs.

Co-designing hardware and software creates massive operational efficiency, but it doesn't mean creating a closed ecosystem. We remain deeply committed to open source and open standards across every layer of the stack including frameworks like llm-d for distributed inference, benchmarking for open models with GKE Prism, and eliminating hardware lock-in with TorchTPU for PyTorch compatibility across TPUs and GPUs. You get the full power of a vertically co-designed stack while maintaining complete freedom across models, frameworks, and silicon. Combining all of these deeply integrated components means building an AI Hypercomputer, capable of exceptional scale, performance, and efficiency. This is the same infrastructure foundation chosen by nine of the top ten AI labs globally. 

2. Scale quickly and economically with a dynamic infrastructure

Today, demand for AI resources is skyrocketing. Internally at Google, our data centers now process 3.2 quadrillion tokens monthly, roughly 7x more than last year1. For enterprise leaders navigating this shift, scaling AI systems are notoriously difficult to architect, resource-intensive, and bursty, which can lead to scaling bottlenecks and large pools of underutilized compute. 

Organizations need a dynamic infrastructure to automate capacity management, modernize business applications at their own pace, and securely connect data, apps, and agents to drive optimized global experiences. 

To thrive at an agentic scale, you need infrastructure capable of operating as a single system. With Google Cloud, you can:

  • Kick-start your AI transformation using Gemini-powered tools to intelligently map and modernize core apps, turning static legacy systems into dynamic foundations for AI and agents.

  • Choose from a wide range of workload-optimized compute types and configurations. You can combine predefined and custom CPU shapes, NVIDIA GPUs, and Google custom silicon (TPUs and Axion CPUs), which are designed to deliver exceptional performance-per-dollar for AI and Enterprise workloads.

  • Connect your enterprise and AI infrastructure on a single, flexible control plane with Google Kubernetes Engine. This includes capacity management capabilities like Dynamic Workload Scheduler to preschedule capacity for planned events and dynamic resource allocation to define advanced rules that dictate how resources are consumed, helping to maximize utilization and reduce costs.

  • Simplify day two operations using Gemini Cloud Assist to proactively identify, troubleshoot, and resolve operational issues for your new agent-based workflows.

  • And finally, run your workloads across hybrid and multicloud environments with Cross-Cloud Interconnect, leveraging Google’s 10+ million kilometer private fiber backbone to deliver up to 40% higher performance than public internet routing and automated delivery in minutes2.

By adopting a unified foundation of dynamic infrastructure, adaptive applications, and responsive systems, organizations can establish the resilient, high-performance infrastructure necessary to lead in this new technological frontier.

3. Embrace digital sovereignty with more choice and security

You shouldn’t have to compromise between modernization, frontier AI capabilities, and regulatory control. Sovereign Cloud from Google gives you access to Gemini and open-weight models across sovereign platforms with three flexible deployment options:

  • Data sovereignty and control: Retain complete control over your data’s location and cryptographic authority with Google Cloud Data Boundary. Manage your encryption keys outside Google infrastructure using External Key Management (EKM) with Key Access Justifications (KAJ), while enforcing precise geographic processing and storage boundaries across both Google Cloud and Google Workspace.

  • Local compliance and regional operations: Run your applications on physically and logically separated regional clouds operated exclusively by local partners, built on Google Cloud dedicated for European customers. In France, S3NS delivers PREMI3NS, providing a standalone sovereign cloud that has achieved the SecNumCloud 3.2 qualification from the French National Agency for the Security of Information Systems (ANSSI). Dedicated sovereign cloud operations operated by Thales are also coming soon to Germany.

  • On-premises and air-gapped flexibility: Bring Google Cloud capabilities directly to your on-premises environment via Google Distributed Cloud (GDC). GDC offers two distinct deployment modes: air-gapped, a fully-managed, self-contained environment operating with zero connectivity to the public internet for public sector, defense, and regulated enterprise workloads; and connected, allowing you to run workloads on your hardware locally while leveraging Google Cloud’s centralized control plane for unified management. 

Accelerate your Cloud journey

Whether you're modernizing core enterprise systems, managing complex compliance requirements, or deploying autonomous AI agents, Google Cloud gives you the performance, scale, and freedom of choice to succeed.

Read the full 2026 Gartner Magic Quadrant for Strategic Cloud Platform Services or explore our AI Hypercomputer page to learn more.


1. Pichai, Sundar. "I/O 2026: Welcome to the Agentic Gemini Era." The Keyword, Google, 19 May 2026
2. During testing, network latency was more than 40% lower when traffic to a target traveled over the Cross-Cloud Network compared to when traffic to the same target traveled across the public internet.

  •  

Dynamic capacity management for AI infrastructure

The internet connected billions of people and mobile devices, putting computers in every hand. Now, we’re in the middle of the next big technology shift, deploying millions of autonomous AI agents to work alongside employees and end users. Today, we announced new FinOps controls for Gemini Enterprise to help organizations manage project-level AI spend and eliminate token shock. But the sheer scale of the agentic era is placing new constraints at every layer of the stack, including infrastructure. AI workloads are notoriously difficult to architect, resource-intensive, and bursty, which can also lead to scaling bottlenecks and large pools of underutilized — or misutilized — compute resources. 

Organizations need insights to help them extract more value from their infrastructure investments. In this blog, we outline best practices for dynamic capacity management — scheduling and utilization strategies to help you run enterprise and AI applications on a single, flexible foundation with predictable cost and performance. These capabilities are designed to augment our on-demand, Spot and committed use discount (CUD) consumption models, which provide flexible pricing and discounting for your workloads. Let’s jump in.

Here's a quick summary

Three ways you can implement dynamic capacity management:

  1. Schedule capacity for planned events. Schedule mission-critical resources (GPUs, TPUs and select VM families) ahead of planned events using calendar mode, or optimize costs for batch jobs with flexible start times using flex-start mode in Dynamic Workload Scheduler. Once you obtain the capacity, those resources are guaranteed for the specified duration.

  2. Maintain service continuity by creating a fallback plan for every application. Define automated, prioritized hardware fallback lists using managed instance groups (MIGs) so your apps automatically pivot to the next approved compute option when your preferred option isn’t available.

  3. Automate your entire capacity management lifecycle on a single, adaptive control plane. Google Kubernetes Engine (GKE) provides an agent-native environment to orchestrate the entire process — from fallback lists using Custom ComputeClasses, to granular hardware slicing with dynamic resource allocation, so agents can rapidly spin up in secure sandboxes and containers while it dynamically reallocating resources on the fly.

Why architectural flexibility matters

Ninety percent of enterprises want to deploy agents within the next three years, but only 17% of IT leaders feel confident their current IT setup can handle the load. Because these workloads have unique performance needs, organizations are racing to adopt specialized infrastructure, including accelerators (GPUs, TPUs) and CPUs with customized compute, memory, and storage ratios. However, agents also require access to enterprise applications and databases — often at a volume and scale that vastly exceeds typical human usage. Handling the intense demands of both agents and the applications they interact with requires a dynamic infrastructure. Infrastructure teams can leverage custom-designed processors like Google’s Axion to meet these needs, but hardware isn’t a complete solution. They also need ways to use that infrastructure wisely, solving execution inefficiencies to enable more flexibility across the stack.

How to overcome infrastructure constraints

Achieving this kind of flexibility requires a two-pronged approach: securing resources for the demand you can predict, and building automation to respond to the demand you can't. Combining the two, you can preschedule capacity for planned events and your infrastructure can adapt to unexpected changes without manual intervention.

1. Schedule capacity for planned events

You can secure mission-critical capacity ahead of scheduled milestones, offline training, or anticipated demand surges using Dynamic Workload Scheduler. By scheduling the resources you need up front, you optimize your spend and ensure you get access to the compute resources you need. Dynamic Workload Scheduler supports hardware accelerators (TPUs and GPUs) and select CPUs with two distinct modes:

  • Flex-start mode: Use this for latency-tolerant workloads like batch processing, model training, or offline fine-tuning. Instead of requiring resources immediately, you submit a defined duration request and the system intelligently queues your job, provisioning the resources as soon as capacity becomes available. This maximizes cost-efficiency and drastically improves your ability to obtain high-demand accelerators.

  • Calendar mode: Use this for mission-critical, time-bound events like a major product launch, a scheduled migration, or a seasonal traffic surge. By specifying the exact start and end dates of your event, you create a future reservation. This guarantees the requested capacity will be available when the event begins.

1

2. Maintain service continuity by creating a fallback plan for every application

Not every spike in traffic is predictable. You also need to plan for unexpected traffic from, say, a breaking news cycle or a sudden market shift that drives a surge in user activity. To help your services get the resources they need without interruption, you need a fallback plan — an automated, prioritized sequence of acceptable hardware configurations. This strategy:

  • Decouples your workloads from a single VM shape, size, or configuration. This allows them to run without manual intervention if your preferred option is unavailable

  • Allows you to execute a progressive tech refresh by adopting the newest VM generations as your primary choice while keeping older generations as an automatic fallback option.

If you run non-containerized workloads on Google Compute Engine, you can dynamically manage capacity with instance flexibility in managed instance groups (MIGs) and bulk VM creation. Instance flexibility lets you specify multiple machine types for your VM instances rather than being limited to a single machine type.

How it works: If your preferred machine type is temporarily unavailable, the MIG automatically provisions a compatible alternative from your list based on real-time capacity. When combined with location flexibility — by specifying multiple zones your MIGs can search within a region — you can drastically improve your provisioning success rate. If your MIGs use Spot VMs, Compute Engine automatically integrates with Spot capacity signals to prioritize machine types that offer longer estimated uptimes and lower risk of pre-emption.

2

You can also extend instance flexibility to your block storage layer by setting baseline disk defaults and configuring disk overrides so your storage adapts when a VM falls back to a different machine type. 

How it works: Most of the time you can simply rely on our default options, omitting ‘disk type’ from the instance template entirely. However, for data disks that will outlive their associated VMs, it’s possible to enable a fast, durable Hyperdisk across multiple VM generations.

While Compute Engine provides instance flexibility for organizations working with virtual machines, GKE goes a step further and automates the entire capacity lifecycle from a single control plane. With GKE custom ComputeClasses, platform teams can design multi-dimensional fallback lists, automatically combine different VM machine families, sizes, and ratios, scale across multiple zones, and shift between on-demand and Spot VMs. By using Dynamic Workload Scheduler as a capacity target, and custom ComputeClasses to define the policy and priority, you can fully automate the capacity management lifecycle.

How it works: Once you’ve set up ComputeClasses, GKE automatically detects when a preferred node configuration is unavailable and falls back to your pre-approved alternative options in order of priority. When active migration is enabled, GKE gracefully migrates workloads back to higher-priority node configurations as capacity becomes available. For short-lived disks such as boot disks, GKE dynamically picks the right defaults based on the instance family. However, for long-term disks that will outlive the VM, you can use Hyperdisk.

3

Another GKE feature, dynamic resource allocation, helps eliminate wasteful, all-or-nothing hardware assignments by letting developers define advanced rules that dictate how resources are consumed.

How it works: Instead of claiming an entire GPU or TPU, your application specifies its exact parameters — such as total memory or number of cores — and the system allocates the perfect slice of hardware, helping to maximize utilization and reduce costs. 

 

4

Take the next step toward dynamic infrastructure

Scaling AI shouldn’t mean linearly scaling your infrastructure budget or accumulating more tech debt. As these examples show, the right tools can help you overcome constraints and dramatically alter the value you get from your compute investments. Here are three steps to get started:

  1. Audit your workloads for immediate cost-savings: Identify any applications currently tightly coupled to a single VM family, machine type, or availability zone, and map out viable alternative hardware shapes. Look beyond your existing configurations to evaluate new compute options that might better serve or act as alternatives based on your workload-level objectives. Then use Compute Engine MIGs, bulk VM creation or GKE Custom ComputeClasses to adopt them automatically, integrating them into your fallback lists.

  2. Commit to a minimum spend for deeply discounted prices: Receive automatic discounts for sustained use, or up to 63% off when you sign up for Compute flexible committed use discounts, where your discount is tied to the resources you use regardless of the specific machine type or location.

  3. Engage your account team: Reach out to your Google Cloud account team to craft a tailored capacity management strategy and configure your automated fallback lists.

  •  

PQC in Plaintext: Google Cloud’s post-quantum cryptography roadmap

Securing infrastructure and services against a future cryptographically-relevant quantum computer has been a goal for Google for a decade, and we’ve dedicated ourselves to help developers by advancing open standards that can benefit everyone. As post-quantum cryptography (PQC) has matured, we’ve been rolling it out in our infrastructure for internal and customer-facing services. 

Today, we're sharing our updated Google Cloud roadmap to migrate to PQC by 2029.

Our strategy: Secure by design

We’ve based our PQC migration strategy on the Google Quantum Threat Model, prioritizing protection across three key domains:

  • Mitigating Store Now, Decrypt Later (SNDL) risks: Protecting today's encrypted data from being harvested and decrypted by a future quantum computer.

  • Ensuring integrity against forgery: Strengthening digital signatures to prevent attackers from falsifying data and identity.

  • Enhancing foundational capabilities for cryptographic agility: Building flexible systems that can easily adopt new cryptographic standards with minimal engineering effort as cryptographic standards evolve.

We’re actively transitioning internal infrastructure and customer-facing services to PQC algorithms far ahead of regulatory deadlines.

We are also deploying PQC solutions across our Sovereign Cloud initiatives, such as Google Cloud Dedicated (GCD) and Google Distributed Cloud (GDC), in collaboration with our partners. Similarly, our strategy allows us to progress on integrating post-quantum protections across our AI services to secure the next generation of cloud workloads. 

These efforts are fundamental pillars of our overarching strategy to achieve full post-quantum readiness across Google Cloud. As this landscape evolves, we will continue to refine and update our deployment schedules.

1

Visualization of our Google Cloud PQC roadmap. Our efforts converge in 2029, and extend beyond it.

We plan to achieve full PQC readiness by 2029, when our efforts converge. We anticipate continuing those efforts into the 2030s to support broader industry guidance and evolving global standards. These standards include CNSA 2.0 and the transition paths defined in NIST IR 8547, which anticipate the final deprecation of legacy, quantum-vulnerable algorithms between 2030 and 2035.

Immediate progress: 2026 milestones

Leadership in the quantum era requires deployment at global scale. We have achieved foundational milestones that provide immediate protection for our customers:

  • API endpoint readiness: Google Cloud API endpoints now offer quantum-safe key exchange, protecting incoming traffic from future decryption. These endpoints include google.com and *.googleapis.com, and both have implemented NIST-standardized ML-KEM (FIPS 203) in hybrid mode.

  • Load balancers PQC support: Application and proxy load balancers now support quantum-safe hybrid key exchange (X25519MLKEM768) for TLS 1.3. Initially available on an opt-in basis, this allows our customers to perform validation, while minimizing impacts to their existing applications.  

  • Quantum-safe certificate experimentation at scale: We’re collaborating with the IETF PLANTS Working Group to produce a public key infrastructure (PKI) standard that minimizes impact to your operations teams. Chrome and Cloudflare have started experimenting with Merkle Tree Certificates to address challenges using PQC signatures for WebPKI, and we have been sharing insights with the standards working group.

  • Cloud KMS PQC algorithms: NIST standardized PQC algorithms (ML-KEM, ML-DSA, SLH-DSA) for your encryption and signing keys are now generally available. 

The roadmap to 2029

We’ve established specific customer-centered journeys for Google Cloud to achieve quantum readiness that allow us to prioritize our quantum-safety initiatives. By adopting this risk-based approach, we focus on the core journeys our security experts have identified as most vulnerable to the potential impacts of quantum computing.

2

Risk-based prioritization for core quantum readiness user journeys.

These scenarios offer diverse platform perspectives to ensure global enablement across our services to meet you where you are.

For each risk domain, we provide a roadmap for key products and services organized by domain, although the services highlighted are not exhaustive lists. We project most services will meet their respective domain's target completion date, though specific product timelines may be adjusted if necessary to account for evolving engineering requirements and any third-party dependencies.

Domain 1: Store Now Decrypt Later (SNDL) mitigation

This domain focuses on addressing vulnerabilities in asymmetric encryption where a future cryptographically-relevant quantum computer (CRQC) could decrypt data captured today.

We’re enabling incremental progress for our customers based on their typical journeys.

  • Securing your customer workloads: Offer quantum-confidential TLS 1.3 handshakes for your Google Cloud services and configured load balancers to protect user sessions.

  • Securing administrator and developer flows: Protect the admin pathways used to manage your cloud environment against SNDL. This includes services such as Cloud VPN and Interconnect. For developers, these include client libraries, SDKs, and Tink, our open-source cryptographic library.

  • Securing data pipelines: Safeguard the confidentiality of data transfers for our analytics and storage platforms. PQC is essential to ensure that sensitive intellectual property and customer data flowing through these systems cannot be captured today and decrypted by a future quantum-capable adversary.

Roadmap
We are targeting these changes for 2027. 

Journey

Benefits

Representative services

Store now decrypt later mitigation (End of 2027)

Securing your customer workloads

Quantum-safe ingress: Protects your cloud perimeter using standardized post-quantum algorithms.

Application and proxy load balancing
[2026 (Completed)]

Securing admin and developer flows

Secure operations: Validates that your management and deployment stack meets emerging cryptographic standards.

Quantum-confidential ALTS
[2025 (Completed)]

API endpoints
[2026 (Completed)]

Cloud VPN, Cloud Interconnect, GCE OS Login, Cloud SDK, gCloud CLI, GKE service mesh, and client libraries
[2026/2027]

Securing data pipelines

Confidential data transfers: Protect sensitive intellectual property and customer data against quantum attackers.

Cloud Storage SDK, Storage Transfer Service, BigQuery CLI, Data Transfer Service
[2026/2027]

Domain 2: Integrity and non-repudiation

This domain addresses quantum-proofing of digital signatures and attestations to safeguard against forgery that could compromise data integrity and authenticity.

  • Securing the software supply chain: Ensure that only trusted, untampered images with quantum-resistant attestations run in production to help prevent a quantum attacker from altering builds. This includes services like Binary Authorization, Cloud Build, and Assured Open Source Software.

  • Issuing quantum-safe certificates: Transition the public key infrastructure (PKI) including our internal and external certificate authorities (CAs) to support ML-DSA certificates and where meaningful, SLH-DSA certificates. This transition will follow Internet Engineering Task Force (IETF) standardization efforts that are currently in development. We’re actively contributing to these efforts, and we’re also conducting several large-scale experiments:

    • Address large PQC signature sizes that can impact the performance of certificate chain validations through novel approaches like Merkle Tree Certificates for Web PKI. 

    • Support ML-DSA/SLH-DSA (pure PQC)-based certificates in private CA solutions such as Certificate Authority Service (CAS). 

    • Add quantum-authentication in addition to our internal traffic protocol ALTS that already supports PQC for confidentiality. You can learn more technical details on our approach to digital signatures and Public Key Infrastructure here.

  • Protecting identity and access: Ensure authentication mechanisms like service account keys and tokens (JWT/OAuth) are resistant to quantum forgery.

Roadmap
We are targeting completion of these milestones by 2028. We also are mindful of ongoing standardization efforts particularly in the field of certificates. Google is actively contributing to quantum-safe certificate standards, and we are committed to help the industry overall meet those deadlines.

Domain / Journey

Benefit

Representative services

Integrity and non-repudiation (End of 2028)

Securing the software supply chain and signature services

Quantum-safe software attestations: Prevents unauthorized build tampering by ensuring only trusted images run in production.

Binary Authorization, Access Approval (AXA)
[2026]

Assured OSS
[2027]

Issuing quantum-safe certificates

Quantum-safe standardized trust: Safeguards the authenticity of your internal and external communications against quantum-calculated certificate forgery.

Quantum-authentic ALTS
[2026/2027]

Private CA (Certificate Authority Service)
[2027]

Google Trust Service: Merkle Tree Certificates
[2028]

Roll out of PQC certificates across Google Cloud products and infrastructure
[2027/2028]

Protecting identity and access

Governed identity: Eliminates the risk of adversarial credential fabrication with NIST-standardized signatures for auditable integrity.

Cloud IAM
[2028]

Infra-wide rollout of quantum-safe authentication and access
[2027/2028]

Domain 3: Foundations and key management

Cryptographic agility is the foundation of our PQC migration. Our ongoing investment in this area drives our end-to-end strategy for key management, libraries, and infrastructure changes.

  • Foundational key management and libraries: Enable NIST-approved algorithms through Cloud KMS and libraries like BoringSSL and Tink. 
      • Note that Cloud KMS achieved general availability for the NIST standardized PQC algorithms (ML-KEM, ML-DSA, SLH-DSA), and is in the process of enabling quantum-safe key import.

  • Hardware-backed cryptographic services: Secure physical foundations using quantum-resistant roots of trust. This includes PQC as part of our Confidential Computing offerings and Cloud Hardware Security Module (HSM).
  • Key sovereignty and partner solutions: Enable PQC orchestration for Google Workspace Client-side Encryption (CSE) and External Key Managers (EKM). Collaborate with partners to support PQC on-premises key providers.

Roadmap
We are targeting completion of these milestones by 2028. 

Domain / Journey

Benefit

Representative services

Foundations and key management (End of 2028)

Foundational key management and libraries

Standardized quantum-safe keys: Provides the NIST-approved building blocks to help migrate your applications.

ML-DSA and SLH-DSA in KMS, ML-KEM and Hybrids in KMS
[2025 (completed)]

Quantum-safe Key Import (BYOK)
[2026]

Hardware-backed cryptographic services

Silicon rooted hardware: Anchors your security in quantum-safe hardware roots of trust.

Confidential Compute (including attestation and vTPM)
[2028]

Quantum-Safe Cloud HSM (FIPS 140-3 L3)
[2028]

Key sovereignty and partner solutions

Cryptographic provenance: Provides control and provenance of your keys where you need them.

External Key Management
[2028]

Partner enablement (key providers and sovereignty solutions)
[2028]

 

A shared responsibility for quantum safety

Security has long been a collaborative partnership with our customers. 

Google’s responsibility — Security of the cloud: We manage the transition to a quantum-safe infrastructure, including our network and encryption in-transit, global front-ends, and the ALTS protocol. 

This responsibility encompasses the end-to-end PQC transition of our servers, ensuring that the underlying hardware and operating systems are secured against quantum threats. We maintain hardware integrity through quantum-safe, open-source silicon foundations such as Caliptra v2.1, TPM 2.0 v185, and OpenTitan. The latter is the first open-source silicon root of trust and already supports quantum-secure boot. 

While we are working toward our 2029 target, hardware transition to PQC involves both active replacement, where feasible, and natural equipment replacement cycles. Our phased approach ensures stability, though the timeline for some physical components may extend beyond 2029.

Customer’s responsibility — Security in the cloud: Organizations must manage their own applications, including updating client-side software to negotiate PQC handshakes and managing the lifecycle of your asymmetric keys.  

In addition, you should update your Google Cloud service configurations with quantum-safe settings and policies.

Our path forward, together

Building momentum toward quantum readiness requires immediate, practical action. We recommend starting with these three steps:

  1. Inventory: Identify your cryptographic resources (such as keys and certificates) using Cloud Asset Inventory and solutions such as Wiz’s cryptography and PQC readiness. When you map cryptographic resource usage across your organization, you can more accurately define and prioritize your migration backlog.

  2. Update: Ensure your development and Site Reliability Engineering teams are using software that supports PQC algorithms such as BoringSSL, Chrome, and SDKs. This update ensures your internal workflows are prepared to negotiate quantum-safe connections by default as we enable them at the edge.

  3. Validate: Test the behaviors of your existing application using our quantum-safe APIs and load balancers. Validating your workflows today will identify architectural bottlenecks before they impact your primary production environments.

Google Cloud is committed to managing the complexity of this transition so you can achieve your regulatory and compliance commitments, while focusing on innovation. We are just beginning to share our progress as we work to empower our customers to lead in the post-quantum landscape.

To learn more about our PQC approach, please visit our post-quantum cryptography (PQC) hub.

  •  

What’s new in AI infrastructure and orchestration in August

Welcome back to What’s new in AI infrastructure and orchestration this month, a collection of product updates, how-tos, customer stories, research and other resources about all the AI compute, networks, storage, frameworks, and orchestration software that you can find at Google Cloud. To be honest, we thought August would be a slow month, but nothing could be further from the truth. Read on and you’ll see what we mean.

August 2026

Product, technology, and tools updates

  • Product update: Filestore, Google Cloud’s first-party, secure, scalable NFS file service, has emerged as a popular storage platform for AI and agentic workflows, and now, it’s even better suited to the task, with a new backend storage layer built directly on Colossus, Google’s foundational distributed storage system. This new backend lets you provision IOPS independently from storage capacity, and is deeply integrated with GKE. In AI environments, this can help you service so-called agentic swarms — large groups of agents that need to read and write to a common dataset — without a drop off in performance. For more, check out the blog post. 

  • New feature: gVisor sandboxes are now available in distributed Ray clusters on GKE. In partnership with Anyscale, we introduced an experimental library for Ray that brings gVisor, Google’s open-source application kernel, directly into distributed Ray clusters. gVisor provides lightweight environments with stronger isolation than ordinary containers, plus fast startup times and low memory overhead. To try out these sandboxing capabilities on GKE, head over to the Ray sandboxing User Guide.

  • Product update: Looking for high-performance, easy-to-use infrastructure on which to run a personal AI agent, but don’t want to spend a lot of money? New Cloud Run instances are dedicated, singleton compute runtimes on Cloud Run that won’t shut down when the agent is idle. Better yet, the cost to run a Cloud Run instance with 1 vCPU and 1 GiB of memory continuously for 30 days is just $5.70.  

Practitioner guides, documentation and how-tos

  • How-to guide: Big news in Model Context Protocol (MCP) land: As of the 2026-07-28 specification, the protocol core is “completely stateless. The handshake is gone. The initialize / initialized handshake (SEP-2575) and the logical Mcp-Session-Id header (SEP-2567) have been removed entirely. Instead, every request is now self-describing and independent.” Whoa. Learn more about the changes that the latest MCP specification brings, and more importantly, how to implement them, in this Google Developers blog.  
  • Guide: Real-time AI systems make a mess of traditional network load balancing techniques. “Instead of handling isolated requests, the backend has to manage a continuous, live bidirectional stream. You’re dealing with a constant stream of audio chunks, transcripts, model outputs, and synthesized speech flowing back and forth simultaneously.” Things only get worse when the user gets involved. “The server has to immediately halt its current speech generation, pivot to update the context, maybe trigger a new tool, and start drafting a different response; this must be done without dropping the connection.” For a new approach to managing load in the AI era, read Scaling real-time AI agents with session-aware load balancing.
  • How-to: Learn how to build an elastic, scalable LLM inference platform on GKE, even with a mix of different GPU accelerators. The proposed architecture combines Capacity Advisor and Compute Advisor, plus high-performance storage like RunAI:model streamer or GCPFuse with parallel downloads. Get all the details here.
  • Documentation: The thing about hosts with GPUs or TPUs is that you can’t use live migration to update them, setting up a maintenance challenge. In this new docs page, learn how to update accelerator-equipped hosts according to your tolerance for downtime for your training and inference workloads.    
  • Documentation: Advanced Compute Images, or ACIs, are standardized image stacks for AI/ML and HPC infrastructure, so you don’t need to manually build your own custom images. In this new docs page, learn how to create an ACI image using the Google Cloud CLI, console, or SchedMD's Slurm workload manager. 
  • Guide: AI workloads are notoriously difficult to architect, resource-intensive, and bursty, which can also lead to scaling bottlenecks and large pools of underutilized — or misutilized — compute resources. A new blog outlines the three main ways to achieve dynamic capacity management in Google Cloud: 1) scheduling capacity for planned downtime; 2) maintaining automated fallback capacity for unplanned downtime; and 3) relying on GKE’s core orchestration capabilities to automate resource allocation. 

Customer and partner updates

  • Business orchestration software provider UiPath was dealing with spiky workloads, and wanted more predictable costs. To get there, it re-architected its infrastructure, moving from isolated clusters to a shared Google Cloud GPU fleet that included both A3 VM instances (NVIDIA H100 GPUs) for training with G4 VM instances (NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs) for inference. You can read more about their architecture here. 

  • Mirendil, an frontier AI lab focused on accelerating AI development, announced that it is using AI Hypercomputer with both TPUs and NVIDIA GPUs to support its model pre-training and post-training applications. 

  • Replenit, a retail CRM provider, built its AI decision engine in Google Cloud, using BigQuery, Gemini Enterprise Agent Platform, and open-source Gemma models that it runs on Cloud TPUs. This latter combination provided Replenit with 90% lower pipeline costs than their previous cloud provider, the company reports. Read the full case study for more. 

  • Malachyte architected its AI-powered e-commerce recommendation platform on top of Bigtable, Managed Service for Apache Kafka, Pub/Sub, Compute Engine, and last but not least, GKE. See how it all comes together in this blog.


July 2026

Product, technology, and tools updates

  • Product update: Google Cloud Managed Lustre is now GA, and available in four distinct performance tiers that deliver throughput ranging from 125 MB/s, 250 MB/s, 500 MB/s, to 1000 MB/s per TiB of capacity — with the ability to scale up to 8 PB of storage capacity. The Managed Lustre solution is powered by DDN’s EXAScaler, combining DDN's decades of leadership in high-performance storage with Google Cloud's expertise in cloud infrastructure.

  • Product update: C4N network and storage optimized VMs are now GA. C4N is our first network- and block-storage-optimized VM series built to eliminate data-transfer bottlenecks. Powered by 5th Gen Intel Xeon Scalable processors and built on Google's Titanium offloading hardware, it achieves 400 Gbps network bandwidth, 95 million packets per second (MPPS), and up to 25 GiB/s of block storage throughput when paired with Hyperdisk Extreme.

  • New feature: GKE Dataplane V2 up to 15K Nodes with Network Policies (GA). This capability enables standard GKE clusters to scale up to 15,000 nodes while maintaining full active Network Policy enforcement, supporting the massive infrastructure needs of large enterprise and AI/ML customers.

  • New feature: Co-operative time-slicing in llm-d. If you’re running reinforcement learning (RL) workloads, you can now interleave independent RL jobs onto shared physical hardware, increasing aggregate accelerator duty cycles from a ~40% baseline up to 70% without impacting model convergence or accuracy. 

  • New AI security tool: Looking to secure your AI supply chain on GKE, deploy AI workloads safely, and cut down on shadow AI? We open-sourced k8s-aibom, a lightweight, unprivileged Kubernetes controller that continuously monitors container clusters to automatically detect running AI runtimes (like vLLM and Triton) and generate standard CycloneDX Machine Learning Bill of Materials (ML-BOMs). Check out the k8s-aibom project and get involved.

Practitioner guides and how-tos

  • How-to guide: On July 27, Google announced Day 0 support for Moonshot AI’s Kimi K3 2.8-trillion-parameter open-weight model, the day weights were released. Whichever your preferred deployment path — via Model Garden, custom orchestration, or GKE with llm-d recipes — this guide offers detailed step-by-step instructions to help you evaluate and pilot Kimi K3 in Google Cloud. 

  • How-to guide: Google Kubernetes Engine (GKE) managed DRANET supports both GPUs and TPUs. There are several configurations to use this implementation, including standard cluster (where you have full control) and autopilot cluster (where Google does the heavy configs for you). Take a deeper dive in the hands-on lab, GKE Autopilot clusters with TPUs, GKE managed DRANET and Gemma 4.

  • How-to guide: Learn to run Ray on TPUs, not GPUs. In Part 1 of this two-part series, we discuss TPU slices (hint: Ray thinks of them as just another accelerator on which to schedule), then walk through Ray’s various AI libraries (Part 2).

  • How-to guide: Evaluate TPUs for sample workloads using a new microbenchmark suite that helps you accurately assess whether a device is achieving its theoretical performance specifications, and to identify specific performance gaps or architecture-specific bottlenecks. Dive in here. 

  • How-to guide: Scale your agents without killing your budget. Learn how GKE orchestration can help you safely pack more agents onto a fixed compute footprint with GKE Agent Sandbox and Pod snapshots. Whether your goal is performance or cost optimization, we teach you how to turn the right dials for optimal agent efficiency. 

  • Technical blueprint: Inside the optimization of Mistral 3 large inference on Ironwood. This blog outlines how one Google team optimized Mistral 3 large MoE model inference on Google’s Ironwood (TPU v7x), achieving a 1.5x performance gain. They did so with hybrid sharding, replacing linear VPU summations with tree reductions, optimizing GMM/MLA kernels, and adopting asynchronous scheduling. As a result, they boosted throughput by up to 48% while maintaining benchmark accuracy neutrality. Read the full blog here.

Research, reports and deep-dives

  • Report: Google was named a Leader in the inaugural GartnerⓇ Magic Quadrant™ for AI Infrastructure, positioned highest for ‘Ability to Execute’ and furthest for ‘Completeness of Vision’. Gartner called out Google’s proprietary scalable compute, integrated AI Hypercomputer architecture, and the scale of our AI compute capacity as key strengths. Download a copy here.

  • Report: We recently surveyed more than 1,400 senior IT leaders for our State of AI Infrastructure report, and a resounding pattern emerged: The gap between AI ambition and infrastructure reality is widening. In fact, 83% of organizations say they require infrastructure upgrades to support production-grade agentic AI. Read the accompanying blog to understand how adapting your infrastructure to meet the demands that agentic applications place on your systems will help you move from pilot to production.


June 2026

Product, technology and tool updates

Practitioner guides and how-tos

  • How-to guide: Learn how to build high availability into an AI inference workload running on GKE Inference Gateway with TPUs, Cloud Storage FUSE and Dynamic Resource Allocation (DRA). This blog provides an overview, or you can get all the technical details in the hands-on codelab.

  • How-to guide: Did you know you can connect your AI agents to unstructured data in Cloud Storage via Model Context Protocol (MCP)? In this blog, learn about why would want to do that from three customer examples, then how to do it, choosing either a fully managed service, or a self-managed local server for more customization and control. 

Research, reports and deep-dives

  • Report: According to an independent benchmark report, GKE Inference Gateway outperforms the next leading managed Kubernetes service with 15.7% higher throughput, 92.8% shorter wait times, and 62.6% lower inter-token latency. This performance can be attributed to its use of prefix caching, which optimizes LLM performance by storing the KV cache (activation states) of long, repetitive prompt prefixes. Learn more in the blog. 

  • Architecture deep dive: A closer look at the cold start problem, this time for TPUs and GKE, and how the Run:ai Model Streamer can help change the dynamic. 

Customer and partner updates


May 2026

Product, technology and tool updates

  • Product update: GKE Agent Sandbox is now generally available.

  • New open-source project: Agent Substrate is a new open-source project aimed at continuing to push the limits of agentic infrastructure density

  • New feature: Google AI Edge Portal, a solution for testing and benchmarking on-device machine learning (ML) at scale, now supports benchmarking and debugging on-device LLMs. Read more here. 

  • Product deep dive: We went into depth about Cloud Storage Rapid, a new family of high-performance storage offerings for AI workloads. At launch, offerings include Rapid Bucket (formerly Rapid Storage), a high-performance zonal object storage offering, and Rapid Cache (formerly Anywhere Cache), which accelerates reads on-demand and colocates compute and data for workloads in existing buckets. 

Research, reports and deep dives

Customer and partner updates

  •  

Accelerating automotive innovation with C4A-metal and Panasonic Automotive vSkipGen

As the automotive landscape accelerates toward software-defined vehicles, Cockpit Domain Controllers (CDCs) are becoming the core of next-generation in-cabin experiences. The ability to rapidly develop, test, and validate CDC software in a flexible, hardware-independent environment is critical for innovation and time-to-market. However, physical hardware constraints and the requirement for high-performance graphics present significant challenges for global development teams. 

Panasonic Automotive’s vSkipGen™ addresses these challenges as a next-generation CDC virtualization platform, now validated on Google Cloud’s C4A-metal, our Axion bare-metal offering. By integrating Panasonic Automotive’s advanced Unified HMI™ remote GPU offload technology with support for Android Automotive OS (AAOS) and Android SDV, vSkipGen delivers a robust, cloud-native solution for cockpit software development and validation — empowering teams to innovate without hardware limitations.

At Google Cloud, we provide workload-optimized infrastructure to help ensure the right resources for every task. Similar to the entire Axion virtual machine family, C4A-metal instances are built on Google Cloud’s custom Arm-based Axion architecture. C4A-metal offers 96 vCPUs, two DDR5 memory configurations (384GB and 768GB), and up to 100Gbps of networking bandwidth. It also provides full support for Google Cloud Hyperdisk, including Balanced, Extreme, Throughput, and ML types. And like the rest of the bare metal portfolio, C4A-metal is powered by Titanium, a key component for multi-tier offloads and security that is foundational to our infrastructure. 

High performance for demanding workloads

C4A-metal is particularly well-suited for complex tasks such as creating digital twins of vehicle cockpits where performance must accurately mirror real-world behavior. Traditionally, the transition to software-defined vehicles has relied on expensive and scarce physical prototypes; C4A-metal overcomes this by offering the high performance and hardware-level access of bare metal with the scalability of the cloud. 

Panasonic Automotive leverages C4A-metal to bypass traditional hardware bottlenecks, enabling their teams to execute complex virtualization tasks and accelerate the development of next-generation cockpit software.

"Google Cloud’s Axion Bare Metal has been a game-changer for our vSkipGen™ platform. By providing scalable, high-performance Arm-based infrastructure, C4A-metal allows our teams to develop and test production-intent software in the cloud with behavior that closely matches target automotive hardware. This cloud-to-car bit parity reduces dependence on costly physical prototypes, improves validation efficiency, increases test coverage and accelerates time-to-market for next-generation cockpit platforms.” - Andrew Poliak, CTO, Panasonic Automotive Systems America.

By leveraging vSkipGen and Unified HMI on C4A-metal, automotive manufacturers can now build, test, and validate full AAOS stacks in a hardware-independent, cloud-native environment, moving from physical dependency to scalable digital twins.

1

Figure 1: Unified HMI solution overview

How vSkipGen works: Virtualizing the cockpit with Cuttlefish

Panasonic Automotive’s vSkipGen acts as a digital twin for physical CDC hardware. To provide a hardware-agnostic environment for Android virtual machines, vSkipGen uses components of Android Cuttlefish. At its core, vSkipGen leverages a cloud-optimized Virtual Machine Monitor (VMM) built on crosvm (the open-source, security-focused VMM originally developed for Chrome OS) which utilizes Linux KVM (Kernel-based Virtual Machine) for hardware-assisted virtualization. The VMM backend is implemented in Rust for enhanced security, scalability, and performance. By running the full stack on C4A-metal (see Figure 2), Panasonic lets developers boot a full AAOS image in the cloud, which behaves exactly like the software running in a physical vehicle.

The platform virtualizes all essential peripherals, such as the audio, GPU, sensors, cameras, Controller Area Network (CAN), Bluetooth, and Wi-Fi, using the VirtIO standard. This VirtIO-native approach allows developers to interact with the virtual devices exactly as they would with the physical hardware. Furthermore, vSkipGen seamlessly connects with automotive simulators and software-in-the-loop (SiL) environments for comprehensive scenario and edge-case validation, enabling teams to conduct software validation and run automated test suites without needing early access to physical prototypes.

2

Figure 2: vSkipGen Cockpit virtualization architecture

3

Figure 3: Remote GPU rendering flow

Accelerating graphics on the go with Unified HMI

High-performance graphics is central to the modern driving experience, but rendering it in a virtual environment can be challenging. Panasonic’s Unified HMI solves this by decoupling HMI rendering from specific hardware. A lightweight Unified HMI component operates outside the VM to offload OpenGL ES commands (the specific data being rendered) from the Cuttlefish instance to GPU-equipped compute resources on Google Cloud, which handle the workloads with hardware acceleration. The rendered UI is then streamed to any standard browser using low-latency WebRTC. This helps ensure that global development teams can experience high-fidelity visuals in real time, regardless of their location.

Unified HMI establishes a unified virtual display layer across multiple Electronic Control Units (ECUs) and virtual machines, allowing applications to render to different displays from anywhere within the system.

Benefits for software-defined vehicle development

With C4A-metal and Panasonic Automotive’s vSkipGen, manufacturers building software-defined vehicles can: 

  • Build and validate AAOS-based software in the cloud using Cuttlefish before physical hardware is available

  • Use industry-standard VirtIO to emulate critical CDC devices for robust, production-grade validation

  • Stream interactive cockpit experiences to any browser to support distributed engineering teams

  • Run multiple isolated CDC instances in parallel to support large-scale automated testing and CI/CD pipelines

  • Reduce cost and environmental impact by minimizing the need for expensive physical hardware prototypes, supporting sustainable development practices

  • Enjoy a future-ready architecture built on open standards, crosvm, and Rust for enhanced security, performance, and long-term adaptability

Get started

C4A-metal is generally available worldwide; please refer to the public documentation for additional information. Panasonic Automotive’s vSkipGen with Unified HMI for Google Cloud will soon be available for evaluation access. To explore how these solutions can help you speed up cockpit software development, contact the team at vSkipGenSupport@panasonicautomotive.com. 


Disclaimer
All trademarks, tradenames, and service marks used herein are the property of their respective owners.

  •  

Three lessons in accelerating foundation model upgrades

Have you run into problems migrating your products from one model to the next?

Upgrading to the latest AI models is rarely simple. For engineering teams, model updates whether migrating to an entirely new model or updating to a newer checkpoint within the same model family, like moving from an earlier Gemini version to Gemini 3.5 — often require a slow and costly process of testing, proving quality, and manually evaluating new responses. For most engineering teams, upgrading to a new model checkpoint means months of manual toil to verify performance. And the industry is moving at breakneck pace – since 2023, we’ve announced six major model evolutions, bringing us to Gemini 3.5 today. 

Our team at Google Cloud, Applied ML, has a goal to deliver transformative infrastructure and services that benefit both Google and our customers globally. As part of that, our team built an agentic workflow that completes model upgrades in hours instead of months. 

In this blog, we’ll show you our approach and three lessons you can apply to accelerate your own foundation model upgrades using Gemini Enterprise Agent Platform — our new, comprehensive platform to build, scale, govern, and optimize agents – and Google Antigravity, our primary solution for developers using AI for coding and agent orchestration.

Three lessons in building a flexible agent system

To support different team needs, we had to rethink traditional automation and learned three key lessons along the way: 

  1. Lesson 1: Start with hands-on discovery. First, our engineers worked closely with product teams on real migration problems. This hands-on work helped us identify complex requirements and build our first guidelines for prompt optimization.
  2. Lesson 2: Beware the rigidity of traditional automation. We turned these guidelines into a standard, automated workflow. While this version gave us some quick wins, we soon found that traditional automation was too rigid to handle different data formats and unique edge cases.
  3. Lesson 3: Pivot to a flexible agent architecture. The real progress came when we rebuilt the tool using a flexible agent. Instead of forcing teams into a rigid process, the agent adapted to specific project needs, helping analyze data and test prompts dynamically with a high degree of adaptability.

How our partner teams cut migration time while boosting quality

Our partner team, which manages video translation and dubbing services, had an interesting challenge: their workflow required rewriting translated text so that the spoken duration matched the original video's pacing exactly, without altering the meaning. Historically, this strict constraint required maintaining a fine-tuned model. Their goal was to migrate to the latest out-of-the-box foundation model, guided purely by prompt engineering.

Using this agentic framework, the team provided their ground-truth dataset and baseline prompt. The system autonomously hill-climbed the prompt quality, migrating the service away from the custom stack

Make your own migration workflow with Agent Platform and Google Antigravity

These learnings can be applied by any engineering team looking to accelerate their own model upgrades. If your organization is struggling to keep pace with new foundational models, replacing manual toil with intelligent automation requires treating migration as an agentic workflow.

To build your own automated migration pipeline, follow these steps:

  1. Deploy Autoraters: Pivot from manual human review to model-based Autoraters to evaluate the quality of a new checkpoint at scale and in a fraction of the time.

  2. Build an agentic loop: You can use the Agent Development Kit within Gemini Enterprise Agent Platform to create your agent. 

  3. Automate the orchestration: To make the process even easier, leverage Antigravity to automate the underlying coding and agent orchestration and add in features such as loss reporting or headroom reports. 

By shifting away from a manual, line-by-line engineering task, organizations can reduce infrastructural tech debt and confidently keep pace with the frontier of AI.


This work is the result of collaboration across Google. We thank key contributors: Anthony Green, Chris Lamb, Chungyen Li, Connie Huang, Elaine Han, Elena Erbiceanu Tener, Eugene Ie, Francesca Ciacchella, Igor Karpov, Jeanie Jung, Jose Menendez, Kiam Choo, Lina Sanders-Self, Longfei Shen, Martin Nikoltchev, Mason Ng, Matt Mancini, Paul Zhou, Pedram Oskouie, Samuel Smith, Tom Lawrie, Ye Tian, Zhen Lin

  •