❌

Vue normale

Reçu avant avant-hierCloud Blog

How Google Cloud Networking Supports Your Fluid Compute Choices for AI Workloads

24 septembre 2026 à 15:00

The availability of resources for AI workloads can be challenging across the industry, especially accelerators. This can slow your AI workload deployment if it’s built around a specific type of accelerator. The concept of fluid compute allows you to design your AI deployment with several options based on available resources that can fit your use case.

In this blog, we will explore how Google Cloud networking supports your AI workloads and considerations that are relevant to your choice of accelerator (GPU or TPU), as the backend networking component configuration is not exactly the same.

The resource options

After deciding the type of work you want to achieve with your AI deployment, another important component is the actual hardware to get this done. In this case, we want to run inference for a private LLM, and the target is the NVIDIA B200 GPU family which is available in the A4 VMs (a4-highgpu-8g).

Now we have identified what we want to get done and a possible compute option, but the challenge is: is this available?

To get access to resources, there are several options which include:

  • Dynamic Workload Scheduler (Flex-start VM): Queues workloads until all required accelerator nodes are available at the same time, provisioning them together and running non-preemptibly for up to seven days.
  • Dynamic Workload Scheduler (calendar mode): Enables reserving accelerator capacity 1 to 90 days in advance with guaranteed start and end times, ideal for scheduled pre-training runs and benchmarking.
  • Future reservations: Guarantees access to committed hardware in a specified zone beginning at a specific future date.
  • Flex reservations: Offers short-term commitment windows to secure scarce accelerator nodes without multi-year lock-in.
  • Dynamic node auto-provisioning and ComputeClasses: In Google Kubernetes Engine (GKE), defining multi-family fallback lists within ComputeClasses allows the cluster to automatically attempt provisioning alternative accelerator types if primary pools face regional constraints.
  • Spot VMs: Delivers surplus compute at substantial discounts for fault-tolerant, checkpointed batch jobs.

Read more on this in the blog Never Run Out of Compute: A Practical Guide to GKE Resource Obtainability.

Networking your choices

The networking component of the accelerator varies based on your choice, so let's explore four configurations: standard networking, accelerated GPU networking (TCPX/TCPXO and RoCEv2), TPU networking, and Cloud Run.

Standard networking

  • Supported accelerators: NVIDIA T4 (N1 series), NVIDIA L4 (G2 series), NVIDIA A100 (A2 machine series single-node and multi-node), Cloud TPU v3, and Cloud TPU v5e (single-host/standalone slices).
  • Architecture: Nodes communicate over the primary Virtual Private Cloud (VPC) network using the Google Virtual NIC (gVNIC) over standard TCP/IP.
  • Workload fit: Provides straightforward portability across Google Cloud compute environments, supporting distributed data preprocessing, decoupled pipeline stages, independent inference replicas, and computer vision workloads using standard VPC routing and network policies.

Accelerated GPU Networking (TCPX/TCPXO and RoCEv2)

Distributed training and multi-node inference require specialized multi-rail network fabrics to handle massive parameter exchanges and collective communications.

GPUDirect-TCPX and TCPXO Fabrics

  • Supported accelerators: NVIDIA H100 (A3 High VMs with 4 rails) and NVIDIA H100 Mega (A3 Mega VMs with 8 rails).
  • Architecture: Uses custom GPUDirect-TCPX (4 dedicated VPCs) and GPUDirect-TCPXO (8 dedicated VPCs) offload engines to achieve high-throughput multi-rail GPU communication over standard Ethernet infrastructure without requiring native RDMA hardware.
  • Deployment blueprints: These multi-VPC topologies can be deployed in many ways including using pre-built blueprints from the Cluster Toolkit.

RoCEv2 Fabrics (VM and Bare Metal)

  • Supported accelerators: NVIDIA H200 (A3 Ultra VMs), NVIDIA B200 (A4 VMs), NVIDIA GB200 NVL72 (A4X VMs), and NVIDIA GB300 (A4X Max Bare Metal).
  • Zonal network profiles: RoCEv2 operates over a dedicated RDMA VPC attached to a specialized zonal network profile: VM instances (A3 Ultra, A4, A4X) use the ZONE-vpc-roce profile, while Bare Metal instances (such as A4X Max) utilize the dedicated ZONE-vpc-roce-metal bare-metal profile.
  • Rail-aligned fabrics: This dedicated VPC is isolated strictly for GPU communication and contains subnets mapped directly to the accelerator NICs. The backend is rail-aligned, with support for Jumbo Frames (MTU 8896), delivering non-blocking multi-terabit bandwidth with minimal cross-rail interference.
  • Automated plumbing with GKE Dynamic Resource Allocation Network (DRANET): When deploying these GPUs on GKE, the GKE managed DRANET can be used to automatically provision additional networks and assign drivers that map the RDMA network interfaces to the GPU. These can then be assigned and consumed directly in your workload pods using standard Kubernetes resource claims.
  • Turnkey deployment: You can deploy this entire end-to-end stack—including RDMA VPCs, MTU tuning, and DRA drivers—using automated blueprints from the Cluster Toolkit.

TPU Networking

  • Supported accelerators: Cloud TPU v4, Cloud TPU v5p, Cloud TPU v5e (multi-host Pod slices), Cloud TPU v6e (Trillium), and TPU7x (Ironwood).
  • Inter-chip interconnect (ICI): Inside a TPU Pod or slice, chips communicate directly over dedicated, ultra-low-latency optical links organized in 2D or 3D torus meshes, bypassing traditional network stacks entirely.
  • Optical circuit switches (OCS): In TPU v4 and TPU v5p SuperPods, software-reconfigurable OCS units dynamically change physical network topologies, route around faulty trays, and provision custom-sized accelerator slices without manual recabling.
  • Multi-NIC architecture (TPU v6e and Higher): While earlier TPU generations relied on ICI within a slice and single-NIC for host traffic, Cloud TPU v6e (Trillium) and TPU7x introduce a native multi-NIC architecture where worker nodes isolate standard Kubernetes management traffic onto a primary VPC while using secondary dedicated VPCs configured for high-throughput TPU data and cross-slice communication.
  • DRANET for TPU deployments: When deploying these TPUs on GKE, the GKE managed DRANET can be used to automatically provision additional networks and assign drivers for TPU communication. These can then be assigned and consumed directly in your workload pods using standard Kubernetes resource claims.
  • Data-center network (DCN) Multislice: For models scaling beyond an individual TPU slice, Cloud TPU Multislice connects multiple independent ICI meshes over Google's high-speed Jupiter Data Center Network utilizing these dedicated multi-NIC paths.

Cloud Run

  • Supported accelerators: NVIDIA L4 (G2 series) and NVIDIA RTX PRO 6000 (Blackwell) on Cloud Run GPU services.
  • Direct VPC egress: Binds serverless containers directly to your private VPC network using sub-minute IP allocation via Direct VPC Egress, enabling secure, low-latency access to internal data lakes, databases, and private APIs without traversing the public internet or requiring legacy connector VMs.
how-google-cloud-networking-supports-your-fluid-compute-choices-networks

Summary

Google Cloud networking options support various accelerator types. When using fluid compute you can adjust your network setup to support the best design to optimise your workloads performance.

Next Steps

Take a deeper dive into Google Cloud AI infrastructure and networking architectures with these resources:

Want to ask a question, find out more, or share a thought? Please connect with me on LinkedIn.

Google's subsea fiber optics, explained

6 avril 2022 à 19:00

Fiber optic networks are a foundation of the modern internet. In fact, subsea cables carry 99% of international network traffic, and yet we are barely aware that they exist. What you might not know is that the first subsea cable was deployed in 1858 for telegraph messages between Europe and North America. A message took over 17 hours to deliver, at 2 minutes and 5 seconds per letter by Morse code. 

Today, a single cable can deliver a whopping 340 Tbps capacity; that’s more than 25 million times faster than the average home internet connection. Over the years at Google, we have worked with partners around the world to engineer more capacity into the fiber optics that can be underground or laid at the bottom of the ocean. It takes an impressive combination of physics, marine technology, and engineering, which is why I set out to make a video about what it takes to plan and implement a global network in a world of ballooning network demand.

When I started making this video, there were a few questions that I wanted to answer. After 25+ hours of research—from interviewing optical network engineers to exploring our network design documents—I finally started to scratch the surface of a sea of information, so let’s dive into my findings (puns intended)!

How does Google Cloud predict network traffic and plan for capacity?

Physical infrastructure permitting, identifying power sources, and installing cooling and hardware. The entire process for one project can take multiple years to plan and implement. As a result, capacity planning must be done far in advance. It’s difficult to predict capacity needs; a typical trend line analysis won’t work. Given these long lead times and a 20+ year typical life of a cable, our forecasting and asset acquisition decision analysis looks at demand forecasts across a longer time horizon of multiple years rather than months. 

One factor is certain: Cloud is a big growth driver of Google’s network demand, with Gartner predicting the world’s cloud spending to increase to $917B by 2025. Google Cloud has pushed our need to increase the availability and speed of our network and services. We need to handle traffic surges that can stem from Google Cloud customers. We also need to plan for higher network capacity when we add new regions, with redundant pathways to those new locations. Because our Global Networking team wants to deliver capacity when Google Cloud customers need it, we design our network with reliability in mind, consider multiple points of failure, and provide fast failover. 

Forecasting in parts

To forecast capacity needs, we predict demand five years out, three years out, and 3-6 months out using sensitivity analyses. We then determine the size of the cable investment that meets an optimal point on the cost curve–one that balances capacity and cost, while meeting Google Cloud requirements like latency. To help with forecasting, we break our network down into three categories:

  • Inter-metro network – pathways connecting major metropolitan areas, both within a continent and across continents. 

  • Regional – pathways connecting data centers within a metropolitan region

  • Edge – pathways connecting Google’s network to internet service providers (ISPs)

Network Expansion

Engineering teams are obsessed with network resilience 

Dozens of engineering teams work to forecast and design the network. They determine:

  1. The bandwidth needs of individual Google services (for example, a Google Cloud service, Search, YouTube). 

  2. The shape of the network topology to optimize for performance (how various nodes, devices, and connections are physically or logically arranged in relation to each other). 

  3. The number of routes (that is, the number of circuits to each location) needed for high availability. We look for routes that are fully disjointed and diverse to prevent any single points of failure. 

Testing optical fiber to improve capacity planning

The Optical Network Engineering team tests our fiber optic cables to understand how they will perform in the Google Cloud network. These tests play a significant role in capacity planning because the results help us predict what we can deliver.

The goal is to achieve the highest signal to noise ratio. As light travels through fiber over long distances, the signal it carries gets distorted. While we can’t house an entire cable in a lab, we do have dozens of spools of fiber that we daisy chain together to replicate the Google Cloud network. Using an optical spectrum analyzer we check the quality of the signal as we pulse lasers, pushing 1.2 Tbps of 0s and 1s through the cable! 

Spectrum Analyzer

But fiber is not just about lasers, cables, and the laws of physics. Our response to changes and issues in hardware requires robust automation through software. We build automation pipelines to enable us to deploy new fiber to connect our regions. If there are disruptions in the fiber, automation enables us to pinpoint the issue with an accuracy of a few meters and respond rapidly. 

What technology have we developed to increase the reliability and scale of the network?

Space Division Multiplexing

The Grace Hopper cable is breaking records by using space division multiplexing (SDM) to fit sixteen fiber pairs into the cable, instead of the usual six or eight. Once delivered, that cable will be able to transmit 340 Tbps–enough to stream my video in 4K 4.5 million times simultaneously! SDM increases cable capacity in a cost-effective manner with additional fiber pairs while taking advantage of power-optimized repeater designs. With the Topaz cable, we are working with partners to use SDM technology and sixteen fiber pairs to give it a design capacity of  240 Tbps. The 19th century electrochemical scientist Michael Faraday would be proud of us knowing that we could send over 2.4 trillion bits every second across the Pacific Ocean in a single cable.

SDM

Wavelength Selective Switching


Topaz will also use wavelength selective switching (WSS). WSS can be used to dynamically route signals between optical fibers based on wavelength. This greatly simplifies the allocation of capacity, giving the cable system the flexibility to add and reallocate it in different locations as needed. Here’s how it works:

  1. Branching Units are used to split a portion of the cable to land at a different location. Branching can be either at the fiber pair or wavelength level.

  2.  WSS splits the signal between the main trunk and the branch path based on its wavelength. This allows the signals from different paths to share the same fiber instead of installing dedicated fiber pairs for each link.

  3. The cable system can then carve up the spectrum on an optical fiber pair and apply capacity to different locations using a single fiber pair, giving us the ability to redirect traffic on the fly.

WSS

WSS for resilient and dynamic paths was first sketched on a Google whiteboard over four years ago. Now this innovation is being adopted across the industry. 

Why is it so rare for Google Cloud customers to notice when a cable is affected? 

While fiber optic cables are protected, they aren’t immune to damage. Fishing vessels and ships dragging anchors account for two-thirds of all subsea cable faults. 

Fiber

Though the risks are unavoidable, it’s important to remember that fiber optic cables are part of the network backbone of the internet that links data centers and thousands of computers together. That’s why Google maintains an intense focus on building and operating a resilient global network while we continue to advance its breadth, reliability, and availability. For example, it can sometimes take weeks to repair a cable that has been physically damaged. So, to ensure that services aren’t affected in such a situation, we design the network with extra capacity: each cross section has multiple cables and no single point of failure. 

Our philosophy is to create enough concurrent network paths at the metro, regional, and global level—coupled with a scalable software control plane—to support traffic redistribution while minimizing network congestion. When service disruptions occur, we’re still able to serve people around the world because other network paths exist to reroute traffic seamlessly. When a link between the US and Chile becomes disrupted, for example, Google Cloud can reroute traffic for customers on our additional diverse paths. 

How can Google Cloud customers make the most of these advancements?

Network planning, operations, and monitoring are linchpins at Google.

Premium Tier network is the gateway to Google’s high-speed network


To understand the Google Cloud network backbone, it is useful to have an  understanding of how the public internet works. Dozens of large ISPs interconnect at network access points in various cities. The typical agreement between providers involves something called hot potato routing. An ISP hands off traffic to a downstream ISP as quickly as it can to minimize the amount of work that the ISP's network needs to do. This can mean more hops between networks and routers before traffic arrives at its destination. This kind of routing is available on Google Cloud with our Standard Tier network.
Standard Diagram
Click to enlarge

Google Cloud’s Premium Tier network, in contrast, uses cold potato routing. It keeps traffic on its private network backbone, requiring fewer hops between ISPs. It offloads traffic to ISPs at the last possible moment, when the data is closest to the end user. 

Diagram
Click to enlarge

Let’s put this in perspective with an example. If you’re a company using traditional networks, your traffic from your own private data center in the US to its destination in Chile will first traverse your local ISP. That local ISP most likely uses a fiber supplier and passes that traffic off to another ISP. Between the many hot-potato hops, you may face higher latency and limited bandwidth capacity. 

With Google Cloud, you can use our Premium Tier network to achieve 1.4X higher throughput than the Standard Tier, as well as Cloud CDN to cache content closest to your end users. 

The beauty of vertical integration

Let’s not forget the software stack that sits on top of the physical network. The network topology and software-defined controller for traffic routing is built for fault tolerance. Our data center network fabric is made up of a closed hierarchical switching fabric that we designed called Jupiter, which connects hundreds of thousands of machines across data centers, providing 1 Pbps of bisection bandwidth.

Jupiter is able to provide such high bandwidth because it’s nonblocking, which means it can handle routing a request to any free output port without interfering with other traffic. This means it can scale or burst with extremely low latency, and is fault-tolerant. If something in the fabric breaks, it is built to handle disruptions.

Jupiter

Google Cloud is underpinned by Andromeda, a virtualized software-defined network built on top of Jupiter, giving you your very own slice of our massive global switching fabric. Andromeda enables you to deploy a global virtual private cloud network–with the aim of providing you both functional and performance isolation, as well as a high degree of security. With its global control plane, high-speed on-host virtual switch, and packet processors, you can burst thousands of stateful machines online in minutes or deploy firewall rules across thousands of machines immediately without chokepoints. Google Cloud’s global Virtual Private Cloud (VPC), for example, gives you the ability to have a single VPC that can span multiple regions without communicating across the public internet. Because various microservices may be separated and talk to each other through the network, they can scale independently so you get virtually unlimited storage and stateless, resilient compute, along with features like live migration.

Andromeda

Whether you’re processing petabytes of data in seconds using BigQuery, running consistent databases across regions using Spanner, or autoscaling GKE clusters across zones, Google's global network backbone provides the capacity to get the job done. 

It’s been a great journey to deep dive into how Google plans and builds its fiber optic cable network. Every engineer, project manager, and public policy expert I talked to exuded a passion to extend the global connectivity of the internet to the entire world (with Google Cloud being a major catalyst in this endeavor). 

Be sure to check out more ways to catch a ride on the Google Cloud Premium Tier network.

Interested in championing Google Cloud technology with a chance to access exclusive events? Join Google Cloud Innovators. 

Have thoughts about this article? Give me a shout @stephr_wong.

How Airtel delivered its flawless Indian Premiere League 2026 cricket broadcasts

9 septembre 2026 à 18:00

For the millions of fervent fans of the Indian Premiere League (IPL), being able to count on a flawless live streaming cricket experience is never up for debate. For Airtel, producing league TV broadcasts with some of the world's most massive concurrent viewership, dropped packets and buffering are simply not options.

During the IPL 2026 season, Airtel partnered with Google Cloud to manage this digital delivery. Across 74 matches, the streaming infrastructure delivered several hundred petabytes of egress data. The final match alone processed tens of billions of requests, hitting a peak egress of several Tbps.

Delivering video under these concurrency spikes requires an edge architecture designed strictly around localization, paired with proactive operational monitoring. 

Our goal for IPL 2026 was to deliver an uninterrupted, stadium-grade viewing experience to cricket fans across India, regardless of concurrency surges or network conditions. Partnering with Google Cloud and using Media CDN gave us deep local edge proximity and excellent cache efficiency. Combined with proactive match-day real-time monitoring, we delivered a reliable broadcast experience from start to finish.

Architecting for concurrency and edge efficiency

IPL-BLog-Architecture

One of the primary challenges in live sports broadcasting is seamlessly handling large traffic spikes and never degrading stream performance or overwhelming backend origins. That’s especially important when millions of viewers simultaneously tune in during a final over because every millisecond counts.

To accelerate content delivery across India’s diverse ISP landscape, Airtel leveraged Google Cloud’s Media CDN. By utilizing Google’s extensive global edge network, Airtel was able to serve viewer requests from edge locations that were physically close to end users. This deep localization was a cornerstone of the broadcast's success, with 99.9% of all tournament traffic being served locally from within India.

This efficient architecture minimized network hops and reduced transit congestion, translating into remarkable infrastructure and viewer experience metrics throughout the 74 matches:

  • Superior caching efficiency: Airtel saw an overall cache hit ratio exceeding 98%. By effectively absorbing massive viewer traffic load at the edge, origin server/video platform demands remained minimal even during peak playoff viewership.

  • Consistent ultra-low latency: Airtel maintained a p99 latency of < 300 ms during the tournament, which supported fast stream start times and minimized buffering risk during critical game moments.

Proactive strategies for operational readiness

While maintaining an intelligent backend architecture was vital to Airtel’s IPL streaming strategy, it was  only half the equation. Executing high-stakes live broadcasts across 74 consecutive matches also demanded meticulous operational preparation and proactive match-day execution.

Because Airtel and Google Cloud recognized that potential bottlenecks had to be identified long before the first ball, they established a deeply integrated operational support model:

  1. Pre-tournament support readiness reviews: Well ahead of the opening match, joint engineering teams conducted comprehensive support readiness reviews. By auditing traffic projections, reviewing manifest configurations, and validating failover mechanisms early, the teams supported robust client readiness, resulting in low operational friction during the tournament.

  2. Monitoring as a service (MaaS): The teams maintained continuous, proactive telemetry monitoring through MaaS on Media CDN, and real-time observability enabled early detection and mitigation of network shifts before anomalies could impact viewer playback.

  3. Dedicated match-day and weekend support: Live sports don't play by the rules of  standard business hours, so Airtel established comprehensive monitoring protocols for every match. During critical weekend fixtures and the high-stakes playoff stage, Google Cloud’s Technical Account Management and MaaS teams worked hand-in-hand with Airtel engineering to provide dedicated, real-time event support.

A blueprint for live broadcast excellence

Airtel’s successful streaming of IPL 2026 demonstrates that handling extreme concurrency is only possible with an integrated strategy across architecture, edge localization, and operational governance. By combining a 98%+ cache hit ratio with 99.9% local delivery and proactive match-day monitoring, Airtel hit a benchmark for live sports broadcasting at scale.

This deployment provides an overview of the technical architecture and operational strategies involved in scaling live media delivery for high-concurrency events. To learn more about optimizing live broadcasts and edge delivery, review the Media CDN developer documentation.

Simplify your resilience testing strategy with Fault Injection Testing

26 août 2026 à 18:00

When databases fail and network paths falter, you still need your mission-critical cloud services to stay online. Yet guaranteeing high availability has become increasingly difficult because of the complexity of modern distributed systems. 

To help you maintain availability and reliability during adverse events, we’re announcing Fault Injection Testing in preview. Fault Injection Testing is designed to help developers and architects automate failure testing to ensure predictable behavior during disruptions. 

By deliberately introducing faults into your environment, you can verify your safety mechanisms before an actual outage impacts your customers.

Why native resilience testing matters

Unlike in self-hosted data centers, cloud applications offer less direct access to underlying infrastructure to facilitate failover testing.

Without native tools to prove your application can survive a failure, you risk a critical gap in your reliability strategy that exposes you to several risks:

  • Damaged trust and reputation: Frequent failures or poor performance lead to customer dissatisfaction and long-term damage to your brand's image.

  • Compliance and regulatory penalties: For many industries, particularly financial institutions, failing to prove disaster recovery capabilities can lead to non-compliance, audits, and fines.

  • Migration delays: Large-scale migrations often stop when teams cannot verify that critical applications will remain stable during a zone failure.

How Fault Injection Testing works

Fault Injection Testing allows you to run experiments by creating experiment templates. These templates act as blueprints, defining the specific fault to be injected and the resources that will be targeted for the experiment.

In this public preview, you can test two primary failure scenarios:

  • Failover Cloud SQL: This fault triggers a failover of a high availability Cloud SQL instance from the primary zone to a standby zone.

  • Degrade application traffic: This allows you to selectively add latency and HTTP error codes through an Application Load Balancer. 

Before any fault is injected, Fault Injection Testing performs an automated dry run. This read-only simulation checks your permissions and provides an up-to-date list of every resource that will be affected. 

Once you verify the scope, you can manually start the injection. The duration you defined in the template will run its course, and the faults will be reverted at the expiration of the timer.  

During the experiment, you can verify that your application is behaving as you planned.  If things do not go as planned, you can use the stop and revert capability to immediately halt the experiment and begin restoring resources to their normal state.

During preview, we recommend as a best practice to use Fault Injection Testing (FIT) in a non-production environment. Preview is an opportunity to get early access to learn how the service fits and complements your existing testing practices, and to provide us with your feedback to improve the product as well!

Built for the enterprise

Partners like KeyBank and Servier are already using Fault Injection Testing to validate their deployments. By using native fault injection, these organizations can approximate demanding failure scenarios — such as zonal outages — to help ensure their services remain stable.

Get started with Fault Injection Testing

Fault Injection Testing is available through the Google Cloud console, the gcloud CLI, and REST APIs.

  1. Request preview access: Talk to your Google Cloud Account Team to add your project to the preview.

  2. Enable the API: Search for "Fault Testing API" in your Google Cloud console and select enable.

  3. Assign roles: Ensure your team has the roles/faulttesting.operator role to configure and run experiments.

  4. Run your first dry run: Create a template for a Cloud SQL or load balancer resource in a non-production environment and execute a dry run to see the potential impact.

For more details on implementation, talk to your account team, or view the User Guide for Fault Injection Testing.

How Uber improves network reliability while unblocking cloud migration

26 août 2026 à 18:00

Uber has a lot in common with the cities it serves. Both are always changing and growing, both must carefully manage the resulting traffic to prevent congestion and sprawl.

Uber has continuously evolved its technical strategies to manage its expanding network, and this careful planning and constant evolution helps ensure that application traffic across its entire platform runs smoothly. Ultimately, maintaining a reliable, high-scale platform that operates seamlessly at any given time is key to preserving user trust.

One important solution in this effort has been application awareness on Cloud Interconnect. An industry-first tool for application prioritization across hybrid networks, application awareness on Cloud Interconnect has helped Uber prioritize critical traffic to ensure business continuity during potential network congestion events. 

Uber acted as an early design partner for application awareness on Cloud Interconnect, helping ensure that this capability met the demands of Uber’s global-scale operations. It not only improved Uber’s daily operations, it also gave Uber the confidence to move forward with a Google Cloud migration, with confidence that there would be less risk of service interruptions during switchovers. 

In this post, we’ll explain the features Uber most sought and why, the inner workings of application awareness on Cloud Interconnect, and how it can help other organizations as well.

Prioritizing critical traffic

When migrating distributed, hybrid, or multicloud applications at a global scale, network reliability becomes a primary concern. Even the most worthwhile migrations may not seem worth it if such migrations interrupt ongoing service. For organizations like Uber, moving vast amounts of data to support large data analytics workload — including emerging AI use cases — can saturate network links, resulting in increased reliability risk for their critical application traffic. 

With standard cloud interconnect approaches, enterprises typically apply simple bandwidth overprovisioning to meet extreme infrastructure needs. But with today's hybrid cloud demands, and given the size of an organization like Uber, overprovisioning network capacity for peak usage is often too costly and unreliable. 

The shortcomings of overprovisioning only become magnified with the integration of cutting-edge AI innovations. Uber needs systems in place that can take on massive data transfers without congesting its network and protecting the performance of business-critical applications.

With the benefit of application awareness on Cloud Interconnect, including the four major features of application awareness — traffic handling, congestion response, latency management, and cost efficiency — Uber was able to achieve the networking optimization its modern tech stack requires.

aai concept value prop with_without picture

Starting with a private preview, Uber deployed this feature across its infrastructure, beginning with Google Cloud Interconnect deployments in Phoenix, Arizona, and Ashburn, Virginia. Application awareness on Cloud Interconnect allows Uber to classify and prioritize end-user application traffic over less time-sensitive data using DSCP marking and configured queuing profiles.

In the following chart, we look at the four key features of application awareness on Cloud Interconnect, how they differ from legacy approaches, and how they help provide better operational continuity for organizations like Uber. 

Feature

Standard interconnect solutions

Application awareness on Cloud Interconnect

Traffic handling

All traffic treated equally (first-in, first-out)

Traffic classified into six distinct traffic classes

Congestion response

High-priority application traffic may be dropped during bursts

Business-critical traffic is protected via strict priority or bandwidth sharing policies

Latency management

Unpredictable latency for high priority applications

Predictable and consistent low-latency for time-sensitive workloads

Cost efficiency

Requires expensive overprovisioning to absorb peaks

Efficient bandwidth utilization and lower TCO

Uber's key takeaways

For Uber, the business value of being able to prioritize business-critical traffic on its networks by deploying application awareness on Cloud Interconnect was immediate. And in doing so, Uber has also created a blueprint that other enterprises with similar hybrid cloud challenges can replicate. The core elements of that blueprint include:

  • Ensuring business continuity: Uber can decide in real time which application traffic to prioritize during major, high-traffic events. This means that mission critical applications stay up and running during even extreme events (both planned and unplanned). Uber leadership has called application awareness on Cloud Interconnect important for its global operations. 

  • Efficient bandwidth utilization: Instead of blindly overprovisioning bandwidth to prevent congestion, application awareness allows Uber to better utilize their existing Cloud Interconnect capacity aligned with their expected network bandwidth needs. The result is lower total cost of ownership for network infrastructure.

  • Unblocked workload migration: By protecting critical applications from network congestion, Uber was able to migrate significant workloads to Google Cloud and, in the process, dramatically reduce operational overhead.

"Application awareness on Cloud Interconnect was the key that unlocked our ability to migrate more strategic workloads to Google Cloud and is critical for maintaining service reliability during peak global demand. By allowing us to intelligently prioritize traffic, it helps us ensure that we can protect our higher priority services and make our infrastructure more efficient, lowering our total cost of ownership. This wasn't just a feature deployment; it was a deep engineering partnership that delivered a solution critical to our business." – Harry Liu, Director of Engineering, Uber

Securing network reliability for AI and beyond

As more enterprises integrate cloud-based AI models, distributed applications, and data analytics, it's becoming a business imperative to be ready to handle the massive data transfers that follow. But in doing so, they also have to ensure they never compromise the reliability of their critical applications. 

With application awareness on Cloud Interconnect, Uber demonstrated that moving beyond simple bandwidth overprovisioning to protect business-critical traffic was an essential step to building the stability required to embrace modern hybrid and multicloud strategies.

You can read our blog about the potential of Cloud Interconnect across industries to learn more about what the service can bring to your organization, and if you’re ready to explore more, our team of networking and industry experts are ready to help.

ClusterNetworkPolicy in GKE: Balancing control and autonomy for your microservices

10 août 2026 à 18:00

Managing network security in a multi-tenant Kubernetes environment typically requires balancing two distinct needs: developers need their microservices to communicate effectively, while platform and security teams must maintain compliance, prevent lateral movement, and establish cluster-wide guardrails.

Historically, the standard Kubernetes NetworkPolicy has been the primary tool for this. While effective for single-namespace isolation, standard NetworkPolicy is scoped strictly to individual namespaces and designed around developer self-service. When cluster administrators attempt to use it for global security enforcement, it can lead to policy conflicts and operational challenges.

To address this, we introduced ClusterNetworkPolicy (CNP), an open-source standard developed by the Kubernetes SIG-Policy Working Group (WG), to Google Kubernetes Engine (GKE). Designed for scale, CNP is a cluster-wide resource that allows administrators to manage network security centrally, providing a mechanism for those responsible for global security to implement consistent, non-bypassable policies.

Read on for technical details about CNP, some common use cases, an example policy, and how to get started. 

Structuring policies with tiers

A core capability of ClusterNetworkPolicy is its hierarchical tier system. Rather than attempting to reconcile flat, conflicting peer rules simultaneously, CNP establishes a deterministic, top-to-bottom evaluation hierarchy:

  1. The admin tier: The highest precedence level. Rules here are enforced before any other policies.

  2. The network policy tier: The standard namespace level, where developers manage their specific application policies.

  3. The baseline tier: The lowest precedence, establishing the cluster’s default behavior when no other policies apply. This can be overridden using namespace scoped policies.

tiers

This tiered structure helps align network security with organizational roles. Using standard role-based access control (RBAC), you can manage the admin tier to enforce compliance mandates, while platform teams can use the baseline tier to set a default "deny-all" zero-trust posture across the cluster. At the same time, developers can write standard network policies for their applications without overriding core security mandates.

This deterministic, top-to-bottom evaluation method resolves conflicts between different teams' policies. The admin tier introduces an explicit Pass action. This allows security teams to inspect traffic against global rules and then delegate the final Accept or Deny decision down to the developer's namespace policy, facilitating both central oversight and distributed management.

Common network security scenarios

This tiered architecture translates complex security requirements into centralized rules. Here are common scenarios where ClusterNetworkPolicy provides a practical solution:

  • Isolating sensitive workloads: You can apply an admin-tier global deny rule to isolate specific namespaces — such as those used for payment processing or compliance data — from the rest of the cluster. This action overrides any permissive developer policies that might otherwise expose these environments.

  • Protecting core services: To prevent configurations that might disrupt internal operations, administrators can create an admin-tier global allow rule for critical services like kube-dns. This allows these services to remain accessible regardless of any misconfigured namespace policies.

  • Managing external egress: By utilizing IP address range matching, egress traffic can be controlled at the cluster level. This functionality allows you to explicitly restrict or permit access to corporate intranets or external IP ranges, serving as a safeguard against unauthorized data exfiltration.

Example scenario

Consider a common enterprise requirement: Application workloads across all namespaces must be permitted to reach central platform infrastructure (such as shared authentication and telemetry services), while access to sensitive environments — like a restricted vault namespace — is strictly prohibited. Meanwhile, routine microservice traffic is delegated to developer-managed, namespace-scoped policies.

ClusterNetworkPolicy makes this straightforward. A platform administrator simply defines an admin-tier guardrail centrally:

code_block
<ListValue: [StructValue([('code', 'apiVersion: policy.networking.k8s.io/v1alpha2\r\nkind: ClusterNetworkPolicy\r\nmetadata:\r\n name: platform-isolation-guardrail\r\nspec:\r\n tier: Admin\r\n priority: 10\r\n subject:\r\n # Target all application tenant namespaces, excluding system and core infrastructure\r\n namespaces:\r\n matchExpressions:\r\n - key: kubernetes.io/metadata.name\r\n operator: NotIn\r\n values: ["kube-system", "shared-services", "restricted-vault"]\r\n egress:\r\n # 1. Mandate access to central shared platform services\r\n - name: allow-shared-services\r\n action: Accept\r\n to:\r\n - namespaces:\r\n matchLabels:\r\n kubernetes.io/metadata.name: shared-services\r\n\r\n # 2. Enforce strict block on accessing the restricted vault namespace\r\n - name: block-restricted-vault\r\n action: Deny\r\n to:\r\n - namespaces:\r\n matchLabels:\r\n kubernetes.io/metadata.name: restricted-vault\r\n\r\n # 3. Explicitly delegate all remaining traffic to developer namespace policies\r\n - name: delegate-remaining-egress\r\n action: Pass\r\n to:\r\n - namespaces: {}\r\n - networks:\r\n - 0.0.0.0/0\r\n - ::/0'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f2488c1ef50>)])]>

Extending open-source foundations

Instead of building this functionality as proprietary extensions, we worked with the Kubernetes community to design the ClusterNetworkPolicy API (policy.networking.k8s.io), distinguishing it from the namespace-scoped NetworkPolicy API (networking.k8s.io). Furthermore, we collaborated closely with the Cilium community to build its implementation of the API.

Because it is built on open-source standards, GKE helps ensure that security configurations remain portable across different environments. The ClusterNetworkPolicy API natively supports tier selection, enabling clear and deterministic policy evaluation. This approach lets administrators enforce robust security guardrails while maintaining the operational flexibility that development teams depend on.

ClusterNetworkPolicy on GKE elevates workload network security — shifting operations from namespace-scoped rules to unified, cluster-wide governance. It is currently in preview in version 1.36 and later. To learn more and get started, check out:

GOL! How TelevisaUnivision streamed the FIFA World Cup to millions with Google Cloud

7 août 2026 à 18:00

Live sports broadcasting represents the ultimate stress test for digital media infrastructure, where operational success or failure is measured in milliseconds and observed live by millions of viewers simultaneously. During the 2026 FIFA World Cup, the stakes reached a high for TelevisaUnivision, the leading Spanish-language media conglomerate. With Mexico serving as both a primary host nation and a core contender on home soil, fan engagement created unprecedented demand across Latin America and TelevisaUnivision's ViX streaming platform.

For a marquee broadcaster like TelevisaUnivision, high-stakes events carry direct, long-term brand and reputational implications. Audiences demand uninterrupted, pristine access to every critical moment of play. Playback interruptions during a key goal, login latency at kickoff, or degraded stream resolutions immediately impact customer satisfaction, risking subscriber churn and brand dilution. When streaming tier-1 global sports events, technical execution directly impacts consumer trust, requiring an absolute commitment to zero-downtime availability and flawless performance on the part of the provider.

Why TelevisaUnivision selected Google Cloud  

Navigating Latin America's complex networking ecosystem, which is marked by heavy ISP fragmentation and cross-border transit bottlenecks, required more than a standard vendor relationship. TelevisaUnivision needed a strategic partner willing to make joint investments in network capacity, infrastructure resiliency, and custom feature development. TelevisaUnivision selected Google Cloud's Media CDN based on two foundational differentiators: its architecture, and Google Cloud’s customer focus.

Platform architecture 

Media CDN provided a resilient, globally distributed infrastructure built specifically to absorb massive live-stream traffic spikes while protecting origin infrastructure. Core architectural advantages included:

  • In-ISP deep edge caching: Media CDN embedded cache nodes deep within local ISP networks across Mexico and Central and South America, placing video segments within a single network hop of viewers.

  • Direct ISP peering: By establishing direct peering connections with major regional telecommunications operators such as América Móvil and Telefônica, the architecture completely bypassed congested international transit routes.

  • Dedicated capacity reservations: TelevisaUnivision reserved live event capacity with dedicated allocated headroom in-region. This isolated livestream traffic from "noisy neighbor" risks and comfortably absorbed peak traffic surges.

  • Sub-millisecond sessions with Valkey 9.0: To handle massive traffic spikes during the World Cup, TelevisaUnivision migrated its session store to Memorystore for Valkey 9.0, operating as serverless microservices on edge compute. This architecture delivered sub-millisecond response times for critical authentication and entitlement checks, while providing automatic scaling to process peak API traffic without the need for manual capacity reservations.

Worldcup diagram

Obsession for customer success

Beyond technical capabilities, TelevisaUnivision chose Google Cloud for its joint co-engineering model and deep operational alignment.

  • Joint 24/7 war rooms: For all 104 matches, TelevisaUnivision engineers and Google Cloud specialists operated side-by-side in unified command centers.

  • Proactive monitoring as a service: Google Cloud’s Customer Reliability Engineering teams provided round-the-clock proactive monitoring and automated alerting.

  • Joint operational authority: Combined teams performed extensive pre-tournament stress tests and simulated failovers. During live matches, unified telemetry empowered joint leads to dynamically route traffic and adjust CDN configurations instantly as regional ISP congestion emerged.

Summary

The strategic partnership between TelevisaUnivision and Google Cloud during the 2026 FIFA World Cup established a new benchmark for global sports broadcasting. Across 39 consecutive days of tournament execution, TelevisaUnivision reported that the joint infrastructure delivered:

Total matches broadcast

104 live matches

Platform availability

100% platform availability (0 downtime)

Cumulative viewership

675 million views across TelevisaUnivision & ViX

By uniting localized edge delivery, sub-millisecond serverless compute, and dedicated operational co-engineering, TelevisaUnivision and Google Cloud solidified a battle-tested blueprint for executing marquee live streaming events at record global scale.

What’s new in AI infrastructure and orchestration in August

31 août 2026 à 18:00

Welcome back to What’s new in AI infrastructure and orchestration this month, a collection of product updates, how-tos, customer stories, research and other resources about all the AI compute, networks, storage, frameworks, and orchestration software that you can find at Google Cloud. To be honest, we thought August would be a slow month, but nothing could be further from the truth. Read on and you’ll see what we mean.

August 2026

Product, technology, and tools updates

  • Product update: Filestore, Google Cloud’s first-party, secure, scalable NFS file service, has emerged as a popular storage platform for AI and agentic workflows, and now, it’s even better suited to the task, with a new backend storage layer built directly on Colossus, Google’s foundational distributed storage system. This new backend lets you provision IOPS independently from storage capacity, and is deeply integrated with GKE. In AI environments, this can help you service so-called agentic swarms — large groups of agents that need to read and write to a common dataset — without a drop off in performance. For more, check out the blog post. 

  • New feature: gVisor sandboxes are now available in distributed Ray clusters on GKE. In partnership with Anyscale, we introduced an experimental library for Ray that brings gVisor, Google’s open-source application kernel, directly into distributed Ray clusters. gVisor provides lightweight environments with stronger isolation than ordinary containers, plus fast startup times and low memory overhead. To try out these sandboxing capabilities on GKE, head over to the Ray sandboxing User Guide.

  • Product update: Looking for high-performance, easy-to-use infrastructure on which to run a personal AI agent, but don’t want to spend a lot of money? New Cloud Run instances are dedicated, singleton compute runtimes on Cloud Run that won’t shut down when the agent is idle. Better yet, the cost to run a Cloud Run instance with 1 vCPU and 1 GiB of memory continuously for 30 days is just $5.70.  

Practitioner guides, documentation and how-tos

  • How-to guide: Big news in Model Context Protocol (MCP) land: As of the 2026-07-28 specification, the protocol core is “completely stateless. The handshake is gone. The initialize / initialized handshake (SEP-2575) and the logical Mcp-Session-Id header (SEP-2567) have been removed entirely. Instead, every request is now self-describing and independent.” Whoa. Learn more about the changes that the latest MCP specification brings, and more importantly, how to implement them, in this Google Developers blog.  
  • Guide: Real-time AI systems make a mess of traditional network load balancing techniques. “Instead of handling isolated requests, the backend has to manage a continuous, live bidirectional stream. You’re dealing with a constant stream of audio chunks, transcripts, model outputs, and synthesized speech flowing back and forth simultaneously.” Things only get worse when the user gets involved. “The server has to immediately halt its current speech generation, pivot to update the context, maybe trigger a new tool, and start drafting a different response; this must be done without dropping the connection.” For a new approach to managing load in the AI era, read Scaling real-time AI agents with session-aware load balancing.
  • How-to: Learn how to build an elastic, scalable LLM inference platform on GKE, even with a mix of different GPU accelerators. The proposed architecture combines Capacity Advisor and Compute Advisor, plus high-performance storage like RunAI:model streamer or GCPFuse with parallel downloads. Get all the details here.
  • Documentation: The thing about hosts with GPUs or TPUs is that you can’t use live migration to update them, setting up a maintenance challenge. In this new docs page, learn how to update accelerator-equipped hosts according to your tolerance for downtime for your training and inference workloads.    
  • Documentation: Advanced Compute Images, or ACIs, are standardized image stacks for AI/ML and HPC infrastructure, so you don’t need to manually build your own custom images. In this new docs page, learn how to create an ACI image using the Google Cloud CLI, console, or SchedMD's Slurm workload manager. 
  • Guide: AI workloads are notoriously difficult to architect, resource-intensive, and bursty, which can also lead to scaling bottlenecks and large pools of underutilized — or misutilized — compute resources. A new blog outlines the three main ways to achieve dynamic capacity management in Google Cloud: 1) scheduling capacity for planned downtime; 2) maintaining automated fallback capacity for unplanned downtime; and 3) relying on GKE’s core orchestration capabilities to automate resource allocation. 

Customer and partner updates

  • Business orchestration software provider UiPath was dealing with spiky workloads, and wanted more predictable costs. To get there, it re-architected its infrastructure, moving from isolated clusters to a shared Google Cloud GPU fleet that included both A3 VM instances (NVIDIA H100 GPUs) for training with G4 VM instances (NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs) for inference. You can read more about their architecture here. 

  • Mirendil, an frontier AI lab focused on accelerating AI development, announced that it is using AI Hypercomputer with both TPUs and NVIDIA GPUs to support its model pre-training and post-training applications. 

  • Replenit, a retail CRM provider, built its AI decision engine in Google Cloud, using BigQuery, Gemini Enterprise Agent Platform, and open-source Gemma models that it runs on Cloud TPUs. This latter combination provided Replenit with 90% lower pipeline costs than their previous cloud provider, the company reports. Read the full case study for more. 

  • Malachyte architected its AI-powered e-commerce recommendation platform on top of Bigtable, Managed Service for Apache Kafka, Pub/Sub, Compute Engine, and last but not least, GKE. See how it all comes together in this blog.


July 2026

Product, technology, and tools updates

  • Product update: Google Cloud Managed Lustre is now GA, and available in four distinct performance tiers that deliver throughput ranging from 125 MB/s, 250 MB/s, 500 MB/s, to 1000 MB/s per TiB of capacity — with the ability to scale up to 8 PB of storage capacity. The Managed Lustre solution is powered by DDN’s EXAScaler, combining DDN's decades of leadership in high-performance storage with Google Cloud's expertise in cloud infrastructure.

  • Product update: C4N network and storage optimized VMs are now GA. C4N is our first network- and block-storage-optimized VM series built to eliminate data-transfer bottlenecks. Powered by 5th Gen Intel Xeon Scalable processors and built on Google's Titanium offloading hardware, it achieves 400 Gbps network bandwidth, 95 million packets per second (MPPS), and up to 25 GiB/s of block storage throughput when paired with Hyperdisk Extreme.

  • New feature: GKE Dataplane V2 up to 15K Nodes with Network Policies (GA). This capability enables standard GKE clusters to scale up to 15,000 nodes while maintaining full active Network Policy enforcement, supporting the massive infrastructure needs of large enterprise and AI/ML customers.

  • New feature: Co-operative time-slicing in llm-d. If you’re running reinforcement learning (RL) workloads, you can now interleave independent RL jobs onto shared physical hardware, increasing aggregate accelerator duty cycles from a ~40% baseline up to 70% without impacting model convergence or accuracy. 

  • New AI security tool: Looking to secure your AI supply chain on GKE, deploy AI workloads safely, and cut down on shadow AI? We open-sourced k8s-aibom, a lightweight, unprivileged Kubernetes controller that continuously monitors container clusters to automatically detect running AI runtimes (like vLLM and Triton) and generate standard CycloneDX Machine Learning Bill of Materials (ML-BOMs). Check out the k8s-aibom project and get involved.

Practitioner guides and how-tos

  • How-to guide: On July 27, Google announced Day 0 support for Moonshot AI’s Kimi K3 2.8-trillion-parameter open-weight model, the day weights were released. Whichever your preferred deployment path — via Model Garden, custom orchestration, or GKE with llm-d recipes — this guide offers detailed step-by-step instructions to help you evaluate and pilot Kimi K3 in Google Cloud. 

  • How-to guide: Google Kubernetes Engine (GKE) managed DRANET supports both GPUs and TPUs. There are several configurations to use this implementation, including standard cluster (where you have full control) and autopilot cluster (where Google does the heavy configs for you). Take a deeper dive in the hands-on lab, GKE Autopilot clusters with TPUs, GKE managed DRANET and Gemma 4.

  • How-to guide: Learn to run Ray on TPUs, not GPUs. In Part 1 of this two-part series, we discuss TPU slices (hint: Ray thinks of them as just another accelerator on which to schedule), then walk through Ray’s various AI libraries (Part 2).

  • How-to guide: Evaluate TPUs for sample workloads using a new microbenchmark suite that helps you accurately assess whether a device is achieving its theoretical performance specifications, and to identify specific performance gaps or architecture-specific bottlenecks. Dive in here. 

  • How-to guide: Scale your agents without killing your budget. Learn how GKE orchestration can help you safely pack more agents onto a fixed compute footprint with GKE Agent Sandbox and Pod snapshots. Whether your goal is performance or cost optimization, we teach you how to turn the right dials for optimal agent efficiency. 

  • Technical blueprint: Inside the optimization of Mistral 3 large inference on Ironwood. This blog outlines how one Google team optimized Mistral 3 large MoE model inference on Google’s Ironwood (TPU v7x), achieving a 1.5x performance gain. They did so with hybrid sharding, replacing linear VPU summations with tree reductions, optimizing GMM/MLA kernels, and adopting asynchronous scheduling. As a result, they boosted throughput by up to 48% while maintaining benchmark accuracy neutrality. Read the full blog here.

Research, reports and deep-dives

  • Report: Google was named a Leader in the inaugural GartnerⓇ Magic Quadrant™ for AI Infrastructure, positioned highest for ‘Ability to Execute’ and furthest for ‘Completeness of Vision’. Gartner called out Google’s proprietary scalable compute, integrated AI Hypercomputer architecture, and the scale of our AI compute capacity as key strengths. Download a copy here.

  • Report: We recently surveyed more than 1,400 senior IT leaders for our State of AI Infrastructure report, and a resounding pattern emerged: The gap between AI ambition and infrastructure reality is widening. In fact, 83% of organizations say they require infrastructure upgrades to support production-grade agentic AI. Read the accompanying blog to understand how adapting your infrastructure to meet the demands that agentic applications place on your systems will help you move from pilot to production.


June 2026

Product, technology and tool updates

Practitioner guides and how-tos

  • How-to guide: Learn how to build high availability into an AI inference workload running on GKE Inference Gateway with TPUs, Cloud Storage FUSE and Dynamic Resource Allocation (DRA). This blog provides an overview, or you can get all the technical details in the hands-on codelab.

  • How-to guide: Did you know you can connect your AI agents to unstructured data in Cloud Storage via Model Context Protocol (MCP)? In this blog, learn about why would want to do that from three customer examples, then how to do it, choosing either a fully managed service, or a self-managed local server for more customization and control. 

Research, reports and deep-dives

  • Report: According to an independent benchmark report, GKE Inference Gateway outperforms the next leading managed Kubernetes service with 15.7% higher throughput, 92.8% shorter wait times, and 62.6% lower inter-token latency. This performance can be attributed to its use of prefix caching, which optimizes LLM performance by storing the KV cache (activation states) of long, repetitive prompt prefixes. Learn more in the blog. 

  • Architecture deep dive: A closer look at the cold start problem, this time for TPUs and GKE, and how the Run:ai Model Streamer can help change the dynamic. 

Customer and partner updates


May 2026

Product, technology and tool updates

  • Product update: GKE Agent Sandbox is now generally available.

  • New open-source project: Agent Substrate is a new open-source project aimed at continuing to push the limits of agentic infrastructure density

  • New feature: Google AI Edge Portal, a solution for testing and benchmarking on-device machine learning (ML) at scale, now supports benchmarking and debugging on-device LLMs. Read more here. 

  • Product deep dive: We went into depth about Cloud Storage Rapid, a new family of high-performance storage offerings for AI workloads. At launch, offerings include Rapid Bucket (formerly Rapid Storage), a high-performance zonal object storage offering, and Rapid Cache (formerly Anywhere Cache), which accelerates reads on-demand and colocates compute and data for workloads in existing buckets. 

Research, reports and deep dives

Customer and partner updates

IDC: Why the right networking approach is foundational to agentic AI

15 juillet 2026 à 18:00

Editor’s note: Today we hear from IDC on the results of its 2026 AI in Networking Special Report Survey exploring the enterprises' concerns about networking infrastructure to support the rise of agentic AI in their organizations. The survey was sponsored by Google Cloud.


Enterprises are moving quickly on AI pilots, but the move from pilot to production remains uneven. While AI models remain important, IDC research indicates that the pilot-to-production bottleneck is primarily infrastructure-centric, with core networking concerns emerging as one of the leading drivers of AI project delays and abandonment. In IDC's 2026 AI in Networking Special Report Survey:

  • 32.6% of respondents cite security concerns: As AI workflows become more distributed and autonomous, enforcing consistent security and governance becomes more difficult.

  • 26.8% of respondents cite challenges in automation: Manual operations and fragmented controls can slow deployment and make AI environments harder to scale.

  • 24.7% of respondents cite staff time and talent restrictions: Limited skills and operational bandwidth can constrain an organization's ability to move AI initiatives into production. 

Agentic AI specifically heightens these concerns by introducing more distributed and dynamic interactions across applications, services, APIs, tools, and data sources. In production environments, these interactions often span different agent frameworks, model providers, clouds, open-source tools, SaaS APIs, and internal applications, expanding both the operational scope and the security and governance surface area. 

Networking for operational control, security, and governance at scale

Networking is the primary enabler of agentic interactions and plays a foundational role for intracloud and intercloud network- and services-layer connectivity, end-to-end security, and consistent governance. In agentic systems, networking increasingly extends into tighter service-centric controls that govern how distributed services identify one another, communicate, and exchange data securely. While AI workloads in general are increasing east-west traffic demands, agentic AI adds an additional layer of complexity by creating dynamic interactions that require tighter policy, visibility, and control closer to the application workflow.

From an infrastructure perspective, networking is much more than just a connectivity function. It is part of the infrastructure platform control plane that applies policy-based controls, supports observability, and helps maintain consistent security and governance across an AI agent's activity. This is significant because framework-level controls alone become insufficient in environments where agents and services span different runtimes, clouds, deployment models, and operating domains.

That is why an infrastructure-level approach becomes key. It does not replace application frameworks or orchestration environments, but it provides broader and more consistent policy implementation across a complex architectural landscape. As agentic AI becomes more autonomous and distributed, organizations need these controls built in as part of the infrastructure to reduce fragmented observability, inconsistent policy application, and unmanaged shadow agent activities. From a cloud infrastructure standpoint, this is where cloud network services become strategically important.

Balancing act: A platform vs. best-of-breed approach

Agentic AI systems are inherently fragmented because of underlying distributed workflows. Enterprises are already navigating a rapidly evolving landscape of business requirements, open-source components, emerging protocol standards, and new architecture patterns. In this context, choices between best-of-breed point solutions and platform-based approaches should be strategic rather than ideological.

Best-of-breed capabilities may be necessary to address specific technical requirements. But it is also true that point solutions introduced across a distributed agentic AI landscape can create inconsistent policies, operational complexity, and governance gaps. IDC research reflects this tension. In IDC’s 2026 AI in Networking Special Report Survey, organizations remained divided between platform and best-of-breed preferences for AI workloads; among respondents who favored platforms, the main reasons cited were stronger security (32.9%), reduced complexity (27.7%), and faster deployment (24.2%).

In IDC's view, a balance is important. Platforms can provide a consistent operational and policy foundation for AI deployments, but at the same time, they need to be modular and extensible to allow the inclusion of best-of-breed functionality as part of the platform toolset. The right platform for agentic AI should be open, flexible, and able to evolve. It should support integration with third-party and open-source tools, allow insertions of needed security and observability functions, and adapt without complete architectural rework.

This is a period of technology disruption. Businesses must meet their AI objectives while carefully managing dynamic agentic AI systems. In this environment, networking not only remains a connectivity piece of the AI infrastructure but becomes foundational to how organizations establish operational control, apply policy consistently, and maintain end-to-end trust across agentic workflows. 

As agentic AI systems continue to evolve, the demands they place are unlikely to be addressed through best-of-breed point solutions alone. Operationalizing agentic AI at scale will require organizations to leverage the right networking approach, supported by infrastructure platforms that are open, flexible, and extensible, enabling a cohesive and adaptable security and governance framework.

Message from the sponsor
The autonomous and non-deterministic communications of agentic applications pose challenges for which the infrastructure and governance models of the cloud-native era are not prepared. In the agent-native era, an infrastructure-led approach is required to enable agentic applications at scale in production with effective governance and observability. An extensible platform based on open standards is critical in enabling the agentic journey today and through its maturity. Learn about the infrastructure imperatives and open standards that make a viable agentic infrastructure here.

❌