❌

Vue normale

Reçu avant avant-hier

Storage Intelligence advisor: Know what changed in your storage estate and act on it

25 septembre 2026 à 18:00

The volume of data being generated today brings both opportunity and massive operational complexity. Most teams that operate at scale don't discover issues until they appear on an invoice — and by the time an unusual access pattern shows up as a line item, it has often been running for weeks. Understanding what happened means exporting inventory, joining it against access logs, and hoping someone still remembers which service account belongs to which job.

That workflow was manageable in the past, but today’s AI training and inference pipelines create data faster than governance systems can classify it, and read data in patterns that shift from week to week. 

Today we're announcing two new features for Google Cloud Storage: the general availability of Storage Intelligence advisor along with expanded capabilities in storage batch operations. Advisor tells you what changed in your storage estate and what to do about it. Batch operations can carry that decision out across millions of objects. These features are available now to all Storage Intelligence customers.

image1

Storage Intelligence advisor in cloud console.

Storage Intelligence advisor makes reporting easy

For the last decade, answering "what’s in my buckets?" has been a data engineering project. Export your inventory, load it somewhere queryable, join it against usage, build dashboards, and then maintain them. Storage Intelligence delivers visibility without the engineering overhead. Teams are voting with their workloads: the number of customers using Storage Intelligence to analyze datasets of over 1 billion objects has more than doubled this year. 

There are two ways to run a large storage estate. Teams can leverage daily activity data and metadata snapshots to build exactly the pipelines they need –Storage Intelligence still gives you that option – but most teams would prefer not to build pipelines if they don’t have to. They want to be told what changed in their storage environment and what to do about it. Storage Intelligence advisor is for them.

What Advisor gives you on day one

Storage Intelligence advisor brings visibility into your storage without having to perform any setup. Advisor starts from a curated set of findings. There's no schema to design, no pipeline to manage, and no dashboard to assemble. Enable Storage Intelligence on an organization, folder, or project, and charts and findings appear for the buckets in that scope. 

Shipt can now more quickly detect anomalies with Storage Intelligence advisor:

"Before Storage Intelligence advisor, tracking critical usage metrics and catching anomalies [in Google Cloud Storage] required heavy engineering and complex data pipelines. Now, with native, out-of-the-box dashboards, we can instantly identify usage spikes and drill down into the details. Having the visibility to immediately remediate unintended usage — without any configuration — has turned what used to be a major effort into a simple, self-service task." - Charley King, DataOps-DevOps Engineer, Shipt (a subsidiary of Target.com)

Once it’s installed, Advisor immediately starts analyzing the Cloud Storage estate, scanning for anomalies and optimization opportunities including:

  • A spike in Class A or B operations against Coldline or Archive data. Cold storage is cheap to use but expensive to access.

  • A spike in 429 errors. Where a request pattern is outrunning limits, timeouts follow.

  • A spike in cross-region egress.

  • Total consumption rising above a long-term trend.

Each finding is baselined from your project's own activity and metadata and works from daily snapshots of your storage usage, so a spike on one day is surfaced within 24 hours, not a line item you discover at the end of the month. In the last 30 days, over 6,000 findings have been generated across hundreds of customers. 

Take a runaway analytics job that issues millions of daily reads against Archive storage. Without Storage Intelligence advisor, this surfaces as a retrieval-fee weeks later on a bill. 

Advisor identifies the anomaly against your project’s baseline, attributes it to the responsible bucket, prefix, and service account, and points at the controls that apply: bulk-transition the affected objects to Cloud Storage Standard to stop retrieval charges, enable Autoclass so tiering follows real access patterns, or tighten access with Managed Folders so the job can’t reach data it was never meant to access. 

Act on findings with storage batch operations

Most storage recommendations go unactioned because carrying them out is a lot of work. Updating retention policies or storage classes across billions of objects means handling throttling, partial failures, and retries. Storage batch operations removes that work. Execution is fully managed and serverless, with progress tracking and automatic retries built in, so a recommendation becomes a policy-driven job rather than a project. 

Palo Alto Networks had this to say about batch operations:

"Object retention locks were essential for our security guardrails, but managing them across billions of objects was once a non-starter. Storage Intelligence changed that. Today, our team uses storage batch operations to seamlessly update retention policies on demand across our entire fleet." - Kurtis Nusbaum, Senior Principal Software Engineer, Palo Alto Networks

Because Storage Intelligence advisor and batch operations are part of the same Storage Intelligence subscription so customers can now quickly identify issues with Advisor and easily remediate those issues with batch operations.

Batch operations enables the following:

  • Remediating operational spikes: Bulk-transition high-traffic Archive or Coldline objects to Standard as soon as the pattern is detected, curbing retrieval and operation charges immediately.

  • Containing runaway growth. Mass-delete stale or temporary data across specific prefixes when the advisor flags above-trend storage growth.

  • Enforcing fleet-wide consistency. Apply metadata, tagging, retention, or encryption changes uniformly across massive object sets, with no dedicated compute to provision.

We also expanded and enhanced the existing capabilities of batch operations, making it easier to execute actions at scale:

  • Multi-bucket processing: Run a single job across up to a thousand buckets per project, rather than executing it bucket-by-bucket.

  • Dry-run validation. Simulate your transformations using dry-run mode before modifying live data. A dry run helps you safely preview a job's impact (including affected object counts, total size, and potential errors) before committing to permanent changes.

  • Advanced filters powered by Storage Insights datasets: Use Common Expression Language (CEL) expressions to select objects directly by specifying conditions that match fields in Insights datasets. For example, you can filter objects across your buckets by storage class, object size, creation date, or custom attributes.

Below is a CLI example demonstrating how to create a batch operations job using advanced filters. This job deletes all temporary objects belonging to the Standard storage class present in a user's "analytics" buckets.

code_block
<ListValue: [StructValue([('code', 'gcloud storage batch-operations jobs create bulk-delete-temp-objects \\\r\n --description="Bulk delete temporary objects in analytics buckets" \\\r\n --target-project="my-project-id" \\\r\n--insights-dataset-config="projects/my-project-id/locations/us-central1/datasetConfigs/my-dataset" \\\r\n --bucket-filters="name.startsWith(\'analytics-\')" \\\r\n --object-filters="storageClass == \'STANDARD\' && name.endsWith(\'.temp\')" \\\r\n --delete-object'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f522ded1990>)])]>

The evolution of storage management

Storage management shouldn’t be a reactive effort reserved for quarterly reviews and post-incident fire drills. It should be continuous, proactive, and contextual.

Storage Intelligence advisor and batch operations help to surface what changed and enable insights and action at scale. As Storage Intelligence gets better at recognizing which findings matter, Cloud Storage can carry more of the operating load for teams that need to manage storage at scale.

Storage Intelligence advisor and enhanced storage batch operations are generally available today.

To get started, enable Storage Intelligence on a project or org. If you haven't used Storage Intelligence before, a 30-day trial is available at no cost.

Introducing Filestore agent volumes: fully managed storage for agent workspaces

15 septembre 2026 à 18:00

From running build tools, to data analysis pipelines, to collaborative research, executing data-driven tasks is essential for any enterprise agent. 

Today, platform teams often stitch together custom workarounds to address agent storage requirements, which could include shuttling state back and forth between agent sandboxes and centralized storage or manually managing local disks and/or self-hosted file systems. However, as agent fleets scale, these approaches force difficult trade-offs between cold-start latency, operational complexity, and the cost of idle, pre-allocated storage.

As organizations scale agent sandboxes to thousands or even millions of concurrent sessions, storage must evolve to overcome these trade-offs and meet the needs of these dynamic workloads, which require strict workspace isolation, instant session resumption, elastic pay-per-use economics, and fluid multi-agent collaboration.

To meet these emerging demands, we’re expanding our AI storage portfolio and announcing availability of Filestore agent volumes, a new, fully managed capability purpose-built to deliver high-performance, elastic file storage for scaling agentic workloads on Google Cloud.

Purpose-built storage for AI agent workspaces

Autonomous agents require isolated runtime environments to safely execute dynamic code, install third-party packages, and run tools without putting host infrastructure or tenant data at risk. While Agent Substrate on GKE and GKE Agent Sandbox provide the dedicated compute environments needed to run high-density agent fleets, those sandboxes also need dedicated persistent workspaces to operate on.

Filestore agent volumes within Google Cloud Filestore, give you purpose-built agentic storage to complement your agentic compute via a dynamic provisioning architecture designed specifically for the scale and elasticity of AI agent fleets. Co-designed with Agent Substrate to support agentic fleets at scale, Filestore agent volumes provide GKE sandboxes with instantaneous access to isolated, persistent file storage. When configured to leverage Filestore, GKE storage management happens behind the scenes: Every time GKE launches a sandbox for a new agent task, Filestore automatically allocates and attaches a dedicated, isolated file workspace to that environment in milliseconds. Platform teams don't need to manually create, attach, or tear down storage volumes for individual agent runs; instead, the system handles the entire volume lifecycle automatically as your agent fleet scales up and down. The result is an efficient, end-to-end infrastructure solution for cost-effective agent management that provides: 

  • Granular isolation and enterprise guardrails: Agent platforms face security and data leakage risks when running untrusted, autonomous code. Filestore agent volumes enforce strict boundary controls and granular access permissions per workspace, ensuring agents operate exclusively within their designated directories and keeping dynamic toolchains strictly isolated across tenants.

  • Sub-second session resumption: Traditional storage provisioning approaches can introduce cold-start latency that stalls interactive agent sessions. Agent volumes attach and detach in milliseconds, making it possible for orchestrators to aggressively suspend idle sandboxes to save compute costs, and resume instantly when new tasks or user inputs arrive.

  • Smart lifecycle economics and pay-per-use pricing: Pre-allocating fixed-size, high-performance storage for thousands of short-lived or intermittent agent tasks can create massive storage waste. With agent volumes, platforms pay only for the storage capacity consumed and benefit from automatic lifecycle tiering. This means you get high performance without wasted spend: When your agents aren’t actively reading/modifying code or analyzing datasets, you can automatically shift idle workspace state to lower-cost storage.

  • Multi-agent collaboration: Coordinating multi-agent swarms can result in brittle data-passing pipelines and risk of file collisions. Built with native Read-Write-Many (RWX) support and POSIX file locking, agent volumes allow orchestrators to attach a single shared workspace across multiple agents. Collaborating agents can safely co-author, test, and review project files concurrently with file-level consistency and protection against write conflicts.

Powering next-generation agentic workloads

By providing an elastic, high-performance, and isolated file tier, Filestore agent volumes unlock a wide spectrum of agentic workloads and use cases in production:

  • Software engineering and coding sandboxes: Agentic coding platforms can spin up thousands of isolated workspaces where agents safely install libraries, write multi-file patches, run build tools, and execute unit tests, all leveraging standard POSIX file semantics with no need for storage-specific customization.

  • Collaborative multi-agent swarms: Complex workflows, such as a lead orchestrator delegating tasks to dedicated research, code generation, and validation sub-agents, can directly share a unified file tree. RWX support allows agents to co-author and review project files concurrently without write conflicts.

  • Interactive long-horizon workflows: For user-in-the-loop applications (such as agents that require asynchronous user approval or run multi-hour data analysis pipelines), platforms can suspend idle agent sandboxes to minimize compute waste, then resume execution on demand with sub-second responsiveness.

Get started today 

If you are building an Agent-as-a-Service platform, scaling coding assistants, or deploying enterprise agent fleets, your storage tier should accelerate your innovation — not hinder it.

Filestore agent volumes are now available to all Google Cloud customers for non-production workloads. GA support for production workloads is available via allowlist. This new offering features out-of-the-box integrations with Agent Substrate on GKE and GKE Agent Sandbox to help you build responsive, scalable, and cost-efficient agent platforms today.

To request access to Filestore agent volumes, submit this form and visit the Filestore documentation and GKE documentation to learn more.

Unlocking the future of shared storage: Filestore on Colossus

5 août 2026 à 15:00

Today,  enterprise storage must be as agile, elastic, and responsive as the workloads it supports. Filestore, Google Cloud’s first-party, secure, scalable NFS file service, can service a wide-range of enterprise use cases as well as cutting-edge AI and agentic workflows. Today, we’re sharing a major enhancement: Filestore now incorporates a cloud-native backend storage layer built directly on Colossus, Google’s foundational distributed storage system. Leveraging Colossus, Filestore can deliver even greater flexibility and scalability to support the most demanding modern workloads.

Colossus: A foundation of global scale

As a first-party service, Filestore is positioned to leverage the best of Google’s infrastructure-level innovation. By leveraging Colossus, Filestore now utilizes the same infrastructure DNA that powers Google’s biggest global services, including YouTube, Gmail, and Gemini. This platform-level upgrade enables superior scalability and operational efficiency compared to legacy VM-based architectures, providing a robust foundation for your data.

This also allows Filestore to decouple storage capacity and performance. Now you can provision IOPS independently to precisely meet workload demands without having to over-provision capacity, supporting a broad range of workloads — from small developer environments to massive datasets.

This decoupled scale is especially powerful when applied to containerized environments. Through the Filestore CSI driver, Filestore delivers persistent, high-performance storage for GKE workloads with enterprise-grade reliability and availability. To further optimize resource utilization at scale, Filestore multishares for GKE lets you segment a single Filestore instance into many smaller shares, starting at just 10 GiB. This means AI teams can scale out their GKE clusters efficiently, carving up large high-performance instances into smaller, project-specific shares to support thousands of concurrent containers, without sacrificing performance or increasing TCO.

Shared storage for agentic swarms

The combination of decoupled performance and deep GKE integration enables a critical new use case: high-concurrency workspaces for AI agent swarms, where multiple specialized, autonomous AI agents collaborate in parallel to accomplish complex goals. In an agentic workflow, multiple agents often need to read from and write to a common dataset simultaneously to maintain state and share context. Filestore file shares act as these common workspaces leveraging the NFS protocol to provide consistent file system access across the swarm.

Backed by Colossus, Filestore can support millions of agents, working independently or securely collaborating with shared data. By utilizing NFS file locking, Filestore helps ensure strict data consistency, preventing conflicts even as swarms grow in size and complexity. This allows AI agents to maintain a unified view of their environment, enabling more complex reasoning and collaboration.

Operational agility and enterprise-grade control

Beyond performance, this enhancement improves operational agility. Now, you can independently scale IOPS via Custom Performance settings, optimizing costs in real-time without having to rebuild clusters. Furthermore, the distributed storage layer provides faster failure recovery and zero-downtime capacity changes compared to traditional monolithic storage architectures. Filestore is also providing security at scale via deep integration with Google Cloud IAM, NFS User IDs and Group IDs (UIDs/GIDs), and IP Access Control Lists (ACLs).

Building for the next era of data

This Filestore update reflects our commitment to meeting the changing needs of AI-era workflows and lays the foundation for future enhancements. By providing a first-party, deeply integrated service that evolves with your needs, we are helping you maximize your business potential flexibly and cost-efficiently. Get started today by visiting the Filestore page.

What’s new in AI infrastructure and orchestration in August

31 août 2026 à 18:00

Welcome back to What’s new in AI infrastructure and orchestration this month, a collection of product updates, how-tos, customer stories, research and other resources about all the AI compute, networks, storage, frameworks, and orchestration software that you can find at Google Cloud. To be honest, we thought August would be a slow month, but nothing could be further from the truth. Read on and you’ll see what we mean.

August 2026

Product, technology, and tools updates

  • Product update: Filestore, Google Cloud’s first-party, secure, scalable NFS file service, has emerged as a popular storage platform for AI and agentic workflows, and now, it’s even better suited to the task, with a new backend storage layer built directly on Colossus, Google’s foundational distributed storage system. This new backend lets you provision IOPS independently from storage capacity, and is deeply integrated with GKE. In AI environments, this can help you service so-called agentic swarms — large groups of agents that need to read and write to a common dataset — without a drop off in performance. For more, check out the blog post. 

  • New feature: gVisor sandboxes are now available in distributed Ray clusters on GKE. In partnership with Anyscale, we introduced an experimental library for Ray that brings gVisor, Google’s open-source application kernel, directly into distributed Ray clusters. gVisor provides lightweight environments with stronger isolation than ordinary containers, plus fast startup times and low memory overhead. To try out these sandboxing capabilities on GKE, head over to the Ray sandboxing User Guide.

  • Product update: Looking for high-performance, easy-to-use infrastructure on which to run a personal AI agent, but don’t want to spend a lot of money? New Cloud Run instances are dedicated, singleton compute runtimes on Cloud Run that won’t shut down when the agent is idle. Better yet, the cost to run a Cloud Run instance with 1 vCPU and 1 GiB of memory continuously for 30 days is just $5.70.  

Practitioner guides, documentation and how-tos

  • How-to guide: Big news in Model Context Protocol (MCP) land: As of the 2026-07-28 specification, the protocol core is “completely stateless. The handshake is gone. The initialize / initialized handshake (SEP-2575) and the logical Mcp-Session-Id header (SEP-2567) have been removed entirely. Instead, every request is now self-describing and independent.” Whoa. Learn more about the changes that the latest MCP specification brings, and more importantly, how to implement them, in this Google Developers blog.  
  • Guide: Real-time AI systems make a mess of traditional network load balancing techniques. “Instead of handling isolated requests, the backend has to manage a continuous, live bidirectional stream. You’re dealing with a constant stream of audio chunks, transcripts, model outputs, and synthesized speech flowing back and forth simultaneously.” Things only get worse when the user gets involved. “The server has to immediately halt its current speech generation, pivot to update the context, maybe trigger a new tool, and start drafting a different response; this must be done without dropping the connection.” For a new approach to managing load in the AI era, read Scaling real-time AI agents with session-aware load balancing.
  • How-to: Learn how to build an elastic, scalable LLM inference platform on GKE, even with a mix of different GPU accelerators. The proposed architecture combines Capacity Advisor and Compute Advisor, plus high-performance storage like RunAI:model streamer or GCPFuse with parallel downloads. Get all the details here.
  • Documentation: The thing about hosts with GPUs or TPUs is that you can’t use live migration to update them, setting up a maintenance challenge. In this new docs page, learn how to update accelerator-equipped hosts according to your tolerance for downtime for your training and inference workloads.    
  • Documentation: Advanced Compute Images, or ACIs, are standardized image stacks for AI/ML and HPC infrastructure, so you don’t need to manually build your own custom images. In this new docs page, learn how to create an ACI image using the Google Cloud CLI, console, or SchedMD's Slurm workload manager. 
  • Guide: AI workloads are notoriously difficult to architect, resource-intensive, and bursty, which can also lead to scaling bottlenecks and large pools of underutilized — or misutilized — compute resources. A new blog outlines the three main ways to achieve dynamic capacity management in Google Cloud: 1) scheduling capacity for planned downtime; 2) maintaining automated fallback capacity for unplanned downtime; and 3) relying on GKE’s core orchestration capabilities to automate resource allocation. 

Customer and partner updates

  • Business orchestration software provider UiPath was dealing with spiky workloads, and wanted more predictable costs. To get there, it re-architected its infrastructure, moving from isolated clusters to a shared Google Cloud GPU fleet that included both A3 VM instances (NVIDIA H100 GPUs) for training with G4 VM instances (NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs) for inference. You can read more about their architecture here. 

  • Mirendil, an frontier AI lab focused on accelerating AI development, announced that it is using AI Hypercomputer with both TPUs and NVIDIA GPUs to support its model pre-training and post-training applications. 

  • Replenit, a retail CRM provider, built its AI decision engine in Google Cloud, using BigQuery, Gemini Enterprise Agent Platform, and open-source Gemma models that it runs on Cloud TPUs. This latter combination provided Replenit with 90% lower pipeline costs than their previous cloud provider, the company reports. Read the full case study for more. 

  • Malachyte architected its AI-powered e-commerce recommendation platform on top of Bigtable, Managed Service for Apache Kafka, Pub/Sub, Compute Engine, and last but not least, GKE. See how it all comes together in this blog.


July 2026

Product, technology, and tools updates

  • Product update: Google Cloud Managed Lustre is now GA, and available in four distinct performance tiers that deliver throughput ranging from 125 MB/s, 250 MB/s, 500 MB/s, to 1000 MB/s per TiB of capacity — with the ability to scale up to 8 PB of storage capacity. The Managed Lustre solution is powered by DDN’s EXAScaler, combining DDN's decades of leadership in high-performance storage with Google Cloud's expertise in cloud infrastructure.

  • Product update: C4N network and storage optimized VMs are now GA. C4N is our first network- and block-storage-optimized VM series built to eliminate data-transfer bottlenecks. Powered by 5th Gen Intel Xeon Scalable processors and built on Google's Titanium offloading hardware, it achieves 400 Gbps network bandwidth, 95 million packets per second (MPPS), and up to 25 GiB/s of block storage throughput when paired with Hyperdisk Extreme.

  • New feature: GKE Dataplane V2 up to 15K Nodes with Network Policies (GA). This capability enables standard GKE clusters to scale up to 15,000 nodes while maintaining full active Network Policy enforcement, supporting the massive infrastructure needs of large enterprise and AI/ML customers.

  • New feature: Co-operative time-slicing in llm-d. If you’re running reinforcement learning (RL) workloads, you can now interleave independent RL jobs onto shared physical hardware, increasing aggregate accelerator duty cycles from a ~40% baseline up to 70% without impacting model convergence or accuracy. 

  • New AI security tool: Looking to secure your AI supply chain on GKE, deploy AI workloads safely, and cut down on shadow AI? We open-sourced k8s-aibom, a lightweight, unprivileged Kubernetes controller that continuously monitors container clusters to automatically detect running AI runtimes (like vLLM and Triton) and generate standard CycloneDX Machine Learning Bill of Materials (ML-BOMs). Check out the k8s-aibom project and get involved.

Practitioner guides and how-tos

  • How-to guide: On July 27, Google announced Day 0 support for Moonshot AI’s Kimi K3 2.8-trillion-parameter open-weight model, the day weights were released. Whichever your preferred deployment path — via Model Garden, custom orchestration, or GKE with llm-d recipes — this guide offers detailed step-by-step instructions to help you evaluate and pilot Kimi K3 in Google Cloud. 

  • How-to guide: Google Kubernetes Engine (GKE) managed DRANET supports both GPUs and TPUs. There are several configurations to use this implementation, including standard cluster (where you have full control) and autopilot cluster (where Google does the heavy configs for you). Take a deeper dive in the hands-on lab, GKE Autopilot clusters with TPUs, GKE managed DRANET and Gemma 4.

  • How-to guide: Learn to run Ray on TPUs, not GPUs. In Part 1 of this two-part series, we discuss TPU slices (hint: Ray thinks of them as just another accelerator on which to schedule), then walk through Ray’s various AI libraries (Part 2).

  • How-to guide: Evaluate TPUs for sample workloads using a new microbenchmark suite that helps you accurately assess whether a device is achieving its theoretical performance specifications, and to identify specific performance gaps or architecture-specific bottlenecks. Dive in here. 

  • How-to guide: Scale your agents without killing your budget. Learn how GKE orchestration can help you safely pack more agents onto a fixed compute footprint with GKE Agent Sandbox and Pod snapshots. Whether your goal is performance or cost optimization, we teach you how to turn the right dials for optimal agent efficiency. 

  • Technical blueprint: Inside the optimization of Mistral 3 large inference on Ironwood. This blog outlines how one Google team optimized Mistral 3 large MoE model inference on Google’s Ironwood (TPU v7x), achieving a 1.5x performance gain. They did so with hybrid sharding, replacing linear VPU summations with tree reductions, optimizing GMM/MLA kernels, and adopting asynchronous scheduling. As a result, they boosted throughput by up to 48% while maintaining benchmark accuracy neutrality. Read the full blog here.

Research, reports and deep-dives

  • Report: Google was named a Leader in the inaugural GartnerⓇ Magic Quadrant™ for AI Infrastructure, positioned highest for ‘Ability to Execute’ and furthest for ‘Completeness of Vision’. Gartner called out Google’s proprietary scalable compute, integrated AI Hypercomputer architecture, and the scale of our AI compute capacity as key strengths. Download a copy here.

  • Report: We recently surveyed more than 1,400 senior IT leaders for our State of AI Infrastructure report, and a resounding pattern emerged: The gap between AI ambition and infrastructure reality is widening. In fact, 83% of organizations say they require infrastructure upgrades to support production-grade agentic AI. Read the accompanying blog to understand how adapting your infrastructure to meet the demands that agentic applications place on your systems will help you move from pilot to production.


June 2026

Product, technology and tool updates

Practitioner guides and how-tos

  • How-to guide: Learn how to build high availability into an AI inference workload running on GKE Inference Gateway with TPUs, Cloud Storage FUSE and Dynamic Resource Allocation (DRA). This blog provides an overview, or you can get all the technical details in the hands-on codelab.

  • How-to guide: Did you know you can connect your AI agents to unstructured data in Cloud Storage via Model Context Protocol (MCP)? In this blog, learn about why would want to do that from three customer examples, then how to do it, choosing either a fully managed service, or a self-managed local server for more customization and control. 

Research, reports and deep-dives

  • Report: According to an independent benchmark report, GKE Inference Gateway outperforms the next leading managed Kubernetes service with 15.7% higher throughput, 92.8% shorter wait times, and 62.6% lower inter-token latency. This performance can be attributed to its use of prefix caching, which optimizes LLM performance by storing the KV cache (activation states) of long, repetitive prompt prefixes. Learn more in the blog. 

  • Architecture deep dive: A closer look at the cold start problem, this time for TPUs and GKE, and how the Run:ai Model Streamer can help change the dynamic. 

Customer and partner updates


May 2026

Product, technology and tool updates

  • Product update: GKE Agent Sandbox is now generally available.

  • New open-source project: Agent Substrate is a new open-source project aimed at continuing to push the limits of agentic infrastructure density

  • New feature: Google AI Edge Portal, a solution for testing and benchmarking on-device machine learning (ML) at scale, now supports benchmarking and debugging on-device LLMs. Read more here. 

  • Product deep dive: We went into depth about Cloud Storage Rapid, a new family of high-performance storage offerings for AI workloads. At launch, offerings include Rapid Bucket (formerly Rapid Storage), a high-performance zonal object storage offering, and Rapid Cache (formerly Anywhere Cache), which accelerates reads on-demand and colocates compute and data for workloads in existing buckets. 

Research, reports and deep dives

Customer and partner updates

❌