❌

Vue normale

Reçu avant avant-hierCloud Blog

Scribd, Inc. classifies more than 400 million documents with Gemini batch inference on Gemini Enterprise

24 septembre 2026 à 18:00

Scribd, Inc. is home to one of the world's largest collections of human-created content. 

Scribd’s  products leverage one of the world's largest collections of human-created content and intelligent tools to help people move from information access to real understanding and application.

This past year, Scribd used Gemini's native PDF understanding and Gemini Enterprise batch prediction to run trust and safety classification across its entire user-generated content corpus of more than 400 million documents, spanning over 12 billion pages, in a matter of months.

Here were the results: 

  • Classified 400M+ user-uploaded documents (12B+ pages of text and images) across Scribd and Slideshare

  • Completed the corpus-wide backfill in a matter of months, with Google Cloud scaling batch throughput to meet the timeline

  • Native PDF input meant more than 99% of the corpus was processed as-is, with no OCR, rendering, or screenshotting pipeline to build

  • Gemini Enterprise’s batch prediction at a 50% discount to interactive pricing made LLM classification viable at corpus scale

Trust and safety at the scale of an entire corpus

Scribd, Inc. is the parent company to four distinct products: Scribd, Slideshare, Everand, and Fable. Across Scribd and Slideshare, hundreds of millions of user-uploaded PDFs, presentations, and documents help people find information, build understanding, and finish projects. With that scale comes responsibility. We aim to balance access with protecting our communities. We leverage a mix of human and automated methods to review and best ensure the content on our platforms complies with our community rules. As the corpus continues to grow and technology evolves, this challenge requires even more resources.

Understanding a document requires reading its text and its images together, in context. Classification has to work across all possible use cases, all possible languages, all possible contexts. There is no single solution that can translate cleanly across all of it. And each policy area traditionally demanded its own specialized detection model, which meant either years of in-house engineering effort or specialized vendor solutions that don't fit the economics of a 400-million-document backfill. The team evaluated several off-the-shelf moderation tools and open models, but none delivered the quality they needed at their scale.

“This is a genuinely hard problem that we have been working on for a long time. Every category of content behaves differently, and historically each one required its own specialized solution. Gemini collapsed all of that into one model, one prompt, and one pipeline.” – Sachin Sebastian, Senior Engineering Manager, Scribd, Inc.

Why Gemini: PDFs are a first-class input

The turning point was realizing that Gemini treats Scribd's corpus the way it actually exists: as PDFs. Gemini accepts PDF input natively and reads each page as both text and image, so a single multimodal model could evaluate everything from dense text documents to image-heavy presentations, with no OCR pipeline, page rendering, or screenshot infrastructure in between. Because Gemini processes each PDF page at a fixed, predictable token count, costs scale linearly and stay low even across 12 billion pages.

After benchmarking model families and versions, the team selected Gemini 2.5 Flash Lite as the classification workhorse, with Gemini 2.5 Pro serving as an LLM judge in a full second consistency pass over the corpus to validate output quality. In the team's evaluations, Gemini's multimodal understanding caught visual policy signals that text-only moderation endpoints routinely missed.

“Gemini's peculiar advantage is that it meets our content in its native format. It reads the text, layout, and images of a PDF directly. More than 99% of our corpus went in exactly as it lives on our site without any pre-processing” – Sachin Sebastian, Senior Engineering Manager, Scribd, Inc.

Batch prediction, simple enough to bet the corpus on

The execution model was deliberately simple. Documents were staged in Cloud Storage, submitted to Gemini Enterprise batch prediction, and the results flowed back into the team's data platform for downstream analysis. There was no serving infrastructure to operate, no rate-limiting logic to write, and no GPU capacity to manage.

Batch pricing, at 50% below interactive rates, is what made the economics work at corpus scale. The team later layered on Gemini Enterprise’s implicit prefix caching, restructuring prompts so the static policy text hit the cache, which pushed efficiency further with no loss in classification quality.

A partnership measured in throughput

Processing 400 million documents is ultimately a throughput problem, and this is where the partnership with Google Cloud mattered most. Scribd's team connected directly with Google Cloud engineering and product to plan the backfill, advise on region strategy, and make sure the right capacity was in place ahead of launch.

As the backfill ramped up, Google Cloud worked closely with the team to scale throughput to the demands of the project. The effect was dramatic: batch jobs began completing far faster than projected, and for much of the run Gemini Enterprise was not the bottleneck. Scribd's own upstream pipeline was.

“Google Cloud didn't just answer support tickets. They partnered with us on the backfill, and there were stretches where Gemini Enterprise finished work faster than our own systems could produce it. That is a good problem to have.” – Sachin Sebastian, Senior Engineering Manager, Scribd, Inc.

What's next

The backfill is now the foundation of an ongoing program: newly uploaded content flows through the same Gemini classification pipeline, keeping the corpus continuously evaluated rather than periodically cleaned. And because the pattern of PDFs in Cloud Storage, Gemini batch prediction, and results in the lakehouse proved so operationally simple, the team is applying it to a growing set of content-understanding workloads across its platforms.

“This project changed how we think about our roadmap. Work we had classified as multi-year, multi-team efforts is now a prompt, a batch pipeline, and a few weeks of runtime.” – Sachin Sebastian, Senior Engineering Manager, Scribd, Inc.


This work was a collaboration between Google Cloud and Scribd. We'd like to thank everyone involved for their support throughout this project:

  • Scribd Engineering: Anish Kumar, Jeanie Lam, James Watkins, Hima Alladi
  • Scribd Applied Research: Rafael Pedrosa Lacerda de Melo, Kara Killough, Eric Chang
  • Scribd Product: Seyoon Kim, Nicole Pauls
  • Google Cloud AI Batch Inference team: James Liu, Digvijay Singh, Wei-chung Wang, Yan Wang, Kun Shi 
  • Google Cloud Customer Engineer: Jennifer Liang

How growing Latin American midsize businesses are building in the AI era

24 septembre 2026 à 16:30

Latin America’s small and medium-sized businesses are the heartbeat of the region's economy — accounting for more than 60% of total employment in the region, according to United Nations estimates. And just like their enterprise peers, everywhere you look, ambitious teams are moving fast to embrace AI. 

Many have already transitioned from experimenting with generative tools and agentic workflows to using them every day to work smarter, save time, and deliver exceptional customer experiences. These growing businesses are particularly focused on maximizing the benefit they get from their investments in AI, whether that’s using a fast, low-cost model to summarize daily emails or deploying an advanced model for complex data analysis, teams can match the right AI capability to their exact task and budget. 

It’s this range of options, and a familiarity with the broader suite of Google business, media, and advertising tools that has led many SMBs to choose Google Cloud, and Gemini Enterprise in particular, as their AI platform of choice. By doing so, they’re able to build custom AI agents, streamline daily tasks and paperwork, and offer customers instant support with the speed and reach needed to compete on a global scale. 

With our unique front row seat, we’ve seen the benefit SMBs are getting from leveraging Gemini Enterprise, not only for generative AI, but as a catalyst for adopting other essential cloud tools like Google Kubernetes Engine and BigQuery for complete end-to-end modernization. The number of Latin American-based small and medium businesses using Google Cloud AI tools has grown 8x year-over-year and the number of Brazil based small and medium businesses using Google Cloud AI tools has grown 9x year-over-year. This rapid adoption spans our Gemini models, Gemini Enterprise, and core Cloud infrastructure, and are helping businesses to:

  • Roll out better customer support systems to help escalate and resolve customer support calls more quickly.

  • Automate repetitive actions in areas like payroll and accounting.

  • Help more employees understand and leverage data at work — even those not trained as data analysts.

  • Rapidly create and implement new designs for marketing collateral.

  • Help more people build their own AI agents to help them in their everyday jobs.

As we head into today’s Google Cloud Summit in Brazil, we were proud to showcase nearly 20 of our newest Latin American SMB customers using Google AI to reduce busywork, serve their customers faster, and grow their businesses.

Announcing new Latin American customers putting Google AI to work

  • AdGoat, an Argentina-based adtech company processing more than 10 billion annual ad requests across more than 100 global websites. It uses Cloud Run, the Gemini API, and Gemini Enterprise to automate content analysis, ad bidding, and audience targeting to help e-commerce brands drive higher campaign returns.

  • Angelus, a Brazilian dental and healthcare manufacturing company, uses Gemini Enterprise to streamline project management across its research and development department. This enables its teams to automatically pull technical project data into pre-approved templates aligned with the company’s brand identity and regulatory requirements.

  • BunkerDB, a marketing science company operating across Latin America, uses Gemini Enterprise, Cloud Run, and Cloud Storage to power an AI platform that organizes marketing assets, checks brand compliance, generates or adapts multimodal content, and predicts ad performance before launch. All of this helps it reduce creative turnaround times from weeks to hours and cut cost per lead by up to 25%.

  • Caffeine Army, a Brazilian wellness and high-performance company that connects people with solutions in nutrition, sports, and well-being, deployed BigQuery and Gemini Enterprise on Google Cloud to unify customer purchase insights, enabling faster creative campaign turnarounds and boosting team productivity across the organization.

  • Convert, a Brazilian marketing and analytics provider, uses Looker, BigQuery, and Cloud Run to power five specialized AI agents that answer complex business questions in natural language, speeding up report deliveries by 65% and reducing operational costs by 32%.

  • Growth Digital, a Google Ad sales rep operating across 13 Latin American countries, used BigQuery and Gemini Enterprise to build over 113 AI agents, enabling teams to build proposals 5x faster, cut campaign reporting time by 80%, and reduce financial error rates to under 0.01%.

  • GrupoTusMaquinas.com, an equipment management platform based in Chile, deployed Google Cloud AI tools and Gemini models to create digital tracking profiles for trucks and machinery, allowing businesses to query fleet status in plain language and manage vehicles regardless of brand or location.

  • HealthAtom, a healthcare technology company, uses the Gemini API, Firestore, and Cloud Functions to power AI assistants across its clinical platforms, automating appointment scheduling and medical record reviews while supporting 80 million annual patient interactions.

  • KLog.co, a Chilean logistics technology company digitizing freight forwarding across Latin America, uses Gemini Enterprise, BigQuery, and Google Workspace to automate cargo tracking and shipping paperwork, cutting manual data entry errors by over 90% and increasing document processing capacity tenfold.

  • NEEOH, a leading Brazilian out-of-home advertising communication platform, uses Gemini Enterprise to standardize secure AI usage across its organization, enabling teams to generate campaign copy and build pitch proposals faster while keeping corporate client data secure.

  • Luxia Agro, an Argentinian foreign trade supplier of crop protection products, uses Gemini 3.5 Flash and Gemini Enterprise to automatically pull key details from complicated shipping emails and update their central business systems. This allows it to automate 80% of foreign trade operations and cut manual processing errors in half.

  • Macal, a Chilean auction company, uses the Gemini Enterprise, Cloud Run, BigQuery, and Security Command Center to automatically verify property records and modernize its technology systems, cutting software development times from weeks to days and lowering infrastructure costs by up to 30%.

  • Ninecon, a Brazilian tech consulting firm, deployed Gemini Enterprise to integrate AI directly into employee workflows, allowing managers to track usage patterns and optimize project turnaround times with real-time insights.

  • Nova Gestões, a customer service and operations provider in Brazil, uses Cloud Speech-to-Text and Gemini Enterprise to translate and analyze 100% of customer calls in real time, reducing post-call manual data entry and boosting team productivity by 30%.

  • Romi, a Brazilian industrial machinery manufacturer, uses the Gemini API and Gemini Enterprise to power an interactive chat assistant directly on CNC machine HMI (human machine iInterface), giving factory operators instant answers grounded in official manuals and generating QR codes for step-by-step instructional videos.

  • Supermercados El Dorado, a leading supermarket chain in Uruguay, leverages Google Compute Engine and Gemini Enterprise to modernize legacy testing infrastructure and connect custom AI agents within their daily workflows, boosting team productivity across departments.

  • Tryvia, a Brazilian IT and business solutions provider, uses Google Cloud, Looker, and Gemini Enterprise to move off legacy physical servers, giving teams real-time reporting dashboards and AI tools that speed up software development and daily tasks.

  • Via Cristais, a major highway operator in Brazil, leverages Google Contact Center as a Service to speed up emergency routing for highway accidents, reducing caller wait times, improving driver satisfaction, and mitigating the impact of call center staff turnover.

  • WeSpeak, an AI conversational platform for the hospitality industry in Latin America, uses Cloud Run, Gemini Pro, and Gemini Flash to automate end-to-end guest interactions across messaging channels like WhatsApp and Instagram. This has helped it achieve an 85% resolution rate and a 2x increase in overall sales volume for hotel clients.

Helping your team build AI skills

To help growing teams get the absolute most out of AI, we’ve created easy, no-cost learning programs that anyone can use:

  • Programs for small and medium businesses: Explore beginner-friendly training paths or join specialized programs to learn how to build custom AI assistants for your day-to-day work.

  • Google skills for organizations: Access thousands of free, on-demand AI courses and hands-on practice labs designed by experts at Google Cloud and Google DeepMind.

  • Get certified: Help your staff gain industry-recognized AI certificates through guided courses, expert mentoring, and skill badges.

By offering easy-to-use tools and free training — from everyday office apps in Workspace to advanced AI on Google Cloud — Google is here to help Latin American businesses thrive today and in the future.

The future of orchestration: Pine59’s journey to Airflow 3 on Google Cloud

17 septembre 2026 à 19:00

Operating large data pipelines requires an orchestration layer that scales smoothly as workloads expand. When your pipelines process millions of complex data points every day to feed predictive models, staying up-to-date with your technology stack is a strategic necessity.

Pine59 provides location intelligence data through data pipelines that produce analytical metrics on cadences ranging from hourly to quarterly. One of the company’s most data-intensive metrics, Daily Foot Traffic, computes data for as many as 14 million distinct locations in a single job. To handle this massive volume, Pine59’s system runs entirely on Google Cloud, with the heavy lifting in BigQuery and all of it orchestrated by Managed Service for Apache Airflow (formerly Cloud Composer) running Apache Airflow 3.

As the company’s volume of data and number of machine learning workloads scaled up, Pine59 decided to modernize its monorepo, which contains hundreds of directed acyclic graphs (DAGs). Here is a look at how that transition improved Pine59’s MLOps capabilities, developer workflow, and pipeline speed.

Proactive modernization for growth

Pine59 has long relied on a shared monorepo with code and tooling spanning multiple projects to run its metric production pipelines. As it considered its infrastructure’s future, the company wanted to help its data pipelines run faster and more reliably.

That’s why it decided to stress-test production workloads against the newly available Managed Airflow (Gen 3) architecture running Airflow 3. The initial results were unambiguous: the Gen 3 environment delivered immediate and significant processing speed, task scheduling, and overall stability improvements. Recognizing the clear potential for performance gains, Pine59 initiated a full transition to the new environment.

1 - Pine59 Google Cloud Architecture Vertical Version

Orchestrating advanced MLOps

Pine59’s pipelines don’t just move data; they drive complex ML models, so a core aspect of its migration was optimizing the orchestration of its ML inference workloads.

Previously, Pine59 had used standard Kubernetes operators for these tasks. By moving to Managed Airflow (Gen 3), which features a highly optimized and abstracted infrastructure layer, the company’s engineering team refined its MLOps architecture. They did so by setting up a dedicated Google Kubernetes Engine (GKE) cluster that was specifically optimized for model inference and integrated it into the Pine59 pipelines.

This clear separation of orchestration and heavy ML execution compute allows data processing and model inference to run efficiently, showcasing Managed Airflow as a resilient, scalable backbone for enterprise MLOps.

Supporting developers with custom extensibility

Beyond infrastructure improvements, Pine59 was also able to immediately capitalize on Airflow 3’s delivery of a vastly improved developer workflow and user interface. Indeed, managing hundreds of interconnected DAGs requires excellent observability, and Pine59 found Airflow 3’s plugin authoring system remarkably easy to use.

To improve internal developer velocity, the company quickly built a number of custom plugins that it integrated directly into its new Airflow UI:

  • BigQuery Auto-linkify: A tool that automatically detects internal BigQuery table references within the Airflow Logs and XCom tabs, dynamically generating direct links to BigQuery Studio for faster debugging (available as a public GitHub gist)

  • DAG Run Configuration Search: A custom search form added directly to the DAG overview page. It allows Pine59 engineers to query specific key-value pairs within DAG run payloads (configs) and instantly surface matching runs. This in turn drastically reduces troubleshooting time.

In addition, the team also deployed a compatibility shim layer within its monorepo. This “compat” module dynamically abstracts logic between Airflow versions, streamlining operator migration across versions.

Faster, more reliable pipelines

For Pine59, migrating to Managed Airflow (Gen 3) with Airflow 3 has yielded clear, quantifiable results.

The most important improvement was the speed of its DAG runs. In the company’s previous setup, tasks often got stuck in a queued state during peak processing surges. With Gen 3, queue latency has dropped dramatically, allowing tasks to start running almost immediately.

Consider the comparison below of total aggregated “queued” & “running” time of more than 300 runs of the same DAG between Managed Airflow (Gen2) with Airflow 2.11 vs. Managed Airflow (Gen3) with Airflow 3.1 below. As we can readily see, the difference in queued time is significant.

image2

Coupled with internal DAG optimizations made during the transition, the performance gains are also highly tangible. For example, the Daily Foot Traffic pipeline previously took nearly 38 minutes to complete. With the new instance, the same workload now takes less than 26 minutes —nearly 32% less processing time.

Today, Pine59 processes all its production workloads on its new Managed Airflow (Gen 3) instance. By moving to this next generation orchestration, the company improved its MLOps capabilities, equipped its developers with better tools, and built a faster, more resilient foundation for future workloads.

If your engineering team spends more time managing infrastructure than delivering value, consider a similar transition and discover how it can help you move from maintaining servers to building the future of your data and AI pipelines today.


Special thanks to the following contributor to this post: Alexandre Crespo-Perez

How a solo founder runs a five-continent tender platform on AlloyDB and MCP

17 septembre 2026 à 18:00

Editor's note: Lucius AI, a tender-intelligence startup covering markets across five continents, runs its entire data platform on AlloyDB for PostgreSQL with a single operator. By migrating semantic search to a ScaNN index and managing database operations through Model Context Protocol (MCP), query latency dropped by 47x while automating day-to-day administrative tasks via MCP.


Executive summary

  • Lucius AI runs a global tender platform spanning more than 210,000 tenders across the UK, EU, India, and Australia, requiring minimal operational overhead for a solo founder.

  • Lucius AI deployed AlloyDB for PostgreSQL to consolidate its relational catalog, audit logs, and vector embeddings into a single managed database engine.

  • Migrating semantic search to a ScaNN index lowered query latency from 1.14 seconds to 24 milliseconds — a 47x speedup on a representative production query.

  • Connecting an AI agent to AlloyDB using the Model Context Protocol (MCP) helps Lucius AI automate query analysis, data freshness checks, and incident forensics under strict least-privilege permissions.

Making tender intelligence work as a company of one

Lucius AI helps businesses bidding on public contracts evaluate opportunities across global markets. The platform ingests public procurement notices from the UK, the EU, the US and Canada, Australia and New Zealand, India and Singapore, alongside World Bank donor-funded notices across Africa and Asia. Lucius AI analyzes tender documents using Gemini to generate compliance matrices, bid recommendations, and draft responses citing original source pages. For small and mid-sized suppliers, this replaces days of manual document reviews and costly external consulting.

Running a platform of this scope requires extensive operational coordination:

  • Nightly ingestion from thirteen public procurement sources

  • A catalog of more than 210,000 tenders, including tens of thousands open for active bidding

  • Two production regions on Cloud Run: Europe, and an Australian deployment on its own AlloyDB cluster with customer-managed encryption keys (CMEK) for defense-adjacent customers

  • Ongoing analytics, performance tuning, data validation, and incident response

Managing these responsibilities without dedicated data engineering or database administration teams requires offloading operational maintenance. Lucius AI addressed this challenge on two fronts: using AlloyDB for PostgreSQL as the core system of record, and connecting an AI agent through the Model Context Protocol (MCP) to safely execute database operations.

1

Consolidating systems into AlloyDB

Rather than deploying separate relational databases, vector databases, and log stores, Lucius AI houses all core data in AlloyDB for PostgreSQL. The relational tender catalog, document metadata, audit logs, and vector embeddings reside in the same database engine. Storing vector embeddings alongside relational rows avoids managing separate vector stores, establishes a unified backup schedule, and centralizes identity management.

Authentication relies strictly on Cloud IAM. Services connect using dedicated Google Cloud service accounts mapped to database roles scoped to specific access requirements, without storing database passwords in application environments. Database reliability is managed natively by AlloyDB through automated backups and point-in-time recovery, avoiding custom disaster recovery procedures.

In production, this consolidated architecture supports:

  • More than 210,000 tenders in the catalog, with embeddings stored directly alongside them

  • Rebuilding the semantic index embedded 115,820 records in 10.6 minutes with the Gemini embedding model, for around three dollars in API spend; AlloyDB auto embeddings now keep those vectors current.

  • Retrieval reranking executed directly inside the database using the ai.rank function — with mean latency of 77-milliseconds - returning the most relevant results for search queries without requiring a standalone reranking microservice

Accelerating semantic search by 47x

Semantic search across the tender catalog initially relied on unindexed vector comparisons, where a representative query took 1.14 seconds. Migrating this workload to a ScaNN index in AlloyDB reduced query latency to 24 milliseconds — a 47x improvement.

The index recommendation originated from the AI agent during an automated performance audit, where it benchmarked the query plan before preparing the index migration.

2

Automating database operations with MCP

To delegate routine administrative tasks, Lucius AI configured the open-source MCP Toolbox for Databases using the prebuilt alloydb-postgres server.

Operational delegation requires strict access controls. The agent connects using a dedicated PostgreSQL role granted SELECT across the schema and UPDATE on a single operational table. Destructive commands (DROP, DELETE, TRUNCATE) are omitted, restricting agent actions to authorized operational boundaries.

Under this configuration, the AI agent performs regular database operations across four key areas:

  • On-demand analytics: Compiles retention cohorts, activation funnels, and catalog coverage by country via ad hoc SQL queries, removing the need to build and maintain manual dashboards or complex analytical pipelines.

  • Performance optimization: Performs query-plan inspections and index analysis, such as identifying the ScaNN indexing strategy.

  • Incident forensics: In response to an external security probe, the agent parsed audit logs to reconstruct the request timeline in minutes, verifying that tenant isolation remained intact.

  • Automated data-quality checks: Evaluates ingestion watermarks and freshness across all thirteen procurement sources every morning.

3

For teams adopting this architecture, establishing a progressive permission structure provides clear guardrails: start with read-only access, expand permissions as requirements dictate, and keep destructive operations restricted to human administrators.

Looking ahead

Lucius AI is planning three technical initiatives to further reduce operational overhead:

  1. Automated vector embeddings in AlloyDB AI: After validating ai.initialize_embeddings across the full catalog, a weekly maintenance job uses ai.refresh_embeddings to update vectors.

  2. Columnar engine acceleration: Having enabled AlloyDB’s columnar engine with auto-columnarization, the database identified and stored 40 frequently queried columns across four tables in memory within a day, accelerating reporting queries without a separate analytical store.

  3. Managed Remote MCP Server: Transitioning from self-hosted Toolbox processes to Google Cloud's fully managed Remote MCP Server for AlloyDB will offload MCP server hosting and maintenance.

By anchoring core data in AlloyDB and managing routine operations through MCP, Lucius AI demonstrates how a single engineer can build and operate a resilient, multi-region procurement platform.

To explore Lucius AI, visit ailucius.com. To evaluate AlloyDB for PostgreSQL, deploy an AlloyDB cluster to test performance against your own workloads.

For SeaVerse, GKE Agent Sandbox reduces infrastructure costs by 60%

16 septembre 2026 à 18:00

Editor’s note: Today we hear from SeaVerse, a gaming startup from SeaArt that is building a platform for playable AI experiences, where users can open lightweight games, character chats, and interactive apps, or create their own experiences from a prompt. To support that creative loop, SeaVerse needed infrastructure that could run dynamic, multi-tenant sandbox workloads with strong isolation, low latency, better observability, and more flexible costs. Google Kubernetes Engine (GKE) and GKE Agent Sandbox gave SeaVerse the managed foundation from which to execute these AI workloads, helping the team reduce their infrastructure costs by up to 60%, while giving creators a faster path from idea to playable experiences. Read on to learn more.


What if AI were a playground? Welcome to SeaVerse, a creation-first platform for playable AI experiences. Here, an AI creation can be as peaceful as drawing a path for a snake to follow, or as chaotic as a music-backed stickman simulation. Some people come to play lightweight games. Others come to chat with AI characters, try interactive apps, create visual patterns, share what they made, or remix an idea into something new. 

We built SeaVerse around a simple promise: Every experience should feel immediate and easy to share. A creator should be able to describe an idea in plain language, refine the result, and publish it in moments, without a traditional coding workflow. 

Delivering that simplicity requires serious infrastructure. Every creation that users make moves through the same chain: generate, run, preview, debug, publish, remix. If any part of that chain is slow, unstable, or poorly isolated, users feel it immediately. That’s why we turned to GKE and GKE Agent Sandbox. 

The infrastructure challenge of instant interaction 

What looks effortless to a user is anything but on our end. Every creation on SeaVerse runs as a distinct workload and is expected to behave reliably from the first interaction. 

Because each workload runs in its own environment, we needed clear security boundaries between users, creations, and sandboxes. But overly strict isolation could slow the very creative loop we were trying to protect, and when something went wrong, diagnosing it was costly. Our engineers had to trace problems across multiple parts of the execution chain with little visibility into what was happening inside the environment. 

We explored existing sandbox approaches, but needed deeper kernel-level isolation and native observability at scale to support fast diagnosis across multi-tenant environments. Something had to change.

Building on GKE and GKE Agent Sandbox 

We chose GKE because we needed a reliable, secure way to operate Kubernetes without turning our engineering team into a cluster maintenance team. GKE brought together the proven ecosystem and operational tooling we needed, freeing us to focus on building the platform rather than managing the infrastructure beneath it. 

As a Kubernetes primitive designed for agent code execution and computer use, GKE Agent Sandbox addressed our requirement for strong isolation, enforcing strong security boundaries without slowing down the creation experience. By utilizing GKE Agent Sandbox with Kata Containers+Cloudhypervisor (microVM), we’ve achieved the perfect balance of multi-cloud flexibility and robust security, option to switch isolation runtime between microVM and gVisor, running our AI sandboxes safely. GKE empowers us to scale toward our long-term vision of supporting over a million sandboxes. Built on gVisor, it provides kernel-level isolation for dynamic sandbox workloads while preserving the Kubernetes orchestration model, so that they can be managed through the same scheduling, monitoring, and operations as the rest of the cluster.  

With SeaVerse, users can generate interactive experiences from a single prompt. After an experience is generated, GKE Agent Sandbox supports the run, test, integration, and verification steps needed to make it ready to preview, refine, and publish. At general availability, it supports allocating up to 300 sandboxes per second, per cluster, with 90% of allocations completing in 200 milliseconds. Together, GKE and GKE Agent Sandbox gave us a reliable foundation for AI-generated interactive workloads that helped keep our team focused on the product experience. 

From black box to glass box 

Before GKE Agent Sandbox, a failed sandbox workload could feel like flying blind. We could often see that something had gone wrong, but didn’t have enough runtime status, metrics, or failure signals to understand why. 

Now, Google Cloud’s native logging and monitoring reach directly into those sandboxed environments, giving us a clearer view of workload behavior, faster issue resolution, and a stronger foundation for managing multi-tenant workloads. 

That visibility matters to developers, but it also matters to the platform’s users: A creator never sees the logs, the cluster, or the orchestration layer. They see whether an experience opens quickly, whether it responds when they draw, click, chat, or share, and whether they can keep building without friction. 

Flexibility that translates to savings

GKE Agent Sandbox also changed how we think about cost. Previously, running secure sandboxed environments meant stronger dependencies on specific server types, which limited how precisely we could match resources to each workload. With GKE Agent Sandbox, we can run secure, isolated workloads on appropriately sized cloud VMs. This gives us greater flexibility in resource allocation and helped us cut our infrastructure costs by up to 60%. 

That same flexibility extended to storage. Not all SeaVerse creations are built in a single session. Some evolve over time as creators return to refine them, build on earlier ideas, or invite others to remix what they’ve made. Our previous architecture didn’t support the persistent file-system capabilities those more complex use cases demanded, but that gap is gone now. We can attach persistent storage where workloads require it while maintaining the isolation boundaries that multi-tenant AI experiences need. For creators, that means experiences that are fast to open and easier to refine, revisit, and build on over time. 

The next remix 

Supporting creations that can evolve and deepen is central to what we’re building. It’s still early in what playable AI can become. As the platform grows, we need to keep strengthening what matters most: stability, observability, elastic scaling, and cost efficiency, all in service of a creator experience that stays fast, reliable, and expressive. 

We’re also exploring additional Google Cloud tools to support smarter analytics and creation assistance. Gemini and agent models could help operators and creators better understand how experiences perform. BigQuery AI and ML capabilities can support use cases such as churn prediction, LTV and ROI prediction, and user segmentation. Multimodal tools such as Imagen and Veo on Gemini Enterprise Agent Platform open up new possibilities for material analysis, creative generation, and AI interactive content production. 

Our goal is to make AI experiences feel immediate, expressive, and connected. With GKE and GKE Agent Sandbox, we have a stronger foundation for the next generation of playable AI.

How Orange built FinOps accountability, and why agents are next

16 septembre 2026 à 18:00

At Orange, the leading France-based multinational telecom provider, there are days when engineering teams set aside their delivery backlogs and spend the day cleaning up cloud spend together. There's a leaderboard. There are goodies on the line. Experienced practitioners guide the newcomers, so people learn the work while doing it. By the end of the day, sponsors can see the results.

Orange calls these FinOps Clean Days. Together with gamified hackathons, they've earned the company's 100-plus person FinOps community a Net Promoter Score within the organization that’s above 70.

Those numbers point at something the wider industry is wrestling with. Recent State of FinOps reports identify getting engineers to take action as one of the top challenges organizations face. Moving from awareness to action means finding ways to build FinOps accountability, and to get teams to genuinely care.

That makes FinOps a business change problem. And business change problems have known solutions. We spoke with Camille Marini, the FinOps lead at Orange, to get a deeper understanding of how the company overcame these hurdles to accelerate AI adoption and ROI, and how your organization might follow the same course.

Why the Clean Days work

Orange has held two principles since it set up its FinOps team. First, Cloud FinOps is a shared responsibility, with every stakeholder in a project involved in their own way. And the only path to that shared responsibility runs through communication and a deliberate change effort. 

“We insisted on the concept of shared responsibility across the organization for our FinOps practices,” Marini told us. “It’s very similar to how we approach cloud security. We needed to make teams understand that every single stakeholder in a project is involved in FinOps, each in their own way, if we are going to achieve responsible and impactful AI spending and usage.”

Those principles led Orange to create a FinOps Community of Practice, with support from Google Cloud Consulting. The team ran it on standardized communication channels so the methodology reached well beyond the central group, and kept the meetings actionable, sharing optimizations and billing updates so every session provided value.

The Clean Days came from a clear-eyed reading of how agile teams actually operate. In agile environments with deployment running constantly, optimization work rarely wins against the sprint. Delivery priorities, backlogs, and daily operations take the available time first. So Orange created protected time, made it collaborative, and made it fun.

McKinsey's four building blocks of change explain why this approach lands. Any large organizational change, the framework holds, requires action across four areas:

  • Conviction and understanding: "I know what is expected of me and I agree with it."

  • Formal mechanisms: "The structures, processes, and systems reinforce the change."

  • Role modeling: "I see my leaders and colleagues behaving differently."

  • Talent and skills: "I have the skills and opportunities to behave in a new way."

Map Orange's practice onto those blocks and the pattern is visible. Gamification and rewards give engineers colleagues to emulate: The leaderboard makes different behavior visible, and sponsors see the quick wins for themselves. Experienced practitioners guiding novices builds talent and skills through the community itself. The regular sessions, sharing optimizations and billing updates, build the conviction that comes from knowing where the money goes.

image2

FinOps activities mapped to the four building blocks of change, with the points where AI agents can reinforce them.

What happens beyond 100 people

A community of 100 engaged people is an achievement. But in an organization with thousands of engineers, no central FinOps team can reach everyone directly. The question for leaders is how to extend what a community like Orange's creates — the awareness, the shared ownership, the habit of acting — to people the FinOps team will never meet.

This is where AI agents extend the capabilities of a FinOps team with two core benefits. They take on complex, time-intensive activities that previously needed a human, and they reduce friction around FinOps for individuals across the business.

Getting teams to adopt them takes a strategy aimed at your own organization's pain points, which often come from high cognitive load, unclear accountability, or competing priorities. Start by finding where engagement drops off in your FinOps lifecycle:

  • An awareness gap: If teams are unsure of their spend impact, an insight agent can push real-time cost data into their daily tools.

  • A bandwidth gap: If engineers are too busy with backlogs, a remediation agent can identify quick wins and present them as ready-to-merge code changes.

  • A complexity gap: If reporting feels like a manual chore, an orchestration agent can gather the data and simplify the process.

Start with trust, then add autonomy

The sensible path runs in sequence. Establish the community practice, the way Orange did. Then introduce read-only agents that inform and suggest. Only once those are established across the community should you build agents that execute changes. Direct action carries operational risk, so manage it carefully. It's also where significant wins often sit.

How you build depends on who's building. For teams that want to deploy quickly with minimal code, the Gemini Enterprise App provides a no-code environment for creating agents. For developers who need granular control, the Gemini Enterprise Agent Platform (formerly Vertex AI) offers advanced tools for launching and governing agents built with frameworks like the Agent Development Kit (ADK).

Cloud FinOps is moving beyond centralized reporting toward action that happens where the work does. The organizations getting there start with the culture, then use agents to carry it further than any one team could reach. 

Orange's numbers came out of the community work. Building that foundation is the part worth copying first. When you're ready to extend it, Google Cloud Consulting can help you shape the community practice, and the Gemini Enterprise App is a low-lift way to put your first read-only agent in front of your teams.

Celebrating our tech and startup customers

20 avril 2022 à 22:00

Our tech and startup customers are disrupting industries, driving innovation and changing how people do things. We’re proud of their success and want to showcase what they’re up to! You’ll hear about their new products, their businesses reaching new milestones and their ability to get things done faster and easier using Google Cloud’s app development, data analytics and AI/ML services.

Congrats to Impact Analytics for Closing PVH
With the COVID-19 pandemic, the rise of e-commerce, and supply chain crisis, Impact Analytics had to quickly offer enhancements on their platform that gave retailers access to intelligent, automated, and edge-aware solutions. Google Cloud's best in class AI and ML solutions and highly performant infrastructure gave Impact Analytics the scalability and building blocks to create Ada, a robust predictive algorithm to give PVH and other retailers the tools to enhance their inventory planning capabilities. And now, Impact Analytics just closed a strategic deal with PVH (parent company of luxury brands Tommy Hilfiger, Calvin Klein, True & Co) to build out AI solutions for assortment planning and pricing optimization. Impact Analytics' cutting edge AI and ML guided forecasting engine is built entirely on Google Cloud! Read more

Podimetrics raises $45M Series C round
Podimetrics, creator of the FDA-cleared SmartMat and integrated clinical care services team, is dedicated to early detection and prevention of diabetic amputations, one of the most debilitating and costly complications of diabetes. Its clinical care services platform leverages Google Cloud services to engage with patients, by helping save limbs, lives, and money - all while keeping vulnerable populations healthy in their own homes. Read more about their Series C funding round.

Anvyl raises +$15M in an oversubscribed Series B funding round
Congrats to Anvyl for raising a hugely successful Series B as they modernize & transform the supply chain technology market and more than doubled revenue in the last year. Read more.

Helios has kicked off 2022 in a big way
The audio tone analysis platform, Comprehend: Elite, that they provide to Wall Street quantitative hedge funds now covers all US equities and is fully available here. It’s entirely powered by Google Cloud!

Dapper Labs uses Google Cloud for performance, reliability and decentralization
In case you missed it before the holidays, Dapper Labs is working with Google Cloud as its hyperscale cloud partner to ensure performance, reliability and decentralization for the next wave of mainstream users on Flow, without needing to compromise on decentralization or sustainability. Find out more.

Geotab’s Intelligent Transportation Systems (Geotab ITS) is built on Google Cloud.
Geotab uses GKE, BigQuery, Dataflow and Cloud Composer to build an innovative solution combining analytics and access to massive data volumes so municipalities can make better transportation planning decisions. The sheer volume of information that it handles, along with a need for highly scalable and flexible tools to manage, store, and analyze that data, led Geotab to invest in Google Cloud technology. Read more.

Mux CEO shares advice for getting started with video
Mux CEO, Jon Dahl, sat down with Google Cloud Director, Nirav Sheth, to share best practices and strategies for getting started with video, along with insights and advice from his learnings as a startup founder. Listen to what he has to say.

Google Cloud is proud to support Unstoppable Women of Web3
Unstoppable Women of Web3 (UWOW3) is an action oriented community made of industry leaders supporting education & opportunities for girls, women, and minorities in this burgeoning industry. This International Women’s Day, March 8th, you can catch live interviews with Tech and Web3 leaders from all over the world, covering topics such as how to build communities, how to learn more about Web3, developing technology on the blockchain, how to talk about complex ideas with kids, and more! How you can engage:

Puppet CTO increases development speed
Hear Puppet CTO Deepak Giridharagopal discuss how they managed to build Puppet's first Saas product, Relay, fast while also ensuring they would be able to remain agile if growth was to happen quickly. Watch video.

Vimeo builds a fully responsive video platform on Google Cloud
The video platform @Vimeo leverages managed database services from Google Cloud to serve up billions of views around the world each day. Read how it uses Cloud Spanner to deliver a consistent and reliable experience to its users no matter where they are. Find out more. 

Nylas improved price-performance by 40%
You don't have to choose between price-performance and x86 compatibility. Hear from David Ting, SVP of Engineering and CISO at @nylas, to learn how Google's x86-based Tau VMs delivered 40% better price-performance than competing Arm-based VMs. Watch now.

Optimizely partners with Google Cloud on experimentation solutions 
Build the next big thing with @Optimizely Experimentation on Google Cloud - driving innovation and next-gen experimentation for enterprise companies and marketers. Check it out.

How Airtel delivered its flawless Indian Premiere League 2026 cricket broadcasts

9 septembre 2026 à 18:00

For the millions of fervent fans of the Indian Premiere League (IPL), being able to count on a flawless live streaming cricket experience is never up for debate. For Airtel, producing league TV broadcasts with some of the world's most massive concurrent viewership, dropped packets and buffering are simply not options.

During the IPL 2026 season, Airtel partnered with Google Cloud to manage this digital delivery. Across 74 matches, the streaming infrastructure delivered several hundred petabytes of egress data. The final match alone processed tens of billions of requests, hitting a peak egress of several Tbps.

Delivering video under these concurrency spikes requires an edge architecture designed strictly around localization, paired with proactive operational monitoring. 

Our goal for IPL 2026 was to deliver an uninterrupted, stadium-grade viewing experience to cricket fans across India, regardless of concurrency surges or network conditions. Partnering with Google Cloud and using Media CDN gave us deep local edge proximity and excellent cache efficiency. Combined with proactive match-day real-time monitoring, we delivered a reliable broadcast experience from start to finish.

Architecting for concurrency and edge efficiency

IPL-BLog-Architecture

One of the primary challenges in live sports broadcasting is seamlessly handling large traffic spikes and never degrading stream performance or overwhelming backend origins. That’s especially important when millions of viewers simultaneously tune in during a final over because every millisecond counts.

To accelerate content delivery across India’s diverse ISP landscape, Airtel leveraged Google Cloud’s Media CDN. By utilizing Google’s extensive global edge network, Airtel was able to serve viewer requests from edge locations that were physically close to end users. This deep localization was a cornerstone of the broadcast's success, with 99.9% of all tournament traffic being served locally from within India.

This efficient architecture minimized network hops and reduced transit congestion, translating into remarkable infrastructure and viewer experience metrics throughout the 74 matches:

  • Superior caching efficiency: Airtel saw an overall cache hit ratio exceeding 98%. By effectively absorbing massive viewer traffic load at the edge, origin server/video platform demands remained minimal even during peak playoff viewership.

  • Consistent ultra-low latency: Airtel maintained a p99 latency of < 300 ms during the tournament, which supported fast stream start times and minimized buffering risk during critical game moments.

Proactive strategies for operational readiness

While maintaining an intelligent backend architecture was vital to Airtel’s IPL streaming strategy, it was  only half the equation. Executing high-stakes live broadcasts across 74 consecutive matches also demanded meticulous operational preparation and proactive match-day execution.

Because Airtel and Google Cloud recognized that potential bottlenecks had to be identified long before the first ball, they established a deeply integrated operational support model:

  1. Pre-tournament support readiness reviews: Well ahead of the opening match, joint engineering teams conducted comprehensive support readiness reviews. By auditing traffic projections, reviewing manifest configurations, and validating failover mechanisms early, the teams supported robust client readiness, resulting in low operational friction during the tournament.

  2. Monitoring as a service (MaaS): The teams maintained continuous, proactive telemetry monitoring through MaaS on Media CDN, and real-time observability enabled early detection and mitigation of network shifts before anomalies could impact viewer playback.

  3. Dedicated match-day and weekend support: Live sports don't play by the rules of  standard business hours, so Airtel established comprehensive monitoring protocols for every match. During critical weekend fixtures and the high-stakes playoff stage, Google Cloud’s Technical Account Management and MaaS teams worked hand-in-hand with Airtel engineering to provide dedicated, real-time event support.

A blueprint for live broadcast excellence

Airtel’s successful streaming of IPL 2026 demonstrates that handling extreme concurrency is only possible with an integrated strategy across architecture, edge localization, and operational governance. By combining a 98%+ cache hit ratio with 99.9% local delivery and proactive match-day monitoring, Airtel hit a benchmark for live sports broadcasting at scale.

This deployment provides an overview of the technical architecture and operational strategies involved in scaling live media delivery for high-concurrency events. To learn more about optimizing live broadcasts and edge delivery, review the Media CDN developer documentation.

How KDDI built Buffmee, a faster, reliable consumer RAG app

8 septembre 2026 à 18:00

When building consumer-facing generative AI applications,  balancing high generation quality with fast response times across diverse media types, can be challenging. KDDI, a major telecommunications carrier in Japan, tackled this challenge head-on when they developed Buffmee, their consumer Retrieval-Augmented Generation (RAG) app.  

Buffmee is an interactive AI service built on the concept of 'AI that helps you grow.' By grounding responses in over 100 sources — including books, magazines, and web media — it helps users search for information, summarize key points, and explore personalized learning and hobby interests. By citing sources, Buffmee alleviates concerns about information reliability, allowing users to safely deepen their knowledge.  To achieve this, KDDI collaborated closely with their development partner KDDI iret, Google Cloud Consulting and our specialized AI engineers.

As part of their app launch, the engineer team needed to ground a massive variety of proprietary content, including books and magazines. However, they struggled with latency issues that prevented them from meeting their target response times, and they needed a reliable way to ensure hallucination-free results.

image1

Buffmee App Description and Images

To meet these performance targets, organizations need a systematic approach to AI evaluation and real-time bottleneck identification. That is why we are sharing the automated evaluation framework and performance optimization techniques that helped KDDI successfully launch their application. 

The results were inspiring: KDDI reduced total application response latency by 38%, successfully hitting their target response performance. They also achieved a nearly 18% improvement in TTFT.

"Our vision hinged on a platform where content, once ingested, would instantly function as a working RAG system. Google's careful, hands-on guidance made that a reality — we're sincerely grateful for their support." — Shunya Onoda, AI Product Department, KDDI.

With these performance and accuracy improvements, Buffmee now empowers users to safely explore their favorite media through interactive Q&A and deep-dive analysis, delivering a highly personalized experience while maintaining strict trust and compliance for content providers.

Let’s deep dive into how they achieved these results. 

Establish automated evaluation for diverse content

Traditional manual testing requires immense effort and cannot scale to accommodate a large content library. To solve this, the development team designed a systematic AI evaluation process using Gemini Enterprise Agent Platform Evaluation Service.

By implementing automated evaluation frameworks like LLM-as-a-Judge and the Rule of Hundreds, the team replaced labor-intensive manual testing with a data-driven process. They ingested their extensive document corpus, constructed hundreds of automated evaluation tests, and built a comprehensive benchmark dataset to measure the reliability of answers for each use case. As a result, the team improved their groundedness scores by 25%, helping deliver highly accurate and reliable outputs.

image2

KDDI's automated evaluation loop: AI generates questions and scores answers, while humans calibrate thresholds and analyze edge-case failures.

Identify bottlenecks and optimize performance with an agentic loop

To improve response speeds, the team implemented BigQuery Agent Analytics and the Agent Development Kit (ADK) log analysis agent. By analyzing actual production logs, they visualized how skill division and prompt bloat—especially with highly complex, multi-page system prompts — impacted the Time To First Token (TTFT).

The team optimized the system prompt, including the inline integration of skills, and reviewed the sub-agent routing. This allowed them to identify and resolve deep-stack bottlenecks in real time without sacrificing response accuracy.

Four core principles for reliable evaluation 

To achieve these results, the team implemented four core technical practices:

  1. Transitioning to binary evaluation: By selectively moving away from ambiguous 1–5 ratings to a binary "pass (1) / fail (0)" system for critical metrics, the team minimized variance and noise, helping improve automation accuracy.

  2. Strategic content sampling: Rather than attempting to evaluate every single document, the team classified their entire corpus along a two-dimensional grid: File Format (Web articles, EPUBs, PDFs, structured data) and Media Composition (Text-heavy, image-heavy, or mixed). By selecting representative samples from each cell of this difficulty grid, they reduced the evaluation workload by 75% while maintaining comprehensive test coverage.

  3. Thresholds grounded in product judgment: Instead of relying solely on default tool parameters, the product owner reviewed randomly sampled answers alongside their automated scores to calibrate and establish what "good enough to ship" actually meant for the user experience.

  4. Modular splitting of massive prompts into ADK Skills: Because massive system prompts exceeding 800 lines can cause LLM attention drift and latency degradation, the team split prompts by function into Agent Development Kit (ADK) Skills, dynamically loading only the required logic to optimize response times.

Get started

Building scalable, reliable generative AI applications requires both automated evaluation and deep performance analytics. To apply these techniques to your own applications:

How Yahoo optimizes resources with flexible VMs in Managed Service for Apache Spark

4 septembre 2026 à 18:00

As a global media and technology company connecting hundreds of millions of users to finance, sports, and entertainment platforms, Yahoo operates a massive data infrastructure where analytics workloads must run continuously at high speed. In deadline-driven data environments, relying on fixed virtual machine (VM) configurations creates a brittle system; if a specific machine shape faces a regional capacity constraint, cluster provisioning in Managed Service for Apache Spark (formerly Dataproc) can experience delays and stall critical data pipelines.

Yahoo utilizes flexible VMs in Managed Service for Apache Spark clusters to automatically absorb these resource fluctuations by defining a ranked list of acceptable VM shapes. This allows the system to dynamically search regional zones and maintain pipeline execution without manual intervention. To search for capacity across a region, teams must also enable Auto-Zone placement.

This optimization builds on Yahoo's broader data modernization journey, which involved migrating on-premises Hadoop and big data estates directly to Google Cloud. By transitioning those legacy workloads, the team established a cloud foundation capable of running high-scale batch and streaming analytics with dynamic resource flexibility.

This post provides a technical blueprint for configuring flexible VM instance rankings in Managed Service for Apache Spark to automatically manage capacity constraints and maintain pipeline execution.

Operational trade-offs of static configurations

Configuring clusters with a single, fixed machine type in a specific zone introduces constraints when regional zonal capacity fluctuations occur, potentially impacting cluster provisioning. Rather than manage these capacity variations through custom retry logic or manual intervention, using flexible configurations allows your infrastructure to automatically adapt. By accepting multiple VM shapes and searching across zones in the selected region, flexible configurations help streamline provisioning to better support high-scale analytics workloads.

Rules for configuring flexible clusters

Deploying flexible configurations requires aligning several connected design choices:

  • Enable auto-zone placement: You must pass a region(--region=${REGION}) or an empty zone string (--zone="") so Managed Spark can search for available capacity across the entire region.

  • Maintain core and memory symmetry: If your Managed Spark cluster uses autoscaling, all machine types in your flexible list must share a similar core count and memory size, even if they come from different VM families. A uniform CPU-to-memory ratio across primary and secondary workers prevents performance degradation, as the smallest ratio determines your effective container sizing.

  • Align component properties: Managed Spark calculates system properties based on VM cores and memory. When mixing machine shapes, you may need explicit property overrides to keep YARN and Spark resource allocations aligned with your expected worker behavior.

Two ways flexible VMs support massive workloads

For large-scale data environments, flexible configurations support operations in two ways:

  1. Higher cluster creation success: Instead of failing when a preferred VM type is out of stock, Managed Spark selects from a ranked list to keep provisioning moving.

  2. Better regional resource use: Auto-zone placement searches the entire region to find capacity, which reduces provisioning friction during high-demand periods.

gcloud example

code_block
<ListValue: [StructValue([('code', 'gcloud dataproc clusters create analytics-cluster \\\r\n --region=us-central1 \\\r\n --zone="" \\\r\n --num-workers=10 \\\r\n --master-instance-selection=\'{"machineTypes":["e2-standard-8"],"rank":0}\' \\\r\n --master-instance-selection=\'{"machineTypes":["n2-standard-8"],"rank":1}\' \\\r\n --worker-instance-selection=\'{"machineTypes":["e2-standard-8"],"rank":0}\' \\\r\n --worker-instance-selection=\'{"machineTypes":["n2-standard-8"],"rank":1}'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f8e31ac9450>)])]>

API example

You can also build this capacity policy into your automated pipelines or Managed Service for Apache Airflow DAGS using the instanceFlexibilityPolicy field in the ‘Dataproc’ API:

code_block
<ListValue: [StructValue([('code', '{\r\n "projectId": "PROJECT_ID",\r\n "clusterName": "analytics-cluster",\r\n "config": {\r\n "gceClusterConfig": {\r\n "zoneUri": ""\r\n },\r\n "secondaryWorkerConfig": {\r\n "numInstances": 8,\r\n "instanceFlexibilityPolicy": {\r\n "instanceSelectionList": [\r\n {\r\n "machineTypes": ["n2-standard-8"],\r\n "rank": 0\r\n },\r\n {\r\n "machineTypes": ["e2-standard-8", "t2d-standard-8"],\r\n "rank": 1\r\n }\r\n ]\r\n }\r\n }\r\n }\r\n}'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f8e32253250>)])]>

This API policy achieves the same goal: it establishes your preferred shape, documents valid fallbacks, and lets Managed Spark resolve resource constraints without breaking your automation scripts.

Establishing an infrastructure policy

Managing data at this scale requires standardizing a clear resource policy rather than relying on a single rigid machine type. Your configuration standards should outline:

  • Preferred and fallback VM families for secondary workers.

  • Default auto-zone placement to enable flexible provisioning.

  • Identical core and memory configurations when using autoscaling.

  • Uniform CPU-to-memory ratios across all worker groups to maintain predictable container sizing.

  • Explicit YARN or Spark property overrides to guarantee consistent runtime behavior across different machine lines.

  • Shuffle-safe patterns for Spark workloads running on Spot or highly elastic capacity.

By adopting flexible configurations, you turn infrastructure scarcity into a predictable fallback plan, keeping your critical data pipelines up and running.

Yahoo impact and results

By implementing flexible VMs in Managed Service for Apache Spark, Yahoo successfully reduced cluster provisioning failures by 85% which were caused by regional capacity stockouts. This flexible configuration allows their data infrastructure to automatically handle capacity constraints and successfully provision resources without requiring manual intervention. As a result, Yahoo ensures continuous workload execution and prevents downstream processing delays across their massive data pipelines.

"Managing high-scale data analytics at Yahoo requires resilient, automated infrastructure. Moving to flexible VMs in Managed Service for Apache Spark has transformed our approach; instead of stalling when a specific machine shape faces capacity constraints, our clusters now automatically pivot to our ranked fallback options. This has helped us reduce provisioning failures by 85%, providing the reliability we need to keep our global media platforms running smoothly." - Akshay Jain, Senior Software Developer Engineer, Yahoo!

Strategic benefits of flexible infrastructure

Adopting a flexible compute stack transforms your environment into a dynamic pool of resources that adapts to your operational needs. By moving away from rigid, single-machine type configurations, you ensure that your workloads reliably access the compute they need, regardless of supply fluctuations. This shift not only maximizes workload obtainability and reliability but also facilitates seamless hardware modernization by allowing you to prioritize newer VM generations while maintaining older types as reliable fallback options.

Build your resilient data pipeline

Transitioning to a fluid compute strategy ensures your critical analytics remain operational despite regional resource shifts. Here is how you can begin optimizing your infrastructure today:

  1. Audit your workloads: Identify applications tightly coupled to specific VM families or zones and map out viable alternative hardware shapes.

  2. Standardize resource policies: Explore the documentation for Managed Spark flexible VMs to establish your preferred and fallback VM families.

  3. Align financial strategy: Utilize Flexible Committed Use Discounts (Flex CUDs) to maintain cost predictability when workloads dynamically pivot to alternative machine types.

  4. Claim your credits: New customers may be eligible for $300 in credits to try Managed Service for Apache Spark and other Google Cloud products at no cost.

How BlackLine simplifies perimeter policy intelligence with VPC Service Controls

1 septembre 2026 à 18:00

Establishing network-level perimeters with VPC Service Controls (VPC-SC) is a critical step that can help you protect your cloud environment against data exfiltration, compromised accounts, and insider threats.

Today, Google Cloud is excited to share new policy intelligence capabilities in VPC-SC that can help drive even greater operational simplicity. With our latest release of the VPC-SC violation analyzer and violation dashboard, we have simplified policy management and troubleshooting, to make managing and optimizing your security perimeter more efficient and straightforward than ever. 

How BlackLine streamlines incident response

BlackLine, a leader in financial operations management, adopted the VPC-SC policy intelligence solution to maintain strict security perimeters. Chosen by over half of Fortune 500 companies, BlackLine uses Google Cloud's full suite of managed services and built-in security capabilities to protect sensitive customer financial data.

VPC Service Controls are the foundation of BlackLine's preventative compliance and security controls in our Google Cloud environment, helping us to mitigate data exfiltration risks and ensure clear separation between our higher and lower environments by establishing strong security perimeters.

Managing these complex perimeters is a continuous process. VPC Service Controls violation analyzer helps BlackLine cloud infrastructure administrators adapt to changing API connection requirements of the business by adjusting security perimeters through approved access levels, ingress policies, and egress policies. 

With only the troubleshooting token or unique ID from any VPC-SC violation error message, we can produce a detailed report identifying the principals and target resources involved in a failed API request, and explaining why and how that API request violated BlackLine's service perimeters. We don’t need to write a Cloud Logging SQL query to extract the data.

The clear access context and actionable insights in the violation details report are an invaluable starting point as we collaborate to resolve violations, significantly reducing our mean-time-to-resolution (MTTR) for service perimeter issues, and helping BlackLine maintain our focus on our customers and continue to innovate on their behalf.

Streamlining the perimeter operations lifecycle

Our new policy intelligence tools — the VPC-SC Violation analyzer and Violation dashboard — simplify real-time monitoring and active incident response. These tools provide clear, actionable insights in the Google Cloud Console, offering greater speed and automation to help you confidently enforce least-privilege perimeters, and quickly resolve access denials.

Violation Dashboard aggregates and visualizes all service perimeter violations across your entire Google Cloud organization in a single pane of glass, helping your team identify trends, spot spikes in access denials, and shareable filters on violations by specific perimeters, projects, or identities.

Violation Analyzer streamlines investigating violations, eliminating the need to query Cloud Logging and manually piece together the details. When you click a troubleshooting token from the dashboard (or input a unique denial ID), the analyzer maps out the identity, source, target, and VPC-SC rule triggered, creating a report telling you why that specific request was blocked. This helps your team more quickly take action to determine whether to modify existing policy rules or create a new one, and resolve incidents more quickly.

Together, the new VPC Service Controls policy intelligence tools go beyond automated log analysis to provide unified visibility of violations and actionable insights to investigate them, making your perimeter deployment and management simpler and lower-risk.

1

Streamlining the VPC Service Controls lifecycle, from deployment to policy refinement.

With the new VPC-SC troubleshooting tools you can more easily:

  1. Test new perimeters (deployment): Use the violation dashboard to visualize the impact of a service perimeter during your initial dry run phase, helping to verify that enforcement is accurate and predictable before it affects production traffic. Filter violations to track and resolve with prebuilt contextual filters for principals, service perimeters, enforcement type, and more.

  2. Track perimeter denials (monitor): The violation dashboard offers a unified view of your perimeter health, allowing your security operations team to monitor status in real time, including dynamic agentic access denials.

  3. Triage an event (investigate): Violation analyzer provides the identity, source, target, and operations for any violation. It cross-references identity and access management (IAM) permissions, resource ancestry, and context evaluation to identify which rule was triggered, reducing manual effort.

  4. Fix the rule (refine policy): Instead of searching through configuration files, violation analyzer maps violations directly to the relevant line in your VPC-SC policy, allowing you to make updates more quickly and with less manual overhead.

output_hq

The VPC Service Controls violation dashboard produces detailed reports to jump-start perimeter access investigations that are simplified using the violation analyzer.

Core VPC-SC operations: Simple perimeter enforcement

Our new troubleshooting capabilities build on VPC Service Controls’ foundational simplicity for designing, enforcing, and managing strong perimeters. 

By using dry run mode, your teams can build precise, contextual ingress and egress rules based on observed traffic — without disrupting vital business workflows. Once you validate these access patterns, moving to full enforcement becomes a more confident, data-driven process. To keep perimeter maintenance more efficient and straightforward, scoped policies allow you to delegate management directly to project-level administrators, empowering the teams closest to the workload.

Getting started

Simplify data security with VPC Service Controls. With the new Violation Analyzer and Violation dashboard, you can spend less time investigating incidents and more time safely scaling your cloud initiatives. Your data is your most valuable asset — protect it with a perimeter that’s as simple to manage as it is effective in enforcing controls.

Learn more and get started with the VPC-SC violation analyzer and violation dashboard in our documentation.

Reimagining work: How Pythian’s internal AI playbook delivers customer ROI

27 août 2026 à 18:00

When Pythian rolled out Google Cloud’s Gemini Enterprise across our 500-person company in 27 countries, the goal was simple: use our own company as a proving ground to discover how enterprise AI actually delivers ROI.

What we found changed our strategy entirely.

Since the rollout of Gemini Enterprise and our previous enterprise AI deployments, Pythian observed firsthand why so many enterprise AI initiatives stall out or fail. 

Most organizations trap themselves in a tool-centric mindset — buying licenses, making tools broadly available, and assuming value will naturally follow. They get stuck chasing "nickel and dime" micro-efficiencies (like saving 5 minutes per user) while missing structural, high-ROI workflow transformations. Compounding the problem, even when custom agents are built, they frequently stall in pilot mode or break down in production because teams lack the operational capability to manage AI model drift, agent lifecycles, and ongoing observability.

To solve this, we engineered the Pythian AI Operating Model — a multifaceted, end-to-end framework designed to take enterprise AI from high-level strategy all the way into sustained production. While our dual center of excellence (COE) serves as the core execution muscle, it is the application of the entire framework, from Field CTO strategy and tooling deployment to the dual COE and XOps, that consistently unlocks million-dollar outcomes.

By proving this complete model internally first, Pythian drove a 3x surge in active user engagement and cut our database incident resolution times by 80%.

The four pillars of the Pythian AI operating model

To move past the common failure points of enterprise AI, our framework consolidates strategy, execution, and operations into a single continuous loop:

Field CTO strategy  ──>  tooling deployment  ──>  dual COE execution  ──>  production XOps

  1. Field CTO strategy and governance: Generative AI is arguably the most academically challenging architectural shift in IT history. Led by former C-suite tech leaders, our Field CTO practice provides executive advisory to establish steering committees and clear value metrics. The team audits operations using 16 horizontal agentic patterns (like automated document processing and runbook creation) to build a prioritized backlog of high-ROI use cases before development starts.

  2. Tooling and platform deployment: The team establishes a secure, production-grade foundation on platforms like Gemini Enterprise and connects AI directly into CRMs, ERPs, and database estates to ground models in real corporate context.

  3. The dualCOE: This execution muscle is split into two specialized engines:

  • People productivity COE: This group handles adoption and change management. Instead of expecting non-technical teams (like HR or Procurement) to build its own agents, this COE builds no-code agents for them, focusing entirely on enablement.

  • Process productivity COE: This team engineers deep, custom-coded AI agents and complex agentic workflows that integrate into core data platforms for autonomous operations.

  • XOps (AI production management): While deploying an agent is 20% of the journey,  maintaining accuracy in production is 80%. Because AI models and prompt structures naturally drift over time, this XOps practice provides the continuous monitoring, prompt tuning, and model observability needed to keep agents performing without breaking core workflows.

  • The difference between chasing minor, scattered efficiencies and driving structural enterprise ROI comes down to how you align your operating strategy:

    Alignment element

    Tool-centric approach

    Pythian AI operating model

    Primary metric

    Individual minutes saved per user

    High-impact workflow reimagination and ROI

    Operational focus

    Broad, unguided tool availability

    Prioritized backlog via 16 agentic patterns

    Execution muscle

    Ad-hoc user experimentation

    Dual COE (people and process productivity)

    Production lifecycle

    Unmonitored static deployments

    Active XOps (Continuous accuracy and drift management)

    Real-world impact: from database ops to global supply chains

    Whether managing 70 manufacturing plants or 30,000 enterprise databases, AI succeeds when tied to structural, high-value workflows:

    • Pythian “as a customer:” Across 15,000 monthly database tickets, our Process COE deployed an agentic workflow that reads tickets, searches knowledge bases, and auto-generates mini runbooks before an engineer touches them. The result was slashed mean time to resolution by 80% and tripled active user engagement.

    • Knowledge management customer: We deployed autonomous IT support agents across 10,000 consultants. As a result, we were able to automate 10% of 20,000 annual IT tickets into "no-touch" resolutions, saving 1,000,000+ operational hours.

    • Supply chain customer: By building custom agentic supply chain tools on Gemini Enterprise, we compressed forecast-matching cycles from weeks down to 2–3 days across 70 global manufacturing sites.

    • Retail customer: We combined Gemini Agentic AI and computer vision to automate store product onboarding. As a result, we transformed a 20-minute manual task into a multi-second flow.

    Ready to build your AI operating model?

    Scaling AI demands more than tool-level experimentation. It also requires an end-to-end AI operating model. Learn how Pythian pairs with Google Cloud to operationalize strategy, streamline XOps, and fast-track your Gemini Enterprise journey.

    How Uber improves network reliability while unblocking cloud migration

    26 août 2026 à 18:00

    Uber has a lot in common with the cities it serves. Both are always changing and growing, both must carefully manage the resulting traffic to prevent congestion and sprawl.

    Uber has continuously evolved its technical strategies to manage its expanding network, and this careful planning and constant evolution helps ensure that application traffic across its entire platform runs smoothly. Ultimately, maintaining a reliable, high-scale platform that operates seamlessly at any given time is key to preserving user trust.

    One important solution in this effort has been application awareness on Cloud Interconnect. An industry-first tool for application prioritization across hybrid networks, application awareness on Cloud Interconnect has helped Uber prioritize critical traffic to ensure business continuity during potential network congestion events. 

    Uber acted as an early design partner for application awareness on Cloud Interconnect, helping ensure that this capability met the demands of Uber’s global-scale operations. It not only improved Uber’s daily operations, it also gave Uber the confidence to move forward with a Google Cloud migration, with confidence that there would be less risk of service interruptions during switchovers. 

    In this post, we’ll explain the features Uber most sought and why, the inner workings of application awareness on Cloud Interconnect, and how it can help other organizations as well.

    Prioritizing critical traffic

    When migrating distributed, hybrid, or multicloud applications at a global scale, network reliability becomes a primary concern. Even the most worthwhile migrations may not seem worth it if such migrations interrupt ongoing service. For organizations like Uber, moving vast amounts of data to support large data analytics workload — including emerging AI use cases — can saturate network links, resulting in increased reliability risk for their critical application traffic. 

    With standard cloud interconnect approaches, enterprises typically apply simple bandwidth overprovisioning to meet extreme infrastructure needs. But with today's hybrid cloud demands, and given the size of an organization like Uber, overprovisioning network capacity for peak usage is often too costly and unreliable. 

    The shortcomings of overprovisioning only become magnified with the integration of cutting-edge AI innovations. Uber needs systems in place that can take on massive data transfers without congesting its network and protecting the performance of business-critical applications.

    With the benefit of application awareness on Cloud Interconnect, including the four major features of application awareness — traffic handling, congestion response, latency management, and cost efficiency — Uber was able to achieve the networking optimization its modern tech stack requires.

    aai concept value prop with_without picture

    Starting with a private preview, Uber deployed this feature across its infrastructure, beginning with Google Cloud Interconnect deployments in Phoenix, Arizona, and Ashburn, Virginia. Application awareness on Cloud Interconnect allows Uber to classify and prioritize end-user application traffic over less time-sensitive data using DSCP marking and configured queuing profiles.

    In the following chart, we look at the four key features of application awareness on Cloud Interconnect, how they differ from legacy approaches, and how they help provide better operational continuity for organizations like Uber. 

    Feature

    Standard interconnect solutions

    Application awareness on Cloud Interconnect

    Traffic handling

    All traffic treated equally (first-in, first-out)

    Traffic classified into six distinct traffic classes

    Congestion response

    High-priority application traffic may be dropped during bursts

    Business-critical traffic is protected via strict priority or bandwidth sharing policies

    Latency management

    Unpredictable latency for high priority applications

    Predictable and consistent low-latency for time-sensitive workloads

    Cost efficiency

    Requires expensive overprovisioning to absorb peaks

    Efficient bandwidth utilization and lower TCO

    Uber's key takeaways

    For Uber, the business value of being able to prioritize business-critical traffic on its networks by deploying application awareness on Cloud Interconnect was immediate. And in doing so, Uber has also created a blueprint that other enterprises with similar hybrid cloud challenges can replicate. The core elements of that blueprint include:

    • Ensuring business continuity: Uber can decide in real time which application traffic to prioritize during major, high-traffic events. This means that mission critical applications stay up and running during even extreme events (both planned and unplanned). Uber leadership has called application awareness on Cloud Interconnect important for its global operations. 

    • Efficient bandwidth utilization: Instead of blindly overprovisioning bandwidth to prevent congestion, application awareness allows Uber to better utilize their existing Cloud Interconnect capacity aligned with their expected network bandwidth needs. The result is lower total cost of ownership for network infrastructure.

    • Unblocked workload migration: By protecting critical applications from network congestion, Uber was able to migrate significant workloads to Google Cloud and, in the process, dramatically reduce operational overhead.

    "Application awareness on Cloud Interconnect was the key that unlocked our ability to migrate more strategic workloads to Google Cloud and is critical for maintaining service reliability during peak global demand. By allowing us to intelligently prioritize traffic, it helps us ensure that we can protect our higher priority services and make our infrastructure more efficient, lowering our total cost of ownership. This wasn't just a feature deployment; it was a deep engineering partnership that delivered a solution critical to our business." – Harry Liu, Director of Engineering, Uber

    Securing network reliability for AI and beyond

    As more enterprises integrate cloud-based AI models, distributed applications, and data analytics, it's becoming a business imperative to be ready to handle the massive data transfers that follow. But in doing so, they also have to ensure they never compromise the reliability of their critical applications. 

    With application awareness on Cloud Interconnect, Uber demonstrated that moving beyond simple bandwidth overprovisioning to protect business-critical traffic was an essential step to building the stability required to embrace modern hybrid and multicloud strategies.

    You can read our blog about the potential of Cloud Interconnect across industries to learn more about what the service can bring to your organization, and if you’re ready to explore more, our team of networking and industry experts are ready to help.

    How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2

    18 août 2026 à 18:00

    Enterprise content management is experiencing its biggest architectural shift since the cloud migration era. 

    For years, enterprises have stored trillions of gigabytes of critical data in Box: financial models, clinical trial protocols, M&A due diligence rooms, engineering schematics, and legal compliance playbooks. Up to this point, text-based search and retrieval-augmented generation (RAG) have successfully unlocked the vast narrative knowledge within these repositories, establishing a powerful and highly effective baseline for enterprise AI intelligence.

    Traditional RAG architectures have mastered text processing, but the agentic era demands more. The next logical evolution is to extend this framework to capture the inherently multimodal, deeply spatial, and highly structured elements that exist alongside text. While text embeddings excel at indexing prose, multimodal architectures unlock a major new capability: For example, they preserve the strict row-column semantics of financial tables, interpret visual evidence like clinical data, and map the logic of multi-page flowcharts without losing their spatial layout.

    To deliver next-generation capabilities that can handle the vast universe of digital content, Google Cloud and Box are integrating advanced multimodal capabilities into Box's Agentic Platform, powered by Gemini Multimodal Embeddings 2 merging Box’s industry-leading Intelligent Content Management platform with Google Cloud’s advanced AI embeddings.

    Benefits of improved embedding: Extending the dimensions of document content

    1. Preserving visual and spatial geometry: Complex document elements like multi-column tables or financial matrices rely on their spatial layout to convey meaning. Converting these elements into a flat string of text can disassociate column headers from their corresponding data points. Multimodal embeddings allow systems to interpret the document exactly as a human does, maintaining the integrity of spatial relationships.

    2. Illuminating the visual modality: Enterprise documents are filled with visual indicators: technical charts, process flowcharts, branding assets, and product photography. Multimodal capabilities ensure that these elements are no longer invisible to search systems, allowing users to query images and text simultaneously.

    3. Connecting hybrid file formats: Real-world business workflows rarely live in a single document format. An agent may need to cross-reference a PDF policy, a spreadsheet tracking log, and a presentation deck. Extending RAG with multimodal embeddings creates a unified understanding across these varied formats.

    The Architectural Solution: Gemini Multimodal Embeddings 2

    Google Cloud’s Gemini Multimodal Embeddings 2 introduces a unified, multimodal vector space capable of embedding text, raster images, document pages, rendered spreadsheet tables, and visual charts into the same semantic representation space.

    GIF_1

    Key product capabilities unlocked by gemini-embeddings-2:

    • Crossmodal retrieval (text-to-visual / visual-to-text): Enables natural language queries to retrieve highly specific visual components, such as locating a target chart or diagram within a massive library of slides, without requiring manual tagging.

    • Layout-aware document embedding: Rather than breaking files into arbitrary text blocks, the system can embed document page renderings directly, preserving visual hierarchies, callout boxes, and structural context.

    • Heterogeneous format bridging: Native support for seamlessly bridging content across .docx, .xlsx, .pdf, .pptx, .png, and .csv without losing modality-specific structural information.

    Three core patterns of multimodal enterprise agents

    By leveraging multimodal embeddings within Box, we have identified three uniqueprimary design patterns that illustrate how organizations can extend traditional RAG to support complex, visual workflows.

    Pattern 1: Complex financial & analytical reporting

    The challenge

    Corporate finance, research, and audit teams analyze highly structured documents where vital data resides in embedded tables, growth charts, and footnote annotations. Text-only indexing can separate these numbers from their context, making automated analysis challenging.

    The multimodal advantage

    • Structural alignment: The embedding model captures the physical structure of tables and charts, allowing financial agents to understand that a column header applies to a specific row of metrics.

    • Visual trend analysis: Agents can cross-reference written summaries with visual trends in accompanying bar or line charts, identifying and pointing out discrepancies between written claims and source data.

    • Contextual sourcing: Users can query complex portfolios and instantly retrieve the exact page, table, or chart supporting a specific metric.

    2

    Pattern 2: Multimodal clinical decision support & assisted diagnosis

    The challenge

    In healthcare and clinical environments, critical patient data is fragmented across vastly different, unstructured visual and textual formats — ranging from external physical photos (visual evidence) and microscopic pathology slides (lab reports) to structured risk matrices (triage grids). Traditional text-based systems or isolated analysis tools cannot synthesize these cross-modal relationships simultaneously, which can delay critical diagnoses or risk missing immediate, life-threatening procedural complications.

    The multimodal advantage

    • Cross-modal clinical synthesis: Evaluates physical symptoms alongside cellular-level laboratory evidence simultaneously by indexing clinical photos, histopathology imagery, and triage grids into a single space.

    • Granular anomaly identification: Connects niche visual patterns under a microscope (like parasitic cyst walls) with medical knowledge to rapidly isolate rare conditions.

    • Risk-aware decision support: Cross-references findings against triage frameworks to deliver instant warnings about immediate patient risks, such as life-threatening anaphylactic shock.

    3

    Pattern 3: Cross-document multimodal synthesis & data reconciliation

    The challenge

    Enterprise information is fragmented across disconnected files and formats (e.g., PDF minutes, Excel charts, PNG flyers, and email threads). Traditional tools analyze these files in isolation, failing to connect the dots when verifying details or resolving data contradictions across independent documents.

    The multimodal advantage

    • Cross-file synthesis: Connects information across entirely different formats (PDFs, spreadsheets, images, emails) simultaneously to answer complex business queries.

    • Conflict resolution: Flags and resolves contradictions between assets, such as catching outdated pricing on an image by cross-checking it against the latest financial spreadsheets.

    • Visual-to-text auditing: Audits visual or scanned files against text-based records (e.g., verifying a signed PDF contract against a legal review email) to catch missing clauses or changes.

    4

    The future of agentic enterprise content management

    The integration of gemini-embeddings-2 into Box’s Agentic Platform is an important new capability to improve the next era of content intelligence. Multimodal embeddings help Box to move beyond basic search to active, intelligent collaboration.Box's Intelligent Content Management platform represents a fundamental shift in enterprise AI infrastructure — moving beyond passive document storage to deliver a governed, semantically indexed reasoning layer where AI agents can interrogate, cross-reference, and act on content with full compliance and security controls already in place. 

    Powered by multimodal embeddings and a suite of native AI agents spanning search, metadata extraction, research, analysis, and composition, Box enables organizations to proactively surface insights such as flagging stale pricing data, expiring contract clauses, or cross-document contradictions before they become business risks. For high-complexity industries like financial services, life sciences, and legal operations, Box's ability to reason across text, tables, charts, and images makes multimodal understanding a competitive requirement. 

    Designed to interoperate with the broader enterprise AI ecosystem, Box serves as the single governed content foundation that ensures every AI-driven workflow is grounded in authorized, auditable enterprise data.

    When you think about it, the enterprise data landscape was always multimodal. Now we have the technology to make the most of it. By integrating gemini-embeddings-2, Box helps its users unlock unprecedented value from unstructured enterprise content. Product leaders who embrace multimodal-first architectures, rigorous precision benchmarking, and audit-ready grounding will lead the next wave of enterprise productivity and innovation.

    The team would like to thank Ken Ikeda, Afshaan Mazagonwalla, and Samip Thakkar for their work on this project.

    Building operational resilience with agentic AI in financial services

    18 août 2026 à 16:00

    For financial institutions, operational resilience has long been embedded in regulatory and supervisory expectations — to say nothing of the high expectations of consumers. With the implementation of the European Union’s Digital Operational Resiliency Act (DORA), those expectations have become even more stringent, with more explicit, harmonized, and evidence-driven requirements. Firms must now demonstrate that their critical business services and supporting digital infrastructures can withstand disruption, support coordinated response, and recover with control.

    To meet these conditions, Deutsche Bank developed an AI-powered agentic resilience platform that modernized its regulatory tabletop resilience exercises at scale and turned manual preparation into context-aware and evidence-ready simulations grounded in actual operational data. The platform builds enterprise context from architecture, data flows, logs, incident history, alerting signals, and operational telemetry to generate scenarios, simulated operational evidence, structured session records, and regulator-ready artifacts.

    At many large banks with operations that span interdependent applications, data flows, and third-party services, this is a critical and even existential shift. Across financial services, supervisory expectations are evolving and as they do, banks’ tabletop exercises must reflect their production dependencies, real operating conditions, and compliance with consistent evidence standards more directly.

    As Deutsche Bank considered how to successfully and efficiently make this shift at scale, it looked to its long-time partner, Google Cloud, and its growing suite of agentic AI tools.

    From tabletop exercises to resilience intelligence

    With its agentic resilience platform, DB has been able to transform its tabletop exercises from manual preparation to a continuous intelligence model. And it’s been able to extend the same agentic layer to root-cause analysis when real operational context is needed.

    This means that every scenario it runs is based on real enterprise signals. The platform can then reflect true system dependencies, failure patterns, and business impact instead of relying on static inputs that are more likely to return assumptions than real-time insights.

    By using Gemini Enterprise Agent Platform, DB has been able to migrate this operational context into structured scenarios with clear timelines, decision points, and expected responses. This has ensured that each exercise is grounded in real system behavior that produces consistent, audit-ready evidence that meets regulatory expectations.

    Dual orchestration for control and flexibility

    In order to deliver both regulator-grade control and operational flexibility, DB’s platform introduced a dual-orchestration architecture that separates workflows into two complementary execution models.

    First, for regulator-aligned execution, the bank is using LangGraph to ensure that it generates every scenario through a traceable, deterministic process — with clear lineage from input context to output — that supports the auditability required for supervisory review.

    Next, for its adaptive and investigative scenarios, DB is using Google Agent Development Kit (ADK) to enable agent-driven coordination. This approach allows the bank’s platform to dynamically analyze conditions and generate responses without predefined execution paths.

    image1

    Figure 1. Architecture for context assembly, orchestration, and scenario generation.

    With this architectural separation, the platform can combine governed execution with adaptive investigation while preserving a common intelligence layer. The same agents and tools can reason over architecture, data-flow diagrams, logs, and code artifacts across tabletop scenario generation and related incident-analysis workflows. Importantly, this supports a consistent resilience model across both planned exercises and real operational events.

    Deutsche Bank’s objective with this platform was to engineer a resilience model for critical financial systems that meets regulatory expectations — even within highly complex, distributed environments. By linking dynamically generated scenarios to real business context and combining governed orchestration with adaptive analysis, the platform has given us an intelligent, continuously adaptive model for operational resilience.” – Sanjay Tripathi, Managing Director, Global Head of Surveillance Technology & Compliance Cloud & AI Transformation Lead, Deutsche Bank

    Powering generation and governance with Google Cloud

    Google Cloud’s suite of agentic tools is providing the foundation for scaling Deutsche Bank’s platform across its many governed, enterprise-grade resilience workflows. Here’s how:

    • Cloud Run supports elastic execution of scenario and evidence-generation services. 

    • Gemini Enterprise Agent Platform transforms operational context into structured resilience scenarios.

    • Google ADK enables adaptive agent coordination.

    • Cloud SQL provides durable persistence for scenarios, session artifacts, and review records.

    Collectively, these services give DB support for the traceable generation, controlled execution, and persistent evidence record required for compliance review and continuous improvement.

    Scalable, evidence-ready resilience testing

    Every scenario generated by Deutsche Bank’s platform drives a structured tabletop session for the teams that run response, escalation, and recovery. Because these exercises are grounded in real enterprise context, they reflect operational reality while also strengthening consistency across teams and creating audit-ready evidence that meets regulatory expectations. For institutions that operate under DORA or similar frameworks, this makes it easier to demonstrate controlled, coordinated, and disciplined response at scale. 

    This model is now being applied across multiple DB portfolios, which is helping the bank establish more consistent and scalable resilience paradigms and a replicable blueprint for the broader financial sector.

    In this model, root-cause analysis acts as the feedback loop between real incidents and future resilience testing. The resulting insights from production events can inform future tabletop scenarios, while exercise outcomes can strengthen response playbooks, escalation paths, and recovery readiness.

    All of this extends the platform’s value from planned resilience exercises to real operational events while keeping scenario-based resilience testing as the primary use case.

    As adoption expands, this platform brings consistency by embedding Google Cloud’s methodology for context-aware resilience. It eliminates fragmented manual approaches and establishes a cross-functional, AI-informed operating model across the bank.

    Toward resilience intelligence

    The bank’s next step is to extend this approach into a broader resilience intelligence layer, which is possible because it can deploy the same patterns to support playbook refinement, recovery-readiness assessments, and continuous validation of controls against evolving system conditions.

    For financial institutions, this is a strategic shift. As systems become more distributed and regulatory expectations more demanding, banks must move from periodic resilience testing to continuous, intelligence-driven capabilities. At Deutsche Bank, Google Cloud is making that transition simple across the organization.

    Learn more about Google Cloud’s methodology for context-aware resilience in this article.

    How WPP operationalizes platform and data engineering for AI marketing

    10 août 2026 à 18:00

    Between chaotic levels of market fragmentation and economic volatility, marketing and communications agencies can no longer rely on the human intuition they’ve traditionally used to win clients and optimize their ad spend. WPP is replacing that guesswork with an AI-powered view of shifting market dynamics, giving brands predictive certainty that lets them invest with confidence while moving at the speed of the market. That’s the value of WPP Open, its agentic marketing system.

    But before it could begin applying sophisticated AI models to power those insights, WPP had to overcome a critical engineering challenge: the marketing data that made up the models was fragmented across hundreds of global agencies. While this dynamic made it nearly impossible to deploy AI tools efficiently and securely, access to models was only part of the equation. And  until it built a reliable way to ingest, clean, and serve data to those models, WPP couldn’t unlock the true potential of generative AI.

    To solve this, WPP partnered with Google Cloud to construct a unified data backbone and  custom platform engineering path. Now, by standardizing its serverless compute patterns and data processing workflows, WPP is able to  securely deploy targeted marketing campaigns in days instead of months.

    Architecting a centralized, service-based data foundation 

    An important part of this effort was accelerating data availability and centralizing management. To do this, WPP adopted a service-based project structure for its current production environment. Rather than isolating every workload into separate silos, its engineering team centralized Google Cloud Storage (GCS) and BigQuery into dedicated, shared data projects, while also segregating the compute and processing workloads into distinct processing projects.

    This structure simplified the core team’s user experience and ensured that all data consumers interacted with a unified source of truth. Because data from WPP’s various product lines lives in shared infrastructure, it was essential that security be strictly enforced at a granular level. By directly applying identity and access management (IAM) controls at the individual GCS bucket and BigQuery dataset levels, the company’s teams only see the data they’re  authorized to access.

    At the same time, raw data from various partners lands in dedicated GCS buckets in order to keep the raw inputs organized and isolated. From there, Managed Service for Apache Spark executes custom Apache Spark jobs to cleanse, normalize, and canonicalize information into standardized cohort definitions (SCDs). By utilizing a serverless architecture combined with Kubeflow for pipeline orchestration, WPP’s data engineering team avoided the overhead that often results from managing cluster infrastructure. This allowed them to focus entirely on the data transformation logic fueling the downstream GCS and BigQuery layers  that ultimately feed the company’s audience & performance AI models.

    1 WPP Data Pipeline Architecture

    What made our collaboration with Google Cloud successful was the balance they struck between uncompromising professionalism when it comes to best practices and timely delivery of incredibly pragmatic, real-world solutions.
    - Jonas Dahlbaek
    Senior Data Engineering Lead, WPP

    Standardizing data into unified cohorts

     When raw data enters WPP’s processing zone, its platform converts it into SCDs that become core concepts used throughout the framework for keying purposes. These are based on five keys: age, gender, geo, product, and interest. But these underlying data definitions are fluid and continuously canonicalized to reflect evolving marketing concepts. As a result, this uniform structure allows WPP to join and aggregate data on a global scale without exposing sensitive underlying particulars or relying on shared identifiers.

    The platform's core processing engine was built in type-safe Scala to ensure comprehensive visibility and compliance This custom framework tightly controls how data is transformed, and it inherently supports full source traceability while guaranteeing that every data point within the curated datasets can be traced back to its origin. This is a crucial level of traceability when building enterprise AI applications, as data scientists and auditors must understand exactly what information feeds into the models, even as WPP concurrently prepares to transition to Google Cloud Knowledge Catalog for automated, enterprise-wide data governance in the future.

    Working with Google Cloud has been instrumental in accelerating and standardizing our engineering efforts. In a world where massive volumes of fragmented data present a daily challenge, having the right infrastructure is paramount to thriving in the AI age and helps our developers and AI marketers alike.
    -Suleman Khan
    Product Manager for OI & Google Partnerships, WPP

    Standardizing the enterprise software lifecycle

    For WPP, even with all these steps in place, processing data is only half the battle. To serve applications and manage the underlying infrastructure, the company’s platform engineering team developed a suite of reusable and centralized GitLab continuous integration and continuous deployment (CI/CD) templates. With this, WPP reduced the cognitive load on individual development teams and ensured that all deployments met strict corporate security standards.

    These templates manage various enterprise workloads autonomously. The suite includes universal Cloud Run templates for full-stack web applications and  batch data processing and scheduled pipelines. It also includes a deploy-only template for multi-stage workflows and a Cloud Run functions deployment template for event-driven microservices.

    2 WPP Cloud Platform Engineering

    Implementing zero-rebuild promotion

    Rebuilding container images in a production environment can introduce unnecessary risk and the potential for configuration drift. In order to maintain environmental consistency, WPP embraced a "build once, deploy many" methodology that applied cross-project IAM logic and Google Cloud Artifact Registry configurations.

    As part of this process, developers build and test container images in the development environment. Once those exact, immutable container images are validated, they’re promote  directly to production. This zero-rebuild promotion ensures total parity across deployment stages and eliminates unexpected production behaviors. The CI/CD templates also facilitate progressive traffic migration, which allowed teams to route a small percentage of traffic to new revisions before initiating a full rollout.

    Immutable deployments. Traceable data. Unshakable trust. When you know exactly what goes into your AI, you can ship at the speed of light.
    - Ranjith K Poldas
    Associate Director , Devops (I&P), WPP Media

    Automating security and intelligent networking

    With this modern architecture, enterprise security acts as a foundational enabler for WPP, so it integrated Wiz security scanning directly into the pre-push phase of the CI/CD pipeline to catch vulnerabilities before code merges. The company also utilized Google Cloud Identity-Aware Proxy to enforce zero-trust access across its  internal applications.

    To further simplify operations, WPP adopted templates with intelligent virtual private cloud (VPC) logic. This configuration automatically identifies and resolves networking conflicts between legacy VPC connectors and modern Direct VPC access. This automated networking prevents deployment failures and accelerates the release cycle.

    Monitoring operational health and driving ROI

    Because a resilient platform foundation requires deep observability, WPP’s engineering team now monitors strict operational metrics instead of relying solely on deployment frequency. The team tracks request latency across p50, p95, and p99 percentiles, alongside 4xx and 5xx error rates. It  also monitors container startup times to mitigate cold starts, while tracking overall CPU and memory utilization. This granularity ensures that both data pipelines and serverless infrastructure always remain highly available.

    "Navigating a transformation of this scale across multiple complex workstreams—spanning data engineering, platform infrastructure, and AI integration—required more than just alignment; it demanded deep, mutual trust. Working as true partners, Google Cloud and WPP moved in lockstep to deliver production-ready platform capabilities on time."
    Yang Yue , Program Manager , Google Cloud

    For WPP, operationalizing its data and AI stacks at this velocity provided the necessary infrastructure for its advanced workloads, and the business impact was clear and quantifiable. By building this dual foundation, the company reduced creative and strategy time from four weeks to just three hours. It also saw a 70% gain in production efficiency, a 33x increase in content volume, and  a 2.8x increase in campaign return on investment. In short, by partnering with Google Cloud and implementing a broad suite of products and tools, WPP was able to quickly realize a significant ROI and boost productivity, efficiency, reliability, and security across the company.

    How Malachyte solves retail’s cold-start problem with managed real-time AI

    10 août 2026 à 18:00

    What’s the best way to recommend products to little-known users? 

    We’ve spent our careers trying to solve this problem for major companies like Spotify and Priceline, and it’s why Sidd founded Malachyte, an AI-powered ecommerce recommendation platform. These days, consumers have come to expect content that feels personalized and relevant, and online services competing for their attention have no choice but to do this exceptionally well.  

    Malachyte was inspired by some unique insights into how advanced AI models, and large language models in particular, could be applied in new ways to old challenges like personalization and recommendations. 

    As Malachyte set out to win potential customers’ business, we needed secure, scalable, reliable and, above all, leading-edge AI infrastructure to continue building the personalization algorithm we had always envisioned. By utilizing Google Cloud tools like Bigtable and Managed Service for Apache Kafka, Malachyte has been able to help some of its retailers double and sometimes even triple their sales. 

    This is the story of how we built it, and the ways any founder can use services like these to start deploying AI foundation models in new ways.

    How Malachyte lifted sales for their users 

    For Malachyte, the aha moment was discovering that it could use neural networks with attention mechanisms — the same concept powering large language models — to personalize retail search and product pages. This approach is what enables LLMs to derive meaning from the relative order of items in a sequence, in their case the order of words and syllables in a sentence. When it comes to a retail website or app, what Malachyte wanted to capture was the sequence of customer interactions with the site.

    1 - Malachyte blog

    What if we predicted the next thing a user wants on an ecommerce website just like LLMs predict the next word in a sentence?

    A pre-GPT language model might have tried to look at a specific sequence of words or even fragments of words (what we now know of as tokens), but those earlier models wouldn’t examine what happens if the words were in the comparable order but weren’t contiguous or were re-arranged. The breakthrough came — in part through Google’s work on transformers — when LLMs gained the ability to understand complex and long-range dependencies within a sequence of items. 

    This more sophisticated method has delivered dramatic results — both for the proliferation of gen AI in general, and for Malachyte’s application of the technology.

    To make this work in practice, Malachyte creates a vector of everything known about a visitor when they arrive on a site.  Most users are visiting for the first time, so little is known about them. This is what’s known as the  “cold start” problem. The trick is to use every interaction with a user to refine this vector. Each new addition to the vector, like a click or a query, does two things: it drives a prediction about the next thing the user wants, and it provides more information about the user.  

    Malachyte’s platform then updates the user vector and the prediction at the same time. This not only enhances the understanding of the individual user and their preferences, it also improves the overall model with the anonymized user data. With every inference, the context of both the average and the specific shopper grows. 

    The company further innovates by not just using attention-based neural networks but combining that with updating user profiles 100 milliseconds at time.

    2 - Malachyte Blog

    Malachyte’s recommendation and search agents populate the next page’s search results or recommendation carousels based on what users clicked on previous pages.

    To be sure, this idea isn’t in itself new. Retailers have long used collaborative filtering recommendations systems to identify similar users and items that required massive sets of interaction history. These models typically required a lot of data, including third-party cookie-based profiles and demographics. 

    By focusing on the sequence of interactions in a session, retailers can achieve far more personalization — with less required data or spend — than by focusing only on a user’s profile. As a bonus, retailers can now offer their users more privacy by not relying on long-term cookie data.

    This works because of the model structure and multimodal vectors that encode everything they know about a user, including browser data, click history and searches. The output, too, is multimodal: The same model can be applied to on-site search product pages, category pages, and add-to-cart carousels.  

    To make this work, each product in the catalog is embedded into the same space as the user vector, which gets updated and subsequently moves the vector closer to relevant products and further from those that aren’t. The neural network computing the embedding is being continuously trained across retailers who work with Malachyte, improving the quality for everyone. The system effectively becomes a data cooperative with each retailer's user helping make the model smarter for everyone.

    3 - Malachyte blog vector space - high res

    A user session represented as a vector in a space of products.

    To make this delivery for every user at every inference in 100 milliseconds, Malachite found real benefits in building onGoogle Cloud’s real-time AI stack.   

    With this system, every behavioral event streams into a Managed Service for Apache Kafka cluster. Rather than queuing for a future training job, each event immediately becomes an update to the user’s profile in Bigtable. The Kafka cluster allows the customer’s front-end to persist, so the user session signals quickly with little worry about how they fit into the user vector.  

    Bigtable allows Malachyte’s services to look up and update the right user vectors, and it and Kafka operate at the order of 10 milliseconds per step, which allows the entire recommendation loop to complete with no disruption to the user experience. 

    4 - Malachyte Blog

    The three layer real-time AI architecture: a retailer’s website, Malachyte’s AI models and serving front ends, and context management infrastructure.

    In addition to a fast core, a second layer of product catalog updates, inventory signals, and retailer-specific dimensional data keeps product data up to date. This flows through Cloud Pub/Sub, which offers globally accessible REST APIs that enable connections retailers can use without deep integration work. Malachyte agents run on Google Kubernetes Engine (GKE), with model inference on Google Compute Engine (GCE).

    5 Malachyte blog

    Continuous ingestion of external data, such as product catalog updates, operates through Pub/Sub’s global messaging system.

    With its migration to Google Cloud’s AI architecture, Malachyte demonstrated that production AI inference and training are about more than GPUs and storage. They require real-time continuous learning infrastructure that includes a fast key-value store, a streaming layer, and a managed messaging system, all integrated with the foundation model architecture. 

    This approach also shows that even a small team like Malachyte’s can have a big impact in an industry. It just needs access to powerful infrastructure and core AI managed services.

    Try it for yourself 

    Looking to shake up your industry or stay ahead of the competition like Malachyte? Try Managed Service for Apache Kafka, Cloud Pub/Sub, and Bigtable. New customers can receive $300 in Google Cloud credits.

    GOL! How TelevisaUnivision streamed the FIFA World Cup to millions with Google Cloud

    7 août 2026 à 18:00

    Live sports broadcasting represents the ultimate stress test for digital media infrastructure, where operational success or failure is measured in milliseconds and observed live by millions of viewers simultaneously. During the 2026 FIFA World Cup, the stakes reached a high for TelevisaUnivision, the leading Spanish-language media conglomerate. With Mexico serving as both a primary host nation and a core contender on home soil, fan engagement created unprecedented demand across Latin America and TelevisaUnivision's ViX streaming platform.

    For a marquee broadcaster like TelevisaUnivision, high-stakes events carry direct, long-term brand and reputational implications. Audiences demand uninterrupted, pristine access to every critical moment of play. Playback interruptions during a key goal, login latency at kickoff, or degraded stream resolutions immediately impact customer satisfaction, risking subscriber churn and brand dilution. When streaming tier-1 global sports events, technical execution directly impacts consumer trust, requiring an absolute commitment to zero-downtime availability and flawless performance on the part of the provider.

    Why TelevisaUnivision selected Google Cloud  

    Navigating Latin America's complex networking ecosystem, which is marked by heavy ISP fragmentation and cross-border transit bottlenecks, required more than a standard vendor relationship. TelevisaUnivision needed a strategic partner willing to make joint investments in network capacity, infrastructure resiliency, and custom feature development. TelevisaUnivision selected Google Cloud's Media CDN based on two foundational differentiators: its architecture, and Google Cloud’s customer focus.

    Platform architecture 

    Media CDN provided a resilient, globally distributed infrastructure built specifically to absorb massive live-stream traffic spikes while protecting origin infrastructure. Core architectural advantages included:

    • In-ISP deep edge caching: Media CDN embedded cache nodes deep within local ISP networks across Mexico and Central and South America, placing video segments within a single network hop of viewers.

    • Direct ISP peering: By establishing direct peering connections with major regional telecommunications operators such as América Móvil and Telefônica, the architecture completely bypassed congested international transit routes.

    • Dedicated capacity reservations: TelevisaUnivision reserved live event capacity with dedicated allocated headroom in-region. This isolated livestream traffic from "noisy neighbor" risks and comfortably absorbed peak traffic surges.

    • Sub-millisecond sessions with Valkey 9.0: To handle massive traffic spikes during the World Cup, TelevisaUnivision migrated its session store to Memorystore for Valkey 9.0, operating as serverless microservices on edge compute. This architecture delivered sub-millisecond response times for critical authentication and entitlement checks, while providing automatic scaling to process peak API traffic without the need for manual capacity reservations.

    Worldcup diagram

    Obsession for customer success

    Beyond technical capabilities, TelevisaUnivision chose Google Cloud for its joint co-engineering model and deep operational alignment.

    • Joint 24/7 war rooms: For all 104 matches, TelevisaUnivision engineers and Google Cloud specialists operated side-by-side in unified command centers.

    • Proactive monitoring as a service: Google Cloud’s Customer Reliability Engineering teams provided round-the-clock proactive monitoring and automated alerting.

    • Joint operational authority: Combined teams performed extensive pre-tournament stress tests and simulated failovers. During live matches, unified telemetry empowered joint leads to dynamically route traffic and adjust CDN configurations instantly as regional ISP congestion emerged.

    Summary

    The strategic partnership between TelevisaUnivision and Google Cloud during the 2026 FIFA World Cup established a new benchmark for global sports broadcasting. Across 39 consecutive days of tournament execution, TelevisaUnivision reported that the joint infrastructure delivered:

    Total matches broadcast

    104 live matches

    Platform availability

    100% platform availability (0 downtime)

    Cumulative viewership

    675 million views across TelevisaUnivision & ViX

    By uniting localized edge delivery, sub-millisecond serverless compute, and dedicated operational co-engineering, TelevisaUnivision and Google Cloud solidified a battle-tested blueprint for executing marquee live streaming events at record global scale.

    Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applications

    6 août 2026 à 15:00

    Nearly every major AI lab uses Google Cloud infrastructure, including for training of models, inference for agents, and new frontier research. Google Cloud also continues to be the platform of choice for new, high-growth AI startups who are driving much of the industry’s research and innovation.

    Today, we’re announcing that Mirendil, an exciting frontier AI lab focused on accelerating AI development, will also utilize Google Cloud’s AI Hypercomputer. This includes using a mix of Google’s TPU AI accelerators and full-stack NVIDIA AI infrastructure running on Google Cloud; this purpose-built AI infrastructure will support model pre-training and post-training applications for Mirendil. 

    The Mirendil team is building new AI systems that can help accelerate and democratize AI research and development. This means managing complex, end-to-end training workflows from initial model pre-training through post-training, and powering reinforcement learning on a massive scale. The ability to choose a mix of both TPU and NVIDIA’s full-stack accelerated computing platform through Google Cloud meant that Mirendil could access critical compute very quickly, and continue to match its workloads to the architecture best-suited to it over time.

    We closely partnered with Mirendil on end-to-end design and deployment of combined TPU and NVIDIA AI infrastructure across compute, storage, networking, and control planes. We also collaborated on a system that uses managed training clusters running in Gemini Enterprise Agent Platform, which effectively streamlines the provisioning and management of both TPU and GPU environments for Mirendil. Mirendil is already live with a cluster of TPU v5P chips, with NVIDIA AI accelerated computing systems coming online soon.

    "Progress in AI has been bounded by how fast humans can run the research loop - designing experiments, evaluating results, and iterating," said Behnam Neyshabur, cofounder and CEO of Mirendil. "We're building AI systems that can accelerate and improve that loop itself. Expanding on Google Cloud gives us the scale and flexibility to push those systems further and put frontier AI research capabilities in the hands of many more scientists and engineers to run that loop faster and at a greater scale."

    You can read more about our partnership on Mirendil’s blog.

    Scaling agentic AI: How UiPath built its high-performance GPU platform on AI Hypercomputer

    5 août 2026 à 18:00

    As a market leader in enterprise agentic automation and business orchestration, UiPath is helping to pioneer an industry shift toward agentic AI. With it, the company is deploying autonomous agents to actively reason, make decisions, and execute complex business processes across its disparate systems. 

    This transition from simple task automation to cognitive decision-making agents requires a massive surge in computational power and powerful infrastructure that’s reliable enough for the needs of the world's largest enterprises.

    Being able to orchestrate hundreds of GPUs in perfect harmony can be what makes the difference between just running a research experiment and building a global AI platform. Such orchestration requires balancing massive training jobs with real-time inference, all without letting costs spiral or latency spike.

    1

    To do so, UiPath re-architected its infrastructure to support high-scale intelligent document processing (IDP) using UiPath IXP and moved from isolated clusters to a shared Google Cloud GPU fleet, balancing A3 VM instances (NVIDIA H100 GPUs) for training with G4 VM instances (NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs) for inference. This architecture lets UiPath solve its “spiky workload” problem and count on predictable costs and open-source patterns that the company’s engineering teams can use to replicate this architecture themselves.

    "Realizing the full potential of enterprise agentic AI requires an infrastructure that matches our ambition. Google Cloud provides the scale and flexibility we need to train specialized models and deploy them globally. This partnership allows us to deliver high-precision intelligent document processing and autonomous agents that don't just chat, but actively drive business outcomes for our customers."  – Raghu Malpani, Chief Technology Officer, UiPath

    The context: heavy-duty math

    UiPath has run its full-stack automation platform on Google Cloud for years, but as its agentic AI initiatives expanded, it faced a series of new infrastructure challenges.

    Core capabilities like IDP, computer vision, and LLM-powered reasoning require heavy-duty math, so UiPath’s engineering team utilizes LLAMA model grounding that allows its robots to "see" interfaces with human-like clarity. And with specialized document models built on the Qwen architecture, the team can extract valuable data from messy, real-world paperwork.

    These models live on the UiPath cloud infrastructure, where cutting every possible millisecond of latency is essential. Moving from a "cool demo" to a reliable production tool without exploding costs meant the team had to rethink its underlying silicon.

    The challenge: more demand than supply

    In the past, when a team at UiPath needed to train a new model or run inference, it provisioned GPU nodes on demand and scaled up or down depending on whether the workloads were spiking or slowing.

    This was a functional strategy when cloud capacity was cheap, abundant, and perfectly elastic. But as its AI ambitions grew, UiPath found this approach could no longer keep up with its operational complexity. It now faced three new challenges:

    1. Spiky workloads: To ensure it had sufficient power for peak demand, UiPath  often had to buy extra capacity that sat idle during quieter periods, wasting expensive headroom. The company needed intelligent,  on-demand scaling that didn't require paying for silicon that wasn't crunching numbers.

    2. Supply bottlenecks: For large-scale fine-tuning, the price-to-performance ratio on gold standard high-end A3 VM instances with 8-cluster H100s is unbeatable. But global demand for those  chips has outstripped  supply, making it nearly impossible to scale training efforts at the speed UiPath desired just by adding nodes.

    3. Operational overhead: UiPath was also struggling with geographical inefficiency because stable inference demand still meant maintaining dedicated clusters in multiple regions to ensure low latency for international customers. Further, managing GPU infrastructure for both training and inference added inefficient layers of operational overhead.

    The solution: a shared GPU fleet

    With all of that in mind, UiPath decided to treat its GPUs as a shared strategic resource instead of a product-centric elastic infrastructure.

    As a result,  its engineering team designed a platform-level shared GPU fleet managed by its  machine learning services (MLS) platform, which prioritizes work across teams and time windows while balancing demand across workflows. During the day, the fleet serves real-time inference and latency-sensitive workloads, and at night or during off-peak hours, it automatically switches to batch training and long-running jobs.

    By coordinating workloads at the fleet level, MLS lets UiPath maximize utilization while reducing contention, all without relying on per-instance elasticity. It also enables the company to schedule capacity in advance, which improves predictability for both research and production use cases.

    Why Google Cloud: AI Hypercomputer architecture

    To support its growing scale, UiPath leveraged Google Cloud AI Hypercomputer, which offers a system-level approach integrating performance-optimized hardware, open software, and flexible consumption models into a unified environment. AI Hypercomputer also minimizes the friction between hardware and software layers, which allows engineering teams to focus on model performance rather than infrastructure management.

    2

    Once it settled on a shared fleet model, UiPath needed a cloud partner that could offer reliable GPU availability, competitive pricing, and burst capacity. That’s why it chose to expand its existing Google Cloud footprint with a highly specialized AI stack running on Google Kubernetes Engine . 

    Now, UiPath can take advantage of predictable capacity by leveraging Google Cloud’s Dynamic Workload Scheduler (DWS) to solve its supply bottleneck. The company knew Google Cloud could secure its GPU capacity consistently with notice windows measured in days. DWS allows the engineering team to schedule training runs in advance and secure capacity for short bursts, and it can now plan for capacity rather than having to react to scarcity. Today, UiPath runs all its training and most of its IDP model inference workloads on Google Cloud.

    While UiPath uses A3 VM instances for heavy-duty training and fine-tuning, not all of its tasks require that level of power. That’s why it now deploys Google Cloud G4 VM instances as a net-new optimization for inference workloads. These instances offer a cost-effective balance of performance and price, which allows UiPath to run lighter inference tasks without occupying the high-performance clusters reserved for training.

    "The shift to a shared fleet on Google Cloud transformed our operational model. We moved from reactive provisioning to a predictable, high-performance engine that powers our most advanced IDP and agentic AI workloads. With tools like Dynamic Workload Scheduler and a mix of A3 and G4 instances, we have the flexibility to optimize for both cost and speed. This ensures our engineers spend their time innovating rather than waiting for compute." - Arthur Wilcke, director of AI infrastructure, UiPath

    Practical validation: differentiated models at scale

    With consistent access to Google Cloud GPUs, UiPath can now bring advanced models into production. It can also schedule large training jobs without blocking production inference, letting it balance research experimentation with production reliability.

    This allows UiPath to deliver advanced IDP capabilities that extract data from highly unstructured and variable documents with high accuracy. For example:

    • Omega Healthcare uses UiPath to automate over 100 million  transactions with 99.5% accuracy, a 40% reduction in processing time, and 15,000 less hours of repetitive tasks per month.

    • Thermo Fisher Scientific uses UiPath to extract data from PDFs like invoices and purchase orders and is now able to process 53% of its invoices without human involvement, while cutting processing time by 70%.

    Lessons learned

    UiPath’s most significant wins so far have been increased availability and reliability. As its workloads continue to transition and it decommissions its legacy GPU resources, the company expects to see additional cost improvements.

    For engineering teams looking to build similar platforms, some key takeaways include:

    1. Decouple capacity: Instead of tying hardware to specific products, pool resources to smooth out usage spikes.

    2. Schedule, don't react: Using tools like DWS to book compute in advance guarantees availability and stabilizes costs.

    3. Right-size the silicon: Use A3 VM instances for training, but choose efficient options like G4 VM instances for inference.

    Next steps

    After its recent infrastructure evolution, UiPath is still refining its MLS platform to support the next evolution of AI innovation. To replicate this success in your own organization, use the resources below:

    ❌