❌

Vue normale

Reçu avant avant-hier

Best practices guide for customizing Gemini models via Reinforcement Learning (RL)

25 septembre 2026 à 18:00

Reinforcement learning (RL) has been a keystone of modern LLM post-training, but it demands large training clusters and access to model internals that external customers can't have with proprietary models like Gemini. So here at Google Cloud, we packaged it into a managed RL fine-tuning service (RLFT service) — you bring prompts and a reward function; we handle the infrastructure and the proprietary model internals. 

Now, you can adapt Gemini with the service — teaching the model from a reward signal you define, rather than from a fixed set of labeled answers. This unlocks a class of problems that supervised fine-tuning (SFT) struggles with: tasks that are hard to demonstrate but easy to score.  

In this guide, we will walk through practical best practices for using RL fine-tuning service. We'll start with a short tour of the RL training loop, how to decide if and when to use RL, and introduce how to get the most value from this approach.

What is RLFT? 

RLFT adapts Gemini from a reward signal you define rather than labeled answers. Instead of authoring a large set of gold examples, you write one program that scores a response and the service improves the model against it — unlocking tasks that are hard to demonstrate but easy to verify: you can't hand-write the ideal SQL for every schema, but you can run the query and check the result.

1 - Single-Step RL Training Loop

At each training step the service generates multiple candidate responses to your prompts, scores them with your reward, and improves the model so that higher-scoring responses become more likely while it stays close to the original Gemini. The reinforcement learning that makes this work is fully managed — you never configure it. The one thing you own, and the thing that most determines your results, is the reward.

Three properties define what RLFT can and can't do:

  1. It learns from the model's own outputs: It refines what the model already produces rather than copying an external target, so it tends to disturb unrelated capabilities less than SFT. 

  2. It rewards outcomes, not paths:   Any response that reaches a good result earns reward, which fits open-ended tasks with many valid solutions. 

  3. It amplifies existing competence:   It makes occasional success reliable, but it can't teach a skill the model never demonstrates.

When to use RLFT

2 - RLFT Approaches - Direct RL or SFT Warmup to Continuous RLFT

Prompting and SFT handle most adaptation; exhaust them first. RLFT earns its keep when you can grade a response but can't cheaply author it, when SFT has plateaued on the metric that matters (faithfulness, schema validity, tone), or when the task has many equally valid answers a single reference target would wrongly penalize. 

SFT and RLFT are complementary, not competing:

  • Direct RLFT when the base model already succeeds part of the time — enough for the reward to tell better answers from worse ones.

  • Two-stage SFT → RLFT when you have SFT data or the base success rate is too low for RL to gain traction. Use SFT as a short, cheap warm start — kept light, since over-fitting the demonstrations leaves less room for RL to improve — then continue into RL via Continuous Tuning, which initializes RL from the SFT checkpoint.

Across early adopters, these patterns show where RLFT delivers the most value — each scoring an outcome the business cares about but could never cheaply demonstrate.

Use cases for RLFT

AI-powered NPCs in games

  • What: In-character, on-brand dialogue held across long, multilingual, multi-turn conversations.

  • Problem: Off-the-shelf models break immersion — wrong language, hallucinated items, ignored players, repetitive loops.

  • Objective and reward: A Gemini autorater (LLM-as-a-judge) scores each turn on persona, flow, and game-state syntax, penalizing format and language errors.

  • Results: Loops and language drift disappeared and state syntax held, making shippable in-game characters viable at scale.

Structured entity extraction

  • What: Pulling a set of items from unstructured documents, such as supplier invoices and shipping manifests, into structured records automatically.

  • Problem: The long tail where SFT plateaus — missing required fields (recall) or inventing ones that aren't there (precision).

  • Objective and reward: A rule-based precision/recall reward forces every field to be grounded in the source text, not imitated from one gold answer.

  • Results: Field-level accuracy rose on noisy real-world documents where tuning had stalled, turning a manual review step into an automated one.

Content moderation

  • What: Applying intricate policies and decision trees at scale.

  • Problem: Models hallucinate false positives or reward-hack with invalid formats to dodge evaluation.

  • Objective and reward: A Cloud Run reward pairs format validation with a deterministic grader to enforce multi-step policy adherence.

  • Results: The model handled complex exemption carve-outs, sharply cut false positives, and stopped reward hacking — reducing the human-escalation volume that makes moderation expensive.

Code measured by execution

  • What: SQL or API calls graded on whether they actually run against customer data.

  • Problem: SFT mimics one reference query and breaks on unseen proprietary schemas.

  • Objective & Reward: A code-execution reward runs the code in a secure sandbox and pays out only if it compiles, executes, and returns the correct result.

  • Results: The model produced first-attempt executable queries at closed-frontier quality and lower inference cost, letting non-technical users query proprietary data in natural language.

Presentation slide generation via HTML

  • What: Multi-slide decks authored as HTML/CSS.

  • Problem: Training on text alone is blind to visual quality — overflows, clipped elements, and inconsistent styling slip through unnoticed.

  • Objective & Reward: A code-execution reward renders the slides and scores visual design, layout integrity, structural completeness, and rubric adherence.

  • Results: The model emitted modular, well-styled decks with cohesive themes and no layout overflow.

Where to start?

  • A dataset. A diverse set of prompts with a held-out validation split is enough for a first run — confirm the loop converges and reward moves the right way, then scale. Keep train and eval strictly separated; a contaminated eval hides overfitting.

  • A reward function. Your task specification as code or configs, and the dominant driver of quality. A good reward correlates with human preference, is robust to malformed output (catch the failed parse and return a clearly negative score rather than crashing), and resists reward hacking — ensemble judges, penalize length, floor degenerate outputs, and prefer a verifiable check over a model's opinion. Validate it offline before launch.

The service handles the rest; start from the defaults, watch reward and eval curves in the console, and take the checkpoint where validation reward saturates rather than the last step.

3 - rlft_tutorial

Get started today

What will you build? The tools are ready and waiting. 

Power your agents: Gemini 3.8 Live with Live Avatar is now generally available

24 septembre 2026 à 17:00

Following our announcement of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking last week, we are thrilled to share that Gemini 3.8 Live with Live Avatar is now generally available in Gemini Enterprise. First previewed at Google Cloud Next 2026, the technology is now officially ready for enterprise production.

As enterprise voice AI evolves beyond basic speed and cost metrics, our priority has shifted to making each interaction even higher quality. Gemini 3.8 Live already delivers a native speech-to-speech foundation for fluid, responsive dialogue. The Live Avatar feature brings an interactive visual presence to conversational video agents across web, mobile, and interactive kiosks. Together, you’ll have access to:  

  1. Video avatars: Conversational video with the Live Avatar feature can generate video avatars with synchronized lip-syncing. Note: Custom avatar feature is available via allowlist only.   
  2. Fluid dialogue: Native speech-to-speech means more natural interruption recovery without dropping conversation context or backend transactions.
  3. Tool calling: It executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting while tasks finish in the background.
  4. Breaks language barriers: Gemini 3.8 Live understands and speaks 97 languages, with automatic language detection.
  5. Make it easy for agents to see what your user sees: Live visual understanding can process live camera feeds and screen shares alongside audio, all at the same time. 

Though Gemini 3.8 Live Extended Thinking remains in private preview, Gemini 3.8 Live with Live Avatar is now available with US and EU endpoints, with provisioned throughput, enterprise compliance, and strict data governance. Try the model and its features in Gemini Enterprise, and start building now with API. 

Trust and transparency at its core

To safeguard identity and prevent misuse, customers can deploy from a library of curated, pre-built avatars, while custom avatar creation is gated behind a strict enterprise allowlisting and verification process. Furthermore, all generated audio and video streams carry imperceptible SynthID watermarks, ensuring AI-generated content remains transparent and verifiable.

Three demos of Gemini 3.8 Live with Live Avatar in action

#1: Interactive custom avatar

Create an interactive custom avatar in a step-by-step process 

Watch how Gemini 3.8 Live makes it possible to build a custom avatar by adding system instructions, uploading a single reference photo and audio file sample. 

Demo #2: Live video understanding in a voice-first claims intake

Streamline claims intake with live video understanding using Gemini 3.8 Live

In this demo, you’ll see a common use case come to life using Gemini 3.8 Live: an intake agent. In this scenario, we chose an insurance claims agent. You talk and show the damage on camera, and the claim notebook fills itself in as you go. In the background, an ADK agent team checks the policy, applies the intake rules, and builds the adjuster packet. To dive deeper, check out the open-source code.

#3: A real-time voice AI agent with Google ADK and Gemini Live API

Build a live voice agent using Google ADK and Gemini Live API

See how developers can use Agent Development Kit (ADK) to define agents, manage runners and session memory, and stream real-time audio directly to the Gemini Live API without a traditional speech-to-text pipeline.

How our customers are innovating with Gemini 3.8 Live with Live Avatar

4 autotrader

Cox Automotive built an AI-powered shopping assistant for Autotrader that uses live screen-highlighting and tool-calling capabilities to guide car shoppers through vehicle search, comparison, and financing — in real time, through natural conversation.

“Shoppers increasingly expect to describe what they need in their own words rather than work through filters and menus. Autotrader’s new conversational AI Avatar brings that experience to vehicle discovery by matching natural conversation to the right inventory. It is another step toward our vision of connected intelligence, where every consumer interaction draws on the full depth of Cox Automotive data.” — Marianne Johnson, EVP and Chief Product Officer, Cox Automotive.

6 equal ai

“We're building a personal AI that knows you, speaks your language, and is always on your side. Today, it handles over a million live calls daily across nine Indian languages. Gemini 3.8 Live improved interruption handling, multilingual conversations, and tool-call reliability. This AI doesn't just answer calls; it gets things done for you.” — Akhilesh Damaraju, CEO, Equal AI.

7 salesforce

“We're excited that Gemini 3.8 Live and Agentforce are coming together to reimagine what's possible in intelligent service. This collaboration between Salesforce AI Research and Google combines real-time, multimodal capabilities with agentic AI to explore new ways to create richer, more intuitive customer experiences from first contact to resolution." — Bob Van Osten, VP of Product for Agentforce, Salesforce.

9 specs

“Our team has been very impressed with Gemini 3.8 Live throughout testing and benchmarking! The updates made to Voice Activity Detection and the improvements to overall latency are huge steps forward and further our ability to deliver the highest quality AI Assistant on the SPECS platform.” — Eric Walsh, Software Engineer, Specs.

Start building today

Gemini 3.8 Live with Live Avatar are available now for enterprise customers:

If your application requires specialized, modular audio capabilities, explore our other audio models:  

Reach out to your Google Cloud sales representative to activate provisioned throughput, discuss customized deployment architectures and allowlisting for custom avatar.

How growing Latin American midsize businesses are building in the AI era

24 septembre 2026 à 16:30

Latin America’s small and medium-sized businesses are the heartbeat of the region's economy — accounting for more than 60% of total employment in the region, according to United Nations estimates. And just like their enterprise peers, everywhere you look, ambitious teams are moving fast to embrace AI. 

Many have already transitioned from experimenting with generative tools and agentic workflows to using them every day to work smarter, save time, and deliver exceptional customer experiences. These growing businesses are particularly focused on maximizing the benefit they get from their investments in AI, whether that’s using a fast, low-cost model to summarize daily emails or deploying an advanced model for complex data analysis, teams can match the right AI capability to their exact task and budget. 

It’s this range of options, and a familiarity with the broader suite of Google business, media, and advertising tools that has led many SMBs to choose Google Cloud, and Gemini Enterprise in particular, as their AI platform of choice. By doing so, they’re able to build custom AI agents, streamline daily tasks and paperwork, and offer customers instant support with the speed and reach needed to compete on a global scale. 

With our unique front row seat, we’ve seen the benefit SMBs are getting from leveraging Gemini Enterprise, not only for generative AI, but as a catalyst for adopting other essential cloud tools like Google Kubernetes Engine and BigQuery for complete end-to-end modernization. The number of Latin American-based small and medium businesses using Google Cloud AI tools has grown 8x year-over-year and the number of Brazil based small and medium businesses using Google Cloud AI tools has grown 9x year-over-year. This rapid adoption spans our Gemini models, Gemini Enterprise, and core Cloud infrastructure, and are helping businesses to:

  • Roll out better customer support systems to help escalate and resolve customer support calls more quickly.

  • Automate repetitive actions in areas like payroll and accounting.

  • Help more employees understand and leverage data at work — even those not trained as data analysts.

  • Rapidly create and implement new designs for marketing collateral.

  • Help more people build their own AI agents to help them in their everyday jobs.

As we head into today’s Google Cloud Summit in Brazil, we were proud to showcase nearly 20 of our newest Latin American SMB customers using Google AI to reduce busywork, serve their customers faster, and grow their businesses.

Announcing new Latin American customers putting Google AI to work

  • AdGoat, an Argentina-based adtech company processing more than 10 billion annual ad requests across more than 100 global websites. It uses Cloud Run, the Gemini API, and Gemini Enterprise to automate content analysis, ad bidding, and audience targeting to help e-commerce brands drive higher campaign returns.

  • Angelus, a Brazilian dental and healthcare manufacturing company, uses Gemini Enterprise to streamline project management across its research and development department. This enables its teams to automatically pull technical project data into pre-approved templates aligned with the company’s brand identity and regulatory requirements.

  • BunkerDB, a marketing science company operating across Latin America, uses Gemini Enterprise, Cloud Run, and Cloud Storage to power an AI platform that organizes marketing assets, checks brand compliance, generates or adapts multimodal content, and predicts ad performance before launch. All of this helps it reduce creative turnaround times from weeks to hours and cut cost per lead by up to 25%.

  • Caffeine Army, a Brazilian wellness and high-performance company that connects people with solutions in nutrition, sports, and well-being, deployed BigQuery and Gemini Enterprise on Google Cloud to unify customer purchase insights, enabling faster creative campaign turnarounds and boosting team productivity across the organization.

  • Convert, a Brazilian marketing and analytics provider, uses Looker, BigQuery, and Cloud Run to power five specialized AI agents that answer complex business questions in natural language, speeding up report deliveries by 65% and reducing operational costs by 32%.

  • Growth Digital, a Google Ad sales rep operating across 13 Latin American countries, used BigQuery and Gemini Enterprise to build over 113 AI agents, enabling teams to build proposals 5x faster, cut campaign reporting time by 80%, and reduce financial error rates to under 0.01%.

  • GrupoTusMaquinas.com, an equipment management platform based in Chile, deployed Google Cloud AI tools and Gemini models to create digital tracking profiles for trucks and machinery, allowing businesses to query fleet status in plain language and manage vehicles regardless of brand or location.

  • HealthAtom, a healthcare technology company, uses the Gemini API, Firestore, and Cloud Functions to power AI assistants across its clinical platforms, automating appointment scheduling and medical record reviews while supporting 80 million annual patient interactions.

  • KLog.co, a Chilean logistics technology company digitizing freight forwarding across Latin America, uses Gemini Enterprise, BigQuery, and Google Workspace to automate cargo tracking and shipping paperwork, cutting manual data entry errors by over 90% and increasing document processing capacity tenfold.

  • NEEOH, a leading Brazilian out-of-home advertising communication platform, uses Gemini Enterprise to standardize secure AI usage across its organization, enabling teams to generate campaign copy and build pitch proposals faster while keeping corporate client data secure.

  • Luxia Agro, an Argentinian foreign trade supplier of crop protection products, uses Gemini 3.5 Flash and Gemini Enterprise to automatically pull key details from complicated shipping emails and update their central business systems. This allows it to automate 80% of foreign trade operations and cut manual processing errors in half.

  • Macal, a Chilean auction company, uses the Gemini Enterprise, Cloud Run, BigQuery, and Security Command Center to automatically verify property records and modernize its technology systems, cutting software development times from weeks to days and lowering infrastructure costs by up to 30%.

  • Ninecon, a Brazilian tech consulting firm, deployed Gemini Enterprise to integrate AI directly into employee workflows, allowing managers to track usage patterns and optimize project turnaround times with real-time insights.

  • Nova Gestões, a customer service and operations provider in Brazil, uses Cloud Speech-to-Text and Gemini Enterprise to translate and analyze 100% of customer calls in real time, reducing post-call manual data entry and boosting team productivity by 30%.

  • Romi, a Brazilian industrial machinery manufacturer, uses the Gemini API and Gemini Enterprise to power an interactive chat assistant directly on CNC machine HMI (human machine iInterface), giving factory operators instant answers grounded in official manuals and generating QR codes for step-by-step instructional videos.

  • Supermercados El Dorado, a leading supermarket chain in Uruguay, leverages Google Compute Engine and Gemini Enterprise to modernize legacy testing infrastructure and connect custom AI agents within their daily workflows, boosting team productivity across departments.

  • Tryvia, a Brazilian IT and business solutions provider, uses Google Cloud, Looker, and Gemini Enterprise to move off legacy physical servers, giving teams real-time reporting dashboards and AI tools that speed up software development and daily tasks.

  • Via Cristais, a major highway operator in Brazil, leverages Google Contact Center as a Service to speed up emergency routing for highway accidents, reducing caller wait times, improving driver satisfaction, and mitigating the impact of call center staff turnover.

  • WeSpeak, an AI conversational platform for the hospitality industry in Latin America, uses Cloud Run, Gemini Pro, and Gemini Flash to automate end-to-end guest interactions across messaging channels like WhatsApp and Instagram. This has helped it achieve an 85% resolution rate and a 2x increase in overall sales volume for hotel clients.

Helping your team build AI skills

To help growing teams get the absolute most out of AI, we’ve created easy, no-cost learning programs that anyone can use:

  • Programs for small and medium businesses: Explore beginner-friendly training paths or join specialized programs to learn how to build custom AI assistants for your day-to-day work.

  • Google skills for organizations: Access thousands of free, on-demand AI courses and hands-on practice labs designed by experts at Google Cloud and Google DeepMind.

  • Get certified: Help your staff gain industry-recognized AI certificates through guided courses, expert mentoring, and skill badges.

By offering easy-to-use tools and free training — from everyday office apps in Workspace to advanced AI on Google Cloud — Google is here to help Latin American businesses thrive today and in the future.

The three things today's hottest startups are looking for in their AI stack

24 septembre 2026 à 15:00

Google Cloud has become the platform of choice for startups building AI. 

Our uniquely complete stack — including a choice of first- and third-party compute and models; our platform for building and managing agents; and our products for securing AI workloads — has emerged as the single most important driver of this growth and it is powering AI development for many of the most exciting and innovative startups in the world.

As a result, startups are choosing to build and run on Google Cloud at a higher rate than they were three years ago, at the start of the AI era.

Given how quickly the technology industry moves in the AI era, the choices startups make can be notable. As we’ve worked together and watch many of these leaders scale, we’ve observed  a few important trends emerging over the past several months. We expect these decisions will continue to shape the choices startups make about the platforms and technology they use: 

  1. Gemini Enterprise, which includes our tools for managing TPU and GPU clusters, services for building and managing agents, and APIs to access both first- and third-party models, is growing significantly with startups. And when startups use Gemini Enterprise, they also tend to use our “core cloud” services like Storage, BigQuery, or GKE.

  2. Gemini models — as well as several of the third-party models available through Gemini Enterprise — are providing very strong price-performance for startups. These customers are increasingly deploying both our frontier models and “workhorse” models as their AI to power workloads as diverse as scientific research, generative media creation, and financial analysis.

  3. Access to compute on GPUs and TPUs is critical for AI and the ability to choose one — or both — is unique to Google Cloud. But importantly, startups almost always use additional products from our stack alongside these chips, like models, tools for building agents, or services like BigQuery or GKE. These additional technologies illustrate how the needs of startups are rarely singular, and just how much value they find in having ready access to a strong suite of second-, third-, and fourth-level technologies beyond just compute.

We can see the demand for these technologies first-hand in some of the recent deals we have struck in the past 60 days with a number of leading startups across sectors:

  • Artificial Agency, a startup focused on generative behavior in games, is running critical AI workloads and research on Google Cloud, where it is using NVIDIA GPUs for model training and inference, as well as Gemini models and Cloud Storage.
  • Arya Health is building the AI workforce for healthcare, deploying agentic AI to perform the non-clinical administrative work that limits providers’ ability to deliver and expand care. Arya’s AI agents work across scheduling, intake, recruiting, onboarding, compliance, after-hours operations, and other critical workflows, interacting through voice, text, email, and providers’ existing systems. Arya uses a range of Gemini models across its agentic infrastructure, including Gemini 2.5 Pro, 3.1 Flash, and 3.5 Flash Lite, selecting models based on the reasoning, speed, and cost requirements of each workflow.
  • Casco is a cybersecurity startup whose autonomous agent swarms execute sophisticated, multi-step attacks to uncover vulnerabilities across enterprise applications, cloud environments, and infrastructure. Its architecture combines advanced reasoning models for complex, long-running tasks with fast models such as Gemini 3.5 Flash for focused subagent work. Google Cloud’s model portfolio and infrastructure, including Provisioned Throughput, help Casco match each workload with the right combination of intelligence, speed, and capacity.
  • CodeRabbit has been a pioneer in independent AI code review and has expanded that layer into Agentic Change Management, the control plane for agentic software development. They use our Cloud Run and Storage products to underpin their application, and are now beginning to leverage Gemini 3.1 Pro and other Gemini models to power use cases like analyzing how a single code change impacts an entire project, writing clear and contextual review comments to explain logic bugs, and instantly generating one-click fixes for developers.
  • Comfy offers a platform for creatives to build brand-consistent content and media with generative AI. They are utilizing a mix of NVIDIA systems and Google media generation models, like Veo and Nano Banana, in their platform.
  • MicroAGI, a German AI robotics startup, recently announced they would access NVIDIA Blackwell systems for model training through Google Cloud. They will also use Gemini Enterprise Agent Platform and AI models on Google Cloud to help robotics process multimodal information like video.
  • Ineffable Intelligence, the London-based superintelligence startup, recently announced a partnership with Google Cloud to access NVIDIA Vera Rubin systems as well as high-efficiency AI networking and storage.
  • PEAR Health Labs built and runs its AI health and fitness companion and agentic health platform entirely on Google Cloud, using our Gemini 3.5 Flash model as well as our data cloud products.
  • Roboforce is a physical AI company building scalable robots for industrial environments. They recently signed a new agreement with Google Cloud to utilize our GPU-powered VMs, which will be used to train and serve their custom and post-trained physical AI models.
  • xFigura offers a platform that gives architecture teams a single, secure canvas that brings generative models into one place for AI-driven design. Built on Google Cloud, xFigura uses models like Nano Banana and Omni to help designers generate and refine concepts in plain language, with authorship staying with the architect.
  • SciFin is a startup that helps revenue teams uncover "revenue reality" by converging fragmented sales context across CRM systems, documents, emails, and other sources. They are utilizing Gemini models, infrastructure, and Gemini Enterprise to build and run their platform.

These new and expanding customers represent some of the most exciting names in AI and are demonstrative of the type of successes that startups are having with Google Cloud’s uniquely complete stack for building AI.

To learn more about Google Cloud’s work with leading AI startups, or to get started building with us, visit here.

Changing the game: Using agentic AI to secure infrastructure code

18 septembre 2026 à 18:00

AI is accelerating software development at an unprecedented pace. But as code generation scales, so do the challenges of securing the code, especially emerging AI-based vulnerability exploitations. To meet these challenges, the Google AI and Infrastructure team is transforming how we approach security. In this article, we discuss new AI-native agentic methods that we’ve developed that systematically embed high-precision, pervasive vulnerability scanning and patching directly into Google’s software development lifecycle. By continuously scanning every code change across hundreds of millions of lines of code that we deploy onto our infrastructure, we are preventing hundreds of vulnerabilities per month from ever reaching our code base or production, defending our global network, AI infrastructure and our users. 

Solution architecture and implementation

image1

Pervasive pre-submit agentic scanning: security as part of ongoing software development

Traditionally, the technology industry relies on large one-off security scans that are slow and lack sufficient context. As a result, they often find vulnerabilities too late. Our approach instead focuses on pre-submit scanning, where we evaluate each code check-in (across every layer of the stack) in real-time using AI agents. By integrating the pre-submit scan into the tools developers already use, security becomes a continuous routine process, similar to rule checkers, readability reviews or other software development tools. Also, from an AI perspective, scanning each individual code change requires much less context than performing a large one-off scan, significantly improving the scan’s effectiveness. 

The importance of localized threat models

For this initiative, we evolved Mantis, our open-source multi-agent review harness, to increase the precision of our security agents by matching them with a cohort of robust localized threat models. Rather than relying on static decoupled documents, the threat models use live codebase metadata. The scanning agent improves its accuracy further using a dependence call graph across packages and libraries to expand and refine its threat model context. Making threat models part of our ongoing vulnerability scanning encourages developers to continuously update threats and dependencies, keeping the models up-to-date. Using localized and precise threat model data translates to dramatic accuracy improvements, bringing our false-positive rates down to 3% in some cases.

Specialized triage agents speed up development

Vulnerability scanning as part of code check-in requires it to respond quickly to the developer or agents generating the code, so as not to impede engineering productivity. To get responses with low latency, we run a two-step validation process. First, we run a quick lightweight scan that validates its findings against a specialized triage agent. This agent programmatically checks the actual structure of the code (using abstract syntax tree parsing, call-graph traversal, and pre-indexed domain safety rules) to prove that the vulnerable path is actually reachable by an attacker. This agent gets over 92% precision and completes its work in less than a minute. Then, a post-submit scan as part of nightly integration testing serves as a second layer of defense, using off-peak cycles to test for vulnerabilities that may have been introduced across multiple changes. 

Bug fix agents close the loop

Finding vulnerabilities is only half the battle. The last component of our solution is an automated bug-fix agent that uses the scan results and generated proofs (snippet of code that demonstrates how the vulnerability is exercised) to autonomously construct precise fixes that are consistent with our internal coding standards. The agent submits the fixes for human review as part of the original change request’s review, further reducing the time between detection and resolution. 

Learnings and call to action 

Embedding continuous scanning directly into the software development lifecycle has been a game changer at Google; its suggestions are widely adopted, and it’s prevented a multitude of vulnerabilities from being introduced into the codebase. But any organization wishing to improve security can adopt a similar AI-native approach, following these principles: 

  1. Keep systems separate: To prevent bias, keep the harnesses, rules, and context for each of your development, scanning, triage agents separate. Pair lightweight AI scans with deterministic, structural validation to drive down latency and improve accuracy. 

  2. Use context wisely: Feed your agents your existing threat models. Precise context is the answer to reducing false positives, and up-to-date threat models set a high floor on a team's security posture by improving the rate of true positives in presubmit scanning.

  3. Build a good harness: While the choice of the underlying model is important, using a multi-agent harness can have substantial impact, by helping compensate for variability in model choice. 

  4. Automate the fix: Use agents to also propose human-in-the-loop fixes, to further reduce time-to-resolution. 

If you want to get started on your own AI-native security transformation, Mantis is now available as open source for you to use and benefit from. You can also learn more about the fundamentals of cybersecurity and the other platforms that power this agentic pipeline: Google Cloud, Gemini Enterprise and Gemini models running on Trillium and Ironwood TPUs. And you can get inspiration from how agentic vulnerability scanning and remediation defends Google Cloud customers as an integral part of Google Cloud’s secure software development lifecycle (SDLC) effort.


With special recognition to critical team members who made this delivery possible: Stella Voutsina (Lead Program Manager), Yulong Zhang (Senior Staff Security Engineer, Mantis), and Nick Galloway (Staff Security Engineer, Mantis).

Cloud CISO Perspectives: How Google monitors AI threats and advances AI defenses

16 septembre 2026 à 18:00

Welcome to the first Cloud CISO Perspectives for September 2026. Today, Sandra Joyce shares the latest details on Google’s visibility into how attackers are using AI, and how we’re using AI to stop them.

As with all Cloud CISO Perspectives, the contents of this newsletter are posted to the Google Cloud blog. If you’re reading this on the website and you’d like to receive the email version, you can subscribe here.

aside_block
<ListValue: [StructValue([('title', 'Get vital board insights with Google Cloud'), ('body', <wagtail.rich_text.RichText object at 0x7fe6c8ab0250>), ('btn_text', 'Visit the hub'), ('href', 'https://cloud.google.com/solutions/security/board-of-directors?utm_source=cgc-site&utm_medium=et&utm_campaign=FY26-Q2-GLOBAL-GCP39634-email-dl-dgcsm-CISOP-NL-177159&utm_content=-&utm_term=-'), ('image', <GAEImage: GCAT-replacement-logo-A>)])]>

‘Spellcheck for cybersecurity’ and beyond: How Google monitors AI threats and advances AI defenses

By Sandra Joyce, VP, Google Threat Intelligence

Sandra Joyce

Sandra Joyce, VP, Google Threat Intelligence

Anyone operating in security knows that speculation is a major liability during periods of technological disruption. While there is plenty of hype and understandable concern around how threats might use and target AI, a CISO’s AI security strategy has to be anchored in ground truth.

Google operates at a rare intersection as both a frontier AI lab and a security company with a frontline view of global incidents. This dual vantage point allows us to understand how AI is built, and exactly how AI is being targeted in the wild. To provide the operational realities that security and business leaders need in the AI era, Google Threat Intelligence Group (GTIG) recently released our latest AI Threat Tracker.

When we strip away the noise and look at the telemetry, the real threat landscape boils down to three structural shifts that CISOs must address:

  1. AI is reshaping how software is built. 

  2. AI is expanding the attack surface.

  3. AI is enhancing threat capabilities. 

Today, we’re sharing details on Google’s visibility into these three challenges, and our approach for solving them.

Building securely in the AI era 

AI has fundamentally altered software development velocity. Across the industry, autonomous agents and AI workflows now push code into production at unprecedented speed. This creates exciting opportunities for innovation, yet CISOs are faced with the difficult task of mitigating enterprise risk while maintaining business momentum. 

We’re seeing threat actors turn our greatest engineering shortcut against us by contaminating upstream packages that AI assistants are trained to suggest and trust. GTIG believes that malicious contamination of AI-assisted coding practices has been contributing to the significant growth in large-scale, open-source software supply chain compromises we observed in 2025 and early 2026.

The solution to a machine-speed threat landscape isn't slowing developers down — it’s building security natively into the AI pipeline. Part of this process involves in-editor guardrails for developers that create a real-time 'spellcheck for cybersecurity.'

We’re also monitoring adversaries targeting agents. The financially-motivated threat actor TeamPCP (UNC6780) has implemented more than half a dozen methods to exploit AI tools and open-source software development practices, including hijacking AI toolkits, prompt injection, and blinding AI scanners with toxic prompts to obfuscate malicious payloads.

The solution to a machine-speed threat landscape isn't slowing developers down — it’s building security natively into the AI pipeline. Part of this process involves in-editor guardrails for developers that create a real-time “spellcheck for cybersecurity.” 

Just as word processors underline typos without forcing the writer to stop, security controls must sit natively inside the developer’s editor and agentic workflows, instantly flagging poisoned packages, toxic prompts, and misconfigured toolkits. 

Crucially, this can’t stop at the editor. Traditional security suffers from context blindness: Code editors can’t see cloud configurations, delivery pipelines miss runtime exposure, and production teams can’t easily patch root-cause blueprints. 

Bridging this gap requires an integrated code-to-cloud approach — the exact design principle behind platforms like Wiz Code. The underlying approach is to ensure code is continuously verified against live cloud realities before it ships.

When organizations think about AI-driven code analysis, the default assumption is to pick one frontier model and point it at their repository. However, our research and telemetry show that single-model security creates a dangerous monoculture: No single AI model can discover every vulnerability, and threat actors are already testing inputs that can blind specific LLM safety filters and scanners.

To secure this expanding attack surface, CISOs should avoid the trap of managing AI through disconnected silos... The future of cloud and AI defense needs to be built on a unified and dynamic graph that connects your code, your models, your data lineage, and your runtime identities into a single living map.

To solve this, Google takes a deliberate multi-model approach. By orchestrating several foundation models — including Gemini, commercial, and open-source — we cross-validate findings, strip out false positives, remediate code, and identify complex logic flaws that a single model misses. We’re smarter with more than one “brain.”

Securing AI 

Securing the development lifecycle is only half the battle. We also need to prevent adversaries from exploiting AI attack surfaces and weaponizing over-privileged agents. Threat actors are targeting AI workloads with techniques that include: 

  • LLMJacking: Cybercriminals and state-sponsored groups target GPU access to support running their AI models and agentic workflows. In one notable intrusion Mandiant investigated in April, a threat actor gained initial access to a victim’s cloud environment from an exposed personal access token, and used it to deploy unauthorized AI infrastructure and scale high-performance compute resources, leaving the victim to absorb the hardware and platform costs.

  • Targeting of AI data and access: Cybercriminals now recognize that your custom prompts, agent instructions, and fine-tuned models represent high-value crown jewels. In Q2 2026, Mandiant investigated multiple data theft extortion operations where threat actors stole proprietary AI data, including models, skills, prompts, source code, and related research. Demand is also surging for AI account credentials in underground marketplace forums, with some sellers offering steep discounts for consumer accounts at up to 99% off retail prices.

To secure this expanding attack surface, CISOs should avoid the trap of managing AI through disconnected silos. Don’t treat agent access policies, model inventories (AI-BOMs) and shadow AI as separate challenges because these risks are deeply connected. The future of cloud and AI defense needs to be built on a unified and dynamic graph that connects your code, your models, your data lineage, and your runtime identities into a single living map. 

Pioneered by the Wiz Security Graph, this approach serves as the contextual engine for Google AI Threat Defense (AITD) — our broader autonomous security framework that fuses the reasoning power of Gemini and other frontier models, the contextual risk prioritization of Wiz, the code remediation capabilities of CodeMender, and the frontline expertise of Mandiant to stay ahead of AI-driven attacks. Crucially, this context is not siloed; it directly feeds Google Security Operations, ensuring that security operations teams can continuously identify, prioritize, and sever toxic attack paths at machine speed.

Defending against AI threats

Threat actors are rapidly moving beyond simple prompt generation toward fully-automated, multi-agent attack pipelines.

Security in the AI Era

Cyber Defense Summit 2026: Security in the AI era

In one notable intrusion investigated by Mandiant, a financially-motivated actor compromised an organization's cloud infrastructure and deployed an autonomous agent framework. The threat actor used an AI coding chatbot, a prompt, and a set of agent instructions to plan, build, and execute a mass credential harvesting campaign in less than six hours.

We’re also tracking adversaries using AI as an intelligent orchestrator across the entire attack lifecycle. GTIG recently observed a PRC-nexus espionage group experimenting with a tool called CC Switch to cycle across multiple accounts and swap AI models — like Claude, Codex, and Gemini — picking the best model for specific tasks, such as writing exploit scripts and drafting lures. While the underlying hacking tools aren’t new, AI turned what had been a disjointed manual process into a smooth and automated workflow.

To take advantage of your deep context, it’s imperative to shift from manual, human-scale incident response to machine-speed security operations. We can no longer rely on human analysts manually triaging endless backlogs of static alerts.

While these machine-speed attacks sound daunting, defenders actually hold an asymmetric advantage. Even when armed with autonomous AI, an attacker operates from the outside with limited context — probing in the dark, guessing connections, and hoping a compromised credential leads to a useful asset. 

Defenders, on the other hand, possess deep context that attackers don’t have. You know your code, cloud configurations, user identities, deployment realities, and internal architecture better than anyone. When you feed this rich, multi-dimensional internal observability into security models, AI defense becomes inherently faster and more accurate than AI offense.

To take advantage of your deep context, it’s imperative to shift from manual, human-scale incident response to machine-speed security operations. We can no longer rely on human analysts manually triaging endless backlogs of static alerts. 

By codifying our frontline threat intelligence directly into these AI models, these autonomous agents can continuously monitor for, investigate, prioritize, and remediate attacks.

How Google is helping defend the ecosystem 

As adversaries adopt AI, we have a unique opportunity to disrupt them at the source. As a major security and AI provider, we take this responsibility seriously, using multiple levers to stay ahead.

  • Disabling malicious infrastructure. If you use Google tools to facilitate an attack, you lose access to those tools. We proactively disable the projects, accounts, and assets of known bad actors.

  • Hardening our AI models and classifiers. We operate a continuous feedback loop for our AI models. By feeding threat intelligence directly back into product development, our models learn to recognize and refuse malicious requests before an attack can even be generated.

  • Automating vulnerability hunting and patching also disrupt adversaries. We are moving from manual patching to AI-driven hunting. Tools like CodeMender automatically fix critical vulnerabilities in the code itself.

  • Developing advanced defenses and threat models. Our teams at Google DeepMind are building specialized defenses for generative AI — deploying active monitoring across our entire ecosystem to identify misuse in real-time.

Securing the AI era can’t be achieved with the disconnected, manual tools of the past, and you can only defend against an AI-powered threat with an AI-powered defense. To tip the scales back in favor of defenders, we must transition to a continuous, machine-speed model of protection — and at Google, we are committed to building that secure future alongside you.

To learn more about our approach to securing the AI era, please check out our new Mandiant AI Risk and Resilience report.

aside_block
<ListValue: [StructValue([('title', 'Learn something new'), ('body', <wagtail.rich_text.RichText object at 0x7fe6c8642ed0>), ('btn_text', 'Watch now'), ('href', 'https://x.com/googlecloud/status/2090213589558698309?s=20'), ('image', <GAEImage: Cloud-CISO-Perspectives-logo-A>)])]>

In case you missed it

Here are the latest updates, products, services, and resources from our security teams so far this month:

  • A manufacturing blueprint for secure agentic AI: AI and agents have arrived on the factory floor. Today’s CISOs and business leaders must balance innovation with precision, physical safety, and operational resilience. Read more.
  • Proactive cyber defense for governments and enterprises: Our new Fairwind Program is a limited access program for governments and trusted partners to use our most advanced cyber defense capabilities. Read more.
  • Getting started with the Mantis harness to find and fix bugs: Mantis is part of how Google finds and fixes vulnerabilities at machine-speed. The open-source AI harness creates a more effective repository analysis. Read more.
  • Breaking into Google's GFile for $100,000: Learn about how a vulnerability — that was not exploited and has now been patched — could have allowed attackers to chain unauthenticated, undocumented internal APIs with overly-permissive shared file libraries to achieve unrestricted data access across core infrastructure. Read more.
  • Introducing new session management tools with native, granular controls: New Google Cloud session controls are deeply integrated and a granular feature of Context-Aware Access. Here’s what you need to know. Read more.
  • How Blackline prevents data exfiltration with VPC Service Controls: We’re excited to share new policy intelligence capabilities in VPC-SC that help drive operational simplicity: Violation analyzer and violation dashboard. Read more.
  • Introducing Continuous Vulnerability Assessment: You can detect exposure to new vulnerabilities the moment they’re published with Wiz CVA. Read more.
  • How developers prevent production risk at the source: Fixing security vulnerabilities in code takes seconds, while patching in production creates high operational costs and risk. Discover how empowering developers as your first line of defense eliminates exposure across every phase of your software pipeline. Read more.
  • Wiz achieves GovRAMP High authorization: Delivering unified cloud security and accelerating secure modernization to protect citizen data and critical infrastructure. Read more.

Please visit the Google Cloud blog for more security stories published this month.

aside_block
<ListValue: [StructValue([('title', 'Join the Google Cloud CISO Community'), ('body', <wagtail.rich_text.RichText object at 0x7fe6c8640610>), ('btn_text', 'Learn more'), ('href', 'https://rsvp.withgoogle.com/events/google-cloud-ciso-community-interest-form-2026?utm_source=cgc-blog&utm_medium=blog&utm_campaign=FY25-Q1-global-GCP30328-physicalevent-er-dgcsm-parent-CISO-community-2025&utm_content=cisop_&utm_term=-'), ('image', <GAEImage: GCAT-replacement-logo-A>)])]>

Threat Intelligence news

  • AI Threat Tracker: From prompting to autonomy: In the newest Google Threat Intelligence Group (GTIG) report on the adversarial misuse of AI, we’ve observed adversaries transition from basic prompting to agentic AI workflows and AI-enabled automation, including threat actors compromise a cloud resource, then plan, build, and execute an agent-enabled mass credential harvesting campaign in under six hours. Read more.
  • Financially-motivated threat actor BREEZE COMET targets Brazil: Learn about BREEZE COMET’s tactics and toolkit, and our mitigation recommendations and detections to support organizations in defending against this active and developing threat. Read more.
  • JFrog Artifactory under attack: Wiz Research has identified active, in-the-wild exploitation of three critical and high-severity vulnerabilities impacting JFrog Artifactory. Attackers are chaining these vulnerabilities to bypass authentication and gain administrative control. Read more.

Please visit the Google Cloud blog for more threat intelligence stories published this month.

Now hear this: Podcasts from Google Cloud

  • Cloud Security Podcast: Patching browsers with AI, agents, Rust, and your tabs: Jasika Bawa and Doug Turner of Chrome Security explore how Google Chrome now uses AI agents to autonomously identify and patch security vulnerabilities at an unprecedented scale, significantly accelerating the browser's update cadence. Listen here.
  • Cloud Security Podcast: All about Project Atlas, Wiz's AI vulnerability research: Nir Orfeld, head of vulnerability research, Wiz, discusses how his team uses multi-agent AI systems for discovering high-impact zero-day vulnerabilities in cloud infrastructure. Listen here.
  • Cloud Security Podcast: How Google eliminates classes of vulnerabilities at scale: How do you build the foundations for a secure Google-scale enterprise that stays secure even if an AI is writing the code and nobody has time to review it? Christoph Kern, principal security engineer, Google, explores what secure-by-design really means in the AI era. Listen here.

To have our Cloud CISO Perspectives post delivered twice a month to your inbox, sign up for our newsletter. We’ll be back in a few weeks with more security-related updates from Google Cloud.

How Orange built FinOps accountability, and why agents are next

16 septembre 2026 à 18:00

At Orange, the leading France-based multinational telecom provider, there are days when engineering teams set aside their delivery backlogs and spend the day cleaning up cloud spend together. There's a leaderboard. There are goodies on the line. Experienced practitioners guide the newcomers, so people learn the work while doing it. By the end of the day, sponsors can see the results.

Orange calls these FinOps Clean Days. Together with gamified hackathons, they've earned the company's 100-plus person FinOps community a Net Promoter Score within the organization that’s above 70.

Those numbers point at something the wider industry is wrestling with. Recent State of FinOps reports identify getting engineers to take action as one of the top challenges organizations face. Moving from awareness to action means finding ways to build FinOps accountability, and to get teams to genuinely care.

That makes FinOps a business change problem. And business change problems have known solutions. We spoke with Camille Marini, the FinOps lead at Orange, to get a deeper understanding of how the company overcame these hurdles to accelerate AI adoption and ROI, and how your organization might follow the same course.

Why the Clean Days work

Orange has held two principles since it set up its FinOps team. First, Cloud FinOps is a shared responsibility, with every stakeholder in a project involved in their own way. And the only path to that shared responsibility runs through communication and a deliberate change effort. 

“We insisted on the concept of shared responsibility across the organization for our FinOps practices,” Marini told us. “It’s very similar to how we approach cloud security. We needed to make teams understand that every single stakeholder in a project is involved in FinOps, each in their own way, if we are going to achieve responsible and impactful AI spending and usage.”

Those principles led Orange to create a FinOps Community of Practice, with support from Google Cloud Consulting. The team ran it on standardized communication channels so the methodology reached well beyond the central group, and kept the meetings actionable, sharing optimizations and billing updates so every session provided value.

The Clean Days came from a clear-eyed reading of how agile teams actually operate. In agile environments with deployment running constantly, optimization work rarely wins against the sprint. Delivery priorities, backlogs, and daily operations take the available time first. So Orange created protected time, made it collaborative, and made it fun.

McKinsey's four building blocks of change explain why this approach lands. Any large organizational change, the framework holds, requires action across four areas:

  • Conviction and understanding: "I know what is expected of me and I agree with it."

  • Formal mechanisms: "The structures, processes, and systems reinforce the change."

  • Role modeling: "I see my leaders and colleagues behaving differently."

  • Talent and skills: "I have the skills and opportunities to behave in a new way."

Map Orange's practice onto those blocks and the pattern is visible. Gamification and rewards give engineers colleagues to emulate: The leaderboard makes different behavior visible, and sponsors see the quick wins for themselves. Experienced practitioners guiding novices builds talent and skills through the community itself. The regular sessions, sharing optimizations and billing updates, build the conviction that comes from knowing where the money goes.

image2

FinOps activities mapped to the four building blocks of change, with the points where AI agents can reinforce them.

What happens beyond 100 people

A community of 100 engaged people is an achievement. But in an organization with thousands of engineers, no central FinOps team can reach everyone directly. The question for leaders is how to extend what a community like Orange's creates — the awareness, the shared ownership, the habit of acting — to people the FinOps team will never meet.

This is where AI agents extend the capabilities of a FinOps team with two core benefits. They take on complex, time-intensive activities that previously needed a human, and they reduce friction around FinOps for individuals across the business.

Getting teams to adopt them takes a strategy aimed at your own organization's pain points, which often come from high cognitive load, unclear accountability, or competing priorities. Start by finding where engagement drops off in your FinOps lifecycle:

  • An awareness gap: If teams are unsure of their spend impact, an insight agent can push real-time cost data into their daily tools.

  • A bandwidth gap: If engineers are too busy with backlogs, a remediation agent can identify quick wins and present them as ready-to-merge code changes.

  • A complexity gap: If reporting feels like a manual chore, an orchestration agent can gather the data and simplify the process.

Start with trust, then add autonomy

The sensible path runs in sequence. Establish the community practice, the way Orange did. Then introduce read-only agents that inform and suggest. Only once those are established across the community should you build agents that execute changes. Direct action carries operational risk, so manage it carefully. It's also where significant wins often sit.

How you build depends on who's building. For teams that want to deploy quickly with minimal code, the Gemini Enterprise App provides a no-code environment for creating agents. For developers who need granular control, the Gemini Enterprise Agent Platform (formerly Vertex AI) offers advanced tools for launching and governing agents built with frameworks like the Agent Development Kit (ADK).

Cloud FinOps is moving beyond centralized reporting toward action that happens where the work does. The organizations getting there start with the culture, then use agents to carry it further than any one team could reach. 

Orange's numbers came out of the community work. Building that foundation is the part worth copying first. When you're ready to extend it, Google Cloud Consulting can help you shape the community practice, and the Gemini Enterprise App is a low-lift way to put your first read-only agent in front of your teams.

Celebrating our tech and startup customers

20 avril 2022 à 22:00

Our tech and startup customers are disrupting industries, driving innovation and changing how people do things. We’re proud of their success and want to showcase what they’re up to! You’ll hear about their new products, their businesses reaching new milestones and their ability to get things done faster and easier using Google Cloud’s app development, data analytics and AI/ML services.

Congrats to Impact Analytics for Closing PVH
With the COVID-19 pandemic, the rise of e-commerce, and supply chain crisis, Impact Analytics had to quickly offer enhancements on their platform that gave retailers access to intelligent, automated, and edge-aware solutions. Google Cloud's best in class AI and ML solutions and highly performant infrastructure gave Impact Analytics the scalability and building blocks to create Ada, a robust predictive algorithm to give PVH and other retailers the tools to enhance their inventory planning capabilities. And now, Impact Analytics just closed a strategic deal with PVH (parent company of luxury brands Tommy Hilfiger, Calvin Klein, True & Co) to build out AI solutions for assortment planning and pricing optimization. Impact Analytics' cutting edge AI and ML guided forecasting engine is built entirely on Google Cloud! Read more

Podimetrics raises $45M Series C round
Podimetrics, creator of the FDA-cleared SmartMat and integrated clinical care services team, is dedicated to early detection and prevention of diabetic amputations, one of the most debilitating and costly complications of diabetes. Its clinical care services platform leverages Google Cloud services to engage with patients, by helping save limbs, lives, and money - all while keeping vulnerable populations healthy in their own homes. Read more about their Series C funding round.

Anvyl raises +$15M in an oversubscribed Series B funding round
Congrats to Anvyl for raising a hugely successful Series B as they modernize & transform the supply chain technology market and more than doubled revenue in the last year. Read more.

Helios has kicked off 2022 in a big way
The audio tone analysis platform, Comprehend: Elite, that they provide to Wall Street quantitative hedge funds now covers all US equities and is fully available here. It’s entirely powered by Google Cloud!

Dapper Labs uses Google Cloud for performance, reliability and decentralization
In case you missed it before the holidays, Dapper Labs is working with Google Cloud as its hyperscale cloud partner to ensure performance, reliability and decentralization for the next wave of mainstream users on Flow, without needing to compromise on decentralization or sustainability. Find out more.

Geotab’s Intelligent Transportation Systems (Geotab ITS) is built on Google Cloud.
Geotab uses GKE, BigQuery, Dataflow and Cloud Composer to build an innovative solution combining analytics and access to massive data volumes so municipalities can make better transportation planning decisions. The sheer volume of information that it handles, along with a need for highly scalable and flexible tools to manage, store, and analyze that data, led Geotab to invest in Google Cloud technology. Read more.

Mux CEO shares advice for getting started with video
Mux CEO, Jon Dahl, sat down with Google Cloud Director, Nirav Sheth, to share best practices and strategies for getting started with video, along with insights and advice from his learnings as a startup founder. Listen to what he has to say.

Google Cloud is proud to support Unstoppable Women of Web3
Unstoppable Women of Web3 (UWOW3) is an action oriented community made of industry leaders supporting education & opportunities for girls, women, and minorities in this burgeoning industry. This International Women’s Day, March 8th, you can catch live interviews with Tech and Web3 leaders from all over the world, covering topics such as how to build communities, how to learn more about Web3, developing technology on the blockchain, how to talk about complex ideas with kids, and more! How you can engage:

Puppet CTO increases development speed
Hear Puppet CTO Deepak Giridharagopal discuss how they managed to build Puppet's first Saas product, Relay, fast while also ensuring they would be able to remain agile if growth was to happen quickly. Watch video.

Vimeo builds a fully responsive video platform on Google Cloud
The video platform @Vimeo leverages managed database services from Google Cloud to serve up billions of views around the world each day. Read how it uses Cloud Spanner to deliver a consistent and reliable experience to its users no matter where they are. Find out more. 

Nylas improved price-performance by 40%
You don't have to choose between price-performance and x86 compatibility. Hear from David Ting, SVP of Engineering and CISO at @nylas, to learn how Google's x86-based Tau VMs delivered 40% better price-performance than competing Arm-based VMs. Watch now.

Optimizely partners with Google Cloud on experimentation solutions 
Build the next big thing with @Optimizely Experimentation on Google Cloud - driving innovation and next-gen experimentation for enterprise companies and marketers. Check it out.

Limitless Data. All Workloads. For Everyone

6 avril 2022 à 07:00

Today, data exists in many formats, is provided in real-time streams, and stretches across many different data centers and clouds, all over the world. From analytics, to data engineering, to AI/ML, to data-driven applications, the ways in which we leverage and share data continues to expand. Data has moved beyond the analyst and now impacts every employee, every customer, and every partner. With the dramatic growth in the amount and types of data, workloads, and users, we are at a tipping point where traditional data architectures – even when deployed in the cloud – are unable to unlock its full potential. As a result, the data-to-value gap is growing. 

To address these challenges, we are unveiling several data cloud innovations today that allow our customers to work with limitless data, across all workloads, and extend access to everyone. These announcements include BigLake and Spanner change streams to further unify customer data while ensuring it’s delivered in real-time, as well as Vertex AI Workbench and Model Registry to close the data to AI value gap. And to bring data within reach for anyone, we are announcing a unified business intelligence (BI) experience that includes a new Workspace integration, along with new programs that further enable our data cloud partner ecosystem. 

Removing all data limits 

Today, we are announcing the preview of BigLake, a data lake storage engine, to remove data limits by unifying data lakes and warehouses. Managing data across disparate lakes and warehouses creates silos and increases risk and cost, especially when data needs to be moved. BigLake allows companies to unify their data warehouses and lakes to analyze data without worrying about the underlying storage format or system, which eliminates the need to duplicate or move data from a source and reduces cost and inefficiencies. 

With BigLake, customers gain fine-grained access controls, with an API interface spanning Google Cloud and open file formats like Parquet, along with open-source processing engines like Apache Spark. These capabilities extend a decade’s worth of innovations with BigQuery to data lakes on Google Cloud Storage to enable a flexible and cost-effective open lake house architecture. 

Twitter already uses storage capabilities with BigQuery to remove the limits of data to better understand how people use their platform, and what types of content they might be interested in. As a result, they are able to serve content across trillions of events per day with an ads pipeline that runs more than 3M aggregations per second. 

Another major innovation we’re announcing today is Spanner change streams. Coming soon, this new product will further remove data limits for our customers, allowing them to track changes within their Spanner database in real time in order to unlock new value. Spanner change streams tracks Spanner inserts, updates, and deletes to stream the changes in real time across a customer’s entire Spanner database. This ensures customers always have access to the freshest data as they can easily replicate changes from Spanner to BigQuery for real-time analytics, trigger downstream application behavior using Pub/Sub, or store changes in Google Cloud Storage (GCS) for compliance. With the addition of change streams, Spanner, which currently processes over 2 billion requests per second at peak with up to 99.999% availability, now gives customers endless possibilities to process their data. 

Remove the limits of your data workloads

Our AI portfolio is powered by Vertex AI, a managed platform with every ML tool needed to build, deploy and scale models, and is optimized to work seamlessly with data workloads in BigQuery and beyond. Today, we're announcing new Vertex AI innovations that will provide customers with an even more streamlined experience to get AI models into production faster and make maintenance even easier.

Vertex AI Workbench, which is now generally available, brings data and ML systems into a single interface so that teams have a common toolset across data analytics, data science, and machine learning. With native integrations across BigQuery, Serverless Spark, and Dataproc, Vertex AI Workbench enables teams to build, train and deploy ML models 5X faster than traditional notebooks. In fact, a global retailer was able to drive millions of dollars in incremental sales and deliver 15% faster speed to market with Vertex AI Workbench.

With Vertex AI, customers have the ability to regularly update their models. But managing the sheer number of artifacts involved can quickly get out of hand. To make it easier to manage the overhead of model maintenance, we are announcing new MLOps capabilities with Vertex AI Model Registry. Now in preview, Vertex AI Model Registry provides a central repository for discovering, using, and governing machine learning models, including those in BigQuery ML. This makes it easy for data scientists to share models and application developers to use them, ultimately enabling teams to turn data into real-time decisions, and be more agile in the face of shifting market dynamics.

Extending the reach of your data

Today, we are launching Connected Sheets for Looker, and the ability to access Looker data models within Data Studio. Customers now have the ability to interact with data however they choose, whether it be through Looker Explore, from Google Sheets, or using the drag-and-drop Data Studio interface. This will make it easier for everyone to access and unlock insights from data in order to drive innovation, and to make data-driven decisions with this new unified Google Cloud business intelligence (BI) platform. This unified BI experience makes it easy to tap into governed, trusted enterprise data, to incorporate new data sets and calculations, and to collaborate with peers.

Mercado Libre, the largest online commerce and payments ecosystem in Latin America, has been an early adopter of Connected Sheets for Looker. Using this integration, they have been able to provide broader access to data through a spreadsheet interface that their employees are already familiar with. By lowering the barrier to entry, they have been able to build a data-driven culture in which everyone can inform their decisions with data. 

Doubling down on the data cloud partner ecosystem

Closing the data-to-value gap with these data innovations would not be possible without our incredible partner ecosystem. Today, there are more than 700 software partners powering their applications using Google’s data cloud. Many partners like Bloomreach, Equifax, Exabeam, Quantum Metric, and ZoomInfo, have started using our data cloud capabilities with the Built with BigQuery initiative, which provides access to dedicated engineering teams, co-marketing, and go-to-market support. 

Our customers want partner solutions that are tightly integrated and optimized with products like BigQuery. So today, we’re announcing Google Cloud Ready - BigQuery, a new validation that recognizes partner solutions like those from Fivetran, Informatica and Tableau that meet a core set of functional and interoperability requirements. Today, we already recognize more than 25 partners in this new Google Cloud Ready - BigQuery program that reduces costs for customers associated with evaluating new tools while also adding support for new customer use cases. 

We're also announcing a new Database Migration Program to help our customers efficiently and effectively accelerate the move from on-premise and other clouds to Google’s industry-leading managed database services. This includes tooling, resources, and knowledgeable experience from alliances like Deloitte, as well as incentives from Google to offset the cost of migrating databases.

We remain committed to continued innovation with the leading data and analytics companies where our customers are investing. This week Databricks, Fivetran, MongoDB, Neo4j, and Redis are all announcing significant new capabilities for customers on Google Cloud.

All of these announcements and more will be shared in detail at our Data Cloud Summit. Be sure to watch the data cloud strategy sessions, breakouts, and get access to hands on content. There is no doubt the future of data holds limitless possibilities, and we are thrilled to be on this data cloud journey.

Investing in our data cloud partner ecosystem to accelerate data-driven transformations

6 avril 2022 à 07:00

By 2023, 60% of organizations will use three or more analytics solutions to build business applications to connect insights to actions. These multiple implementations add complexity and challenges with multiple data models, disparate toolsets, and lack of integration and governance. To provide organizations the flexibility, interoperability and agility to accelerate data-driven transformations, we have significantly expanded our data cloud partner ecosystem, and are increasing our partner investment across a number of new areas. 

This week at the Data Cloud Summit, we are announcing a new Data Cloud Alliance, along with the founding partners Accenture, Confluent, Databricks, Dataiku, Deloitte, Elastic, Fivetran, MongoDB, Neo4j, Redis, and Starburst, to make data more portable and accessible across disparate business systems, platforms, and environments—with a goal of ensuring that access to data is never a barrier to digital transformation.

We are also rolling out updates to ensure that organizations can effectively utilize the expertise and power of our data cloud partners, including our new Google Cloud Ready - BigQuery initiative to help customers identify validated partner integrations with BigQuery; a public preview of our Analytics Hub to help partners share and monetize their data; a new Built with BigQuery initiative to highlight partner products that utilize our data cloud capabilities; and several new innovations and launches from our partners.

Helping customers identify validated partner integrations with the Google Cloud Ready - BigQuery initiative 

We strive to give customers the best experience when using partner solutions together with Google’s data cloud products. And as more and more customers deploy partner solutions alongside BigQuery, it’s critical that they are able to identify highly effective, validated, and trusted integrations to get the most out of their data. 

To enable this, we are launching a new Google Cloud Ready - BigQuery initiative. Google Cloud Ready - BigQuery is a validation program whereby Google Cloud engineering teams evaluate and validate BigQuery integrations and connectors using a series of data integration tests and benchmarks. Today, we’re announcing 25 launch partners whose integrations and connectors are validated as Google Cloud Ready - BigQuery:

BigQuery partners.jpg
Google Cloud Ready - BigQuery partners

For example, Google Cloud-validated connectors from Informatica help customers streamline data transformations and rapidly move data from any SaaS application, on-premises database, or big data source into Google BigQuery.

“Google Cloud and Informatica have been strategic cloud partners for the last five years, providing end-to-end, scalable enterprise-class data migration, integration and management solutions for customers. Being recognized as a Google Cloud Ready - BigQuery partner further validates Informatica's ability to help customers be successful in their journey to cloud with Google” said Jitesh Ghai, Chief Product Officer at Informatica.

Google Cloud-validated Fivetran connectors continuously replicate data from key applications, event streams, file stores, and more into BigQuery, helping turn big data into informed business decisions. Customers can keep up-to-date with the performance and health of the connectors through logs and metrics available through Google Cloud Monitoring. 

"Customers are looking to move data reliably and securely into Google BigQuery to meet the needs of their business," said Fraser Harris, VP of Product at Fivetran. "We are proud to announce that we have achieved Google Cloud Ready - BigQuery Designation. This marks another milestone in our long-standing partnership with Google Cloud that provides our customers with further assurance that Fivetran products work seamlessly with BigQuery - today and into the future."

Similarly, Google Cloud-validated BigQuery and Tableau integrations allow customers to analyze billions of rows in seconds without writing a single line of code and with zero server-side management. Organizations can create dashboards in minutes and share insights with users instantaneously.

“Tableau strives to meet customers where they are and for many organizations with large complex data problems, that’s on the Google Cloud Platform,” said Brian Matsubara, Vice President, Global Technology Alliances at Tableau. “Partnering with Google empowers our customers to explore their data in real-time to unlock actionable insights that can transform a business."

If you are already a Google Cloud partner, sign up to get your product integration validated by our experts. To become a Google Cloud partner, click here to enroll.

Helping ISVs build and grow their applications with BigQuery

More than 700 partners power their applications with Google’s data cloud - including companies like ZoomInfo, Equifax, Exabeam, Bloomreach, and Quantum Metric. We’re committed to helping these partners both build effective products and go to market, and this week we’re excited to launch the Built with BigQuery initiative, which helps ISVs get started building applications using data and machine learning products like BigQuery, Looker, Spanner, and VertexAI. The program provides dedicated access to Google Cloud expertise, training and co-marketing support to help partners build capacity and go to market. Furthermore, Google Cloud engineering teams work closely with our partners on product design and optimization, to share architecture patterns and best practices. This allows SaaS companies to harness the full potential of data to drive innovation at scale.

“Built with Google’s data cloud, Exabeam’s limitless-scale cybersecurity platform helps enterprises respond to security threats faster and more accurately” said Sanjay Chaudhary, VP of Products at Exabeam. “We are able to ingest data from over 500 security vendors, convert unstructured data into security events, and create a common platform to store them in a cost effective way. The scale and power of Google’s data cloud enables our customers to search multi-year data and detect threats in seconds”

Click here to learn more about the Built with BigQuery initiative.

Enhancing secure data sharing with Analytics Hub

We are also launching a public preview of Analytics Hub, a fully-managed service built on BigQuery that allows our data sharing partners to efficiently and securely exchange valuable data and analytics assets across any organizational boundary. With unique datasets that are always-synchronized, and bi-directional sharing, partners can create a rich and trusted data ecosystem

“As external data becomes more critical to organizations across industries, the need for a unified experience between data integration and analytics has never been more important. We are proud to be working with Google Cloud to power the launch of Analytics Hub, feeding hundreds of pre-engineered data pipelines from hundreds of external datasets,” said Dan Lynn, SVP Product at Crux. “The sharing capabilities that Analytics Hub delivers will significantly enhance the data mobility requirements of practitioners.”

Click here to join the public preview of Analytics Hub.

New launches from our data cloud partners

We’re excited to highlight several important launches from our partners themselves. At Google Cloud, we’re proud to support the fastest-growing and most innovative data and analytics companies, whether they’re running applications on Google Cloud, launching new integrations or connectors, co-creating entirely new capabilities with BigQuery, or continually tweaking and updating their platforms to provide the best experience for customers.

This week our partners Databricks, Fivetran, MongoDB, Neo4j, and Starburst Data are all announcing new capabilities for customers, including:

  • Databricks SQL will be publicly available for all customers on Google Cloud this month, enabling customers to operate multi cloud lakehouse architectures with performant query execution. Learn more, here.

  • Fivetran, in addition to joining the Cloud Ready - BigQuery initiative, is now a partner for the Google Cloud Cortex Framework. With deep experience in moving data from a variety of SaaS and database sources - including SAP, Fivetran offers Google Cloud customers accelerated time to value in unlocking the Google Cloud Cortex Framework data models, driving real-time analytics and business insights.

  • MongoDB is working to launch real-time integration of operational data from Atlas to Google BigQuery (and vice versa) via Dataflow Templates. This enables customers to cross-reference operational data and leverage BigQuery and it’s Analytics, as well as AI/ML tools to support use cases such as anomaly detection in IoT, product recommendations in Retail and fault detection in Manufacturing and feed these insights back to MongoDB Atlas to power the modern real-time enabled enterprise for Continuous Intelligence. This is targeted to be available in Q3/22. 

  • Neo4j is launching a fully-managed graph technology service for data scientists and developers to build intelligent, algorithm-powered applications with Neo4j Graph Data Science on Google Cloud.

  • Starburst is announcing a packaged offer for customers to enrich their BigQuery data foundation with hybrid, cross-cloud data stores.

The depth and breadth of innovation and support from the Google Cloud ecosystem is a tremendous asset for customers as they accelerate their data-driven digital transformations. Our community of expert services partners and systems integrators are heavily engaged, too - to date, our partners have earned more than 80 Specializations and more than 200 Expertises pertaining to data cloud technologies on Google Cloud. Visit our partner directory to find partners specialized in Google’s data cloud.

If you are already a Google Cloud partner, sign up to get your product integration validated by our experts. If you are looking to build your applications on Google’s data cloud, apply for the Built with BigQuery initiative. To become a Google Cloud partner, click here to enroll.

Announcing BigQuery and BigQuery ML operators for Vertex AI Pipelines

5 avril 2022 à 18:00

Developers (especially ML engineers) looking to orchestrate BigQuery and BigQuery ML (BQML) operations as part of a Vertex AI Pipeline have previously needed to write their own custom components. Today we are excited to announce the release of new BigQuery and BQML components for Vertex AI Pipelines, that help make it easier to operationalize BigQuery and BQML jobs in a Vertex AI Pipeline. As official components created by Google Cloud, using these components will allow you to more easily include BigQuery and BigQuery ML as part of your Vertex AI Pipelines. For example, with Vertex AI Pipelines, you can automate and monitor the entire model life cycle of your BQML models, from training to serving. In addition, using these components as part of a Vertex AI Pipelines provides extra data and model governance, as each time you run a pipeline, Vertex AI Pipelines tracks any artifacts produced automatically.

For BigQuery, the following components are now available:

BigqueryQueryJobOp

Allows users to submit an arbitrary BQ query which will either write to a temporary or permanent table. Launch a BigQuery query job and wait for it to finish.

For BigQuery ML (BQML), the following components are now available:

BigqueryCreateModelJobOp

Allow users to submit a DDL statement to create a BigQuery ML model.

BigqueryEvaluateModelJobOp

Allows users to evaluate a BigQuery ML model.

BigqueryPredictModelJobOp

Allows users to make predictions using a BigQuery ML model.

BigqueryExportModelJobOp

Allows users to export a BigQuery ML model to a Google Cloud Storage bucket 

Learn how to create a simple Vertex AI Pipeline train a BQML model, and deploy the model to Vertex AI for online predictions:

In addition to the notebook above, check out the end-to-end example of using Dataflow, BigQuery and BigQuery ML components to predict the topic label of text documents using BQML and Dataflow.

End-to-end example using BigQuery and BQML components in a Vertex AI Pipeline

In this section, we'll show an end-to-end example of using BigQuery and BQML components in a Vertex AI Pipeline. The pipeline predicts the topic of raw text documents by first converting them into embeddings using the pre-trained Swivel TensorFlow model with BQML. Then, it trains a logistic regression model in BQML using the text embeddings to predict the label, which is the topic of the document. For simplification, the model will just predict if the topic is equal to "acq" (1) or not (0), where "acq'' means that document is related to the topic of "corporate acquisitions'' as defined in the dataset. 

Below you can see a high level picture of the pipeline

BQML Pipeline
The high level picture of the pipeline (Click to enlarge)

From left to right: 

  1. Start with news-related text documents stored in Google Cloud Storage

  2. Create a dataset in BigQuery using BigqueryQueryJobOp

  3. Extract title, content and topic of (HTML) documents using Dataflow and ingest into BigQuery

  4. Using BigQuery ML, apply the Swivel TensorFlow model to generate embeddings for each document’s content

  5. Train a logistic regression model to predict if a document's embedding is related to a pre-defined topic (1 if related, 0 if not)

  6. Evaluate the model 

  7. Apply the model to a dataset to make predictions

Let’s dive into them. 

ETL with DataflowPythonJobOp component

Imagine that you have raw text documents from Reuters stored in Google Cloud Storage (GCS). You want to preprocess them and import the text into BigQuery to train a classification model using BQML.

First, you will need to create a BQ dataset, which you can do by running the SQL query CREATE SCHEMA IF NOT EXISTS mydataset. In a Vertex AI Pipeline you can execute this query within a BigqueryQueryJobOp.
code_block
<ListValue: [StructValue([('code', 'from google_cloud_pipeline_components.v1.bigquery import (\r\n BigqueryQueryJobOp,\r\n BigqueryCreateModelJobOp,\r\n BigqueryEvaluateModelJobOp,\r\n BigqueryPredictModelJobOp) \r\n\r\nBQ_DATASET = "mydataset"\r\n\r\n# create the BQ dataset\r\nbq_dataset_op = BigqueryQueryJobOp(\r\n query=f"CREATE SCHEMA IF NOT EXISTS {BQ_DATASET}",\r\n project=project,\r\n location="US",\r\n )'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fe5237a7dc0>)])]>

For the full code, please see this notebook.

Aftering creating the dataset, you will need to prepare the text documents for model training. Importantly, you would need to parse out the relevant sections ("Title", "Body", "Topics") of text from the raw documents. In our case we use a Beam pipeline on Dataflow which uses the Beautiful Soup (bs4) library in Python to extract the following:

  • Title, the title of the article

  • Body, the full text content of the article

  • Topics, one or more categories that the article belongs to

This preprocessing step is based on the ETL pipeline with Apache Beam and can be wrapped in a DataflowPythonJobOp to be used in a Vertex AI Pipeline. DataflowPythonJobOp components enable you to submit Apache Beam jobs written in Python to Dataflow for execution on Google Cloud. In other words, you can use this component to run your data pre-processing step to extract the relevant sections of text via Dataflow.
code_block
<ListValue: [StructValue([('code', '# Dataflow job\r\ndataflow_python_op = DataflowPythonJobOp(\r\n requirements_file_path=requirements_file_path,\r\n python_module_path=python_file_path,\r\n args=build_dataflow_args_op.output,\r\n project=project,\r\n location=region,\r\n temp_location=temp_location,\r\n ).after(build_dataflow_args_op)\r\n\r\n# Wait for Dataflow job to finish running\r\ndataflow_wait_op = WaitGcpResourcesOp(\r\n gcp_resources=dataflow_python_op.outputs["gcp_resources"]\r\n ).after(dataflow_python_op)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fe5237a7d90>)])]>

For the full code, please see this notebook.

To clarify what's happening in the component above, the requirements_file_path contains the GCS bucket URI to a requirements file, the python_module_path contains the bucket uri to the Apache Beam pipeline python script and args contains a list of arguments that are passed via the Beam Runner to your Apache Beam code. Also notice that WaitGcpResourcesOp is required, which allows the pipeline to wait for the component to finish execution before continuing to the next part of the pipeline.

Feature Engineering with BigqueryQueryJobOp and DataflowPythonJobOp

Once the documents have been pre-processed in Dataflow, the next step is to generate text embeddings from the pre-processed documents. The embeddings can then be used to as training data for the logistic regression model to predict the topic. 

To generate embeddings from text, you can use the Swivel model, which is a pre-trained TensorFlow model publicly available on TensorFlow Hub. How do you use a pre-trained TensorFlow SavedModel on text in BigQuery? You can import it using BQML and apply it to the text of our documents to generate the embeddings. Then you split the dataset to get the training sample you will consume into the text classifier.  

The following shows what the component looks like in the pipeline:

code_block
<ListValue: [StructValue([('code', '# run preprocessing job\r\n bq_preprocess_op = BigqueryQueryJobOp(\r\n query=bq_preprocess_query,\r\n project=project,\r\n location="US",\r\n ).after(dataflow_wait_op)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fe5237a79d0>)])]>

For the full code, please see this notebook.

Where bq_preprocess_query contains the preprocessing query, project and location where the BigQuery job runs. 

For example, within bq_preprocess_query, the first step will be to import the Swivel model using BigQuery ML. And a step thereafter is to retrieve the embeddings, which serves as the training data for the document classifier model in the next section.
code_block
<ListValue: [StructValue([('code', "# Create the embedding model by importing the model from GCS\r\n CREATE OR REPLACE MODEL\r\n mydataset.swivel_model \r\n OPTIONS(model_type='tensorflow',\r\n model_path='{MODEL_PATH}');\r\n\r\n...\r\n\r\n# Retrieve embeddings from text using ML.PREDICT\r\n SELECT\r\n title,\r\n sentences,\r\n output_0 as content_embeddings,\r\n topics\r\n FROM ML.PREDICT(MODEL `{PROJECT_ID}.{BQ_DATASET}.{MODEL_NAME}`,(\r\n SELECT topics, title, content AS sentences\r\n FROM `{PROJECT_ID}.{BQ_DATASET}.{BQ_TABLE}`\r\n )"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fe5237a7d00>)])]>

For the full code, please see this notebook.

Training a document classifier with BigqueryCreateModelJobOp

Now that you have the training data (as a table), you are ready to build a document classifier, using logistic regression to determine if the label of the document is 'acq' or not, where 'acq' means that the document is related to the topic of "corporate acquisitions". You can automate this BQML model creation operation within a Vertex AI Pipeline using the BigqueryCreateModelJobOp. This component allows you to pass the BQML training query and, in case you need, parameterize it and the job associated in order to schedule a model training on BigQuery. The component returns a google.BQMLModel which also tracks the BQML model automatically using Vertex ML Metadata which provides extra insight into model and data lineage.
code_block
<ListValue: [StructValue([('code', 'create_bq_model_query = f"""\r\nCREATE OR REPLACE MODEL `{PROJECT_ID}.{BQ_DATASET}.{CLASSIFICATION_MODEL_NAME}`\r\n OPTIONS (\r\n model_type=\'logistic_reg\',\r\n input_label_cols=[\'label\']) AS\r\n SELECT\r\n label,\r\n feature.*\r\n FROM\r\n `{PROJECT_ID}.{BQ_DATASET}.{PREPROCESSED_TABLE}`\r\n WHERE split = \'TRAIN\';\r\n"""'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fe5237a7880>)])]>

For the full code, please see this notebook.

Evaluate the model using BigqueryEvaluateModelJobOp

Once you train the model, you would probably evaluate it before deploying into production for generating predictions. With the BigqueryEvaluateModelJobOp, you just need to pass the google.BQMLModel output you received from the previous training BigqueryCreateModelJobOp and it will generate various evaluation metrics, based on the type of the model. In our case we get precision, recall, f1_score, log_loss and roc_auc which are available in the lineage of pipeline in the Vertex ML Metadata.
Figure 2 - A view of metrics in Vertex AI Metadata
A view of metrics in Vertex AI Metadata (Click to enlarge)

Notice that, thanks to these components, you can implement conditional logic in order to decide what you want to do with the model downstream. For example, if the performance is above a certain threshold, you can have the pipeline proceed to make predictions directly with BigQuery ML, or register your model into a model registry service, or deploy the model to a staging endpoint. This will really depend on the logic you want to implement with your ML pipeline. 

In this example pipeline with document classification, you'll just be generating batch predictions directly without conditional logic. 

Predict using BigqueryPredictModelJobOp

In order to operationalize BQML model prediction, the bigquery module of google_cloud_pipeline_components library provides you the BigqueryPredictModelJobOp which allows you to launch a BigQuery predict model job by consuming the google.BQMLModel of the training component.

code_block
<ListValue: [StructValue([('code', '#simulate prediction\r\n bq_predict_op = BigqueryPredictModelJobOp(\r\n model=bq_model_op.outputs["model"],\r\n query_statement=create_bq_prediction_query,\r\n job_configuration_query=job_config,\r\n project=project,\r\n location=\'US\'\r\n ).after(bq_evaluate_op)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fe5237a7ca0>)])]>

For the full code,  please see this notebook.

In this component,  you also pass in a job_config in order to define the destination table (project ID, dataset ID and table ID) beyond the query statement to format the columns you want to have in the prediction table. 

Below you can see the visualization of the overall pipeline you get in the Vertex AI Pipelines UI.

Figure 3  - The visualization of the pipeline in the Vertex AI Pipelines UI.
The visualization of the pipeline in the Vertex AI Pipelines UI.

Conclusion

In this blogpost, we described the new BigQuery and BQML components now available for Vertex AI Pipelines. We also showed an end-to-end example of using the components for document classification involving BigQuery ML and Vertex AI Pipelines. 

What’s Next

Are you ready for running your BQML pipeline with Vertex AI Pipelines? Check out the following resources and let give it a try: 

References 

  1. https://google-cloud-pipeline-components.readthedocs.io/en/google-cloud-pipeline-components-0.2.2/google_cloud_pipeline_components.experimental.bigquery.html

  2. https://cloud.google.com/architecture/analyzing-text-semantic-similarity-using-tensorflow-and-cloud-dataflow?hl=en

  3. https://towardsdatascience.com/how-to-do-text-similarity-search-and-document-clustering-in-bigquery-75eb8f45ab65 

Special thanks to Bo Yang, Andrew Ferlitsch, Abhinav Khushraj, Ivan Cheung, and Gus Martins for their contributions to this blogpost.

Google is a Leader in the 2026 Gartner® Magic Quadrant™ for Enterprise AI Assistants

9 septembre 2026 à 20:00

We are excited to share that Gartner has named Google a Leader in its inaugural 2026 Magic Quadrant for Enterprise AI Assistants. In this comprehensive evaluation of top enterprise AI assistant vendors, Gartner placed Google in the Leaders quadrant for its evaluation across both Completeness of Vision and Ability to Execute.

Gemini Enterprise helps organizations bring helpful, secure AI directly into the daily work of their employees. It moves teams past basic chat interactions to automating multi-step, end-to-end workflows with AI agents. It connects with the tools and infrastructure companies already use and scales easily and cost-effectively, all with security and governance in place. 

We see this recognition from Gartner as validation of our goal: creating a unified platform where everyone — from business users to developers — can work alongside AI agents to accomplish more together.

[High Res] Gartner EAIA Magic Quadrant

Our take on Google as a Leader

Amid a complex landscape of standalone AI tools and emerging platforms, Gemini Enterprise emerges as a unified, open agentic platform backed by Google’s full AI stack, with strengths mentioned in the report such as:

  • Unified “AI front door”: Gemini Enterprise unifies enterprise chat and search, first and third-party agents, a no-code agent designer, Google Workspace integration, and third-party connectors all in one platform, eliminating the need for organizations to piece together disparate AI tools.

  • Open connectivity: Gemini Enterprise offers extensive connectivity beyond Google's ecosystem — extending to Microsoft 365, other third-party software, and internal enterprise data sources. This open connectivity lets organizations adopt Gemini Enterprise alongside their existing infrastructure without costly system overhauls.

  • Simple economics: Gemini Enterprise offers a straightforward pricing model, with actions like chat and search included in the base SKU. Organizations can select per-user seat subscriptions, as well as a pay-as-you-go option that lets users run agent workloads without hitting quota limits mid-task.

  • Built-in governance: Gemini Enterprise provides robust agent governance out-of-the- box at no extra cost, enabling enterprises to seamlessly manage users, agents, and data permissions while curbing security risks and agent sprawl.

  • Full-stack depth and scale: Google’s vertically-integrated stack provides global infrastructure, custom silicon, world-class models, and a secure, enterprise-ready foundation — all optimized for security, interoperability, and cost. 

To read the full report, download it here.

Accelerating our vision with the latest Gemini Enterprise advancements

Over the past months, we have accelerated Gemini Enterprise with significant product advancements:

  • Tailored industry solutions: We introduced specialized solutions for legal and financial services last month. These tailored solutions bring pre-built agents, domain-specific skills, secure data connectors, and an open partner ecosystem, allowing for rapid deployment. They are also built on top of Gemini Enterprise’s governed control plane so you can create and manage workflows in these highly regulated industries. 

  • Google Antigravity in Gemini Enterprise: With the introduction of AI developer tools in Gemini Enterprise, enterprises can now easily enable agentic dev tools like Antigravity and Android Studio for their developer teams with eligible Gemini Enterprise licenses, and maintain full governance and observability inside the admin console.

  • FinOps and cost controls: New FinOps and cost-control capabilities were introduced last month, allowing organizations to optimize AI spend through more pricing options, Flexible Savings Plans, and granular spend management.

Gemini Enterprise customers are also seeing the value

Hearing this recognition from Gartner is great, but it's not just them—our customers are seeing this real-world value too, and it's driving a positive impact across their organizations every day. 

Check out what our customers are saying about the recent product advancements.

“Deploying Antigravity in Gemini Enterprise allows Accenture to arm our engineers with Google DeepMind’s premier technology on the secure, trusted foundation of Google Cloud. Abstracting away operational complexity ensures our teams don't have to choose between developer speed and enterprise-grade governance — freeing them to deliver high-velocity engineering and transformative value for our clients.” — Chetna Sehgal, Global Practice Lead, Accenture Google Business Group. Learn more here. 

“Cleary is committed to embedding AI into our workflows in strategic and competitive ways. Using Google’s Gemini Enterprise, which can slot in seamlessly with other daily work tools, we can unlock greater efficiencies for our teams and help them deliver even higher quality work for our clients.” — Jeff Karpf, Managing Partner, Cleary Gottlieb. Learn more here.

“As a design partner for the Financial Research agent, Deutsche Bank has helped shape this capability in view of the realities of a highly regulated industry – from data protection and governance to the workflows our teams use every day,” – Marie-Jeanne Deverdun, Chief Technology, Data and Innovation Officer, and Member of the Deutsche Bank Management Board. Learn more here. 

“We’re thrilled to partner with Google Cloud in the early adoption of Gemini Enterprise for Legal. We look forward to integrating Google’s technology to streamline workflow and further support our litigators in shaping outcomes critical to our clients’ futures.” — Joe Petrosinelli, Chairman, Williams & Connolly. Learn more here.

Get started today

How Airtel delivered its flawless Indian Premiere League 2026 cricket broadcasts

9 septembre 2026 à 18:00

For the millions of fervent fans of the Indian Premiere League (IPL), being able to count on a flawless live streaming cricket experience is never up for debate. For Airtel, producing league TV broadcasts with some of the world's most massive concurrent viewership, dropped packets and buffering are simply not options.

During the IPL 2026 season, Airtel partnered with Google Cloud to manage this digital delivery. Across 74 matches, the streaming infrastructure delivered several hundred petabytes of egress data. The final match alone processed tens of billions of requests, hitting a peak egress of several Tbps.

Delivering video under these concurrency spikes requires an edge architecture designed strictly around localization, paired with proactive operational monitoring. 

Our goal for IPL 2026 was to deliver an uninterrupted, stadium-grade viewing experience to cricket fans across India, regardless of concurrency surges or network conditions. Partnering with Google Cloud and using Media CDN gave us deep local edge proximity and excellent cache efficiency. Combined with proactive match-day real-time monitoring, we delivered a reliable broadcast experience from start to finish.

Architecting for concurrency and edge efficiency

IPL-BLog-Architecture

One of the primary challenges in live sports broadcasting is seamlessly handling large traffic spikes and never degrading stream performance or overwhelming backend origins. That’s especially important when millions of viewers simultaneously tune in during a final over because every millisecond counts.

To accelerate content delivery across India’s diverse ISP landscape, Airtel leveraged Google Cloud’s Media CDN. By utilizing Google’s extensive global edge network, Airtel was able to serve viewer requests from edge locations that were physically close to end users. This deep localization was a cornerstone of the broadcast's success, with 99.9% of all tournament traffic being served locally from within India.

This efficient architecture minimized network hops and reduced transit congestion, translating into remarkable infrastructure and viewer experience metrics throughout the 74 matches:

  • Superior caching efficiency: Airtel saw an overall cache hit ratio exceeding 98%. By effectively absorbing massive viewer traffic load at the edge, origin server/video platform demands remained minimal even during peak playoff viewership.

  • Consistent ultra-low latency: Airtel maintained a p99 latency of < 300 ms during the tournament, which supported fast stream start times and minimized buffering risk during critical game moments.

Proactive strategies for operational readiness

While maintaining an intelligent backend architecture was vital to Airtel’s IPL streaming strategy, it was  only half the equation. Executing high-stakes live broadcasts across 74 consecutive matches also demanded meticulous operational preparation and proactive match-day execution.

Because Airtel and Google Cloud recognized that potential bottlenecks had to be identified long before the first ball, they established a deeply integrated operational support model:

  1. Pre-tournament support readiness reviews: Well ahead of the opening match, joint engineering teams conducted comprehensive support readiness reviews. By auditing traffic projections, reviewing manifest configurations, and validating failover mechanisms early, the teams supported robust client readiness, resulting in low operational friction during the tournament.

  2. Monitoring as a service (MaaS): The teams maintained continuous, proactive telemetry monitoring through MaaS on Media CDN, and real-time observability enabled early detection and mitigation of network shifts before anomalies could impact viewer playback.

  3. Dedicated match-day and weekend support: Live sports don't play by the rules of  standard business hours, so Airtel established comprehensive monitoring protocols for every match. During critical weekend fixtures and the high-stakes playoff stage, Google Cloud’s Technical Account Management and MaaS teams worked hand-in-hand with Airtel engineering to provide dedicated, real-time event support.

A blueprint for live broadcast excellence

Airtel’s successful streaming of IPL 2026 demonstrates that handling extreme concurrency is only possible with an integrated strategy across architecture, edge localization, and operational governance. By combining a 98%+ cache hit ratio with 99.9% local delivery and proactive match-day monitoring, Airtel hit a benchmark for live sports broadcasting at scale.

This deployment provides an overview of the technical architecture and operational strategies involved in scaling live media delivery for high-concurrency events. To learn more about optimizing live broadcasts and edge delivery, review the Media CDN developer documentation.

How KDDI built Buffmee, a faster, reliable consumer RAG app

8 septembre 2026 à 18:00

When building consumer-facing generative AI applications,  balancing high generation quality with fast response times across diverse media types, can be challenging. KDDI, a major telecommunications carrier in Japan, tackled this challenge head-on when they developed Buffmee, their consumer Retrieval-Augmented Generation (RAG) app.  

Buffmee is an interactive AI service built on the concept of 'AI that helps you grow.' By grounding responses in over 100 sources — including books, magazines, and web media — it helps users search for information, summarize key points, and explore personalized learning and hobby interests. By citing sources, Buffmee alleviates concerns about information reliability, allowing users to safely deepen their knowledge.  To achieve this, KDDI collaborated closely with their development partner KDDI iret, Google Cloud Consulting and our specialized AI engineers.

As part of their app launch, the engineer team needed to ground a massive variety of proprietary content, including books and magazines. However, they struggled with latency issues that prevented them from meeting their target response times, and they needed a reliable way to ensure hallucination-free results.

image1

Buffmee App Description and Images

To meet these performance targets, organizations need a systematic approach to AI evaluation and real-time bottleneck identification. That is why we are sharing the automated evaluation framework and performance optimization techniques that helped KDDI successfully launch their application. 

The results were inspiring: KDDI reduced total application response latency by 38%, successfully hitting their target response performance. They also achieved a nearly 18% improvement in TTFT.

"Our vision hinged on a platform where content, once ingested, would instantly function as a working RAG system. Google's careful, hands-on guidance made that a reality — we're sincerely grateful for their support." — Shunya Onoda, AI Product Department, KDDI.

With these performance and accuracy improvements, Buffmee now empowers users to safely explore their favorite media through interactive Q&A and deep-dive analysis, delivering a highly personalized experience while maintaining strict trust and compliance for content providers.

Let’s deep dive into how they achieved these results. 

Establish automated evaluation for diverse content

Traditional manual testing requires immense effort and cannot scale to accommodate a large content library. To solve this, the development team designed a systematic AI evaluation process using Gemini Enterprise Agent Platform Evaluation Service.

By implementing automated evaluation frameworks like LLM-as-a-Judge and the Rule of Hundreds, the team replaced labor-intensive manual testing with a data-driven process. They ingested their extensive document corpus, constructed hundreds of automated evaluation tests, and built a comprehensive benchmark dataset to measure the reliability of answers for each use case. As a result, the team improved their groundedness scores by 25%, helping deliver highly accurate and reliable outputs.

image2

KDDI's automated evaluation loop: AI generates questions and scores answers, while humans calibrate thresholds and analyze edge-case failures.

Identify bottlenecks and optimize performance with an agentic loop

To improve response speeds, the team implemented BigQuery Agent Analytics and the Agent Development Kit (ADK) log analysis agent. By analyzing actual production logs, they visualized how skill division and prompt bloat—especially with highly complex, multi-page system prompts — impacted the Time To First Token (TTFT).

The team optimized the system prompt, including the inline integration of skills, and reviewed the sub-agent routing. This allowed them to identify and resolve deep-stack bottlenecks in real time without sacrificing response accuracy.

Four core principles for reliable evaluation 

To achieve these results, the team implemented four core technical practices:

  1. Transitioning to binary evaluation: By selectively moving away from ambiguous 1–5 ratings to a binary "pass (1) / fail (0)" system for critical metrics, the team minimized variance and noise, helping improve automation accuracy.

  2. Strategic content sampling: Rather than attempting to evaluate every single document, the team classified their entire corpus along a two-dimensional grid: File Format (Web articles, EPUBs, PDFs, structured data) and Media Composition (Text-heavy, image-heavy, or mixed). By selecting representative samples from each cell of this difficulty grid, they reduced the evaluation workload by 75% while maintaining comprehensive test coverage.

  3. Thresholds grounded in product judgment: Instead of relying solely on default tool parameters, the product owner reviewed randomly sampled answers alongside their automated scores to calibrate and establish what "good enough to ship" actually meant for the user experience.

  4. Modular splitting of massive prompts into ADK Skills: Because massive system prompts exceeding 800 lines can cause LLM attention drift and latency degradation, the team split prompts by function into Agent Development Kit (ADK) Skills, dynamically loading only the required logic to optimize response times.

Get started

Building scalable, reliable generative AI applications requires both automated evaluation and deep performance analytics. To apply these techniques to your own applications:

Spanner migrations: Automating dual-write with Antigravity CLI for minimal disruption

4 septembre 2026 à 18:00

When Google's Finance Engineering team needed to modernize their legacy data layer, they chose Spanner, a globally distributed, strongly consistent, multi-model database with high availability capabilities. But migrating to Spanner without taking production services offline was a daunting engineering challenge: As the internal team responsible for the application, we needed to manually rewrite dual-write logic across dozens of Data Access Objects (DAOs), a process that is slow and prone to human error. Further, doing so without disruption would have required implementing multi-phase dual-write architectures across every DAO in our codebase. 

To solve this, we took an alternative approach: We built an automated refactoring pipeline powered by Antigravity CLI in headless mode. This helped us accelerate our migration velocity significantly while maintaining strict data parity in our staging environments as we prepare for production. 

The challenge: Anatomy of a dual-write migration

When migrating high-throughput production services where financial accuracy is essential, simple cutover scripts do not work. You must verify that both the legacy datastore and Spanner receive identical writes simultaneously until all the historical data backfills and verifications are complete.

We structured our migration across three distinct phases:

  • Historical backfill: Copying existing historical records to Spanner while maintaining referential integrity.

  • Dual-write / dual-read implementation: Modifying every DAO to write mutations to both the primary store and Cloud Spanner in parallel during the migration window.

  • Automated API verification and parity checking: Intercepting RPC traffic and verifying end-to-end that every write lands with byte-for-byte equivalence across both stores.

1 - Dual Write Architecture

The architectural pattern is clean, but at our scale, we began to encounter friction. That’s because each DAO requires:

  • A dedicated MutationConverter class mapping complex domain models to Spanner schema columns

  • Dual-write branch handling and rollback or error-reporting logic

  • A suite of unit tests verifying both primary and Spanner writes using fake time sources and test doubles (FakeTimeSource)

Performing these identical, high-precision code changes across 30+ DAOs by hand would have taken months of engineering time.

The solution: Standardized mutation converter patterns

To verify that our automation pipeline could reliably generate clean code, we first standardized our DAO refactoring pattern around a decoupled MutationConverter interface.

Instead of embedding raw Spanner table names and column assignments directly inside core DAO business logic, we isolate Spanner schema translation into dedicated converter units:

code_block
<ListValue: [StructValue([('code', '// Example of the standardized pattern generated by our pipeline\r\n\r\ntype BpcTransferAmountsMutationConverter interface {\r\n ToInsertMutation(entity *model.BpcTransferAmount) (*spanner.Mutation, error)\r\n ToUpdateMutation(entity *model.BpcTransferAmount) (*spanner.Mutation, error)\r\n}\r\n\r\ntype bpcTransferAmountsMutationConverterImpl struct {\r\n tableName string\r\n}\r\n\r\nfunc (c *bpcTransferAmountsMutationConverterImpl) ToInsertMutation(entity *model.BpcTransferAmount) (*spanner.Mutation, error) {\r\n if entity == nil {\r\n return nil, errors.New("entity cannot be nil")\r\n }\r\n \r\n // Map domain fields to Cloud Spanner table schema\r\n cols := []string{"TransferId", "AmountCents", "CurrencyCode", "LastModifiedTimestamp"}\r\n vals := []interface{}{\r\n entity.TransferId,\r\n entity.AmountCents,\r\n entity.CurrencyCode,\r\n spanner.CommitTimestamp, // Use Spanner commit timestamps\r\n }\r\n \r\n return spanner.Insert(c.tableName, cols, vals), nil\r\n}'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f8e31a4f0d0>)])]>

By establishing a rigid, deterministic contract between the DAO and the Spanner SDK (spanner.Mutation), we created an exact target specification that an AI coding agent could reason about and generate reliably.

Why use Antigravity CLI in headless mode?

Interactive AI chat interfaces in IDEs work well for exploratory coding, but they are poorly suited for systematic, multi-file code updates across an entire codebase. When you need to apply repeatable refactoring to dozens of targets without missing edge cases, you need automated workflows.

We addressed this by building an orchestration script (migration_ui.py) that runs Antigravity CLI in headless mode (-p).

Headless mode lets Antigravity run directly inside shell scripts, continuous integration pipelines, and background automation jobs without requiring manual terminal prompts. This approach helped us scale our work in three key ways:

  • Deterministic prompt architectures: We treated our prompts as version-controlled engineering artifacts. We codified precise rules handling common Spanner edge cases — such as timestamp serialization, nullability conversions, mutation ambiguity, and FakeTimeSource test injection — directly into reusable prompt templates.

  • Batch execution and automated verification: Our orchestration script takes a target DAO name as input, retrieves the existing single-write source code and schema, and feeds it to headless Antigravity alongside our structural conventions. Antigravity generates the new converter, the refactored dual-write DAO, and corresponding unit tests. The script then runs blaze test. If a linter error or test assertion fails, the error log feeds directly back into Antigravity for self-correction.

  • Overnight execution at scale: Because the loop runs unattended, engineers can queue up 10 DAOs at the end of the day. By morning, the pipeline generates, tests, and validates 10 clean changelists ready for human code review.

Results and key takeaways for cloud engineers

Combining Spanner's distributed database primitives with Antigravity CLI's headless automation produced clear benefits across our engineering organization:

  • Significant reduction in migration effort: DAO dual-write migrations that previously required extensive manual coding and testing were completed and reviewed in a fraction of the time 

  • Highly reliable data migration: Because every generated DAO adhered to the exact same tested MutationConverter pattern and underwent automated unit testing against Spanner test doubles, we sustained high data fidelity during our extensive migration testing. 

  • Focus on higher-value engineering: Engineers avoided repetitive boilerplate refactoring, giving them time to focus on data modeling, architectural resilience, and performance optimization.

Three tips for your next database migration

  1. Decouple schema translation first: Before writing migration scripts, define a strict interface (like our MutationConverter) that isolates your new cloud database SDK requirements from your existing business logic. AI agents work best when given clear, bounded design patterns.

  2. Move from interactive chat to headless automation: When executing repetitive refactoring across more than three or four files, invest in scripted, headless workflows. Treating prompt inputs and test verifications as automated build steps help maintain quality and consistency.

  3. Let the build system act as your guardrail: Connect your AI generation loop directly to your build and test harness (bazel test or go test). This lets the model fix compile and assertion errors before a developer reviews the code.

Get started

Whether you’re migrating financial systems or building cloud-native applications from scratch, Spanner and Antigravity provide a foundation for scalable software development.

Getting started with Mantis, our open-source bug finding-and-fixing harness

2 septembre 2026 à 18:00

AI models have clearly proven their ability to discover and exploit vulnerabilities without much, if any, human assistance. To help defenders gain the advantage with AI, we built the Mantis harness to automate the discovery, triage, reproduction, and patching of software vulnerabilities. 

Available to all as an open-source framework, Mantis is part of Google’s internal approach to find and fix vulnerabilities at machine-speed. It creates a more effective scalable, context-aware repository analysis. 

While sloppiness in AI code scanning frequently leads to hallucinated bugs and weak true-positive rates under 7%, we designed Mantis to be effective by combining industry-standard agentic techniques like critic and review agents with sandboxed reproduction of vulnerabilities for grounding. 

As we detailed in June, it examines the history of the repository to learn from past security fixes and automatically builds up architectural and threat model documentation, even if these are not provided. 

It constructs a hierarchical security summary tree, condensing individual files into directory and root-level summaries. This technique reduced token overhead by over 85%, while preserving critical structural context across massive repositories.

Mantis distills decades of cybersecurity expertise across a wide spectrum of codebases, and is available on GitHub. Here’s how you can get started using Mantis.

  • First, clone the Mantis repo locally using:

code_block
<ListValue: [StructValue([('code', 'git clone https://github.com/google/mantis.git'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fcc5fa559d0>)])]>
  • Second, open your favorite coding agent and use the prompt, “I would like to use Mantis framework in path/to/mantis to review my code in path/to/your/code, can you help me get started?” 

Internally at Google, this exact prompt has been used to find real vulnerabilities across our many code repositories. As part of the Mantis repository on GitHub, we’ve included sample sandboxing options. You can also implement your own sandbox to match your own workflow.

Mantis is intended to be an easy place to start with vulnerability discovery, true positive filtering, and patching. Once you've got a handle on AI-discovered vulnerabilities, you can use the new mantis-advise skill to make use of the accumulated knowledge and get your coding agents to write secure code the first time.

To get the most out of AI-driven vulnerability discovery and modernize your development practices, we strongly recommend two essential practices:

  1. Feed your tools the right context: While Mantis automatically analyzes commit history and code to build documentation for itself, human-curated knowledge often can dramatically improve the quality of your results. For example, if you would never waste time fixing bugs where the user can crash their own program, this is critical information for a scanning pipeline to ensure that those types of bugs are never surfaced.

  2. Build a cyber sandbox with vulnerability acceptance criteria. Safe, sandboxed environments where you can reproduce vulnerabilities with clear vulnerability-reproduction criteria will give you better results for surfacing only the things you need to know and also for ensuring that your fixes are correct.

You can learn more about Mantis here.

Reimagining work: How Pythian’s internal AI playbook delivers customer ROI

27 août 2026 à 18:00

When Pythian rolled out Google Cloud’s Gemini Enterprise across our 500-person company in 27 countries, the goal was simple: use our own company as a proving ground to discover how enterprise AI actually delivers ROI.

What we found changed our strategy entirely.

Since the rollout of Gemini Enterprise and our previous enterprise AI deployments, Pythian observed firsthand why so many enterprise AI initiatives stall out or fail. 

Most organizations trap themselves in a tool-centric mindset — buying licenses, making tools broadly available, and assuming value will naturally follow. They get stuck chasing "nickel and dime" micro-efficiencies (like saving 5 minutes per user) while missing structural, high-ROI workflow transformations. Compounding the problem, even when custom agents are built, they frequently stall in pilot mode or break down in production because teams lack the operational capability to manage AI model drift, agent lifecycles, and ongoing observability.

To solve this, we engineered the Pythian AI Operating Model — a multifaceted, end-to-end framework designed to take enterprise AI from high-level strategy all the way into sustained production. While our dual center of excellence (COE) serves as the core execution muscle, it is the application of the entire framework, from Field CTO strategy and tooling deployment to the dual COE and XOps, that consistently unlocks million-dollar outcomes.

By proving this complete model internally first, Pythian drove a 3x surge in active user engagement and cut our database incident resolution times by 80%.

The four pillars of the Pythian AI operating model

To move past the common failure points of enterprise AI, our framework consolidates strategy, execution, and operations into a single continuous loop:

Field CTO strategy  ──>  tooling deployment  ──>  dual COE execution  ──>  production XOps

  1. Field CTO strategy and governance: Generative AI is arguably the most academically challenging architectural shift in IT history. Led by former C-suite tech leaders, our Field CTO practice provides executive advisory to establish steering committees and clear value metrics. The team audits operations using 16 horizontal agentic patterns (like automated document processing and runbook creation) to build a prioritized backlog of high-ROI use cases before development starts.

  2. Tooling and platform deployment: The team establishes a secure, production-grade foundation on platforms like Gemini Enterprise and connects AI directly into CRMs, ERPs, and database estates to ground models in real corporate context.

  3. The dualCOE: This execution muscle is split into two specialized engines:

  • People productivity COE: This group handles adoption and change management. Instead of expecting non-technical teams (like HR or Procurement) to build its own agents, this COE builds no-code agents for them, focusing entirely on enablement.

  • Process productivity COE: This team engineers deep, custom-coded AI agents and complex agentic workflows that integrate into core data platforms for autonomous operations.

  • XOps (AI production management): While deploying an agent is 20% of the journey,  maintaining accuracy in production is 80%. Because AI models and prompt structures naturally drift over time, this XOps practice provides the continuous monitoring, prompt tuning, and model observability needed to keep agents performing without breaking core workflows.

  • The difference between chasing minor, scattered efficiencies and driving structural enterprise ROI comes down to how you align your operating strategy:

    Alignment element

    Tool-centric approach

    Pythian AI operating model

    Primary metric

    Individual minutes saved per user

    High-impact workflow reimagination and ROI

    Operational focus

    Broad, unguided tool availability

    Prioritized backlog via 16 agentic patterns

    Execution muscle

    Ad-hoc user experimentation

    Dual COE (people and process productivity)

    Production lifecycle

    Unmonitored static deployments

    Active XOps (Continuous accuracy and drift management)

    Real-world impact: from database ops to global supply chains

    Whether managing 70 manufacturing plants or 30,000 enterprise databases, AI succeeds when tied to structural, high-value workflows:

    • Pythian “as a customer:” Across 15,000 monthly database tickets, our Process COE deployed an agentic workflow that reads tickets, searches knowledge bases, and auto-generates mini runbooks before an engineer touches them. The result was slashed mean time to resolution by 80% and tripled active user engagement.

    • Knowledge management customer: We deployed autonomous IT support agents across 10,000 consultants. As a result, we were able to automate 10% of 20,000 annual IT tickets into "no-touch" resolutions, saving 1,000,000+ operational hours.

    • Supply chain customer: By building custom agentic supply chain tools on Gemini Enterprise, we compressed forecast-matching cycles from weeks down to 2–3 days across 70 global manufacturing sites.

    • Retail customer: We combined Gemini Agentic AI and computer vision to automate store product onboarding. As a result, we transformed a 20-minute manual task into a multi-second flow.

    Ready to build your AI operating model?

    Scaling AI demands more than tool-level experimentation. It also requires an end-to-end AI operating model. Learn how Pythian pairs with Google Cloud to operationalize strategy, streamline XOps, and fast-track your Gemini Enterprise journey.

    FinOps for the AI era: New flexible billing and cost controls for agents

    26 août 2026 à 15:30

    Editor's note: A product image was updated after initial publication.


    As AI takes on more complex work, business leaders face a new challenge: enabling rapid innovation using agents while protecting their margins and budgets. To get a real return on AI, financial operations (FinOps) and cost management must evolve alongside technology, giving you clear visibility, proactive cost controls, and flexible payment models that fit your needs. 

    That’s why today we’re introducing expanded billing flexibility and new cost management tools for agent workloads across Gemini Enterprise and developer tools like Google Antigravity in Gemini Enterprise and Android Studio.

    • Flexible payment options: You can mix our existing, predictable per-user seat subscriptions with a new pay-as-you-go option in Gemini Enterprise app that lets you run agent workloads without hitting quota limits mid-task.

    • Developer access, one place to manage your AI: Google Antigravity and Android Studio AI use is now included in your Gemini Enterprise subscription (available for select customers and rolling out broadly soon), giving your developers more without giving you more to manage. Usage across Antigravity, the platform, and the app rolls up into a single view instead of separate licenses and billing silos.

    • Pay less as your usage grows: If your AI workloads are steady or climbing, Flexible Savings Plans let you commit to a monthly spend you're comfortable with and take 10–20% off your token costs — no minimums, no maximums, and no new billing silo to manage.

    • Consolidated spend guardrails: You can now set hard monthly caps on AI spend and projects, estimate agent runtime costs, and catch sudden budget spikes before they hit your invoice.

    Give your teams flexibility without losing control over spend in Gemini Enterprise

    Every organization operates differently. Even within the same business, no two teams consume AI in the same way. Your business users might rely on steady, everyday productivity tools. Meanwhile, your technical teams might run AI agent workloads in bursts. 

    To help align costs with how work actually gets done, you can combine these payment and licensing choices and features across Gemini Enterprise:

    Option

    How it works

    Why it helps optimize spend

    Gemini Enterprise app per-user seat subscription

    You pay a fixed monthly fee per user, which includes daily quota pools that are shared across your entire project.

    Predictable budgeting. It provides finance teams with a clear, steady monthly baseline for teams with consistent daily productivity needs.

    [New] Gemini Enterprise app pay-as-you-go consumption edition

    *available for select customers and rolling out broadly soon

    There is no upfront commitment or base subscription fee, meaning you pay strictly for the compute and tokens your teams consume at standard model API rates.

    Only pay for what you use. Your spend scales up and down automatically with real usage, ensuring you never pay for empty seats when project demand dips.

    [New for Antigravity in Gemini Enterprise] Consolidated pooled quotas

    Daily usage allowances are pooled project-wide, letting business apps, developer tools, and custom agents draw from the same shared quota. Pooled quota is always exhausted first, and admins can control if overages are allowed, at which point it’s charged at pay-as-you-go rates. 

    Maximized resource usage: Unused daily allowances from business users automatically absorb heavy developer or custom API agent demands, so no quota allowance goes to waste.

    [Coming soon] Deferred execution pricing

    *available for select workloads soon

    Mark eligible agent workloads as deferred, and our intelligent scheduler in the Gemini Enterprise Agent Platform runs them during off-peak capacity windows.

    Substantial discounts for work that can wait: AI workloads can run on separate, off-peak capacity, you pay up to half the inference cost and bypass standard quota limits entirely – letting you run substantially more agentic volume under the same budget.

    Equip developers with advanced agentic tooling under a single Gemini Enterprise subscription

    We’re rolling out access to Google Antigravity in Gemini Enterprise, an agent-first developer platform that brings powerful agentic coding and agent-building capabilities to technical teams, included with Gemini Enterprise subscriptions for eligible customers. In addition, Android developers can leverage the Google Antigravity quota included in their Gemini Enterprise subscriptions natively in Android Studio, the agentic IDE for professional Android development.

    To be more efficient with agentic coding costs, we are pooling developer tools quota included in each Gemini Enterprise subscription and making it available across the whole Google Cloud project so your teams can benefit from the capacity you’re already purchasing. Your developers get access to advanced agentic tools, while you maintain centralized governance and control.

    For a closer look into what’s new with Antigravity in Gemini Enterprise and how customers are putting it to work in production, take a look at our deep-dive.

    Budget smarter with Gemini Enterprise Flexible Savings Plans (FSPs) 

    If your organization has steady or growing AI workloads, Gemini Enterprise Flexible Savings Plans offer a simple, spend-based commitment model across Gemini Enterprise usage. FSPs are designed to lower token costs while keeping budgets flexible:

    • Programmatic savings: Receive 10% off for 1-year or 20% off for 3-year commitments for monthly spending across Gemini Enterprise.

    • Tailored to your pace: With no minimum or maximum spend requirements, you can determine a monthly commitment that fits your current traffic and make adjustments as your usage increases over time. 

    • Enterprise Agreement (EA) friendly: FSP spend seamlessly draws down against your existing Google Cloud EA, giving lines of business dedicated budget control without fragmenting your broader cloud commitments.

    Gemini Enterprise Flexible Savings Plans are already available for self-serve customers and customers on enterprise agreements.

    Give your teams the freedom to build while maintaining financial discipline

    As a leader, your goal isn't to restrict the potential value of AI  – it's to remove the financial and operational risk that you face without managed AI costs. You should be able to give engineering, marketing, and operational teams the freedom to innovate with agents, but you should also have the visibility to trust what those agents are doing and the safety nets to protect your budget.

    To bridge this gap, we've built robust, native governance tooling directly into the Google Cloud Billing Console around three simple goals:

    1. Plan before you scale: The Google Cloud Pricing Calculator lets you estimate anticipated costs in Gemini Enterprise across per-user licenses, developer tools, and background agent runtimes. It gives you the numbers you need to build clear business cases upfront before project work begins.

    2. Enforce boundaries without micromanaging spend: Instead of spending time tracking daily usage variations across project teams, let these tools do the monitoring for you:

    • Early anomaly detection: If a project’s AI spending trends higher than normal, the system flags the deviation with root cause analysis and pinpoints the top 3 SKUs driving the increase so you can see exactly what changed.

    1 Jul22_Anomalies_Image1

    Billing Console showing an Early Anomaly alert with the Root Cause Analysis (RCA) breakdown highlighting the driving SKUs

      • Project-level spend caps: When a project needs defined financial boundaries, you can set a firm monthly spend limit directly in the Google Cloud Billing Console. If a project hits its limit, the agent's API calls temporarily pause – protecting your budget without affecting the rest of your production infrastructure. Automated email alerts at 50%, 80% and 100% of the budget keep you informed of your progress against the spend limit. 
      • Overage controls: If a spend cap triggers, you can choose to resume work with a single click in the console. Alternatively, if your priority is continuous operation, you can turn on overages so excess usage smoothly transitions to consumption rates, which can draw directly against your FSP to keep overage unit costs heavily discounted.
    3 PAYG Overage Enabled

    Enabling overage pay-as-you-go for a project.

    3. Get visibility into business value: Use centralized billing reports paired with the FinOps agent to generate natural-language cost insight summaries of where your budget went, making it simple to show ROI to leadership.

    Cost overview FinOps

    AI spending reporting in Google Cloud Console

    Go deeper with AI cost optimization

    To build a full-stack FinOps strategy that optimizes the cost, latency, and performance of your models and infrastructure, explore our detailed architecture specifications and frameworks:

    • How to outsmart infrastructure constraints with dynamic capacity management: Discover how to optimize your compute investments with capabilities in Google Kubernetes Engine and Google Compute Engine that automatically schedule and reallocate resources to avoid interruptions, over-provisioning, and over-reliance on any one hardware configuration.

    • Expanding Google Antigravity for Enterprise Customers: Read our developer tooling deep-dive to see how technical teams are accelerating software delivery with agent-first workflows.

    • What sports cars can teach us about optimizing AI spend: More tokens doesn't always mean better AI. Read our conversation with Mike Clark, Director of Product Management for Gemini Enterprise Agent Platform, on how to balance horsepower with efficiency and get the highest return out of every dollar you spend on AI. 

    • Protection during usage spikes: Your heavy workloads can surge during peak hours without forcing you to pay for expensive, dedicated infrastructure that sits idle the rest of the time. As your AI usage grows, Gemini models can automatically scale on demand without hitting artificial rate limits – processing up to 50 million tokens per minute.

    Now introducing Gemini Enterprise for Legal

    25 août 2026 à 14:00

    Few professions are as exacting as the practice of law. A team reviewing a contract or building a case works inside strictly privileged information, firm-specific playbooks, and a body of law that changes constantly. The work thrives on nuanced, professional judgment — and the systems supporting it inherit real obligations: ethical walls that cannot be crossed, matter permissions that cannot be flattened, and a duty of confidentiality that does not bend for convenience.

    General-purpose AI, however capable, does not meet that standard on its own. Foundational model intelligence is necessary. For legal work, it is nowhere near sufficient.

    What makes the difference is the system built around the model: skills that enhance a firm's own expertise, connections into the systems where matters actually live, agents that complete work rather than return suggestions, and an open ecosystem to extend all of it — with governance running underneath all four. Each is valuable alone. Only in combination do they produce something a firm or a legal department can put into production and actually rely on.

    Today we're bringing that to legal practice with Gemini Enterprise for Legal, part of our new suite of purpose-built industry solutions.

    Bringing Gemini Enterprise to your legal practice

    Four components of Gemini Enterprise for Legal

    Developed alongside industry leaders, Gemini Enterprise for Legal provides an integrated, fully governed environment configured for rapid deployment across firms and corporate legal departments:

    1. Purpose-built skills for legal work. Skills are reusable packages of instructions and context, designed by domain experts, that teach an agent to run a specialized task while enforcing your firm's playbooks, citation rules, and house style. They cover contract review and redlining, playbook creation, regulatory horizon scanning, legal research, DSAR fulfillment, and more — and they are where a firm's institutional knowledge becomes something the platform can execute rather than something a partner has to re-explain.

    2. Connections to trusted systems and data. Secure MCP connectors link agents to the document management systems, case repositories, research services, and industry applications legal teams already rely on — inheriting each platform's existing user permissions and access controls rather than working around them.

    3. Agents that act within the data. Skills and connections come together in agents that carry work through: pre-built agents from Google and leading legal software providers deploy out of the box. Specialized agents handle legal and policy research, regulatory screening, and contract drafting — bringing deep legal expertise onto a platform with centralized governance.

    4. An open partner ecosystem. Every firm and legal department practices differently. Partnerships with global systems integrators and legal-tech specialists — Accenture, Deloitte, Devoteam, Factor Law, KPMG, Tribe.ai, Valtech, Zazmic, Zencore, and 66degrees — let organizations customize, integrate, and scale across complex enterprise architectures without vendor lock-in.

    Running underneath: a governed control plane. A single dashboard for legal IT and risk teams that natively enforces security policies (VPC, CMEK), maintains private data isolation, and holds every output to verifiable grounding with traceable citations.

    Unlocking high-value workflows with domain-specific skills

    Gemini Enterprise for Legal shifts AI from passive querying to agentic execution, automating high-volume, precision-critical workflows such as:

    • Proactive regulatory horizon scanning: Keeps legal and compliance teams ahead of global mandates by autonomously tracking legislative updates, court dockets, and supervisory bodies. It cross-references emerging changes against enterprise policies to flag exposure gaps and generate updated policy drafts for immediate practitioner review.

    • Automating data discovery and DSAR response: Modernizes privacy workflows by compiling personal data across fragmented enterprise systems in seconds. It eliminates the manual toil of Data Subject Access Requests (DSARs), and allows for adherence to regulatory timelines while minimizing operational risk.

    • Accelerating contract review and negotiation: Compresses turnaround times for inbound vendor agreements, NDAs, and complex M&A documentation by benchmarking terms against enterprise playbooks. It surfaces high-risk clauses and potential exposure, enabling attorneys to focus on strategic negotiation and high-value judgment.

    • Building and updating contracting playbooks: Transforms legacy agreement archives into dynamic, actionable playbooks instantly. It automatically extracts key terms, fallback positions, and institutional knowledge to maintain portfolio-wide term consistency and lower negotiation variance across the enterprise.

    • Redacting documents for motions to seal: Eliminates the manual burden of preparing court filings and redacting legal documents. It intelligently identifies sensitive terms and PII for rapid practitioner confirmation, dramatically accelerating filing timelines while safeguarding confidentiality.

    • Drafting NDA documents: Elevates contract creation through structural fidelity validation that enforces firm standards and logical document hierarchies. It allows legal teams to rapidly generate and evolve non-disclosure agreements with complete formatting confidence and minimal review overhead.

    Open ecosystem of connectors across the legal technology stack

    Legal work is only as good as its sources, and legal data carries permissions that have to travel with it. Gemini Enterprise for Legal connects directly to core legal systems via secure MCP connectors. Crucially, access is bound by existing role-based access controls, document-level permissions, and trusted data controls inherited from document management and ediscovery systems.

    Productivity and collaboration:

    • Google Workspace: Connects seamlessly with Google Docs, Gmail, Drive, and Sheets to analyze matter communications, correspondence, and surface internal files while enforcing enterprise access controls.

    • Microsoft 365: Integrates directly with Word, Outlook, and SharePoint to triage inquiries, redlines, and securely ground work product across emails and matter folders without breaking workflow context. 

    Document management:

    • iManage: Gives Gemini Enterprise for Legal permission-bound, auditable access to governed iManage content, including matter history, documents, and institutional knowledge, eliminating the need for bulk exports or custom integrations.

    • NetDocuments: Enables Gemini Enterprise to search and analyze an organization's knowledge and expertise while preserving each user's existing permissions and ethical walls. Source documents never leave the governed NetDocuments environment. 

    Contract lifecycle and execution:

    • Docusign: Integrates agreement metadata, active approval workflows, and contract repositories to surface obligations, track renewal dates, and streamline drafting-to-execution lifecycles.

    E-discovery and litigation intelligence:

    • Everlaw: Connects Gemini Enterprise to litigation and investigations evidence in Everlaw, allowing legal teams to search and analyze their data, uncover case insights, and build timelines directly in Gemini Enterprise, with responses grounded in the underlying documents and access governed by each user’s existing Everlaw permissions.

    • RelativityOne: Allows legal teams to stand up workspaces, organize case data, and manage operations within a secure perimeter. 

    Primary law, research, and public dockets:

    • Thomson Reuters HighQ: Connects Gemini Enterprise for Legal with HighQ, helping legal teams securely access and reference relevant HighQ content within their workflows. 

    • Free Law Project’s CourtListener.com: Provides access to millions of federal and state court opinions, PACER dockets, judicial profiles, and oral arguments.

    • Courtroom5: Delivers jurisdiction-aware civil litigation datasets, procedural rules, and deadline calculation logic.

    Specialized legal AI and intellectual property:

    • Harvey: Bridges Harvey’s legal reasoning intelligence into Gemini Enterprise, supporting complex legal reasoning and research across Vault projects. 

    • Solve Intelligence: Links Gemini Enterprise to worldwide patent and non-patent literature, SEP technical standards, and prior art databases for patent drafting and claim charting.

    • Legora: Agentic operating system for legal work, supporting lawyers in research, review, and drafting across complex matters 

    Third-party agents and implementation partners 

    Every firm and legal department practices differently. Through our open platform, organizations can deploy pre-built partner agents or collaborate with systems integrators to scale custom capabilities:

    • Deloitte: Contract Summarize Pro Agent that synthesizes complex contracts into clear summaries for rapid insight and informed decision-making. Clause Guard contract redlining agent to accelerate turnaround times, and minimize risk in contract management. 

    • Eudia Knowledge agent: Accelerates high-stakes legal and contracting work by combining institutional intelligence with a suite of agents that execute deep legal research, high-volume document analysis, and regulatory compliance screening.

    • Global systems integrators & tech partners: Strategic partnerships with Accenture, Deloitte, Devoteam, Factor Law, KPMG, Tribe.ai, Valtech, Zazmic, Zencore, and 66degrees ensure legal teams can customize, integrate, and scale these capabilities across complex enterprise architectures without vendor lock-in.

    Gemini Enterprise for Legal offers leading firms a way to manage modern legal work with a secure agentic platform

    Developed alongside leading global law firms

    We are working closely with leading law firms, including Cleary Gottlieb, Freshfields, Weil, and Williams & Connolly, to ensure these capabilities address the realities of sophisticated legal practice.

    “Cleary is committed to embedding AI into our workflows in strategic and competitive ways. Using Google’s Gemini Enterprise, which can slot in seamlessly with other daily work tools, we can unlock greater efficiencies for our teams and help them deliver even higher quality work for our clients.” — Jeff Karpf, Managing Partner, Cleary Gottlieb.

    “The legal sector is entering a period of accelerated change and transformation. For Freshfields the opportunity lies in how effectively we combine frontier technology like Gemini Enterprise for Legal with our expertise, robust governance and institutional knowledge to create value for our clients. Our strategic, multi-year partnership with Google Cloud is helping us accelerate that work and enhance how we deliver legal services.” — Alan Mason, Global Managing Partner, Freshfields.

    “We’re thrilled to partner with Google Cloud in the early adoption of Gemini Enterprise for Legal. We look forward to integrating Google’s technology to streamline workflow and further support our litigators in shaping outcomes critical to our clients’ futures.” — Joe Petrosinelli, Chairman, Williams & Connolly.

    "Our collaboration with Google gives us early access to emerging capabilities while allowing us to help shape the platform based on the realities of sophisticated legal practice. The result is technology that helps us continue to deliver the innovative, high-quality service our clients expect from Weil. We are looking forward to working with Google Cloud engineers and the Gemini Enterprise product team as we further innovate and evolve our AI capabilities." — Ramona Nee, Incoming Executive Partner, Weil.

    Weil scales judicial insights with Benchmark, built on Gemini Enterprise

    Built on an enterprise-grade foundation

    Confidentiality is not a feature of legal technology; it is the precondition for using any at all. Because Gemini Enterprise for Legal runs on Google Cloud infrastructure and the Gemini Enterprise platform, the permissions and access controls your firm already maintains are the boundaries the platform operates within — not settings it asks you to reconstruct.

    Client data, firm-specific playbooks, intellectual property, custom agents, and model outputs stay private to your organization, and are never used to train or fine-tune Google's foundation models. Because we operate the full stack, from infrastructure and models through the application layer, we can tune performance and cost together — so expanding what your teams can take on does not mean expanding spend at the same rate.

    This is just the beginning

    The launch of Gemini Enterprise for Legal represents another defining step in delivering on the promise of Gemini Enterprise: bringing the best of Google AI to every professional, for every workflow, natively tailored to the way they work.

    Gemini Enterprise for Legal is available in preview today, launching alongside our new purpose-built solution for Financial Services. We invite law firms and legal teams to explore how Gemini Enterprise can transform their most critical workflows, with solutions for Healthcare, Life Sciences, and other Professional Services on the horizon.

    Now introducing Gemini Enterprise for Financial Services

    25 août 2026 à 14:00

    Protecting capital in today's markets requires immense speed and precision. A financial analyst preparing a deal memo works across licensed market data, internal models, and confidential client files. General-purpose AI lacks the real-time accuracy, verifiable data lineage, and strict security that financial institutions demand. While model intelligence is necessary, without deep integration into trusted financial systems, it is not sufficient.

    Making AI genuinely useful inside an industry requires four things, together: domain expertise encoded into reusable skills, secure connections to the systems and data the work depends on, agents that can act inside real workflows, and an open ecosystem that extends and scales all of it — with governance running underneath all four. Each is valuable alone. Only together do they produce something an institution can actually put into production and see true return on investment.

    Today, we are delivering on this vision with Gemini Enterprise for Financial Services, bringing Google’s agentic AI directly into the workflows of capital markets and corporate banking.

    Bringing Gemini Enterprise into the workflows of capital markets and corporate banking

    Four components, built for financial work 

    Gemini Enterprise for Financial Services delivers an integrated, secure environment configured for rapid deployment with four core components:

    1. Purpose-built financial skills. Skills are reusable packages of instructions and context that teach an agent to run a specialized task the way your institution runs it — applying custom formatting to a report, pulling a specific data cut, following a defined research methodology. They are available inside the Financial Research agent and to any agent your teams build.

    2. Secure Model Context Protocol (MCP) connectors. Direct integrations, using MCP, into essential financial platforms and licensed data sources, configured inside your own environment. Access stays bound by the entitlements you already maintain — licensed data stays licensed, and permissioned data stays permissioned.

    3. Agents that act. At its core is the Financial Research agent which is a Google-built, Google-managed agent that runs end-to-end research with full explainability. It ships with more than 50 foundational skills and exposes its reasoning through confidence scores, explicit methodologies, data snapshots for auditing, and precise source citations. Analysts can use it directly in the Gemini Enterprise app or wire it into existing agent workflows through Agent-to-Agent (A2A) APIs, and it connects to enterprise data sources over MCP to produce reports and documents in the formats your teams already use. Alongside it, out-of-the-box partner agents cover other workflows and extend the capabilities further.

    4. An open partner ecosystem. Scale with global systems integrators and specialized fintech providers including 66degrees, Accenture, Artefact, Capgemini, Cognizant, Deloitte, Genpact, GFT Technologies, Infosys, KPMG, PwC, Quantiphi, Slalom, Tribe AI, and Zencore to customize and integrate the platform into your own architecture, without vendor lock-in.

    Running underneath: a governed control plane. A single dashboard for IT and risk teams that natively enforces security policies (VPC, CMEK), maintains private data isolation, and holds every output to verifiable grounding with traceable citations.

    Unlocking high-value workflows with domain-specific skills

    Whether used by private equity specialists, wealth managers, or compliance teams, the solution adapts to diverse workflows like credit risk assessment, portfolio monitoring, market news synthesis, and investigative financial research:

    • Elevate advisor insights: Equips relationship managers and advisors with AI-generated insights, personalized recommendations, and tailored artifacts, enabling higher-quality conversations and fostering loyalty. 
    • Deepen Know Your Customer (KYC) research and analysis: Modernizes onboarding and Know Your Customer (KYC) workflows across private banking and prime brokerage by using multi-format ingestion (PDFs, Excel, SEC filings) to map complex corporate hierarchies, evaluate risk personas, and resolve ultimate beneficial owners (UBOs).
    • Enhance portfolio resilience: Helps trading desks deal with sudden macroeconomic shocks. It reduces complex bond portfolio risk exposure analysis to a sub-5-minute execution, complete with automated duration-hedging strategy suggestions.
    • Uncover credit market opportunities: Transforms credit data into actionable trade ideas by identifying and isolating potential mispricings. This enables teams to expand trading volumes while lowering back-office risk and underwriting latency.
    • Accelerate bond issuance: Compresses client pitch presentation timelines from days to minutes so that fixed-income and underwriting teams can proactively target prospects, increase deal capacity, and secure a crucial first-mover advantage to help win more business.

    Open ecosystem of connectors across the financial technology stack

    Gemini Enterprise connects directly to core financial systems via secure MCP connectors. Access is bound by existing role-based controls, ensuring verifiable grounding and precise source citations:

    Productivity and collaboration:

    • Google Workspace: Enables seamless analysis and live artifact generation across Docs, Sheets, and Slides while adhering to enterprise DLP policies.

    • Microsoft 365: Integrates directly with Excel, Word, and PowerPoint to populate financial models, research memos, and client pitch decks.

    Market data and financial fundamentals:

    • Daloopa: Provides the structured, source-linked financial data layer that enables finance professionals and AI tools to produce accurate and auditable results. 

    • FactSet: Enables secure, authorized access to FactSet's multi-asset class financial and non-financial datasets, powering reliable AI-driven workflows with fully auditable, compliant insights. 
    • Finnhub: Provides real-time financial APIs, global fundamentals, and earnings call transcripts for in-depth financial research.

    • Fiscal.ai: Delivers institutional-grade financial data within minutes of earnings, covering financials, news, ownership, segments & KPIs, filings, and earnings call transcripts. 

    • Guidepoint: Connects to primary research insights and expert network transcripts to inform and validate investment theses.

    • LSEG: Provides access to a broad range of trusted financial data, analytical models, indices and news. Enabling customers turnkey access to trusted financial intelligence across every stage of the investment lifecycle.

    • S&P Global: Integrates cited, verifiable S&P Global data for a range of workflows, from financial analysis, to peer benchmarking, industry research, and more.

    Risk, ratings, and private markets:

    • Moody’s: Brings ratings, default risk models, and real time news fused into one lens for counterparty risk assessment. 

    • MSCI: Connects to proprietary indexes, data and models spanning public and private assets and also provides risk analytics and factor exposures. 

    • PitchBook: Provides comprehensive data and research on private equity, venture capital, credit, M&A, and public markets, including, company financials, deal terms, valuations, and fund performance.

    Regulatory and corporate records:

    • SEC Edgar: Delivers instant, verifiable retrieval of statutory filings, 10-Ks, 10-Qs, and 8-Ks with precise citation mapping.

    • Dun & Bradstreet: Accelerates commercial onboarding and KYB verification through direct access to global corporate hierarchy records.

    Digital assets and indices:

    • CoinDesk Data and Indices: Supplies institutional-grade digital asset pricing, benchmark indices, and crypto market intelligence for multi-asset strategies.

    Introducing Gemini Enterprise for Financial Services

    Third-party agents and implementation partners

    Organizations can deploy out-of-the-box partner agents or collaborate with global systems integrators to scale custom capabilities without vendor lock-in:

    • D&B Business Verification agent: Accelerates commercial onboarding and strengthens KYC compliance.

    • FlowX agents: Automate loan pack completeness check, document reconciliation and many other mission critical processes for financial institutions.

    • Obin Financial agent: Helps asset management, commercial lending, and insurance teams accelerate complex financial analyses.

    • S&P Global agents: Data Retrieval Agent for multi-step analysis, report generation, research workflows, and the Horizons Agents that help turn complex energy and sustainability data into fast insights for finance workflows.

    • Global systems integrators and tech partners: Strategic partnerships connect firms with specialist FinTech and leading global systems integrators, including 66degrees, Accenture, Artefact, Capgemini, Cognizant, Deloitte, Genpact, GFT Technologies, Infosys, KPMG, PwC, Quantiphi, Slalom, Tribe AI, and Zencore to manage custom configurations and deploy specialized capabilities at a global scale.

    Developed alongside leading global financial institutions

    We are developing these capabilities in close collaboration with financial institutions, including Deutsche Bank and CME Group, to ensure they reflect the operational realities of the industry.

    “As a design partner for the Financial Research agent, Deutsche Bank has helped shape this capability in view of the realities of a highly regulated industry – from data protection and governance to the workflows our teams use every day,” said Marie-Jeanne Deverdun, Chief Technology, Data and Innovation Officer, and Member of the Deutsche Bank Management Board. “Starting in the Corporate Bank, we see significant potential to reduce manual research effort, improve the consistency and auditability of outputs, and give our teams more time for client conversations. This is an important step in applying AI where it can make a practical difference: safely, responsibly and at scale.”

    This launch builds on the rapidly growing momentum of Gemini Enterprise, with many leading financial institutions like BNY,  Citi Wealth, Lloyds Banking Group, Macquarie Bank, and Signal Iduna using it to equip their workforce with advanced, agentic workflow tools to drive growth and efficiency. 

    Built on an enterprise-grade foundation

    Because Gemini Enterprise for Financial Services runs on Google Cloud infrastructure and the Gemini Enterprise platform, organizations get the security, governance, compliance, and cost-management capabilities they expect from an enterprise platform. 

    Customer data, business rules, intellectual property, custom agents, and model outputs remain private to their organization. Your data is never used to train or fine-tune Google’s foundation models. Furthermore, our full-stack approach - from infrastructure and models to the application layer - allows us to optimize performance and cost, helping organizations maximize the value of their AI investments. 

    This is just the beginning

    Gemini Enterprise for Financial Services is available in preview today, launching alongside our new purpose-built solution for Legal. We invite enterprise leaders in financial institutions to explore how Gemini Enterprise can transform their most critical workflows, with solutions for Healthcare, Life Sciences, and other Professional Services on the horizon.

    ❌