❌

Vue normale

Reçu avant avant-hiercloud

Accelerating the borderless Lakehouse: Announcing preview of cross-cloud caching

18 septembre 2026 à 18:00

Today, we are excited to announce enhancements to the borderless Lakehouse, our answer to how data engineers, data scientists, and increasingly, AI agents, can query governed data directly where it lives.

To reason accurately and automate complex enterprise workflows, agents and data consumers of all types need fast, unified access to an organization's complete data estate, joining customer records, transaction logs, and operational telemetry across clouds. However, modern enterprise data is rarely confined to a single location; data estates often span Amazon S3, Azure Data Lake Storage (ADLS), Google Cloud Storage, operational databases, and SaaS platforms like Salesforce, SAP, and Workday. Historically, uniting these distributed datasets required brittle ETL pipelines, duplicated storage, and prohibitive cross-cloud data transfer costs.

We introduced the borderless Lakehouse earlier this year to let organizations query and activate data in place across clouds. By adopting the Apache Iceberg REST catalog specification, we federate directly to catalogs such as Databricks Unity Catalog, AWS Glue, and Snowflake Horizon. We also introduced Partner Cross-Cloud Interconnect to establish high-bandwidth, private links to other cloud providers, lowering per-gigabyte transfer costs compared to the public internet. 

Today, we are taking multi-cloud efficiency a step further by optimizing how much data needs to be transferred across the wire in the first place.

We are excited to announce two new features to help further reduce costs of querying cross-cloud data.  First, the preview of cross-cloud caching for Lakehouse transparently accelerates cross-cloud queries in BigQuery and cuts remote transfer costs by caching frequently accessed data locally in Google Cloud. Combining standard Iceberg columnar compression with cross-cloud caching means you often only need to transfer under 5% of the data you process across clouds, which helps lower the Total Cost of Ownership (TCO) to make cross-cloud analytics and AI viable at enterprise scale. In addition, BigQuery cross-cloud connections are also available in preview to query non-Iceberg data in other clouds and accelerate workloads.

How cross-cloud caching works

Cross-cloud caching meets enterprise performance and security requirements with no knobs to turn or storage to manage to accelerate your queries. Some of the mechanisms used under the hood are:

  • Sub-file block granularity: Instead of transferring entire multi-gigabyte files across clouds when a query touches only a few columns, cross-cloud caching operates at the sub-file block level for columnar formats like Apache Parquet. BigQuery caches only the specific column chunks and dictionary pages projected by the query. On a cache miss, BigQuery fetches the needed data from the remote cloud to answer the query, and saves a local copy in the cache for future queries, drastically cutting network transfer and latency on repeated workloads.

  • Default encryption at rest: Cached data blocks are encrypted at rest by default using Google-managed encryption keys (GMEK) so that temporary cache storage maintains the same enterprise-grade security posture as native BigQuery storage without extra overhead.

  • Tenant and regional isolation: Cache entries are strictly partitioned by project and catalog boundaries to help prevent cross-tenant data exposure. Lakehouse anchors both the local cache and query execution strictly to the configured Google Cloud region (e.g., us-east4) to support compliance with regional data residency requirements when querying remote clouds.

  • Freshness checks: Multi-cloud caching often forces a trade-off between speed and freshness. To avoid stale reads, BigQuery fetches remote object metadata before using cached data to ensure the data hasn’t changed and the user still has access. Any upstream table modification prompts BigQuery to fetch new files, while unreferenced cached blocks expire automatically, delivering local query speed with single-source-of-truth accuracy.

For more details on caching mechanics, statistics counters, and regional considerations, see the Lakehouse intelligent caching documentation.

Cross-cloud caching in action

So how does this work in day-to-day operations? Consider an e-commerce team querying a 10 TiB Iceberg sales table (aws_lakehouse_catalog.sales.web_sales) in Amazon S3, federated into Lakehouse from Databricks Unity Catalog. During evening promotional drops (8:00–9:00 PM), analysts query historical transactions to identify which storefronts drive peak volume and revenue among high-intent demographics:

code_block
<ListValue: [StructValue([('code', 'SELECT w.web_name, hd.hd_buy_potential, COUNT(*) AS total_transactions, ROUND(SUM(ws.ws_sales_price), 2) AS total_sales\r\nFROM `aws_lakehouse_catalog.sales.web_sales` ws\r\n-- Joins household_demographics, time_dim (8:00-9:00 PM), and web_site.\r\nGROUP BY w.web_name, hd.hd_buy_potential;'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f09257a4190>)])]>

Initial execution: Cold columnar retrieval

On this initial cold run, the local cache is empty (cacheBytesRead: "0"). BigQuery applies partition pruning and column projection to transfer only the required Parquet byte ranges from Amazon S3 over Partner Cross-Cloud Interconnect:

code_block
<ListValue: [StructValue([('code', '{\r\n "totalBytesProcessed": "230343464114",\r\n "objectStorageStats": [\r\n{"cloudProvider": "AWS", \r\n"objectStorageBytesRead": "25834740486", \r\n"cacheBytesRead": "0"}]\r\n}'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f09257a7ad0>)])]>
  • Logical data processed: BigQuery processes 214.5 GiB across the 10 TiB dataset.

  • Standard Iceberg compression efficiency: BigQuery reads 24.1 GiB from S3 thanks to standard Iceberg columnar compression with Zstandard (zstd) — an 8.9:1 compression ratio. As these sub-file Parquet blocks arrive in Google Cloud, BigQuery populates the regional cache.

Follow-on exploration: Adding a dimension

In practice, analysts and agents rarely run the exact same query twice in a row. To drill deeper into fulfillment methods, the analyst modifies the query by adding the shipping method dimension (sm.sm_type):

code_block
<ListValue: [StructValue([('code', 'SELECT w.web_name, sm.sm_type, hd.hd_buy_potential, COUNT(*) AS total_transactions, ROUND(SUM(ws.ws_sales_price), 2) AS total_sales\r\nFROM `aws_lakehouse_catalog.sales.web_sales` ws\r\nJOIN `aws_lakehouse_catalog.sales.ship_mode` sm ON ws.ws_ship_mode_sk = sm.sm_ship_mode_sk\r\n-- Reuses existing joins on household_demographics, time_dim, and web_site.\r\nGROUP BY w.web_name, sm.sm_type, hd.hd_buy_potential;'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f09257a62d0>)])]>

Job statistics for this follow-on query show:

code_block
<ListValue: [StructValue([('code', '{\r\n "totalBytesProcessed": "287928766472",\r\n "objectStorageStats": [\r\n{"cloudProvider": "AWS", \r\n"objectStorageBytesRead": "1426587648", \r\n"cacheBytesRead": "25834740486"}]\r\n}'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f09257a46d0>)])]>
  • 94.8% cache hit rate: BigQuery serves 24.1 GiB of previously queried columns directly from local cache.

  • Granular remote retrieval: BigQuery transfers only 1.33 GiB from S3 for the new ws_ship_mode_sk column and ship_mode table.

  • Sub-file flexibility: Modifying a query reuses cached column chunks and transfers only newly required bytes.

Compounding efficiency at enterprise scale

When thinking about TCO of cross-cloud queries, the top two factors to account for are:

  • Compression ratio: when using default compression algorithms (Zstandard/zstd) on Iceberg, columnar data is highly compressible. If you assume that your data achieves a compression ratio of 8:1, it means every 1 TiB of logical data processed only requires ~128 GiB of data to move over the network.

  • Cache hit rates: when data is retrieved from cache rather than across the network because it was recently accessed, a network transit is avoided. Assuming 80% of your data results in a cache hit it means for every 100 GiB of physical data accessed only 20 GiB moves over the network.

Taking both factors and assumptions into account, for every 1 TiB of data your organization processes, you only need to transfer ~26 GiB across the network (under 3% of total data processed). Combining this reduction with Partner Cross-Cloud Interconnect lowers TCO enough to make cross-cloud analytics and AI cost-effective at petabyte scale.

BigQuery cross-cloud connections now in preview

Alongside cross-cloud caching, the preview of BigQuery cross-cloud connections lets organizations connect BigQuery directly to open-format data in Amazon S3 and Azure Storage. 

Understanding when to use catalog federation versus cross-cloud connections is straightforward:

  • BigQuery cross-cloud connections (for raw files): For standalone files (CSV, JSON, ad-hoc Parquet) without an Iceberg catalog, cross-cloud connections let you create BigQuery external tables referencing remote bucket paths directly.

  • Lakehouse catalog federation (for Iceberg): For Iceberg data managed by catalogs like Databricks Unity, AWS Glue, or Snowflake Horizon, Lakehouse automatically synchronizes schemas and table snapshots to simplify the user experience and ensure users are always querying the latest data.

Cross-cloud connections serve as the modern architectural evolution by using standard BigQuery compute workers in Google Cloud regions rather than compute workers in other clouds. This approach helps unlock global region availability and provides full BigQuery feature parity — including with BigQuery AI and Gemini on remote files.

The cross-cloud caching capabilities for Lakehouse applies to data queried from BigQuery cross-cloud connections as well as Lakehouse catalog federation. To learn how to create connections and query external bucket paths, see the BigQuery cross-cloud connections setup documentation.

Scaling Telco Autonomy: Leveraging GNNs with Distributed GraphFlow

15 septembre 2026 à 18:00

The telecommunications industry is currently undergoing a paradigm shift, moving from traditional manual human-driven operations to fully Autonomous Network Operations. Modern networks have grown increasingly complex, heterogeneous, and large-scale, making handcrafted rules-based methods and traditional Machine Learning (ML) approaches alone insufficient to automate network operations. While ML methods can identify subtle patterns and make fine predictions from large amounts of structured data, they lack the ability to understand, reason about the data and the system it represents, and ultimately make the kind of decision a human operator would.

The growth of AI agents and their ability to reason is a promising solution to this shortcoming. However, in the same way a human operator is not capable of directly ingesting the statistical information spread across the billions of data points created in a large network, AI agents also lack the ability to operate at this scale. To address this challenge, telecommunications companies are adopting Graph Neural Networks (GNNs), a modern form of machine learning designed to operate natively on massive volumes of temporal and relational data. By integrating GNNs with AI agents, operators can combine advanced diagnostics such as root cause analysis, capacity planning, traffic forecasting, what-if simulations, and real-time anomaly detection with the reasoning power required to interpret these insights and execute justified actions. This powerful combination enables networks to safely move towards Level 5 Autonomy as defined by TM Forum, where the system operates autonomously. 

In this post, we present the three components (Data, ML, and AI) that will power Google Cloud’s Autonomous Network Operations framework.

1
2

Google Autonomous Network Operations framework architecture

Foundation: Digital Twin on Spanner Graph

At the heart of Google Cloud’s Autonomous Network Operations framework is the network digital twin: a highly detailed, virtual replica that continuously mirrors its living telecommunications network in real time. Rather than being a static model, it is represented as a dynamic, temporal network graph that captures the evolving state and relations of its components over time. This architectural approach allows operators to "go back" in time to train and evaluate ML models on historical data, while providing AI agents with the foundational operational knowledge required to achieve Level 5 Autonomy. By simulating the impact of proposed network changes within this digital environment, the Digital Twin establishes a critical layer of trust, enabling AI agents to confidently design future states and automatically resolve network issues.

Google Cloud’s Spanner Graph is well suited to host this digital twin:

  • Scalability and Availability: Spanner Graph provides a no compromise foundation for modern applications, offering virtually unlimited scaling that grows as the network grows, along with 0-RPO/0-RTO and five 9s of availability.

  • Multi-Model Support: Supports multiple data models (Relational, Graph, Vector, and Full-Text Search) in a single platform allowing developers to build complex compositions such as graph transversals combined with nearest neighbor vector search.

  • Global Consistency: Spanner provides a globally consistent view of the network, simplifying system development.

The next figure illustrates a network topology with four node types: routers, interfaces (the physical ports), VPNs (L3VPN service instances), and flows (active traffic sessions). These are connected by directed edge types capturing the full network stack: physical containment (router-interface), physical links (interface-interface), control-plane peering (router-router via OSPF/iBGP), service membership (router-VPN), and traffic anchoring (flow-interface, flow-VPN).

3

High Level network topology

The ML layer: Distributed Graph Flow (DGF)

To predict how a network will behave and react, the digital twin leverages an ML layer powered by Distributed Graph Flow (DGF). By training on the vast volumes of structured historical data hosted within Spanner Graph, this layer uncovers critical predictive insights that enable human operators and AI agents to manage networks proactively rather than reactively.

DGF is a recently open-sourced Python library designed to manage the entire end-to-end lifecycle of GNN modeling. Developed by Google CoreML and Google Research, it brings a decade of internal Google-scale tools and expertise directly to Google Cloud enterprise clients. To accommodate different engineering needs, the library offers high-performance, composable, low-level primitives for advanced teams, alongside a simple API for rapid development that requires no prior GNN expertise.

For instance, training and evaluate a GNN model in GraphFlow with the high level API can be as simple as writing 5 lines of code:

code_block
<ListValue: [StructValue([('code', 'import dgf\r\n\r\n# Fetch the data from Spanner Graph\r\ngraph, schema = dgf.io.read_spanner_graph(...)\r\n\r\n# Train a node attribute prediction model\r\nmodel = dgf.learning.train_node_model(graph, schema, target_column="risk_score")\r\n\r\n# Evaluate the model\r\nmodel.evaluate()\r\n# Make predictions\r\nmodel.predict(graph, seed_node_idxs=[0, 1, 2])\r\n\r\n# Save the model for later\r\nmodel.save("/tmp/model")'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fce29db6350>)])]>

The DGF provides high-level concepts that map directly to Autonomous Network Operations requirements:

4

Use cases

By leveraging DGF and GNNs, telcos can move from reactive maintenance to proactive prevention through several advanced use cases:

  • Anomaly detection: GNNs generate node and edge embeddings that encapsulate historical patterns and current health. Any anomalous embeddings are flagged for review before they lead to service degradation.

  • Root cause analysis (RCA): DGF can output specific subgraphs containing only the relevant network instances related to an incident, such as "Attach Failures" in a specific ZIP code. This allows troubleshooting agents to perform high-speed analysis without scanning the entire global network.

  • Predictive maintenance: The system can predict the likelihood of device failures or edge breaks, such as "handover failures" for fast-moving equipment, enabling proactive load balancing or rerouting. Furthermore, by combining agents, remedial actions can be automated by adopting a ‘human-on-the-loop’/’human-in-the-loop’.

  • What-if analysis: GNNs enable Telcos to simulate scenarios like fiber cuts,  or traffic surges or device configuration changes. By modeling topological dependencies, GNNs can predict how these local changes propagate across the entire network, allowing engineers to test resilience and evaluate mitigation strategies in a risk-free digital environment.

Scenario: Root cause analysis with GNNs and DGF

Once you have created a digital twin (example code), a straight-forward 5-step process can be used to implement Root Cause Analysis(RCA) detection using GNNs and DGF. 

  1. Connect to the Digital Twin: Use the DGF Spanner Graph connector (dgf.io.read_spanner_graph) to load the network topology directly from Spanner Graph's Digital Twin into the DGF environment.

  2. Train a Supervised Node (or Edge) Prediction model: Depending on the training data and objective, you will train a supervised node prediction model to predict a target node feature or an edge prediction model to predict an edge between the root cause entity node and the affected entity node. For the given sample data you will use the high-level dgf.learning.train_node_model API to train a supervised node prediction model.

  3. Use the node prediction model to predict root cause node: The node prediction model can be directly used to predict the impact score on the node with the anomaly. Entity nodes affected by the anomaly with highest predicted impact score will be the top candidates for root cause.

  4. Deploy to Gemini Enterprise Agent Platform (formerly Vertex AI): Export the model and host it on a Gemini Enterprise endpoint to enable scalable, low-latency predictions.

  5. Real-time Inference: Make prediction calls to the inference endpoint with the anomaly date as input. The endpoint will return the predicted root cause Entity nodes. 

Get started today

The integration of GNN using Distributed Graph Flow into network operations is more than just a technical upgrade; it is a critical evolution for the telco industry. By moving towards a GNN-powered autonomous framework, operators can significantly shorten outage times, optimize capacity in real-time, and ultimately deliver a superior customer experience through improved operational efficiency. 

To start building your own intelligent network applications, check out the Distributed GraphFlow (DGF) library, which provides the essential primitives for scalable GNN training and inference. For a hands-on experience, follow our step-by-step code sample. You can also explore our recent award-winning Moonshot project on Business-aware GNN-healing networks, and dive deeper into our approach on self-optimizing autonomous networks by reviewing this whitepaper.

Agent-ready analytics: Unlocking insights with BigQuery augmented analytics

14 septembre 2026 à 18:00

BigQuery now features a suite of augmented analytics Table-Valued Functions (TVFs) designed to automate complex data analysis at scale. Augmented analytics combines AI, ML and statistical methods to automate insight discovery and pattern explanation. These functions allow you to diagnose why metrics changed, uncover underlying trends and relationships across the data, and even isolate the true impact of business decisions. 

These TVFs run directly where your data lives, which helps speed up analysis and reduces the need to export data into external tools. In addition, since these functions are compact and yield structured SQL outputs, they can easily be integrated as skills for AI agents, which easily enables automated, conversational data investigation workflows. 

We are introducing six new augmented analytics functions in BigQuery, each created to address a specific analytical challenge:

TVF Function

What It Helps You Find

Real World Question It Answers

AI.KEY_DRIVERS

Identifies the top drivers behind an increase or drop in a metric between two time periods or groups. 

Why did revenue spike this quarter compared to last quarter?

AI.CAUSAL_EFFECT

Quantifies the impact of an action or event by comparing the observed results to an expected baseline.

How much of the revenue lift came from our pricing update rather than organic growth?

ML.CORRELATION

Evaluates the direction and strength of the relationship between pairs of numeric metrics. 

Does increased user session duration correlate with higher lifetime customer value?

ML.DETECT_CHANGE_POINTS

Identifies specific dates or intervals where a metric experiences a shift compared to surrounding patterns.

During which time periods did our platform latency experience persistent, structural shifts?

ML.TREND

Separates the underlying growth or decline from short-term fluctuations or noise. 

What are the underlying trends of my revenue over the past year, abstracting away the outlying spikes and drops?

ML.SEASONALITY

Discovers predicable repeated cycles across hours, days, weeks, months or quarters.  

Which days of the week consistently experience the highest server load?

As we show in the next section, these functions can be easily chained together. The output of one function, such as a detected time window, can directly parameterize the next analytical step.

A step-by-step example of chaining insights

Consider a case where there is a shift in a metric, and you need to diagnose the underlying cause and measure the business lift. 

To diagnose, we can chain ML.DETECT_CHANGE_POINTS, AI.KEY_DRIVERS and AI.CAUSAL_EFFECT using the Austin Bikeshare sample dataset (bigquery-public-data.austin_bikeshare.bikeshare_trips). This dataset contains historical trip volume and demographic data for the city’s bikesharing program. 

Step 1: Detect change points

ML.DETECT_CHANGE_POINTS automatically identifies statistically significant structural shifts or level changes in your time-series data. While this example demonstrates the analysis  in a single aggregate metric, this function is highly scalable and is capable of running across millions of individual time series. 

To find these shifts,  we run the following query across the daily baseline:

code_block
<ListValue: [StructValue([('code', "WITH daily_trips AS (\r\n SELECT\r\n TIMESTAMP_TRUNC(start_time, DAY) AS trip_day,\r\n COUNT(*) AS total_trips\r\n FROM `bigquery-public-data.austin_bikeshare.bikeshare_trips`\r\n GROUP BY 1\r\n)\r\nSELECT\r\n begin_timestamp,\r\n end_timestamp,\r\n metrics.avg AS avg_daily_trips,\r\n metrics.min AS min_daily_trips,\r\n metrics.max AS max_daily_trips,\r\n metrics.count AS duration_days\r\nFROM ML.DETECT_CHANGE_POINTS(\r\n (SELECT * FROM daily_trips),\r\n data_col => 'total_trips',\r\n timestamp_col => 'trip_day'\r\n);"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f313193db90>)])]>

The output identifies the exact time intervals where the baselines have shifted over the company’s history:

1

If we look at the raw daily session counts, this aligns with shifts over time. We highlight the two change points with the longest durations below:

2

The shift in February 2018 aligns with the day the Austin City Council passed the “Dockless Mobility Pilot Program”, to transform the transit ecosystem, integrating shared electric scooters and bikes into the public. 

Step 2: Key drivers attribution

We can input the February 2018 slice found directly to AI.KEY_DRIVERS to determine the particular factors (i.e. bike_type, subscriber_type, etc) driving the surge. AI.KEY_DRIVERS can scan through millions of rows of multi-dimensional data in seconds. 

We define the interest group as the slice of time after the shift occurs and compare it against the time period before the shift as the reference group.

code_block
<ListValue: [StructValue([('code', "WITH daily_segments AS (\r\n SELECT \r\n start_station_name,\r\n end_station_name,\r\n subscriber_type,\r\n bike_type,\r\n 1 AS trip_count,\r\n -- We use the precise breakpoint identified by Change Points\r\n IF(EXTRACT(DATE FROM start_time) >= '2018-02-11', TRUE, FALSE) AS after_shift\r\n FROM `bigquery-public-data.austin_bikeshare.bikeshare_trips`\r\n -- Equidistant ~30 day window around the event\r\n WHERE start_time BETWEEN '2018-01-12' AND '2018-03-13'\r\n)\r\nSELECT \r\n drivers,\r\n metric_interest,\r\n metric_reference,\r\n difference,\r\n relative_difference,\r\n unexpected_difference,\r\n contribution\r\nFROM AI.KEY_DRIVERS(\r\n (SELECT * FROM daily_segments),\r\n metric_col => 'trip_count',\r\n interest_label_col => 'after_shift',\r\n dimension_cols => ['start_station_name', \r\n 'end_station_name', \r\n 'subscriber_type', \r\n 'bike_type'],\r\n top_k => 10\r\n);"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f3130d38f50>)])]>

AI.KEY_DRIVERS isolates the top contributing dimension values.  Each row contains a segment, which represents a slice of data identified by a specific combination of dimension values (e.g., subscriber_type = 'UT Student' and bike_type = 'classic').

3

The analysis reveals that the overall trip count increased +374.7% (+40,159 trips) between the reference and interest time windows. The massive growth was overwhelmingly concentrated in U.T. Student Memberships (+7,167.1%) and trips ending at the 21st & Speedway @PCL station (+20,739.1%).

This aligns with Austin Bikeshare’s response to the Dockless Mobility Pilot Program. In early February, the bikeshare program launched a large promotional partnership with the University of Texas that offered free annual memberships to all UT students.

Step 3: Causal effect

While we know what drove the surge and when it started, we need to isolate the true return on investment over organic expectations. AI.CAUSAL_EFFECT can construct an ARIMA_PLUS counterfactual to measure what the volume would have been had the program never launched.

code_block
<ListValue: [StructValue([('code', "WITH daily_trips AS (\r\n SELECT \r\n TIMESTAMP_TRUNC(start_time, DAY) AS trip_day, \r\n COUNT(*) AS total_trips\r\n FROM `bigquery-public-data.austin_bikeshare.bikeshare_trips`\r\n -- Training on the 6-month baseline leading up to the intervention\r\n WHERE start_time BETWEEN '2017-08-11' AND '2018-04-11'\r\n GROUP BY 1\r\n)\r\nSELECT \r\n *\r\nFROM AI.CAUSAL_EFFECT(\r\n (SELECT * FROM daily_trips),\r\n data_col => 'total_trips',\r\n timestamp_col => 'trip_day',\r\n -- We inject the breakpoint found in Step 1 as our intervention\r\n intervention_timestamp => '2018-02-11 00:00:00',\r\n output_time_series => TRUE\r\n);"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f3130d3bcd0>)])]>

If we graph the predicted and actual trips per day, we can see the surge compared to the counterfactual.

4

If we set the output_time_series => FALSE, we can see a summary of the lift

5

AI.CAUSAL_EFFECT reveals that the program caused a +358% volume surge above organic baseline projections, resulting in an estimated 89,775 incremental trips (with 99.9% probability of causal effect).

Connecting augmented analytics to Conversational Analytics

Conversational Analytics lets you chat with agents about your data using natural language. All new BigQuery augmented analytical functions are now available in Conversational Analytics. Since these TVFs can execute complex analytics at BigQuery-scale in seconds, Conversational Analytics can orchestrate multi-step investigative workflows based on a given prompt. Below we show two examples:

Example 1: Chicago taxi trips

Here is an example using the Chicago Taxi Trips (`bigquery-public-data.chicago_taxi_trips.taxi_trips`).

Prompt: What metric has the strongest correlation with drivers getting tipped? Then run an attribution analysis to tell me which categorical dimensions (like location and payment type) most disproportionately drive that specific metric.

6

The results here used ML.CORRELATION in combination with AI.KEY_DRIVERS.

Credit card payments serve as the primary positive driver of trip distance, adding +1.65M due to longer travel routes and automated digital tip tracking. Trips originating from O'Hare International Airport (Community Area 76) represent another major positive factor, contributing an additional +1.10M miles among tipped credit card rides. In contrast, cash transactions act as a significant negative driver (-652.96K miles), reflecting that cash is predominantly used for shorter journeys rather than extended airport travel.

Example 2: Iowa liquor dataset

Here is an example using the Iowa liquor dataset (`bigquery-public-data.iowa_liquor_sales.sales`) that uses both ML.TREND in combination with ML.SEASONALITY.

Prompt: Find the historical trend for bottles sold. Then, describe the yearly seasonality patterns.

7

The results show that liquor sales in Iowa show persistent long-term growth, rising from 1.3–1.5 million bottles in 2012 before stabilizing around 2.6 million in recent years. There are strong seasonal cycles, particularly during October and December as well as May and June. There is a drop in sales around January and February.  

The skills for these TVFs are now available at the Google Skills Github repository. The BQ AI/ML skills can be found here. 

Take the next step


We would like to extend our sincere thanks to Katelin Amann, Shirley Fu, Chaoyi Shen, Haiyang Qi, Zheng Zhang, Xi Cheng and the wider engineering team for their feedback and contributions of this work.

Simplify pipelines with new BigQuery identity columns

2 septembre 2026 à 18:00

To further empower our customers in their data journey, we are excited to announce the launch of identity columns in BigQuery. This new feature allows users to define columns that automatically generate sequential 64-bit integer values, simplifying the way you manage unique identifiers within your tables.

Data engineers are always looking for ways to make data ingestion smoother and more reliable. BigQuery identity columns offer a powerful, built-in mechanism to automatically generate unique numerical values for your tables. By shifting the responsibility of ID generation to BigQuery, you can significantly reduce the complexity of your data pipelines and focus on delivering insights.

Key benefits for your data pipelines

Implementing identity columns provides several advantages that help streamline the development and maintenance of your data architecture.

  • Streamlined ingestion: You can now ingest data without needing to pre-calculate unique keys in your application logic or ETL tools.

  • Reduced boilerplate: By using auto-generated sequences, your SQL code becomes cleaner and easier to maintain, as the database handles key management natively.

  • Integrated automation: Identity columns work harmoniously with standard DML operations, ensuring that every new row receives a unique identifier automatically.

  • Flexible integration: Whether you are using INSERT or MERGE statements, identity columns adapt to your existing workflow.

How to implement identity columns

Setting up an identity column is simple and can be done directly within your CREATE TABLE statement. You have two primary ways to define how these values are handled.

Definition options

Clause

Description

GENERATED ALWAYS AS IDENTITY

BigQuery automatically manages and ensures the uniqueness of the values.

GENERATED BY DEFAULT AS IDENTITY

Provides an automatic value but still allows for manual overrides when necessary.

Example usage
The following SQL statement demonstrates how to create a table that automatically increments IDs, starting at 1 and increasing by one for each new entry.

code_block
<ListValue: [StructValue([('code', "CREATE TABLE my_project.my_dataset.orders (\r\n order_id INT64 GENERATED ALWAYS AS IDENTITY (START WITH 1 INCREMENT BY 1),\r\n customer_name STRING,\r\n order_date DATE\r\n);\r\n\r\n-- Ingesting data is now simpler:\r\nINSERT INTO my_project.my_dataset.orders (customer_name, order_date)\r\nVALUES ('Joe Doe', CURRENT_DATE());"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fcc84da5c50>)])]>

Get started today

Identity columns represent our ongoing commitment to providing a flexible, high-performance, and standards-compliant data platform. By automating the generation of surrogate keys, we are making it easier for you to build scalable and maintainable data architecture.

To learn more about how to implement this feature in your projects, please visit the BigQuery identity columns documentation.

Introducing TabFM in BigQuery: Predictive analytics reimagined

1 septembre 2026 à 18:00

Historically, enterprise predictive analytics tasks such as predicting churn, purchase intent, or fraud scoring have meant building custom models using libraries like XGBoost, Random Forest, or Deep Neural Networks (DNNs). While effective, the traditional train-tune-deploy-retrain cycle can be complex and time-consuming. Additionally, the overhead of manual feature engineering, hyperparameter tuning, lengthy and expensive training, and the need for specialized data science skills can lead businesses to underutilize predictive models in their decision-making. 

Today, we are announcing the TabFM model in BigQuery. Developed by Google Research, TabFM is a state-of-the-art, pre-trained foundation model for regression and classification on tabular data. It leverages in-context learning (ICL) to deliver highly accurate predictions on your tabular datasets instantly via a single SQL statement, removing the separate training and deployment steps. TabFM on BigQuery is currently in preview. 

Here is what TabFM brings to your BigQuery analytics:

  • Zero-shot predictions: Skip model training, tuning, and artifact deployment. Simply pass your labeled historical data and new prediction tables into a single SQL function to get instant, high-quality predictions.
  • Predictive ML for your agentic applications: Building an agent for your business use? Add predictive powers to it with TabFM plus BigQuery MCP server. No runtimes or infrastructure to manage, just data in and predictions out.
  • State-of-the-art accuracy: Outperforms custom-trained, out-of-the-box traditional models on complex datasets, achieving superior accuracy scores on industry benchmarks.
  • Simple developer experience: Runs natively in BigQuery and is accessible via simple SQL syntax. Automatically handles featurization tasks such as missing values, categorical encoding, etc., with no complex feature engineering pipelines to manage.
  • Scalability: Processes massive inference tables (up to millions of rows) in minutes using BigQuery’s distributed inference architecture.

The leading model for tabular predictions

Google’s TabFM delivers industry-leading accuracy across a wide range of tabular data. In evaluations on the TabArena benchmark, TabFM consistently outperforms both classic machine learning models and other tabular foundation models.

image1

ELO ratings (↑) for the top 10 models across TabArena classification (upper) and regression (lower). (D) = default; (T+E) = tuned + ensemble. Higher scores denote superior performance.

Learn more about the TabFM model here.

Getting started with TabFM in BigQuery

Using TabFM is straightforward. It is exposed directly through new, built-in SQL functions: AI.PREDICT and AI.EVALUATE.

1. Get instant predictions with AI.PREDICT
To make predictions, you write a single query that passes your training  data and prediction data. The model automatically infers whether the task is a classification or regression problem based on the data type of your target label.

code_block
<ListValue: [StructValue([('code', "-- Classifying transactions as fraudulent or not\r\nSELECT *\r\nFROM AI.PREDICT(\r\n TABLE `my_project.my_dataset.historical_transactions`, -- Training data (in-context examples)\r\n TABLE `my_project.my_dataset.new_transactions`, -- Prediction data\r\n label_col => 'is_fraud'-- Target column to predict\r\n);"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fa6f05bb210>)])]>

In this example, the output contains all original columns from your prediction table plus predicted label and probability columns (e.g. predicted_is_fraud). No manual feature engineering or model creation was required.

2. Evaluate models with AI.EVALUATE
You can quickly check prediction performance against a test set using the AI.EVALUATE function. This allows you to generate standard evaluation metrics in a single step.

code_block
<ListValue: [StructValue([('code', "-- Regression Evaluation for Customer Lifetime Value (LTV)\r\nSELECT *\r\nFROM AI.EVALUATE(\r\n TABLE `my_project.my_dataset.historical_customer_ltv`,\r\n TABLE `my_project.my_dataset.test_customer_ltv`,\r\n label_col => 'ltv'\r\n);"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fa6f24557d0>)])]>

AI.EVALUATE returns a robust set of metrics such as r2_score, mean_absolute_error etc. for regression problems and metrics such as precision, recall, and f1 for classification problems.

TabFM in BigQuery under the hood

Traditional machine learning requires fitting model parameters to a training dataset. TabFM, in contrast, uses in-context learning. Similar to how large language models (LLMs) learn a task from few-shot examples in a prompt, TabFM reads your training table as in-context examples and generates predictions for your target table in a single forward pass.

To handle the computational complexity and memory footprint of tabular foundation models, BigQuery performs distributed, parallelized inference on your data. Further, to optimize performance and resource utilization, it uses intelligent training-data sampling as well as distributed execution. This allows BigQuery to handle large input rows for training data while executing predictions quickly and efficiently across millions of rows of inference data.

Choosing the right tool for the job

TabFM introduces groundbreaking zero-shot capabilities to BigQuery, and complements existing offerings such as XGBoost models. Here’s how to choose between TabFM and other models:

  • Use TabFM when you need rapid, high-quality predictive insights without machine learning expertise, when historical datasets are small-to-medium sized, when data changes frequently, and when you need to retrain your models frequently to maintain accuracy. It is also a great fit for conversational or agentic workflows where you need predictive analysis on demand.

  • Use traditional models like XGBoost when you have very large historical datasets, require complete control over custom hyperparameter tuning, have a high number of features that exceed current limits of TabFM, or need feature-importance explainability, i.e., which of the input features contributed most to the prediction.

Predictive machine learning made easy

With TabFM natively integrated into BigQuery, predictive ML is now as easy as running a standard SELECT query. By eliminating the manual overhead of model training, tuning, and management, TabFM lets developers, data scientists and analysts go from raw data to rich predictive insights in seconds.

To get started today, check out the public documentation. For questions or feedback reach out to our team at bqml_feedback@google.com.  We look forward to seeing what you build!

BigQuery Graph is now GA: the knowledge foundation for the agentic era

1 septembre 2026 à 01:00

Many of the questions that matter in enterprise data aren't just about individual rows — they're about how things connect: how two accounts are linked, what path a payment took, what context grounds an AI agent's answer. That’s what a graph is built to solve. Historically, unlocking these insights meant extracting data into standalone graph databases, creating silos and operational overhead. To remove these barriers, we brought native graph capabilities directly to the data warehouse. Today, we are announcing the general availability of BigQuery Graph.

We introduced BigQuery Graph in preview to unify graph and relational analytics. ISO-standard Graph Query Language (GQL) sits alongside SQL, traversals run natively, and there’s no ETL. And because it’s built on BigQuery, BigQuery Graph inherits and expands its capabilities: It reaches petabyte-scale without the memory bottlenecks of a scale-up database, runs under your existing row- and column-level security, and calls BigQuery ML and AI functions in the same query. One engine, two jobs — large-scale graph analytics, and connected context for AI agents.

"BigQuery Graph has been a game-changer for our threat detection pipeline, allowing us to move beyond simple, siloed alerts. By modeling our security signal data as a property graph, we can now perform complex, multi-hop traversals in seconds - something that was previously computationally prohibitive. This graph-centric approach automatically clusters anomalies into coherent attack stories, which, combined with the seamless integration of Gemini models, helps us generate actionable threat narratives. We look forward to integrating native BigQuery Graph algorithms to further streamline our workflows." - Pete Rubio, VP of Global engineering at Thales Cybersecurity Products

Since preview, we saw data teams across industries adopt BigQuery Graph for both analytical and agentic workflows:

  • Threat and fraud detection:  Security and financial organizations correlate signals across event logs to uncover multi-hop attack paths, fraud networks, and suspicious transaction loops.

  • Supply chain digital twins: Manufacturing and logistics organizations map dependencies across suppliers, parts, and distribution routes to simulate disruptions and optimize fulfillment.

  • Identity resolution and Customer 360: Ad-tech and retail platforms stitch fragmented user identifiers and behavioral touchpoints into unified customer profiles across channels.

  • Knowledge graphs and AI agent grounding: Enterprise AI teams build structured knowledge graphs from unstructured documents, providing domain context to ground Gemini models and GraphRAG workflows. 

  • Network lineage and infrastructure management: Telecommunications and enterprise IT teams track complex network topologies, service dependencies, and data lineage across multi-hop paths.

What’s new in BigQuery Graph

Reaching GA is more than a stability milestone. The work fell into two movements: we made the graph engine itself faster and broader, and we built an agentic ecosystem around it — so agents can build a graph, chat with it, and keep an auditable memory on it. Some of what follows is generally available today; some is in preview or rolling out over the coming weeks.

A faster, broader graph engine

“Advertising has spent decades optimizing individual events; the agentic era will optimize the relationships between them. At Yahoo, BigQuery Graph gives our AI agents connected context - campaigns, audiences, exposures, and outcomes, traversable with standard GQL right where our monetization data already lives, with no separate graph engine and no data movement. Our agents don't just read the graph; they reason over it and write their conclusions back as new relationships. That's how monetization moves beyond automation, to autonomous systems we can trust to act.” - Mikul Bhatt, Director of Engineering, Monetization Platform at Yahoo

Borderless graph Lakehouse

Agents are only as good as the context they can reason over, and that context is rarely in one place. With borderless Lakehouse, a single BigQuery Graph can span native BigQuery tables and open Iceberg tables in other clouds — through Databricks Unity Catalog, AWS Glue, or Snowflake — traversed in place, without copying data or building ETL pipelines.

Say a support agent needs to answer, "who supplies the product behind this customer's delayed order, and where are they based?" The customer data sits in an Iceberg lakehouse on Google Cloud, the product and supplier records in a Databricks catalog on AWS. Instead of stitching the sources together per request, the agent traverses one virtual knowledge graph that already connects them — over data that never moved.

1, Virtual Graph

Figure 1: A diagram illustrating a virtual knowledge graph spanning across Google Cloud (blue nodes), AWS (yellow nodes), and other clouds (green nodes) without data movement.

The following DDL statement shows how you can define this virtual graph, mapping your node and edge tables directly across both cloud environments:

code_block
<ListValue: [StructValue([('code', '-- A virtual knowledge graph spanning two clouds - no data movement\r\nCREATE OR REPLACE PROPERTY GRAPH `my_project.retail.virtual_kg`\r\n NODE TABLES (\r\n -- Google Cloud\r\n `my_project.gcs_lake.retail.customers` AS Customer KEY (customer_id),\r\n -- AWS\r\n `my_project.dbx_fed_catalog.retail.products` AS Product KEY (product_id),\r\n `my_project.dbx_fed_catalog.retail.suppliers` AS Supplier KEY (supplier_id)\r\n )\r\n EDGE TABLES (\r\n `my_project.gcs_lake.retail.purchases` AS Bought KEY (purchase_id)\r\n SOURCE KEY (customer_id) REFERENCES Customer (customer_id)\r\n DESTINATION KEY (product_id) REFERENCES Product (product_id),\r\n `my_project.dbx_fed_catalog.retail.products` AS Supplied_By KEY (product_id)\r\n SOURCE KEY (product_id) REFERENCES Product (product_id)\r\n DESTINATION KEY (supplier_id) REFERENCES Supplier (supplier_id)\r\n );'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff407a4d310>)])]>

With that, the agent gets a grounded, multi-hop answer assembled across two clouds in a single traversal:

code_block
<ListValue: [StructValue([('code', "-- Agent grounding: trace a customer to the supplier behind their product, across clouds\r\nGRAPH `my_project.retail.virtual_kg`\r\nMATCH (c:Customer {customer_id: 'C1'})-[:Bought]->\r\n (:Product)-[:Supplied_By]->(s:Supplier)\r\nRETURN s.name AS supplier, s.country AS supplier_country"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff407a4d460>)])]>

Faster and more expressive GQL

BigQuery Graph is built for questions about connection: how two accounts are linked, what path a payment took, which entities sit within a few hops of a flagged one. These are the questions SQL joins struggle to express, and they're where a graph engine earns its place. At GA, we've made them both faster to run and easier to write:

  • Faster execution. GA optimizes path-finding for acyclic and undirected traversals: against public benchmarks, GQL is 2x faster since preview and undirected traversal 100x, with faster, more resource-efficient cycle detection in ACYCLIC and TRAIL path modes. Lower query latency keeps the neighborhood and path lookups that ground an agent's answer responsive under frequent, interactive access.

  • More expressive queries. With the new CALL statement and extended subquery support, you can run a graph subquery for each entity in a result, or invoke a reusable named function, so a complex question breaks into parts instead of one sprawling pattern. The same functions an analyst writes become the building blocks an agent calls as a tool.

Built for the agentic era

“Companies have plenty of workforce data, but very little shared understanding of what their people can do or where they fit. BigQuery Graph lets us turn that scattered information into a reusable property graph and traverse connections across people, roles, capabilities, and evidence at scale, so the same connected workforce context can support thousands of decisions instead of being recreated one decision at a time. That gives AI a stronger foundation for much harder questions about how work should get done.”  - Heiko Roth, Founder & CEO, Workerbee

Chat with your graphs

You don't have to write GQL to explore a graph. BigQuery conversational analytics lets you chat with your graph directly in natural language: it reads the relationships in your schema to translate a question into SQL or GQL, and visualizes the traversal for path-based answers. The agent draws on graph metadata like descriptions and synonyms to keep results grounded — the relationships that make a graph a graph are exactly what cut the ambiguity and hallucination that plague free-form natural language querying. You can also connect Gemini Enterprise to BigQuery Graph through an MCP server, or publish the conversational data agent to it directly.

2, Graph CA Blog V1 2x high res

Build a graph with an agent

Standing up a graph — modeling tables into nodes and edges, then writing GQL against them — is work you can hand to the data agent you already use. We've packaged BigQuery Graph expertise into an agent skill that makes your agent fluent in graph: GQL pattern matching, blending graph and SQL, and schema design that follows our recommended practices. The capabilities are accessible out of the box in your preferred agentic coding tool, such as Antigravity, Visual Studio Code, Claude Code, and Codex, with the Google Cloud Data Agent Kit extension.

The skill is also learning to author, not just advise — a capability rolling out soon. Point it at a dataset, a model document, or an ER diagram and it proposes the nodes and edges, then verifies each relationship against your data before building, showing you the match rates: this one resolves at, say, 98%, that one 56%. You get a graph you can trust from day one.

3, Graph GA Skill Demo V2

Give your agents an auditable memory

Grounding an agent is half the job; the other half is remembering what it did. As agents move from advising to acting, every decision has to be explainable after the fact — which option was chosen, which policy applied, which alternatives were rejected. With context graph in BigQuery Agent Analytics, each action an agent takes is captured and shaped into a context graph: a typed, queryable trace of the agent's reasoning, stored right in BigQuery Graph. Because the trace is itself a graph, "why did the agent do this?" is a single traversal — and the outcomes you join back to those decisions become the data that improves the next one.

Get started with BigQuery Graph today

BigQuery Graph runs graph analytics and grounds AI agents on your data, across clouds. To get started, check out the overview and data model to see how GQL, node tables, and edge tables fit together, then put them to work on your team’s common patterns. Trace suspicious money movement and synthetic identities in the fraud detection codelab, stitch fragmented emails, devices, and cookies into one customer in the identity resolution codelab, or model a supply chain as a digital twin you can query for hidden dependencies when disruption hits.

From there, take it toward agents. The agent context graph codelab turns raw event logs into a graph that audits, explains, and traces what your autonomous agents actually did — the connected memory behind a system you can trust to act. If your workloads span both real-time operational transactions and massive-scale analytics, explore our unified graph solution to see how Spanner Graph and BigQuery Graph work together. And when you are ready to go deeper — our ebook walks the journey end-to-end.

Zero-code, low-cost data ingestion: New BigQuery DTS capabilities

7 août 2026 à 19:00

In a fast-paced digital economy, data is your most critical engine. Yet, many enterprises find themselves trapped in a costly paradox, spending over 100 hours a week building and fixing fragile, in-house ETL pipelines or wrestling with unpredictable third-party tools.

Trusted by thousands of customers every single day, BigQuery Data Transfer Service (DTS) eliminates this engineering burden. As a fully managed, zero-code data movement solution, BigQuery DTS automates data ingestion into BigQuery allowing your teams to transition from pipeline maintenance to strategic data science in minutes.

Expanding the ecosystem: New connectors and capabilities

We are rapidly expanding our integration landscape to eliminate data silos across databases, ads and marketing platforms. Here are the latest additions and enhancements

Open Lakehouse ingestion

  • Direct ingestion into Apache Iceberg managed tables (Preview): You can now ingest data from common sources such as Google Cloud Storage, Amazon S3, and Azure Blob Storage directly into Iceberg managed tables. This enables you to maintain full multi-cloud storage cross-compatibility with other query engines while leveraging BigQuery's top-tier performance tuning.

Next-gen agentic architecture

Enterprise and relational databases (supports both full or incremental transfers)

  • Microsoft SQL Server (Preview): Easily centralize transactional tables, schemas, and operational data directly into your analytical environment in BigQuery. 

  • PostgreSQL(GA) and MySQL (GA): Automates data delivery and simplifies the replication of high-volume web and application workloads into your central data warehouse within minutes. Supports data replication from on-premise environments, CloudSQL, and other clouds.

E-commerce and growth marketing

  • Shopify(Preview): Automates the extraction of granular order histories, inventory logs, and customer profiles straight into your analytical schema.

  • Klaviyo(Preview): Extracts detailed email and SMS engagement logs (such as clicks, sends, and opens) to build precise multi-channel lifecycle attributes.

  • HubSpot(Preview) : Syncs pipeline, contact tracking, and inbound marketing metrics to keep your revenue operations teams aligned.

  • Mailchimp(Preview): Automatically moves campaign performances and audience list attributes directly into your data warehouse.

Migration connectors

  • Snowflake(GA): Migrate your data from Snowflake with features like incremental transfer, auto schema detection, private connectivity and support for migrating data residing on all three major clouds (Google Cloud, AWS, and Azure)

 Enhancement to major connectors

  • ServiceNow, Salesforce, and Oracle: Enhanced with native incremental update support to speed up large-scale pipeline refreshes for enterprise CRM, ITSM, and financial workflows.

Why you should choose BigQuery Data Transfer Service

Moving data across an enterprise architecture shouldn't require complex compromises between cost, management overhead, and pipeline health. BigQuery DTS delivers unique advantages across three pillars:

1. Unbeatable cost efficiency

  • Zero Ingestion Costs for major sources: Data ingestion at no-charge for all first-party Google sources (except Google Play), including Google Ads, Google Analytics 4, Campaign Manager, YouTube and Google Cloud Storage. Ingestion cost are also free for Amazon S3, Azure Blob Storage, Amazon Redshift, and Teradata

  • Low consumption-based rates for third-party SaaS: Ingestion from 3rd-party sources are entirely on a flexible consumption model. Compute fees run less than 6 cents per slot-hour in major regions. Because pricing scales with compute footprint rather than row volume, depending on your source data compression format and available network bandwidth, you can efficiently transfer massive data volumes. 

2. Frictionless security and native management

Eliminate middlemen servers and external security and API configurations. BigQuery DTS fully integrates with Cloud IAM. Data transfers instantly inherit your destination dataset’s Column-Level Security, Row-Level Security, and Customer-Managed Encryption Keys (CMEK) without extra configuration overhead. Your streams flow securely inside the native Google Cloud perimeter.

3. Industry-leading performance and resilience

When you manage data at enterprise scale, downtime means lost business. BigQuery DTS provides a highly resilient ingestion footprint backed by a strict Google Cloud Service Level Agreement (SLA). The system delivers a monthly uptime percentage of >=99.99%, guaranteeing your analytics, automated pipelines, and operational dashboards update reliably.

Ready to transform your data operations?

Stop letting manual ingestion scripts limit your growth. Join the thousands of companies relying on Google's native cloud lakehouse data movement architecture to build a modern, scalable data stack.

Try the platform out today: Navigate directly to the BigQuery Data Transfer Service console, pick your connector, and deploy your first automated transfer in just a few clicks!

What connectors should we build next?

We are constantly expanding our native integration library based on your business needs. What sources are you currently forced to extract manually? Are there specific relational databases, NoSQL engines, or regional SaaS platforms you need to replicate next? Let us know with a feature request.

Please file feature requests via the public issue tracker.

Unifying Structured and Unstructured Data Insights with BQ Search Innovations

7 août 2026 à 18:00

Modern enterprises possess a vast amount of unstructured data, yet they frequently encounter significant challenges in managing and extracting value from it. Historically, unlocking the insights hidden within PDFs, audio files, images, and unstructured text required a fragmented architecture: moving data out of your warehouse, stitching together complex LLM pipelines, and managing disparate search indexes.

BigQuery has worked with many enterprises to make sense of their unstructured data sources. For example, consider an advanced healthcare company managing thousands of clinical trial documents in PDF form. BigQuery helps unlock insights from these documents through a simple, five-step lifecycle: Access, Process, Ground, Relate, and Activate.

In this post, we are highlighting three major milestones focused heavily on the "Ground" phase of this framework:

  1. General Availability (GA) of Autonomous Embedding Generation

  2. General Availability (GA) of AI.SEARCH with massive single-query performance gains

  3. Public Preview of Hybrid Search

Let’s dive into how these features work together to simplify your AI architecture, using a real-world clinical trial research platform as an example.

Simplify Pipelines with Autonomous Embedding Generation (GA)

Building a retrieval-augmented generation (RAG) pipeline or search application usually requires managing complex, asynchronous embedding infrastructure. You have to handle retries, error logging, and pipeline orchestration every time a new record arrives.

With the General Availability of Autonomous Embedding Generation, BigQuery manages this entirely for you. By simply defining a column in your schema, BigQuery asynchronously and continuously generates embeddings as new data is ingested. You have the flexibility to choose external models (like Vertex AI text-embeddings) or natively utilize Gemma embedding models directly within BigQuery.

How it works in practice:

Imagine you are building a research platform analyzing clinical trial PDFs stored in Google Cloud Storage. After extracting the study titles and disease areas into a table, you can automatically embed those titles:

code_block
<ListValue: [StructValue([('code', "CREATE OR REPLACE TABLE mydataset.clinical_trials\r\n(\r\n trial_id STRING,\r\n study_title STRING,\r\n study_embedding STRUCT<result ARRAY<FLOAT64>, status STRING>\r\n GENERATED ALWAYS AS (AI.EMBED(\r\n study_title,\r\n connection_id => 'myconnection',\r\n endpoint => 'text-embedding-005'\r\n ))\r\n STORED\r\n OPTIONS( asynchronous = TRUE )\r\n);"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f6e8690a990>)])]>

BigQuery eliminates the need for complex third-party vector databases by managing enterprise-scale processing with one configuration. This autonomous embedding generation keeps data synchronized automatically as source text changes, removing the need for manual machine learning pipelines. This integrated approach streamlines workflows for dynamic datasets and reduces the operational burden of maintaining custom data scripts.

Finally, with this GA launch, Autonomous Embedding Generation now also supports generating embeddings natively over images using ObjectRefs, unlocking true multimodal search and analytics.

Natural Language Search at Scale with AI.SEARCH (GA)

To truly enable conversational analytics agents and snappier user experiences, your underlying search infrastructure needs to be intuitive and performant.

Once your data is seamlessly embedded, you need an efficient way to query it. Today, we are announcing the General Availability of AI.SEARCH(). This function provides a streamlined, natural-language-focused search experience, allowing you to easily find semantically related records without generating embeddings in your search path. In pairing this with the Autonomous Embedding Generation, we leverage the same embedding model used in your dataset for easier use. 

Furthermore, as part of efficiency investments in the last year, we have heavily optimized AI.SEARCH for single-query execution. For online applications and single-query searches (those most common in agentic searches), we have observed up to a 133x gain in slot efficiency..

This means you can serve highly concurrent, user-facing natural language searches directly out of BigQuery faster and more cost-effectively than ever before.

code_block
<ListValue: [StructValue([('code', 'SELECT base.trial_id, base.study_title, distance\r\nFROM\r\n AI.SEARCH(\r\n TABLE mydataset.clinical_trials,\r\n \'study_title\',\r\n "What treatments are available for advanced tumors?"\r\n );'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f6e86146a90>)])]>

Unifying Keyword and Vector with Hybrid Search (Public Preview)

Semantic (vector) search is incredibly powerful; for example, the query above will successfully return conceptually related terms like "chemotherapy." However, semantic search isn't always enough. What if a researcher is searching for a specific, highly technical immunotherapy drug designation like "MK3475"? Because this alphanumeric string lacks broad semantic meaning, pure vector search might struggle to rank it correctly.

By merging lexical search with existing semantic capabilities, BigQuery's hybrid search allows for data retrieval based on both keyword similarity and underlying meaning. This approach unites the conceptual depth of semantic vector search with the pinpoint accuracy of lexical matching, utilizing algorithms such as Reciprocal Rank Fusion and BM25. The result is a significant boost in search precision and a reduction in LLM hallucination costs through the reranking of results based on keyword frequency and semantic relevance. Users can implement this via the AI.SEARCH and VECTOR_SEARCH functions by employing the hybrid mode or lexical_search_columns parameters. Furthermore, performance can be optimized by extending vector indexes to include keyword data, which accelerates the lexical search process.

You can now perform hybrid searches effortlessly using the AI.SEARCH() function by simply setting the mode to HYBRID:

code_block
<ListValue: [StructValue([('code', 'SELECT base.trial_id, base.study_title, distance\r\nFROM\r\n AI.SEARCH(\r\n TABLE mydataset.clinical_trials,\r\n \'study_title\',\r\n "Cancer treated by MK-3475",\r\n mode => \'HYBRID\'\r\n );'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f6e86146010>)])]>

To speed up these hybrid queries at scale, you can easily extend your CREATE VECTOR INDEX DDL to include the keyword columns you want to use for the lexical portion of the search, natively combining your indexes.

Building an End-to-End Unstructured Data Analytics Platform

The search and embedding features launching today are part of a much broader vision. We are building an end-to-end unstructured data analytics platform. One common type of unstructured data is documents, and BigQuery now provides the complete toolset to manage this workflow from end to end:

image1
  • Access: Enable zero-ETL workflows by querying unstructured PDFs and documents directly where they live in Google Cloud Storage using Object Tables.

  • Process: Utilize embedded AI capabilities like AI.PARSE_DOCUMENT and AI.CHUNK_DOCUMENT (coming soon) for layout-aware chunking (perfect for RAG), AI.GENERATE for entity extraction/summarization, and AI.CLASSIFY to instantly categorize records using foundational models directly in your SQL pipelines.

  • Ground: Build highly accurate context using Autonomous Embeddings and Hybrid Search. Hybrid search combines the conceptual understanding of semantic vector search with the exact precision of lexical keyword matching By reranking results based on both semantic relevance and keyword frequency, you drastically increase search precision and drive down LLM hallucinations.

  • Relate: Uncover hidden multi-hop insights by mapping extracted entities (like Sponsors, Trials, and Drugs) into a BigQuery Graph—no specialized graph database required.

  • Activate: Bring it all together with BigQuery's Conversational Analytics agents. Using functions like AI.AGG, you can chat directly with your complex data, generate visualizations, and perform trend analysis at massive scale.

With unstructured data as a first-class citizen in BigQuery, you can finally bridge the gap between your raw documents and conversational AI.

Ready to get started?

Level Up Your Column-level Security: Using IAM Data Governance Tags in BigQuery

17 juillet 2026 à 18:00

Many BigQuery customers rely on policy tags for protecting their sensitive information in BigQuery. Policy tags were the go-to solution for applying column-level access controls, allowing only users with the right permission to view sensitive columns like personally identifiable information (PII). It was a robust and effective system — for its time.

However, data ecosystems have grown in complexity, and the tools we use to help secure them need to evolve with them. New challenges include creating and managing a taxonomy that supports multiple tags across multiple regions and locations, enabling disaster recovery, and integrating with a broad centralized governance strategy.

To help you meet the needs of today’s data ecosystems, we're excited to introduce the preview of data governance tags in BigQuery. Built on Google Cloud's Identity and Access Manager’s (IAM) Resource Manager infrastructure, data governance tags provide a scalable, and robust method to help you manage access controls and protect your BigQuery column data. 

What are IAM data governance tags?

Data governance tags are a special type of Resource Manager tags.  You can create it by setting the purpose field to DATA_GOVERNANCE when creating a tag key in IAM, you designate it for use in BigQuery column-level security. You can create a hierarchical tree of data governance tags specifically for column-data governance purposes and apply them directly to your BigQuery columns. 

Why use data governance tags for column-level security?

  • Global scope, regional enforcement: Unlike policy tags (which are regional-only), data governance tags are global. You can define a single tag key:value pair (like “data_sensitivity:high”) at the organization level and use it across any project or region in your organization.

  • Managed disaster recovery: Security policies should persist during a failover. Data governance tags and their associated data policies are automatically replicated to secondary regions. If you need to switch regions, your security posture moves with you automatically.

  • Hierarchical security: You can now build a tree of tags up to five levels deep. This allows for inheritance and more granular classification (such as PII > Financial > CreditCardNumber).

  • Decoupled governance: You can tag your data to organize and classify it before you decide to enforce security. Access control only kicks in once you define a data policy for that tag, giving your team more flexibility during data onboarding.

Three steps to column-level security

Step 1: Create the tag key and values

1. Create data governance tag key: First you create an IAM tag key in Console, gcloud CLI, or API. The magic happens when you specify the purpose field as --purpose=DATA_GOVERNANCE for the tag key. This key change tells Google Cloud that this tag will be used for column-level security in BigQuery.

code_block
<ListValue: [StructValue([('code', '# Example: Creating a Data Governance tag key named "data_class"\r\ngcloud resource-manager tags keys create data_class \\\r\n --parent=projects/my-governance-project \\\r\n --purpose=DATA_GOVERNANCE'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fae13044c70>)])]>

2. Create tag values:

Once your data governance tag key has been created, you need to create specific tag values under the key that you will use to categorize/classify your column data.  One of the useful features of data governance tags is the ability to build a hierarchical tree of tag values. The tag-values tree allows you to create broad categories and then drill down into specific categories based on data type. You can go up to five levels deep for granular access control.

code_block
<ListValue: [StructValue([('code', '# Level 1: Create a tag value called "pii"\r\ngcloud resource-manager tags values create pii \\\r\n --parent=my-governance-project/data_class\r\n\r\n\r\n# Level 2: Create a child value under "pii" for "private" data\r\ngcloud resource-manager tags values create private \\\r\n --parent=my-governance-project/data_class/pii\r\n\r\n\r\n# Level 3: Create another child tag value for "email" under "private"\r\n# You can go up to 5 levels deep for granular control\r\ngcloud resource-manager tags values create email \\\r\n --parent=my-governance-project/data_class/private'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fae13044b50>)])]>

Step 2: Attach tags to your columns via JSON schema

1. Export your existing schema

For existing tables, the most efficient way to manage tags is by updating the table schema using a JSON file and using API or BQ CLI because it allows you to tag multiple columns at once.

code_block
<ListValue: [StructValue([('code', '# Save the current table schema to a local JSON file.\r\nbq show --schema --format=prettyjson my_project:my_dataset.my_table > schema.json'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fae130448b0>)])]>

2. Add the dataGovernanceTags to your JSON file

Open schema.json and add the tag mapping to your sensitive columns. Note the use of the namespaced key and the short name for the value.

code_block
<ListValue: [StructValue([('code', '[\r\n {\r\n "name": "user_email",\r\n "type": "STRING",\r\n "dataGovernanceTagsInfo": {\r\n "dataGovernanceTags": {\r\n "my-governance-project/data_class": "email" \r\n }\r\n }\r\n },\r\n {\r\n "name": "phone_number",\r\n "type": "STRING",\r\n "dataGovernanceTagsInfo": {\r\n "dataGovernanceTags": {\r\n "my-governance-project/data_class": "private"\r\n }\r\n }\r\n },\r\n {\r\n "name": "government_id",\r\n "type": "STRING",\r\n "dataGovernanceTagsInfo": {\r\n "dataGovernanceTags": {\r\n "my-governance-project/data_class": "pii"\r\n }\r\n }\r\n }\r\n]'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fae13044970>)])]>

3. Update the table:

Apply the schema to your BigQuery table.

code_block
<ListValue: [StructValue([('code', '# Overwrite the table schema with your newly tagged JSON file.\r\nbq update --project_id=my-data-project --schema=schema.json my_dataset.my_table'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fae130440a0>)])]>

Alternatively you can also use SQL to bind data governance tags to BigQuery table columns.

code_block
<ListValue: [StructValue([('code', "CREATE OR REPLACE TABLE my_dataset.my_table(\r\n user_email STRING\r\n OPTIONS (\r\n data_governance_tags = [('my-governance-project/data_class', 'email')]),\r\n );\r\nALTER TABLE my_dataset.my_table\r\nALTER COLUMN phone_number\r\n SET OPTIONS (\r\n data_governance_tags = [('my-governance-project/data_class', 'private')]);\r\nALTER TABLE my_dataset.my_table\r\nADD COLUMN government_id\r\n STRING\r\n OPTIONS (\r\n data_governance_tags = [('my-governance-project/data_class', 'pii')]);"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fae13044a60>)])]>

You can also remove a column tag by setting it to [], for example:

code_block
<ListValue: [StructValue([('code', 'ALTER TABLE my_dataset.my_table\r\nALTER COLUMN phone_number\r\n SET OPTIONS (\r\n data_governance_tags = []\r\n);'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fae1305eb20>)])]>

You can use information_schema COLUMNS view to see the columns tags:

code_block
<ListValue: [StructValue([('code', "SELECT\r\n column_name,\r\n data_governance_tags[SAFE_OFFSET(0)].key AS tag_key,\r\n data_governance_tags[SAFE_OFFSET(0)].value AS tag_value,\r\nFROM `my_project.my_dataset.INFORMATION_SCHEMA.COLUMNS`\r\nWHERE table_name = 'my_table'"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fae1305ef10>)])]>

The result is similar to the following:

code_block
<ListValue: [StructValue([('code', '+---------------+----------------------------------+-----------+\r\n| column_name | tag_key | tag_value |\r\n+---------------+----------------------------------+-----------+\r\n| user_email | my-governance-project/data_class | email |\r\n| phone_number | my-governance-project/data_class | private |\r\n| government_id | NULL | NULL |\r\n+---------------+----------------------------------+-----------+'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fae1305e400>)])]>

Step 3: Create data policies

Finally, define a BigQuery data policy to govern access to these tagged columns. These policies explicitly reference the tag values you attached previously. Note that, while data governance tags are global, data policies are regional. 

To protect your data, the policy must be created in the same region where your BigQuery table is located. Once the policy is defined, access is only granted to the specified grantees; all others will be denied access to the sensitive column data.

Also, keep in mind that security in BigQuery is layered. For a data policy to be effective, the users (grantees) must first possess base-level access to the table itself (typically via a role like roles/bigquery.dataViewer). Data policy then acts as a second security layer, determining whether they view the raw, sensitive column data or a masked, obfuscated version.

Masking policy for ‘pii’ tagged column-data (SHA256 Masking):

code_block
<ListValue: [StructValue([('code', 'curl --request POST "https://bigquerydatapolicy.googleapis.com/v2/projects/myProject/locations/us-east1/dataPolicies" \\\r\n --header "Authorization: Bearer $(gcloud auth print-access-token)" \\\r\n --header \'Accept: application/json\' \\\r\n --header \'Content-Type: application/json\' \\\r\n --data \'{\r\n "dataPolicy": {\r\n "dataPolicyType": "DATA_MASKING_POLICY",\r\n "dataMaskingPolicy": { "predefinedExpression": "SHA256" },\r\n "grantees": [ "principalSet://goog/group/grp-sales@corp.com" ],\r\n "dataGovernanceTag": { "key": "myProject/data_class", "value": "pii" }\r\n },\r\n "dataPolicyId": "masking_policy_for_data_class_pii"\r\n}\' \\\r\n --compressed'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fae1305e910>)])]>

 Raw access policy for ‘pii’ tagged column-data

code_block
<ListValue: [StructValue([('code', 'curl --request POST "https://bigquerydatapolicy.googleapis.com/v2/projects/myProject/locations/us-east1/dataPolicies" \\\r\n --header "Authorization: Bearer $(gcloud auth print-access-token)" \\\r\n --header \'Accept: application/json\' \\\r\n --header \'Content-Type: application/json\' \\\r\n --data \'{\r\n "dataPolicy": {\r\n "dataPolicyType": "RAW_DATA_ACCESS_POLICY",\r\n "grantees": [ "principal://goog/subject/abc@xyz.com" ],\r\n "dataGovernanceTag": { "key": "myProject/data_class", "value": "pii" }\r\n },\r\n "dataPolicyId": "raw_access_policy_data_class_pii"\r\n}\' \\\r\n --compressed'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fae1305e6d0>)])]>

Masking policy for “private” tagged column data (NULL Masking):

code_block
<ListValue: [StructValue([('code', 'curl --request POST "https://bigquerydatapolicy.googleapis.com/v2/projects/myProject/locations/us-east1/dataPolicies" \\\r\n --header "Authorization: Bearer $(gcloud auth print-access-token)" \\\r\n --header \'Accept: application/json\' \\\r\n --header \'Content-Type: application/json\' \\\r\n --data \'{\r\n "dataPolicy": {\r\n "dataPolicyType": "DATA_MASKING_POLICY",\r\n "dataMaskingPolicy": { "predefinedExpression": "ALWAYS_NULL" },\r\n "grantees": [ "principal://goog/subject/abc@xyz.com" ],\r\n "dataGovernanceTag": { "key": "myProject/data_class", "value": "private" }\r\n },\r\n "dataPolicyId": "null_policy_data_class_private"\r\n}\' \\\r\n --compressed'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fae1305ee80>)])]>

With these three steps, your column data is now protected. The next time a principal queries your BigQuery table, our authorization engine automatically evaluates their identity against your data policies. If the principal is part of the policy, they get to see the masked or raw data as per the policy; if they are not, then they will be denied access. 

Get started today

Data governance tags are a powerful new tool to enhance your data security and governance strategy in BigQuery. We are continuously working to enhance data governance capabilities in BigQuery. Future updates include support for using SQL to create tags and tag based policies, ability to attach multiple tags to a single column,  ability to define policies based on combinations of tags, and deeper integrations with services like Knowledge Catalog.

You can start tagging your columns and defining fine-grained access controls at scale. To learn more, dive into the Data Governance Tags documentation.

How to Analyze and Govern Gemini Enterprise App Usage at Scale with BigQuery

15 juillet 2026 à 18:00

Deploying the Gemini Enterprise app across an organization marks a transformative leap forward in workforce productivity, providing employees with an amazing, high-performance suite of agentic AI tools, search-grounded assistants, and specialized solutions like NotebookLM. As adoption grows to a large scale, it can introduce a critical administrative scale challenge: how to audit, govern, and extract insights from a massive volume of telemetry without getting bogged down in manual overhead. To help administrators succeed, Google Cloud provides comprehensive, out-of-the-box analytics via pre-computed dashboards to track day-to-day adoption, user engagement, and active user metrics. While this provides a product-centric lens to look at Gemini Enterprise app's usage, to understand the impact of agentic AI, administrators might need a more nuanced, organization-centric perspective tailored to their own internal context. This is where using Google BigQuery becomes a crucial tool in the administrator's arsenal to run deep-dive forensics across their organization to analyze and govern the adoption of agentic AI.

Why Gemini Enterprise app + BigQuery is a game-changer

Augmenting the Gemini Enterprise app with BigQuery through log sinks allows a lean administrative team to analyze and govern a large-scale deployment. Specifically, it empowers IT, Data, and Security teams to:

  • Profile nuanced adoption and behaviors: Segment usage patterns by department to see which teams are building custom agents, track NotebookLM utilization, and calculate agent-to-employee ratios.

  • Quantify organizational value: Combine conversational logs with HR or line-of-business datasets to calculate actual employee hours saved, trace value creation, and build executive Looker dashboards.

  • Execute precision compliance audits: Audit grounding queries across Google Drive folders and enterprise directories to prevent data leaks and protect corporate IP.

  • Investigate safety alerts instantly: Query historical logs when security filters flag a prompt, identifying the exact text that triggered a Model Armor block to resolve compliance alerts.

To support these use cases, the telemetry is partitioned into five distinct log tables in BigQuery, capturing unique data fields:

BigQuery Destination Table

Telemetry Captured

Gen AI User Messages

`discoveryengine_googleapis_com_g
en_ai_user_message`

Verbatim prompt inputs typed by users

Gen AI Choices

`discoveryengine_googleapis_com_g
en_ai_choice`

Verbatim model responses, finish reasons, and LLM reasoning steps

User Activity Telemetry

`discoveryengine_googleapis_com_g
emini_enterprise_user_activity`

Corporate identity (IAM emails) and grounding file access paths

Cloud Audit Activity

`cloudaudit_googleapis_com_activity`

Control plane configuration changes and administrative user logs

Cloud Audit Data Access

`cloudaudit_googleapis_com_data_ac
cess`

High-volume data plane interactions and search queries

Aggregate OOB Metrics

(Batch Export Table)

Pre-aggregated seats claimed, seat purchases, and engagement metrics from the past 30 days. To be pulled asynchronously via custom daily batch runs of the analytics:exportMetrics API to build high-level adoption and cost dashboards.

Ingestion pipeline and architecture

To implement scale-ready observability, administrators establish an automated telemetry pipeline. Moving your Gemini Enterprise data to BigQuery does not require complex custom software development; instead, it leverages a continuous Cloud Logging Log Router Sink for conversational logs and an asynchronous batch export API for high-level aggregate seat metrics.

The diagram below illustrates the ingestion pipeline and how telemetry is mapped to BigQuery:

1

Here is your blueprint for connecting Gemini Enterprise to BigQuery to build the ultimate analytics and governance foundation for your organization.

Routing pipelines: Continuous logging and audit sinks

To capture your telemetry, establish log sinks within Cloud Logging to intercept and route runtime events to BigQuery:

  • The streaming pipeline (detailed logs): Streams row-by-row conversational data (user prompts, model choices, and grounding events). Ensure prompt and response logging is enabled in your Gemini Enterprise Admin Console (see Set Up Usage & Audit Logs).

    • Inclusion Filter (replace [PROJECT_ID] with your Google Cloud Project ID):

code_block
<ListValue: [StructValue([('code', 'logName="projects/[PROJECT_ID]/logs/discoveryengine.googleapis.com%2Fgemini_enterprise_user_activity" OR\r\nlogName="projects/[PROJECT_ID]/logs/discoveryengine.googleapis.com%2Fgen_ai.user.message" OR\r\nlogName="projects/[PROJECT_ID]/logs/discoveryengine.googleapis.com%2Fgen_ai.choice"'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f4248bc1430>)])]>
  • The governance pipeline (audit logs): Captures administrative actions (Admin Activity) and data plane operations (Data Access, such as grounding data connector lookups).

    • Inclusion Filter (replace [PROJECT_ID] with your Google Cloud Project ID):

code_block
<ListValue: [StructValue([('code', 'logName:"projects/[PROJECT_ID]/logs/cloudaudit.googleapis.com" AND \r\nprotoPayload.serviceName="discoveryengine.googleapis.com"'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f4248bc10a0>)])]>
    • Admin Activity Logs: Always enabled by default; tracks resource changes (e.g., custom agent creation, updates, deletions).

    • Data Access Logs: Off by default; must be enabled in GCP IAM settings for the Discovery Engine API to log user-level data read/write interactions during chats.

Unlock advanced intelligence in BigQuery

Transform raw telemetry into insights

BigQuery provides AI-powered analysis tools that make understanding and navigating telemetry effortless. By leveraging Gemini in BigQuery, administrators can translate raw log streams into visual insights and clear documentation without manual guesswork.

No-Code Conversational Analytics (BigQuery CA)
Querying nested JSON schemas is made simple with Conversational Analytics in BigQuery (BQ CA). BQ CA acts as an intelligent agent within BigQuery Studio, automatically generating and executing SQL grounded in your schema, business metadata, and verified queries/UDFs to ensure metrics consistency. It also surfaces its "thinking process" alongside the generated code to build administrative trust. 

For example, as shown in the screenshot below, asking "Compare the usage of notebooklm, deep research and custom agents using oob_metrics?" generates the correct SQL, runs the query and outputs the result in seconds:

2

As shown in the screenshot below, BQ CA goes beyond traditional querying and standard SQL generation by allowing users to execute sophisticated AI and machine learning tasks directly within the console. Administrators can leverage these native capabilities to run advanced analysis, such as classification of user prompt sentiment or forecasting future adoption trends, streamlining the governance process.

3

Auto-generated schema documentation and insights

Understanding telemetry fields like useriamprincipal, finish_reason, or groundedContent is crucial for extracting the right insights. BigQuery simplifies this through automated schema documentation and AI-powered context:

  • Automated profiling and metadata: By pairing Knowledge Catalog Data Profiling with Gemini, you can evaluate unique value counts, null rates, and data distributions in raw tables. With a single click, Data Insights generates descriptive metadata for both tables and individual nested columns.

  • Unified data insights: Gemini leverages this rich context to surface insights across your entire data estate. It automatically recommends queries to find anomalies or safety failures within a single table (like gen_ai_user_message). At the dataset level (Preview), it generates an interactive relationship graph to map cross-table join paths and suggests queries that combine data—like user activity and model outputs—to calculate task complexity.

  • Seamless agent integration and glossaries: Table and data insights integrate directly into the BigQuery Conversational Analytics (BQ CA) agent UI, giving agents immediate access to enriched metadata and few-shot examples. To ensure agents accurately interpret domain-specific prompts, BQ CA also supports business glossaries. You can define custom terms directly for your agents or import existing glossaries from Knowledge Catalog to establish a standardized vocabulary.

Administrators can then leverage the profiling, enriched metadata and insights to navigate logged fields, understand the telemetry structure, and catalog data for compliance audits. As shown in the screenshots below, the output of Gemini-powered auto generation of schemas, descriptions and linkages makes it easy to make sense of the complex relationships and telemetry data output by agentic interactions on the Gemini Enterprise app.

4
5

Visualizing with Data Studio dashboards

For executive stakeholders, raw log tables can be transformed into interactive, high-impact business intelligence dashboards. By connecting Data Studio directly to BigQuery, you can build dashboards that monitor:

  • User adoption and seat ROI: Segment usage trends by department, highlighting the ratio of custom agents built relative to employee headcount.

  • Data grounding traffic: Map which enterprise connectors—such as SharePoint, Google Drive, or Gmail—experience the highest utilization.

  • Content safety and violations: Track Model Armor sanitization blocks and sentiment feedback loops over time to maintain safety standards.

  • Share BQ Conversational Analytics agent: Share the BQ CA agents you built via Data Studio to give business users the ability to ask more questions of the data.

Empower your organization with the Gemini Enterprise app and your administrators with BigQuery

  1. Deploy the Gemini Enterprise App: Bring the best of Google AI to every employee.

  2. Enable Prompt & Response Logging: Turn on prompt and response logging in the Admin Console to begin recording user activity telemetry.

  3. Configure Log Router Sinks: Establish sinks to stream telemetry into BigQuery.

  4. Track Metrics & Export Analytics: Access pre-computed, out-of-the-box dashboards on the console and export historical aggregate statistics.

  5. Extract Table-Level & Dataset-Level Insights: Explore unfamiliar log tables and discover relationship join paths automatically.

  6. Query with Conversational Analytics: Build data reasoning agents and leverage natural language querying inside BigQuery Studio.

  7. Visualize with Data Studio: Connect Data Studio to BigQuery to build executive-level dashboards & give access to BQCA agents to business users.

  8. Consult a Google Cloud Customer Engineer for the most cost-effective and secure configuration for the analytics setup described above.

The authors would like to acknowledge and thank the Google Forge team, especially Vicky Falconer, Dharini Chandrashekhar and Adhaar Gupta, for contributing to the core work that led to this article.

Building the AI-defined vehicle with Android, Google Cloud, and Nexus SDV

13 juillet 2026 à 18:00

The automotive industry is moving from building hardware-centric platforms toward building their own sophisticated Software-Defined Vehicle (SDV) architectures. For OEMs, a vehicle is no longer just a way to go from point A to point B, but an intelligent, connected node within an AI-native ecosystem! 

With its partners, Google’s Android and Google Cloud are at the forefront of this transition. Android’s open source Automotive OS (AAOS) SDV implements the AI-defined vehicle while  Google Cloud provides scalable infrastructure including a full suite of AI integration tools, leveraging services like Bigtable for automotive and manufacturing telematics at scale. Valtech, a Google Cloud partner, uses Google technologies as part of its Nexus SDV platform, establishing a full end-to-end connected vehicle system that enables truly agentic mobility, offering automotive OEMs a ready-to-use, end-to-end foundation for the next generation of connected vehicles. Let’s take a look at how this all comes together.

The vehicle side: AAOS SDV

As the foundational in-vehicle platform, Google’s open source AAOS SDV platform abstracts core functions into reusable services independent of physical hardware, establishing a modular Service-Oriented Architecture (SOA). By decoupling non-safety domains like climate control, lighting, and diagnostics from Electronic Control Units (ECUs), the AAOS SDV platform introduces dynamic runtime service discovery. With this, the SDV can easily discover what services are running (e.g., the odometer, HVAC, sunroof, motorized seats, electric windows, etc.) and their status.

To accelerate development, engineering teams leverage the Android Cuttlefish emulator to build digital twins in the cloud, simulating high-frequency sensor streams to validate these decoupled services bit-for-bit before physical silicon is ready. Valtech Nexus SDV utilizes this AAOS SDV middleware layer to discover, map, and manage vehicle resources, structuring and streaming high-frequency telemetry data straight into Bigtable. Compare this to the prior state of affairs, where OEMs outsourced system software to a variety of suppliers, each with their own pipelines, protocols, and data stored in separate silos. 

Crucially, this model decouples services from the heavy main infotainment stack, so they can run independently, even when the vehicle is off and parked. This allows functions like remote vehicle monitoring to remain active even when the primary infotainment system is powered down, ensuring continuous telemetry access without draining the vehicle’s 12V battery or main EV battery pack.

This tight integration between the AAOS SDV platform and Nexus SDV enables a number of agentic AI and innovative first-party solutions. Unlike traditional sandboxed infotainment tools, multimodal AI agents can utilize the service discovery layer to safely interact with the physical car and process complex, intent-based requests. For example, an AI agent could automatically adjust climate zones, window actuators, or interior lighting based on a conversation with the driver, or in response to climate sensors, as in this clip:

1

By linking this on-vehicle service layer managed by Nexus SDV with historical fleet telemetry stored in Bigtable, you deliver deeply integrated experiences that unlock new mobility solutions. Now let’s take a quick look at the Cloud side.

The Google Cloud side: AI-native mobility

Beyond SDV, we are rapidly moving toward AI-defined vehicles, or AIDV, where AI is core to a vehicle's operational logic. To be AI-native means being autonomous by design, with AI embedded at every architectural level. With this level of AI, the system can perceive environments, reason through complex scenarios using engines like Google Gemini, and proactively execute actions. For example, a Gemini-powered vehicle doesn't just warn you that you’re low on power; it analyzes your schedule, traffic, and charger availability to suggest an optimized charging stop that pre-conditions the battery for maximum efficiency. This is the level of contextual understanding and proactive automation that characterizes AIDV.

Compare this to legacy architectures, which weren’t designed to capture the volume and variety of data coming from different systems across the vehicle. This can lead to data silos of isolated maintenance and safety information telematics. Moreover, because this data is fragmented, it can be very difficult to get cohesive value from the data across systems. An AI-native approach can help collapse these silos, providing a unified contextual understanding. This solves a primary OEM pain point: the massive complexity of managing high-bandwidth telemetry from multiple sources like SDV telematics. 

Bigtable: The data backbone for Automotive Telemetry

Bigtable was purpose-built for the massive ingestion rates and sub-millisecond latency requirements, and serves as the data backbone for petabyte-scale automotive and manufacturing telemetry datasets. In fact, Bigtable is already being used to support business critical automotive telemetry solutions. Its flexible, sparse-row schema allows OEMs to evolve their data models without downtime, accommodating diverse sensor arrays — from high-frequency engine metrics to LiDAR point clouds — within a single, unified table structure. Then, by versioning time-series events in a way that is natively optimized for both massive writes and complex, multi-dimensional analytical lookups, Bigtable helps avoid the data overload typical of legacy systems.

Meanwhile, features like Continuous Materialized Views (CMV) allow for pre-calculating key metrics, such as average battery temperature or fleet-wide torque distributions, directly within the storage layer, minimizing computational overhead. Bigtable’s integration with Agent Development Kit (ADK) further bridges the gap between data and action by giving AI agents access to data. This kit combined with Bigtable’s integrations with frameworks like Apache Spark help monitor the "firehose" of live telemetry data and trigger automated workflows in real time, e.g., logging mission-critical alerts, initiating proactive over-the-air (OTA) software adjustments, or pre-ordering replacement parts, the moment specific degradation patterns are detected.

Bring it all together: Nexus-SDV platform

The Nexus SDV platform is built on Google Cloud and integrated with AAOS SDV, supporting the future of connected vehicles. By providing a standardized data foundation, Nexus empowers automotive OEMs to go beyond building infrastructure from scratch and start focusing on unique brand experiences.

Nexus SDV uses Google components like Gemini Enterprise Agent Platform, Bigtable, and BigQuery. Setting up Nexus SDV is quick, automated and transparent. OEMs can create  brand-specific customer experiences in the vehicle, as well as in other customer touch points such as the UI screen, mobile app, or service centers.  The connection to the vehicle is accomplished by leveraging the open source Synadia NATS interface. This integration with the vehicle is facilitated through simple Cloud and vehicle SDKs, for service discovery on both sides. Nexus SDV is optimized for AAOS SDV, but can integrate with any vehicle framework.

2

Security is woven into the Nexus architecture via a "Defense-in-Depth" model. Mutual TLS (mTLS) and Google Cloud Certificate Authority Service (CAS) provide vehicles with a cryptographically secure identity. Network isolation is maintained through Private GKE clusters, while the Secure AI Framework (SAIF) helps ensure data privacy throughout the machine learning lifecycle, protecting sensitive user data and OEM intellectual property.

Together, the quick setup and integration time coupled with a standardized data foundation and built-in state-of-the-art security leads to an immediate and measurable business impact for the car manufacturer.

3

Let’s put it all together and look at a use case in more detail…

Predictive maintenance

By moving from reactive to predictive maintenance, OEMs can reduce warranty costs, improve customer loyalty, and ensure higher vehicle uptime.

The challenge: Traditional scheduled maintenance is often inefficient, leading to unnecessary service visits or unexpected vehicle breakdowns that incur significant costs for both OEMs and owners. By moving to a proactive, AI-driven approach, Nexus SDV,  Bigtable, and ADK transform this experience. The process begins by taking the firehose of vehicle telemetry data —monitoring engine RPM, vibration, fluid levels, brake pressure, and more — ingesting it and storing it directly into Bigtable.

To enable real-time anomaly detection, agentic AI can monitor telemetry streams as they arrive. Bigtable CMVs pre-calculate rolling aggregations such as average engine vibration or sudden fluctuations in battery temperature profiles. AI models consuming these live aggregates can then detect subtle deviations from normal parameters, identifying early signs of engine wear or accelerated battery degradation long before a warning light appears on the dashboard.

Once an anomaly is detected by specialized AI models, the system shifts into the agentic reasoning and action phase. A Gemini-powered engine assesses the severity and context of the data, considering factors like mileage, model, make, service history, and upcoming trips. Based on this intelligent assessment, the system can proactively notify the driver via the AAOS infotainment system, suggests an optimized service appointment at a nearby dealership, or can even trigger an automated parts order to ensure everything is ready upon arrival. The AI model works against false negatives to protect customer sentiment or erosion of confidence, while the solution as a whole ensures higher vehicle uptime, transforming maintenance from a reactive burden into a brand-defining service experience.

Get started

The AI-native Nexus SDV platform with AAOS SDV is available today, providing a sophisticated, end-to-end connected vehicle ecosystem designed to meet the extreme scale and analytical rigors of modern mobility. By adopting this unified, open-source architecture, OEMs can transcend the limitations of legacy infrastructure and redirect their resources toward the development of high-impact, brand-defining features. 

Nexus SDV takes the connected vehicle service into the agentic era, where vehicles are no longer merely connected, but serve as intelligent, proactive partners in the driving experience. Give it a try today.

Learn more

If you’d like to learn more about Nexus SDV platform, AAOS SDV and Bigtable contact us today at nexus-sdv@google.com.

AAOS SDV is available in the Android Automotive 26Q2 release. Nexus SDV documentation can be found here.

Go here to learn more about Bigtable as the time-series database for automotive telemetry.

Thinking about your connected vehicle security, check this out, Shift into high gear with agents: Securing the software-defined vehicle.

❌