Vue normale
Unlock 3x QPS and microsecond latency with Memorystore for Valkey 9.1
At Google Cloud, we are committed to delivering the best managed experience backed by open source software. Today, we’re announcing the general availability of Memorystore for Valkey 9.1, which achieves up to 3x queries per second (QPS) at microsecond latency compared to Memorystore for Redis Cluster.
Our support for Valkey dates back to 2024, when Redis Inc. shifted its licensing away from the permissive open-source BSD license to a dual-license model. In response, Google Cloud, alongside other technology leaders, backed the creation of Valkey, an open-source alternative governed by the Linux Foundation.
Valkey has come a remarkably long way since then, delivering major performance and feature updates that push boundaries far beyond the original fork. Valkey is particularly compelling for organizations scaling AI and microservices to handle millions of concurrent users. Here, backend developers and architects must deliver both massive throughput while also maintaining microsecond latency.
In this blog, let’s take a look at how Valkey 9.1 achieves its performance, new developer capabilities, how to get started, and how customers are using it.
Under the hood: Rethinking thread communication
In high-throughput, in-memory datastores, efficient I/O offloading is critical to keeping the main execution loop unblocked. Previously, Valkey assigned client sockets to I/O threads statically in a round-robin fashion, requiring the main thread to continuously poll lists of pending clients to detect completed work.
Valkey 9.1 replaces list-polling with a lock-free, multi-queue messaging architecture that eliminates cross-thread CPU waste and unlocks dynamic work balancing. It involves three complimentary queues:
-
Main thread to I/O thread queue: Dispatches read and write jobs to a single-producer multi-consumer (SPMC) queue. Free worker threads pull tasks on demand, enabling dynamic work-stealing that prevents thread starvation or hot-spotting.
-
I/O thread to main thread queue: Worker threads push completed tasks into a multi-producer single-consumer (MPSC) queue. The main thread pops completed work instantly, eliminating busy-wait list iteration.
-
I/O thread-specific queues: Dedicated single-producer single-consumer (SPSC) queues handle thread-affine memory cleanup and high-volume epoll offloading.
Valkey 9.1 also replaces static thread thresholds with a two-phase dynamic scaling engine:
-
CPU-driven "ignition": When main-thread CPU usage crosses 30%, the engine automatically activates the first background I/O thread to absorb incoming traffic before queue bottlenecks form.
-
Queue-depth auto-scaling: Once ignited, Valkey dynamically scales the number of active I/O worker threads up or down based on real-time SPMC queue backlog, ensuring extra cores are used only when needed and parked when idle.
New developer capabilities in Valkey 9.1
Beyond raw performance, Valkey 9.1 addresses key feature requests from engineering teams with powerful new commands and enhanced security controls. Here is a look at what you can do with these new capabilities:
1. Granular database-level access control (ACLs)
We recently launched support for access control lists on Memorystore for Valkey to provide more granular key-level and command-level authorization using IAM. This foundational security mechanism is offered at no additional cost and includes the following capabilities:
-
Centralized management: A 1:N mapping approach allows you to define a single ACL policy and attach it across multiple clusters.
-
Secure multi-tenancy: Organizations can easily enforce least privilege and secure multi-tenancy across their database fleets.
-
Enhanced observability: The feature includes versioned policy revisions and comprehensive audit logging.
Previously, ACL rules applied globally across an instance. Valkey 9.1 allows administrators to restrict user access at the specific numeric database level within the ACL framework.
Real-world example: You can configure a staging or service-specific user and isolate their access strictly to non-production databases:
-
productionuser:@all ~* db=0 -
staginguser:@all ~* db=1 -
devuser:@all ~* db=2
Protect against unauthorized data access and guard against application bugs by leveraging database-level access control across multiple databases, all without needing to prefix your keys.
2. CLUSTERSCAN: Efficient cluster-wide key scanning
Previously, scanning keys across a large cluster required querying nodes individually. This approach was not cluster- or failover-aware. Consequently, scans could miss keys, return duplicates, or fail if slot migrations or node failovers occurred during the process.
The CLUSTERSCAN command addresses these limitations by introducing a topology-aware cursor. This cursor encodes the current slot, the fingerprint of the local hashtable, and the local cursor. With this additional encoded information, clients can scan keys across the entire cluster while gracefully handling topology changes and redirections.
CLUSTERSCAN supports two primary scanning strategies:
Use case 1: Sequential full cluster scan (single worker)
This strategy is suitable for simple scripts or background jobs that prioritize simplicity over speed. The client starts with cursor 0 and sequentially traverses all slots in the cluster:
- code_block
- <ListValue: [StructValue([('code', 'CLUSTERSCAN 0 MATCH "user:*" COUNT 10\r\n1) "0B3a21-{06S}-64"\r\n2) 1) "user:101"\r\n2) 2) ...'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f4b3c6768d0>)])]>
To continue the scan, pass the returned cursor to the next call. The cursor automatically transitions to the next slot when the current one is fully scanned.
- code_block
- <ListValue: [StructValue([('code', 'CLUSTERSCAN 0B3a21-{06S}-64 MATCH "user:*" COUNT 10\r\n1) "0B3a21-{07T}-0"\r\n2) 1) "user:102"\r\n2) 2) ...'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f4b3c6f5d90>)])]>
The scan is complete when the command returns a cursor of "0".
Use case 2: Parallelized cluster scan (multiple workers)
This strategy is suitable for high-throughput scans. Using the SLOT argument restricts the scan to a specific slot, allowing you to partition the 16,384 slots across multiple parallel workers.
Worker 1 (Scanning Slot 0):
- code_block
- <ListValue: [StructValue([('code', 'CLUSTERSCAN 0 SLOT 0 MATCH "user:*" COUNT 10\r\n1) "0B3a21-{06S}-64"\r\n2) 1) "user:101"\r\n2) 2) ...'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f4b3c6f7550>)])]>
Worker 2 (Scanning slot 1000 in parallel):
- code_block
- <ListValue: [StructValue([('code', 'CLUSTERSCAN 0 SLOT 1000 MATCH "user:*" COUNT 10\r\n1) "0B3a21-{08X}-32\r\n2) 1) "user:999"\r\n2) 2) ...'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f4b3c6f4650>)])]>
From here, Worker 1 continues to pass SLOT 0 and Worker 2 continues to pass SLOT 1000. Mismatching the slot and the cursor returns an error. Once all 16384 slots have been scanned, the cluster scan is considered complete.
3. More commands for atomicity and expirations
HGETDEL: Atomic fetch and delete
A frequent application pattern involves reading a hash field and deleting it immediately (such as consuming single-use authentication tokens or short-lived session states). Valkey 9.1 introduces HGETDEL, which retrieves the value of a hash field and deletes it atomically in a single network round-trip.
Real-world example:
HSET user:1001 temp_token "abcde"
(integer) 1
HGETDEL user:1001 FIELDS 1 temp_token
1. "abcde"
HGET user:1001 temp_token
(nil)
MSETEX: Shared expiration for multiple keys
To eliminate multi-command pipeline overhead, the new MSETEX command enables setting multiple keys simultaneously with a single, shared expiration time.
Real-world example: Setting up a temporary session state where multiple distinct keys must expire together in 300 seconds:
MSETEX 2 session:auth "ok" session:user_id "1001" EX 300
(integer) 1
TTL session:auth
(integer) 300
Enhanced HSETEX with conditional flags
HSETEX now supports the NX (only set if the field does not exist) and XX (only set if the field exists) conditional flags.
Real-world example: Initializing a rate-limit threshold field with a 1-hour TTL, ensuring you don't overwrite an existing active limit:
HSETEX config:123 NX EX 3600 FIELDS 1 "rate_limit" "100"
(integer) 1
Built on Memorystore for Valkey 9.0
The release of Valkey 9.1 builds upon the major updates we unveiled for Memorystore for Valkey at Google Cloud Next '26:
-
Built-in modules for AI & vector workloads: Native JSON support and Bloom filters enable fast document querying and membership checks.
-
Six new node sizes: To help you manage costs and scale, we added six new node sizes.
-
Small Size Nodes: Custom-Pico (1.25 GB), Custom-Micro (2.5 GB), and Custom-Mini (3.5 GB) for lightweight microservices and dev/test environments. These are only available for cluster mode disabled environments.
-
High CPU and Large SKUs: HighCPU-Medium (8 vCPU/13 GB) and Standard-Large (8 vCPU/26 GB) optimized for CPU-heavy applications.
-
XXL SKU: Highmem-XXLarge with 110 GB RAM and 16 vCPUs per node for massive cluster consolidation to power your most demanding workloads.
(Note: The figures above are based on open-source benchmarks; actual performance improvements will vary depending on your specific workloads.)
Migrating to fully managed Memorystore for Valkey
Having to self-manage your Redis OSS /Valkey caching layers drains valuable engineering bandwidth and creates operational friction during scaling. We are also excited to announce a new migration workflow to Memorystore for Valkey.
With this release, migrating your infrastructure is straightforward, fully managed, and requires a simple configuration change on your application to point to Memorystore for Valkey once your data is migrated. This workflow is generally available.To move off self-managed Redis or Valkey to fully managed Memorystore for Valkey, follow these four steps:
1. Provision the target instance: Deploy a Memorystore for Valkey instance configured with your required shard count, node sizing, and clustered database options.
2. Establish online replication: Initiate continuous, dual-sync online migration directly from your source database to Memorystore.
3. Validate data synchronization: Monitor replication metrics in real time to verify full dataset alignment and low-latency replication health.
4. Execute the cutover: Switch application connection endpoints over to Memorystore for Valkey to start using the new cache.
What Memorystore for Valkey customers are saying
Already, over 95% of the top 100 Google Cloud customers already rely on Google Cloud Memorystore to power demanding, high-throughput workloads, led by increasing numbers of Memorystore for Valkey users.
Consider the fast-paced world of live sports, where delivering a flawless digital experience is of utmost importance. When a game-changing play happens, millions of fans immediately reach for their devices to check real-time stats, watch highlights, and engage with interactive features. These massive, unpredictable traffic spikes require an underlying architecture capable of immense scale. For organizations like Major League Baseball (MLB) , a partner since Valkey’s early days, managing unpredictable traffic spikes without compromising performance is essential.
"We trust Memorystore for Valkey to power the massive scale of live baseball, delivering real-time stats and uninterrupted digital experiences to millions of fans. As we look ahead, we are incredibly excited about the Memorystore for Valkey 9.1 launch. The engine optimizations and latency enhancements will give us even more horsepower to handle the most unpredictable game-day traffic spikes, ensuring fans get the best technology-powered experience the game has to offer." - Rob Engel, SVP of Software Engineering, Major League Baseball
Beyond the stadium, the retail industry faces its own intense scaling challenges, particularly during major shopping holidays or flash sales. Modern e-commerce platforms rely on real-time personalization, dynamic pricing, and instant inventory updates to keep shoppers engaged. A lag of even a few milliseconds can disrupt the customer journey and impact the bottom line. To maintain a competitive edge, leading retailers such as Target require ultra-responsive caching layers to power their most crucial customer-facing platforms.
"By leveraging Google Cloud Memorystore for Valkey, Target delivers ultra-low-latency, resilient caching for personalization services. We look forward to leveraging the performance enhancements in Valkey 9.1 to make our personalization platform even faster, more scalable, and more resilient during periods of peak demand." - Scott Weide and Sumanth Huddar, Senior Engineering Managers, Target
The demand for these ultra-low-latency architectures extends far beyond sports and retail. Across the digital landscape, organizations in banking, AI-native development, digital streaming, and telecommunications all share a common mandate: the need for superfast, highly available caches. Whether it is processing high-frequency financial transactions, serving complex machine learning inferences in real time, delivering seamless global video streams, or routing immense volumes of telecom data, microsecond latency is the new baseline for success.
Make the move to Valkey
Stop letting cache bottlenecks slow down your most demanding applications. Experience the performance, dynamic scalability, and enhanced security of Memorystore for Valkey 9.1 today.
- Start building: Create a Memorystore for Valkey 9.1 instance in the Google Cloud console.
- Dive deeper: Read the technical documentation to view the full list of supported commands, ACL configurations, and detailed capabilities of Memorystore for Valkey.
-
Cloud Blog
- AlloyDB delivers PostgreSQL for agents: Real-time data at agent scale, with full workload isolation
AlloyDB delivers PostgreSQL for agents: Real-time data at agent scale, with full workload isolation
Enterprises rely on mission-critical operational databases where performance slowdowns simply aren’t an option. Yet when even a few agents execute dense reasoning loops, the unpredictable surge in queries can easily overwhelm traditional architectures.
Today, we’re announcing that AlloyDB delivers PostgreSQL for agents (in preview), enabling real-time data access without compromising your mission-critical systems. AlloyDB now scales to dynamic agent bursts by provisioning sandboxed database instances in seconds, enabling full workload isolation. You can cost-effectively run your agents at any scale — from a few agents to millions of agents — and the instances automatically spin down when agents finish. With this announcement:
-
AlloyDB now features an agentic database architecture engineered to scale PostgreSQL to thousands of serverless database instances that have up-to-the-second read-only access to production. These instances remain fully separated from the primary, standby, and read replica instances where production workloads run.
-
Each of these instances accesses real-time data in the database backed by a unified storage layer in Colossus, Google’s exabyte-scale distributed storage system. This helps agents achieve sub-millisecond I/O, and terabit-per-second aggregated scan throughput, supporting over 3 million queries per second.
-
These instances utilize the full AlloyDB PostgreSQL engine, providing access to every index, the full capability of SQL, and comprehensive vector, full-text, and spatial search. Agents can also leverage BigQuery and Spark to run lakehouse analytics without requiring complex ETL pipelines.
-
When agents complete their tasks, these instances scale right back to zero, thus reducing your cloud spend by billing only for active reasoning loops.
With AlloyDB for PostgreSQL, we pioneered the agentic enterprise relational database by integrating advanced vector operations, machine learning inference, and foundation model integrations directly within a 100% PostgreSQL-compatible engine. It protects your data through deep Google Cloud security integrations — replacing static passwords with IAM authentication, isolating traffic via VPC Service Controls, and providing customer-managed encryption and auditing. This functionality, combined with the scalability now provided by our agentic architecture, makes AlloyDB the premier enterprise-grade agentic PostgreSQL offering.
To get started, sign up here for the preview.
Why this matters
In today’s agentic era, we’re swiftly moving from single copilot agent interactions to networks of millions of agents collaborating simultaneously. When these agents query and operate all at once, sudden traffic spikes can overwhelm your core databases, competing with the systems that run your business.
Agents require both fast analytics and low-latency access to real-time production data, utilizing B-tree, vector, text, and spatial indexes to efficiently execute their workflows. Emerging architectures rely on page-caching layers that sit above object storage, but they suffer from scaling and cost challenges that can compromise the stability of production systems.
In addition, they face a challenging trade-off: To unlock production data for analytics, they create performance bottlenecks for operational access, putting mission-critical databases at risk the moment agents are unleashed in production. These approaches attempt to solve the problem using traditional object stores for database storage. While this enables analytical access that can help some agents, the underlying databases are too slow for production workloads, suffering from up to an order of magnitude higher I/O latency. Page caching layers are at best a patch; the caches themselves are often still not fast enough, and they pose a scalability bottleneck that is easily saturated by agentic workloads.
When active multi-agent systems execute dense reasoning cycles, they trigger highly concurrent, unpredictable bursts of queries that overwhelm these caching layers, leaving mission-critical production systems vulnerable to agent-induced outages. Consequently, an entire class of operational use cases is precluded from running on these architectures, locking businesses out of the transformative power of AI on live enterprise data.
AlloyDB’s unique agentic PostgreSQL architecture
We are taking a different approach. AlloyDB delivers an agentic database architecture purpose-built for the AI era, with four key differentiated capabilities:
- Sub-millisecond I/O latency, without artificial choke points: Combining AlloyDB’s industry-leading transaction and query processing with low-latency object storage backed by Google’s planet-scale Colossus storage infrastructure, this architecture provides a large-scale, shared storage plane for agents. It achieves sub-millisecond I/O and over a terabit-per-second of aggregate scan bandwidth, allowing agents to execute intensive read queries and vector searches directly against fresh operational data.
- Fully isolated from production workloads while scaling to meet demand: AlloyDB scales by dynamically provisioning sandboxed database instances in seconds against fresh production data. Unlike traditional architectures where agents compete for operational resources, agentic database compute remains completely isolated from the primary database clusters — allowing agents to execute dense, unpredictable reasoning loops without degrading performance in production. These robust safety guardrails, coupled with enterprise-grade governance and fine-grained access control, allow you to confidently unleash the full, unconstrained power of PostgreSQL on your production data — seamlessly mixing analytical queries, vector queries, and operational point lookups in active agentic loops.
- Pay-as-you-go billing: Most agent activity is characterized by sharp spikes of concurrent queries followed by periods of inactivity. Provisioning dedicated read replicas to absorb these bursts forces you to maintain expensive infrastructure around the clock. To support massive groups of agents cost-effectively, and eliminate the idle compute overhead of provisioned systems, these agentic AlloyDB instances can rapidly scale to handle millions of queries per second, and automatically scale to zero with a flexible, pay-as-you-go pricing model.
- Native lakehouse integration, without ETL: All production data is natively integrated with Google Cloud’s borderless Lakehouse. This allows agents to run federated queries across BigQuery and Lightning Engine for Apache Spark, joining massive lakehouse datasets with up-to-the-second transactional data in AlloyDB. This eliminates the need to build and maintain fragile batch ETL pipelines, giving autonomous agents instant access to both live operational state and historical lakehouse context.
“As supply chains become increasingly autonomous, our platform relies on real-time transactional intelligence to coordinate complex logistics workflows across thousands of facilities. AlloyDB's new PostgreSQL architecture for agents has been a game changer for us. We can now deploy networks of agents collaborating simultaneously to help us analyze inventory and order data with sub-second freshness, while ensuring our core transactional processing remains entirely untouched. It delivers the isolation, speed, and cost efficiency we need to power the next generation of enterprise supply chain AI.” - Sanjeev Siotia, Executive Vice President & Chief Technology Officer, Manhattan Associates
Availability
PostgreSQL for agents in AlloyDB is now available in preview.
To learn more, visit the documentation page, and sign up here to get started.
A new, no-compromises database architecture for the agentic era
Entire database engineering careers have been spent on a single question: How do you scale an OLTP workload without compromising the system of record that owns the data?
Exadata answered the question by offloading queries into a scale-out storage tier beneath the database, removing the network as the bottleneck. Azure SQL Hyperscale did it with shared block servers, scaling out to tens of read replicas. Aurora offloaded log application to distributed storage nodes, scaling reads across tens of PostgreSQL nodes. Meanwhile, emerging architectures persist data in traditional object storage with a provisioned cache tier in front, recovering latency for hot data but leaving a high-latency tail on every cache miss.
Each of these architectures is inherently constrained by at least one of these three properties: scale, latency, and isolation — and sometimes even two. For instance, architectures built on shared block servers compromise scalability, because I/O inevitably bottlenecks on the block server. They also sacrifice isolation, as production workloads get throttled whenever replica traffic spikes.
Some of these trade-offs were actually sound when they were developed; they met the requirements of enterprise database workloads for four decades. However, in the agentic era, these compromises are no longer acceptable. Agentic workloads are generated dynamically and cannot be vetted in advance, making it a business-continuity imperative to isolate them from mission-critical systems. Agentic workloads also require low latency that is only possible with the full power of the database engine and all its indexes, as well as a whole new level of elastic scale that has never been tried with a single database: a burst of agents that demand 1,000 compute nodes over a single database within seconds, and that may finish inside a minute.
Three tenets needed for a truly agentic database architecture
We believe the agentic era demands a new agentic database architecture defined by three fundamental tenets. An agentic database architecture must satisfy all three, or it isn’t really agentic.
-
Tenet: Isolation — isolation by design, but with real-time data access. Agents must read live production data with sub-second freshness over a data path that does not share database components with the primary cluster. Real-time means up-to-the-second, not a stale copy or branch. This is physical separation, not a quota — because shared allocations mean shared fate. The boundary extends straight through the storage layer, eliminating resource contention by design.
-
Tenet: Latency — sub-millisecond baseline I/O. Operational workloads demand sub-millisecond block I/O, and that bar does not drop for agents. While compute nodes leverage DRAM and local SSD for acceleration, cache misses that reach remote storage — whether application or agentic — must complete in under a millisecond. An architecture that degrades into an order-of-magnitude performance cliff is fundamentally unusable by agents.
-
Tenet: Scale — agent-scale compute and I/O. Agent scale is simultaneously instantaneous, volatile, and massive: Database compute nodes must spin up in seconds, scale to thousands, run for short bursts, and automatically spin down to zero when agents are done with them. No one has thus far ever dreamed of expecting a database to scale compute and I/O dynamically to thousands of nodes while leaving production untouched. Due to the dynamic nature of agents, pre-provisioning is a non-starter across the entire stack, whether it’s compute, storage I/O, or any caching tier in between.
Crucially, an agentic architecture must uphold all three tenets at once. And by doing so, the architecture allows agents to work directly against live operational data, i.e., enterprise truth, without compromising production stability. The outcome is transformative:
-
No correlated failures: Total decoupling between the engines running the business and the fleets of agents reasoning over it removes a path for agents to affect production.
-
No capacity guesswork: True elasticity that eliminates the friction of pre-provisioning for unforecastable agent scale.
-
No semantic compromises: Nothing is withheld from agents — they get access to the full power of relational SQL, hybrid search (vector, full-text, spatial), and indexes within every single reasoning step.
AlloyDB's agentic architecture
AlloyDB’s new agentic database architecture is the first system that satisfies all three tenets. We engineered this from the ground up across storage, network, compute and databases to deliver:
-
Isolation, avoiding shared fate by design: The transactional production cluster runs on dedicated, pre-provisioned infrastructure, completely isolated from agent workloads. Agents interface via the Model Context Protocol (MCP) to an independent, ephemeral pool of microVM-based AlloyDB nodes that read directly from dedicated Colossus storage segments, separate from those for production.
-
Predictable sub-millisecond storage I/O: Every storage read is served directly by Google’s Colossus storage system inheriting its baseline sub-millisecond latency, eliminating performance cliffs on cold cache misses.
-
True zero-to-thousands compute scaling: The agent pool scales rapidly from zero to thousands of nodes for bursty agentic activity, and scales back to zero the moment tasks complete.
Agents query production data with sub-second freshness, with the complete PostgreSQL engine — point lookups, index traversals, vector, full-text and spatial search, columnar scans, and federated queries across the lakehouse — at their disposal to power their reasoning loops.
Run agents against production data at any scale by joining the preview of AlloyDB PostgreSQL for agents. You can learn more about its full capabilities in the companion announcement blog.
Why existing architectures can’t satisfy all three tenets
Traditional and emerging operational databases attempt to scale using one of three architectural paradigms. When assessed against the demands of autonomous AI agents, each paradigm exhibits a fundamental structural compromise — none satisfies all three tenets simultaneously.
Independent replicas (shared-nothing storage)
Traditional relational architectures scale reads by streaming replication logs from a primary instance to dedicated replica databases, each with its own local or attached block storage. They meet Tenet: Isolation – replicas share no physical resources with the primary cluster, and continuous log replication maintains near-real-time currency. They meet Tenet: Latency – dedicated local storage guarantees predictable, sub-millisecond read latency. However, they fail Tenet: Scale – scaling requires provisioning a new replica and rehydrating hundreds of gigabytes or terabytes of storage. All this takes hours — an impossible mismatch for agent-reasoning bursts measured in seconds. Furthermore, statically provisioned compute and storage continue to incur idle costs long after the agent completes its run.
Disaggregated shared-storage servers
A second approach decouples stateless compute nodes from a shared, multi-tenant tier of custom storage servers that manage persistence, replication, and that may offload block writes. This approach meets the Tenet: Latency – reads hitting the optimized storage servers resolve with consistent, low operational latency. However, it fails the Tenet: Isolation – because every replica reads from the same servers as the primary, so agent I/O contends directly with production I/O, creating shared fate. It also fails the Tenet: Scale – stateless compute replicas spin up quickly because no data is copied, but total storage I/O bandwidth is fixed to the pre-provisioned storage tier. Adding compute nodes without scaling underlying I/O capacity simply accelerates storage saturation and throttling.
Object storage with shared-block servers
A third emerging approach keeps data durable in general-purpose object storage and serves block reads from a shared tier of block servers. Because a random read from object storage takes tens of milliseconds — an order of magnitude slower than traditional database storage, and slower than an enterprise disk array has been for at least 25 years — the block servers hold hot data in order to serve it at low latency. This approach meets the Tenet: Latency — with one caveat: A block server miss still falls through to object storage at unacceptably high latency. It fails the Tenet: Isolation — because replicas share the block servers with production: Agent I/O and production I/O draw on the same capacity, so when that capacity is exhausted or throttled, production is affected along with the agents. It also fails the Tenet: Scale, for the same reason as shared storage servers: Replicas start quickly, but the block servers do not scale their I/O with the burst.
Some architectures in this family also allow analytical engines like Apache Spark to read the underlying object storage directly, bypassing the database engine. For analytics workloads, that is a valuable and viable path. However, since agents need low-latency retrieval, stripping away indexes, point lookups, and vector search forces brute-force table scans, exploding latency, and therefore breaks the ability for agents to execute their retrieval-reasoning loops.
Evaluating existing architectures
We evaluated a commercially available service that uses the object storage architecture with shared block servers by running concurrent index lookups over a dataset larger than available DRAM, testing both scaling limits and production isolation. Starting with a single reader instance, we scaled the workload by adding up to eight read replicas.
In architectures with shared physical resources, scaling agents via read replicas quickly degrades both replica and primary performance. In our tests as seen in the chart below, adding replicas provided less than a 2x throughput increase, peaking at four replicas before dropping off as the shared block-server bandwidth saturated.
The impact on the primary database was immediate and severe: Primary throughput plummeted by more than 75% as replicas were added.
In short, neither traditional nor emerging architectures can meet the scale that agents demand, and certainly not without jeopardizing the stability of production systems.
Assessing against the Tenets
* Partially meets: Hot data is served at low latency from the block servers, but a block server miss falls through to object storage at tens of milliseconds.
In each case the gap is structural, not just a matter of tuning. Replication isolates by giving each replica its own storage, so it cannot add a replica faster than it can populate that storage. Shared storage servers add compute quickly by sharing storage, so they can neither isolate nor scale I/O. Block servers over object storage recover latency with a provisioned tier, so they can neither isolate nor burst, and every miss still reaches object storage. Each approach solves the problem at one layer and pays for it at another. Meeting all three tenets at once requires rethinking the database architecture across compute, network and storage together.
How we engineered AlloyDB across the stack
AlloyDB's agentic database architecture is vertically integrated across Google's data, AI and infrastructure stack: AI models, the database engine and analytical engines, but also storage, networking and compute infrastructure.
Storage: Colossus as the foundation
At the persistence layer, AlloyDB builds on Colossus, Google's exabyte-scale distributed storage system that underpins Google Search, YouTube, Gmail, Google Drive, Spanner, and Bigtable. A single Colossus cluster scales to exabytes of storage and tens of thousands of machines. With Spanner, we demonstrated that a transactional database engineered directly on Colossus can scale to thousands of nodes. The new AlloyDB architecture applies the same foundation to a new problem: agents.
Colossus has three properties that enable AlloyDB to satisfy the three tenets.
-
Direct, sub-millisecond I/O: Colossus is engineered to minimize read latency. A database node opening a Colossus stream receives a handle that describes where data physically resides. Authorization and metadata resolution happen once, when the stream is created; every subsequent read goes directly to the disks holding the data, over an optimized network protocol. The result is sub-millisecond latency across all of the database's data, with no intermediary to warm and no tier to miss.
-
Massive throughput: Colossus delivers up to 15 TB/s of aggregate throughput and 20 million queries per second to a single AlloyDB database without needing to provision bandwidth and with an unlimited number of concurrent hosts. At Colossus scale, a fleet of AlloyDB agent nodes is not a load the storage must be sized for; it is a fraction of the load the storage already serves!
-
Physical segment partitioning: AlloyDB serves agents from a separate set of Colossus segments, so agent I/O is deliberately spread away from the production data path rather than contending with it.
At no point along the data path — compute, network or storage — can an agent ever share a database component with production.
Network: Scalable bandwidth with Jupiter
Compute and storage are bound together by Jupiter, Google's high-capacity data center network. A single Jupiter fabric connects more than 100,000 servers with 13 petabits per second of bisection bandwidth — enough to carry a video call for every person on Earth.
Because Jupiter provides high bisection bandwidth with predictable low latency across the networking fabric, agent nodes can be scheduled flexibly anywhere in the cluster with consistent access to centralized storage. As the agent pool scales from zero to thousands, the underlying interconnect capacity absorbs the expanding traffic without creating placement bottlenecks.
Compute: Elastic and serverless PostgreSQL and analytics
At the compute layer, agents connect to AlloyDB's agent pool through MCP. The agent pool consists of AlloyDB agent nodes with read-only access to the up-to-the-second state of the database. This layer provides:
-
MicroVM isolation: Each agent node is a fully functional AlloyDB for PostgreSQL database engine running inside a lightweight, secure microVM. These instances are fully isolated from each other and from the dedicated primary cluster.
-
Rapid spin-up and scaling: Agent nodes are provisioned in response to requests from agents and stop automatically when the agents finish. In response to a burst, AlloyDB rapidly provisions thousands of agent nodes, serving millions of concurrent agents, and releases them as the agents finish. Because billing is per second of agent-node activity, a burst that uses a thousand nodes for tens of seconds will only be charged for the resources that the job consumed, and nothing more.
Meanwhile, the production cluster remains as it is today: pre-provisioned, on dedicated infrastructure, sized for the system of record. Agent nodes read from Colossus directly and see a consistent production state with sub-second freshness.
Beyond the agent pool, BigQuery and Spark can read AlloyDB data from Colossus with the same isolation from the production cluster, so agents can use lakehouse federation to join real-time operational data with large-scale lakehouse datasets.
By building on these Google-scale storage, network, and compute layers, AlloyDB’s new agentic database architecture achieves a remarkable goal: Share the data. Share nothing else.
Evaluating AlloyDB’s agentic database architecture
We tested AlloyDB by running concurrent index lookups over a dataset larger than available DRAM, testing scalability across the full stack. We ran the agentic workload starting with a single agent node — an independent database instance in the agent pool, rather than a traditional read replica — and scaled dynamically to thousands of nodes over a single database, measuring both the aggregate agentic throughput as well as any impact on production.
In this test, throughput scaled linearly from 3.9K to 41K QPS when expanding from one to 10 agent nodes. Scaling by two additional orders of magnitude yielded near-linear performance up to 1,000 nodes. We observed:
-
Zero primary degradation: Scaling from 1 to 1,000 agent nodes produced no measurable impact on primary cluster performance.
-
Massive throughput: Aggregate throughput dynamically scaled 773x to 3 million QPS, driving over 8 million IOPS in Colossus across 1,000 compute nodes.
In a similar benchmark running concurrent full table scans across 2,100 agent nodes, aggregate scan throughput exceeded 1 terabit per second.
Because the agent pool shares no physical infrastructure with the production cluster, teams can scale reasoning fleets to thousands of nodes without placing production systems at risk.
Engineering all three tenets by design
The table below shows how AlloyDB’s architecture satisfies each of the three tenets:
Every agentic database architecture will require these three foundational elements: storage with the properties of Colossus, a network that connects compute to that storage without constraint, and compute that can be provisioned and released at agent scale.
Google has spent more than two decades building exactly that, to run Google Search, YouTube, and Gmail. Now it underpins our agentic database architecture.
Give agents live data without impacting production
Every organization building with AI faces the same core dilemma: how to give agents full access to live operational data without putting the systems running the business at risk. Until now, architecture — not application needs — dictated that choice. Giving agents direct access to the database meant exposing mission-critical systems to unforecastable load, severe resource contention, and production outages.
An architecture built on these three tenets removes these compromises entirely. Agents reason over live production data with sub-second freshness. They have the complete engine at their disposal — every index, vector, full-text and spatial search, and the full capability of SQL — at sub-millisecond I/O. The architecture scales dynamically to thousands of isolated nodes when agents need it, then to zero when agents finish. Throughout, core transactional workloads remain untouched: no shared components, no shared quota, no correlated failures. Agents can deliver innovation without conflicting with business continuity.
The same property extends to every other reader of production data. Reporting, analytics and applications can freely read live data without putting production at risk, ending a constraint that has shaped operational databases for five decades.
The data in an enterprise's systems of record is its crown jewels. Built on this foundation, that data can finally be put to work in full.
Databases, unfettered.
To learn more, visit the documentation page, and sign up here to get started.
What managing 150,000 AI agents could look like for database teams
The database administrator of the future will spend considerably less time administering databases.
That sounds contradictory, but AI agents are taking over that work. For decades, DBAs have handled the decidedly hands-on work of keeping databases available, performant, secure, and affordable. They provision capacity, troubleshoot slow queries, manage migrations, and step in when something inevitably goes sideways.
AI is already taking on some of that work. At the same time, it is creating a much bigger data infrastructure fleet to manage.
The result is likely to be a very different kind of DBA: one who spends less time tending individual databases and more time supervising the autonomous systems doing it for them.
Congratulations, you’re managing robots now
This shift starts with a familiar problem: more infrastructure needs managing than the people available to manage it.
Database automation is hardly new, but agents can potentially go further than the scripts and rules DBAs already rely on. Rather than automating one predetermined task, an agent can inspect what is happening, decide what needs attention, use tools to act on it, and check whether its intervention worked.
That changes the DBA’s relationship with the database. A performance problem that once required someone to dig through metrics, identify the troublesome query, and decide how to respond could increasingly be investigated by an agent before a human gets involved.
It doesn’t remove the DBA from the equation. Someone still has to decide what an agent can do, where human approval is required, and what happens when it gets something wrong. But the work moves up a layer. Instead of personally performing every operational task, DBAs start managing the systems carrying them out.
Instead of personally performing every operational task, DBAs start managing the systems carrying them out.
And before anyone gets too comfortable with that idea, the number of those systems could become enormous.
150,000 agents walk into a database…
Gartner predicts that the average global Fortune 500 company will have more than 150,000 AI agents in use by 2028, up from fewer than 15 in 2025. Only 13% of organizations currently believe they have the right governance in place to manage them.
Not every agent will need its own database, but plenty will. They will create state, retrieve data, remember previous interactions, and exchange information with other agents. Many will also behave very differently from the applications DBAs are used to supporting: spinning up quickly, sitting idle for long stretches, and suddenly becoming busy when there is work to do.
Nobody is hiring 150,000 DBAs to manage them.
That is the scale problem Yugabyte is targeting with YugabyteDB AMP, or Agentic Multitenant PostgreSQL. Rather than treating each new agent workload as another database for an administrator to provision and babysit, AMP manages databases as a fleet.
The platform packs hundreds of small Postgres workloads onto shared distributed infrastructure while keeping their databases isolated. Lifecycle operations, including provisioning, branching, scaling, migration, and teardown, can be exposed to agents through MCP. Yugabyte has also built specialized agents for setup, migration, performance tuning, and integrations.
In that model, a DBA is no longer provisioning database number 14,372. The interesting job is setting the rules for how database number 14,372 is provisioned, operated, and fine-tuned without them.
Do more with less (no, really)
Scale is only half of the problem. Someone also has to pay for all this stuff.
Agent workloads make traditional capacity planning particularly awkward because many are bursty and frequently idle. Giving every experimental agent permanently provisioned infrastructure could leave companies paying for many databases that spend much of their lives doing very little.
This is where consolidation becomes as much an economic question as an operational one.
AMP’s approach is serverless multitenancy and scale-to-zero. Multiple small workloads share the underlying distributed infrastructure, while customers pay by CPU minute and idle agents consume no compute. Resource governance can impose CPU limits on individual workloads, preventing a single overeager agent from consuming the capacity intended for its neighbors.
The human equivalent matters too. If routine setup, migrations, tuning and other database operations can increasingly be delegated, a smaller database team can potentially look after a much larger estate.
That doesn’t mean companies get to fire the DBAs and hand the keys to the robots. It means scarce database expertise can be spent on architecture, governance, and genuinely difficult problems instead of repeatedly doing the work that software can handle.
Your 2028 database problem starts now
The harder question is what to build underneath all of this when nobody really knows what the enterprise AI estate will look like in two years.
An agent that begins as an experiment today could disappear next month. Another could suddenly become a production application used across the business. Building one infrastructure stack for cheap experiments and another for serious workloads risks creating a migration problem every time an experiment succeeds.
Yugabyte bets that both ends of that journey should sit on the same foundation.
YugabyteDB AMP lets workloads start on serverless Postgres and transition to fully distributed YugabyteDB as their scale and criticality increase, without rewriting the application or migrating data to a different database platform.
Then there is the problem above the individual database: agents need to remember what happened, and not just in a silo.
That’s where Meko fits into the Yugabyte stack. Meko is an agent-native context engine designed for multi-agent AI systems. It provides persistent memory, shared knowledge, decision traces, and autidability across multiple agents, rather than leaving each agent working from its own isolated context. An agent can pick up information learned by another agent instead of retrieving it again or restarting the reasoning process.
Taken together, it delivers a single data stack for an agent’s entire lifecycle: Meko for the context shared among agents, YugabyteDB AMP for agentically managing fleets of Postgres databases, and distributed Postgres-compatible YugabyteDB for workloads that outgrow their serverless beginnings.
Of course, there’s no guarantee that 2028 will look exactly like today’s forecasts. That’s rather the point. The safest architectural bet may be one that doesn’t require you to know in advance which of today’s tiny AI experiments will become tomorrow’s critical applications.
The DBA is still critical in that world, but the job will look different. The DBA of the future may manage fewer databases directly, while taking responsibility for vastly more of them. Instead, managing the autonomous systems that do the administering.
The post What managing 150,000 AI agents could look like for database teams appeared first on The New Stack.
-
Azure service updates
- [Launched] Generally Available: Logical replication slot sync status metric for Azure PostgreSQL Flexible Server
[Launched] Generally Available: Logical replication slot sync status metric for Azure PostgreSQL Flexible Server
Announcing Native BM25 Ranking in AlloyDB and Cloud SQL
Vector search is a critical component of generative AI, retrieval-augmented generation (RAG), and data agent architectures, but sometimes vector search alone isn't enough. While vector embeddings are incredible at understanding conceptual meaning, they stumble on specific alphanumeric IDs and exact product SKU numbers. To build truly robust search and AI applications, you may need the combination of semantic vector search and traditional exact keyword full-text search — what we call hybrid search.
In search, Best Matching 25, or BM25, is a key algorithm used to estimate how relevant a document is to a given query. Until today, if you wanted BM25 ranking with AlloyDB or Cloud SQL, you needed to add an additional full-text search backend. This introduced data silos, sync lags, and operational complexity. Today, we are eliminating the friction of maintaining a separate full-text search backend altogether, with the preview of the native BM25 index in AlloyDB and Cloud SQL for PostgreSQL 17+, made possible through the open-source pg_textsearch extension created by Tiger Data.
Now, with a unified hybrid search backend, you no longer need to provision, manage, or pay for separate systems to get state-of-the-art full-text retrieval. It all happens directly inside your database, where your operational data lives, delivering:
-
Industry-standard keyword ranking: Powered by Tiger Data's
pg_textsearch, bring lightning-fast, C-optimized BM25 scoring directly to your Postgres tables. -
No complexity, total consistency: Eliminate the data duplication, ETL pipelines, and synchronization lag that you get when you maintain multiple backends for vector and full-text retrieval.
-
Supercharged semantic search (AlloyDB exclusive): Get up to 6x and 10x faster vector search queries (when compared to standard PostgreSQL) with ScaNN and HNSW index types.
Why pg_textsearch?
If you’ve used PostgreSQL's built-in ts_rank for full-text search at any meaningful scale, you already know its limitations. Ranking quality degrades as your corpus grows. There’s no support for inverse document frequency, so common words carry the same weight as rare ones. There’s no term-frequency saturation, so a document that mentions "database" 50 times outranks one that mentions it once.
BM25 is the information retrieval gold standard, providing inverse document frequency (rarer terms matter more), term frequency saturation (repetition doesn't dominate), and document length normalization. You can learn more in this blog post by Tiger Data about how they built a BM25 search engine on PostgreSQL pages.
Full-text search example
Here’s how to get started with BM25 full-text search on both AlloyDB and Cloud SQL. Consider a sample table, cymbal_products, that contains the unique identifier uniq_id, a product_name column, a product_description column containing a text description of each product, and a generated product_embedding column. cymbal_products contains information on various retail products, including indoor and outdoor plants.
Index creation
To use BM25, enable the pg_textsearch extension.
- code_block
- <ListValue: [StructValue([('code', '-- Install pg_textsearch extension\r\nCREATE EXTENSION pg_textsearch;'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f15a7c5b810>)])]>
Create the index on the product_description column from the cymbal_products table.
- code_block
- <ListValue: [StructValue([('code', "-- Create the native BM25 index on the content column\r\nCREATE INDEX idx_docs_bm25 \r\nON cymbal_products \r\nUSING bm25 (product_description) \r\nWITH (text_config='english');"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f15a7c58fd0>)])]>
A BM25 full-text search query can be executed using the <@> special operator. In the snippet below, we search for ‘cherry tree’.
- code_block
- <ListValue: [StructValue([('code', "-- Full text search query\r\nSELECT product_name, product_description <@> 'cherry tree' AS bm25_score \r\nFROM cymbal_products\r\nORDER BY bm25_score \r\nLIMIT 5;"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f15a7c5a850>)])]>
Sample output is shown below. A more negative score indicates a stronger relevance match.
AlloyDB hybrid search example
Setting up a hybrid search system in AlloyDB is simple. You can create both your vector and keyword indexes on the same table and merge the results seamlessly using the hybrid search user-defined function (UDF).
Vector index creation
Here is how to create a ScaNN vector search index:
- code_block
- <ListValue: [StructValue([('code', '-- Install vector extension\r\nCREATE EXTENSION vector;\r\n\r\n-- Install scann extension\r\nCREATE EXTENSION IF NOT EXISTS alloydb_scann;\r\n\r\n-- Create scann vector search index \r\nCREATE INDEX cymbal_products_embeddings_scann ON cymbal_products USING scann(product_embedding cosine);'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f15a7c590d0>)])]>
Hybrid search
AlloyDB provides an out-of-the-box hybrid search UDF that makes it very simple to run hybrid search queries. The UDF merges the ranked results from each search component into a single, unified list using the Reciprocal Rank Fusion (RRF) algorithm. This query utilizes the UDF to perform a vector search for ‘trees that grow taller than houses’ and a keyword search for ‘California’ in the product description.
- code_block
- <ListValue: [StructValue([('code', 'CREATE EXTENSION google_ml_integration;\r\n\r\nSELECT *\r\nFROM ai.hybrid_search(\r\n search_inputs => ARRAY[\r\n \'{\r\n "data_type": "vector",\r\n "weight": 0.5,\r\n "table_name": "cymbal_products",\r\n "key_column": "uniq_id",\r\n "vec_column": "product_embedding",\r\n "distance_operator": "public.<=>",\r\n "limit": 10,\r\n "query_vector": "ai.embedding(\'\'text-embedding-005\'\', \'\'trees that grow taller than houses\'\')::vector"\r\n }\'::JSONB,\r\n \'{\r\n "data_type": "text",\r\n "weight": 0.5,\r\n "table_name": "cymbal_products",\r\n "key_column": "uniq_id",\r\n "text_column": "product_description",\r\n "limit": 10,\r\n "ranking_function": "<@>",\r\n "query_text_input": "California"\r\n }\'::JSONB\r\n ],\r\n);'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f15a7c5b850>)])]>
As shown in the sample output below, results are ranked in descending order of their RRF scores.
Here, hybrid search bridges the gap between semantic intuition and exact keyword matching. While vector embeddings excel at grasping conceptual queries, like "trees that grow taller than houses", traditional full-text search provides the pinpoint precision needed for strict identifiers like "California." By fusing the two, AlloyDB helps ensure your application prioritizes highly specific, locally relevant results like ‘California Sycamore’ right at the top of the list.
Cloud SQL hybrid search example
In Cloud SQL, you can create both your vector and keyword indexes on the same table and merge the results seamlessly using Common Table Expressions (CTEs) and coalescing the RRF score, as shown below.
Vector index creation
Here is how to create an HNSW index in Cloud SQL.
- code_block
- <ListValue: [StructValue([('code', '-- Install vector extension\r\nCREATE EXTENSION vector;\r\n\r\n-- Create an HNSW index on the embedding column for fast approximate nearest neighbor search\r\nCREATE INDEX product_hnsw_idx ON cymbal_products USING hnsw(product_embedding vector_cosine_ops);'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f15a7c5ac50>)])]>
Hybrid search
Here is the hybrid search query.
- code_block
- <ListValue: [StructValue([('code', "CREATE EXTENSION google_ml_integration;\r\n\r\n-- BM25 keyword results\r\nWITH keyword_results AS (\r\n SELECT uniq_id, product_name, \r\n ROW_NUMBER() OVER (ORDER BY product_description <@> 'California') AS rank_kw\r\n FROM cymbal_products\r\n ORDER BY product_description <@> 'California'\r\n LIMIT 10\r\n),\r\n-- Semantic vector results\r\nsemantic_results AS (\r\n SELECT uniq_id, product_name, \r\n ROW_NUMBER() OVER (ORDER BY product_embedding <=> google_ml.embedding('text-embedding-005', 'trees that grow taller than houses')::vector) AS rank_vec\r\n FROM cymbal_products\r\n ORDER BY product_embedding <=> google_ml.embedding('text-embedding-005', 'trees that grow taller than houses')::vector\r\n LIMIT 10\r\n)\r\n-- Reciprocal Rank Fusion (RRF) to merge and score both lists\r\nSELECT COALESCE(k.uniq_id, s.uniq_id) AS uniq_id,\r\n COALESCE(k.product_name, s.product_name) AS product_name,\r\n COALESCE(1.0 / (60 + k.rank_kw), 0) + COALESCE(1.0 / (60 + s.rank_vec), 0) AS rrf_score\r\nFROM keyword_results k\r\nFULL OUTER JOIN semantic_results s ON k.uniq_id = s.uniq_id\r\nORDER BY rrf_score DESC\r\nLIMIT 5;"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f15a7c5b150>)])]>
The resulting output is identical to the AlloyDB hybrid search results shown above.
Watch it in action
Watch how this all comes together in this demo video.
Relevant resources
We are incredibly excited to work with Tiger Data and cannot wait to see how you leverage native BM25 support to build faster, smarter, and simpler AI applications. Turn on the pg_textsearch extension today, and experience the ultimate hybrid search engine experience with AlloyDB and Cloud SQL.
Want to get started? Check out”
-
AlloyDB resources
-
New to AlloyDB? Discover AlloyDB with a 30-day free trial
-
Cloud SQL resources
-
Azure service updates
- [Launched] Generally Available: New and improved troubleshooting guides for Azure Database for PostgreSQL
[Launched] Generally Available: New and improved troubleshooting guides for Azure Database for PostgreSQL
-
Azure service updates
- [Launched] Generally Available: PG18 support for Azure Database for PostgreSQL elastic clusters
[Launched] Generally Available: PG18 support for Azure Database for PostgreSQL elastic clusters
How a solo founder runs a five-continent tender platform on AlloyDB and MCP
Editor's note: Lucius AI, a tender-intelligence startup covering markets across five continents, runs its entire data platform on AlloyDB for PostgreSQL with a single operator. By migrating semantic search to a ScaNN index and managing database operations through Model Context Protocol (MCP), query latency dropped by 47x while automating day-to-day administrative tasks via MCP.
Executive summary
-
Lucius AI runs a global tender platform spanning more than 210,000 tenders across the UK, EU, India, and Australia, requiring minimal operational overhead for a solo founder.
-
Lucius AI deployed AlloyDB for PostgreSQL to consolidate its relational catalog, audit logs, and vector embeddings into a single managed database engine.
-
Migrating semantic search to a ScaNN index lowered query latency from 1.14 seconds to 24 milliseconds — a 47x speedup on a representative production query.
-
Connecting an AI agent to AlloyDB using the Model Context Protocol (MCP) helps Lucius AI automate query analysis, data freshness checks, and incident forensics under strict least-privilege permissions.
Making tender intelligence work as a company of one
Lucius AI helps businesses bidding on public contracts evaluate opportunities across global markets. The platform ingests public procurement notices from the UK, the EU, the US and Canada, Australia and New Zealand, India and Singapore, alongside World Bank donor-funded notices across Africa and Asia. Lucius AI analyzes tender documents using Gemini to generate compliance matrices, bid recommendations, and draft responses citing original source pages. For small and mid-sized suppliers, this replaces days of manual document reviews and costly external consulting.
Running a platform of this scope requires extensive operational coordination:
-
Nightly ingestion from thirteen public procurement sources
-
A catalog of more than 210,000 tenders, including tens of thousands open for active bidding
-
Two production regions on Cloud Run: Europe, and an Australian deployment on its own AlloyDB cluster with customer-managed encryption keys (CMEK) for defense-adjacent customers
-
Ongoing analytics, performance tuning, data validation, and incident response
Managing these responsibilities without dedicated data engineering or database administration teams requires offloading operational maintenance. Lucius AI addressed this challenge on two fronts: using AlloyDB for PostgreSQL as the core system of record, and connecting an AI agent through the Model Context Protocol (MCP) to safely execute database operations.
Consolidating systems into AlloyDB
Rather than deploying separate relational databases, vector databases, and log stores, Lucius AI houses all core data in AlloyDB for PostgreSQL. The relational tender catalog, document metadata, audit logs, and vector embeddings reside in the same database engine. Storing vector embeddings alongside relational rows avoids managing separate vector stores, establishes a unified backup schedule, and centralizes identity management.
Authentication relies strictly on Cloud IAM. Services connect using dedicated Google Cloud service accounts mapped to database roles scoped to specific access requirements, without storing database passwords in application environments. Database reliability is managed natively by AlloyDB through automated backups and point-in-time recovery, avoiding custom disaster recovery procedures.
In production, this consolidated architecture supports:
-
More than 210,000 tenders in the catalog, with embeddings stored directly alongside them
-
Rebuilding the semantic index embedded 115,820 records in 10.6 minutes with the Gemini embedding model, for around three dollars in API spend; AlloyDB auto embeddings now keep those vectors current.
-
Retrieval reranking executed directly inside the database using the
ai.rankfunction — with mean latency of 77-milliseconds - returning the most relevant results for search queries without requiring a standalone reranking microservice
Accelerating semantic search by 47x
Semantic search across the tender catalog initially relied on unindexed vector comparisons, where a representative query took 1.14 seconds. Migrating this workload to a ScaNN index in AlloyDB reduced query latency to 24 milliseconds — a 47x improvement.
The index recommendation originated from the AI agent during an automated performance audit, where it benchmarked the query plan before preparing the index migration.
Automating database operations with MCP
To delegate routine administrative tasks, Lucius AI configured the open-source MCP Toolbox for Databases using the prebuilt alloydb-postgres server.
Operational delegation requires strict access controls. The agent connects using a dedicated PostgreSQL role granted SELECT across the schema and UPDATE on a single operational table. Destructive commands (DROP, DELETE, TRUNCATE) are omitted, restricting agent actions to authorized operational boundaries.
Under this configuration, the AI agent performs regular database operations across four key areas:
-
On-demand analytics: Compiles retention cohorts, activation funnels, and catalog coverage by country via ad hoc SQL queries, removing the need to build and maintain manual dashboards or complex analytical pipelines.
-
Performance optimization: Performs query-plan inspections and index analysis, such as identifying the ScaNN indexing strategy.
-
Incident forensics: In response to an external security probe, the agent parsed audit logs to reconstruct the request timeline in minutes, verifying that tenant isolation remained intact.
-
Automated data-quality checks: Evaluates ingestion watermarks and freshness across all thirteen procurement sources every morning.
For teams adopting this architecture, establishing a progressive permission structure provides clear guardrails: start with read-only access, expand permissions as requirements dictate, and keep destructive operations restricted to human administrators.
Looking ahead
Lucius AI is planning three technical initiatives to further reduce operational overhead:
-
Automated vector embeddings in AlloyDB AI: After validating
ai.initialize_embeddingsacross the full catalog, a weekly maintenance job usesai.refresh_embeddingsto update vectors. -
Columnar engine acceleration: Having enabled AlloyDB’s columnar engine with auto-columnarization, the database identified and stored 40 frequently queried columns across four tables in memory within a day, accelerating reporting queries without a separate analytical store.
-
Managed Remote MCP Server: Transitioning from self-hosted Toolbox processes to Google Cloud's fully managed Remote MCP Server for AlloyDB will offload MCP server hosting and maintenance.
By anchoring core data in AlloyDB and managing routine operations through MCP, Lucius AI demonstrates how a single engineer can build and operate a resilient, multi-region procurement platform.
To explore Lucius AI, visit ailucius.com. To evaluate AlloyDB for PostgreSQL, deploy an AlloyDB cluster to test performance against your own workloads.
Perplexity’s AI agents helped build a database. They weren’t allowed to run it.
Perplexity decided it was paying too much for DynamoDB and wasn’t getting the control it wanted over read performance. So it built its own database: CobbleDB.
Built by two engineers in two months with help from hundreds of persistent coding agents throughout development, CobbleDB is a roughly 40,000-line Rust key-value store that now handles part of Perplexity’s production search traffic. The company measured median batch-read latency at 5.6 milliseconds after the move, compared with 31.4 ms on DynamoDB before the cutover, while p99 went from 123 ms to 24.2 ms.
It’s expected to cost at least 20% less than DynamoDB and plans to open-source the database at some point.
But the database itself is only part of the story. CMU professor Andy Pavlo argued at Percona Live earlier this year that databases are the hardest and most important challenge for AI agents, in part because mistakes involving production data can be difficult or impossible to reverse.
Perplexity went ahead and used hundreds of agents to help build one anyway, but they weren’t given the keys to production.
It’s expected to cost at least 20% less than DynamoDB and plans to open-source the database at some point.
Why DynamoDB couldn’t keep up
Each search requires the serving layer to retrieve pre-chunked passages and vector embeddings, with a single Search API call fetching 100 to 120 page keys in batches of 10 to 20. Each item averages about 50 KB.
DynamoDB gave Perplexity little control over how it handled reads, which meant a slow replica could hold up the entire things. It also charged for the steady flow of large reads and writes generated by search, crawling, and reprocessing, which made cloud costs difficult to justify as traffic and the corpus grew.
That led Perplexity to separate long-term document storage from the database serving live searches.
Three tiers for search data
The storage stack is split into three pieces. Pillar keeps durable document state in YTsaurus on HDDs, including versioned metadata, chunks and embeddings, while Lorry packages updates into partition-specific batches and moves them through S3 to CobbleDB.
Processed page data is spread across three replicas per partition, with hashed URLs as keys and RocksDB keeping often accessed data in memory while the rest stays on local NVMe. Reads stay within the same availability zone when possible, and the router can try another replica if one is slow rather than hold up the batch.
Updates come through S3 and are applied independently, allowing a replica to fall behind and catch up without blocking the others.
Roughly 5X Lower Batch-Read Latency
Perplexity was handling approximately 200,000 requests per second when it measured CobbleDB at 5.6 ms for a median batch read, down from the 31.4 ms it had recorded on DynamoDB. At p99, latency went from 123 ms to 24.2 ms.
In later load testing, CobbleDB reached 500,000 requests per second before performance started to decline.
The comparison comes with an important caveat; DynamoDB and CobbleDB weren’t tested side by side against identical traffic: the DynamoDB figures were recorded before the cutover, and CobbleDB’s afterward. Perplexity separately ran synthetic benchmarks using batches of 10 to 15 keys with values ranging from 100 bytes to 100 KiB.
Its cost model puts CobbleDB at least 20% below DynamoDB across the commitment options evaluated, though that estimate doesn’t include the engineering cost of supporting the database.
In later load testing, CobbleDB reached 500,000 requests per second before performance started to decline.
Agents built it, engineers controlled it
The agents carried context across sessions, catching problems with restore assumptions and runtime configuration while working on fixes and tests. But they weren’t running the database.
The two engineers kept control of the architecture and production system, particularly important given Pavlo’s warning about putting agents near critical production data.
Ownership has long-term costs
Shipping CobbleDB in eight weeks solved Perplexity’s immediate engineering bottleneck, but maintaining a custom datastore could prove considerably harder. The latency results aren’t from a controlled side-by-side benchmark, and the projected savings don’t include the engineers needed to maintain CobbleDB and respond when something breaks.
Like Shopify and Ramp, which built custom coding agents around third-party models, Perplexity kept the cloud infrastructure but replaced a managed service with something built for its own needs. CobbleDB shows how AI-assisted development is changing that calculation, making custom infrastructure more practical for smaller engineering teams.
CobbleDB shows how AI-assisted development is changing that calculation, making custom infrastructure more practical for smaller engineering teams.
The post Perplexity’s AI agents helped build a database. They weren’t allowed to run it. appeared first on The New Stack.
[In preview] Public Preview: Azure SQL updates for mid-September 2026
-
Cloud Blog
- M4N VM family, now GA: Highest per-core IOPS and throughput for I/O and memory-bound workloads
M4N VM family, now GA: Highest per-core IOPS and throughput for I/O and memory-bound workloads
As enterprise organizations scale mission-critical applications, storage I/O and memory access can become severe operational bottlenecks. Whether its Oracle databases, in-memory databases like SAP HANA, or high-throughput SQL Server clusters, EHR systems, and real-time big data analytics, memory-bound databases often force enterprises to over-provision compute cores (vCPUs) to get the RAM capacity and storage bandwidth they need, driving up costly third-party software licensing fees.
Today, we are thrilled to announce the general availability of the M4N machine series in Google Compute Engine, purpose-built for I/O intensive, high-memory workloads, the second offering in our network- and block-storage optimized VM family. Compared to similar offerings from other hyperscalers M4N provides the highest per-core IOPS and throughput for high-memory instances, and over 20% TCO reduction for Oracle databases.
M4N is also the industry’s first instance of network and block storage optimized with higher memory ratios (up to 26:1) and size (6TB). Powered by 5th Gen Intel® Xeon® Scalable processors and built on Google Cloud's custom Titanium offload architecture, M4N instances deliver up to 25,000 MiB/s (25 GiB/s) of aggregate host storage performance and up to 1 million IOPS when paired with Hyperdisk Extreme — doubling the block storage performance of current M4 instances.
M4N targets workloads that demand both extreme high-density RAM and uncompromising I/O performance, complementing our existing memory-optimized families (such as M1, M2, M3, M4, and X4) by solving specific storage and network bottlenecks for high-throughput enterprise applications.
Built for demanding workloads
|
Workload Category |
Typical Applications |
Why M4N Wins |
|---|---|---|
|
Mission-critical enterprise DBs |
Oracle, SAP HANA, SQL Server, IBM DB2, MySQL, PostgreSQL |
Memory-to-core ratios (up to 26.57 GB/vCPU) paired with 25 GiB/s storage for rapid data ingestion, transaction logging, and zero-stall backup cycles. |
|
Generative AI and RAG data layers |
Milvus, Pinecone, Qdrant, Vespa, Redis, In-Memory Context Caching |
Sub-millisecond similarity search across massive vector indexes in RAM, combined with 400 Gbps network bandwidth for distributed model retrieval. |
|
Enterprise healthcare and ERP |
Epic Systems (Operational Database), SAP ECC, SAP S/4HANA |
Sustained I/O headroom that prevents query latency spikes during peak clinical/transactional hours. |
|
Real-time analytics and EDA |
Electronic Design Automation, Genomic Modeling, In-Memory OLAP |
High memory capacity to load massive datasets entirely in RAM with maximum storage bandwidth for checkpoint dumps. |
Optimizing Oracle licensing costs
Enterprise IT departments struggle with the rising cost of core-based software licensing. For workloads like Oracle database, licensing fees are typically calculated based on the number of vCPUs or physical cores assigned to the instance. Historically, this has forced a difficult trade-off: paying for more compute cores than necessary just to obtain the required amount of RAM and storage performance.
M4N changes this paradigm with its industry-leading high memory-to-vCPU ratio. By providing the highest per-core IOPS and throughput for high-memory instances of all the leading hyperscalers, M4N allows database administrators to:
-
Reduce TCO and licensing overhead: Stop over-provisioning of cores while meeting Oracle database performance density requirements, resulting in over 20% TCO reduction compared to similar offerings from leading hyperscalers.
-
Right-size infrastructure: Allocate the exact amount of compute power needed for the workload while still accessing massive memory pools.
-
Improve cache-hit ratios: With more memory available per core, larger portions of the database can reside in the system global area (SGA), reducing expensive I/O operations and further boosting efficiency.
What customers are saying
Early experiences with M4N show that workload-optimized infrastructure is the engine for transformation.
“Before M4N, meeting our demanding I/O requirements on Google Cloud often required over-provisioning our compute to achieve the necessary performance density. The new M4N instances solve this by delivering high throughput across the smaller to larger shapes.” - Sherri Trojan, Sr Principal Solution Architect, Sabre
"We are delighted to see Google Cloud introduce this next-generation high-performance infrastructure for mission-critical database workloads. The new compute platform demonstrates tremendous potential for enterprise Oracle deployments requiring scalability, resiliency, and performance. We are excited about what this innovation means for customers running Oracle workloads on Google Cloud.” - Bala Kuchibhotla, Co-Founder and CEO, Tessell
"With M4N, Google Cloud continues to push the boundaries of platform co-design. By combining 5th Gen Intel Xeon Scalable processors with Google's custom Titanium offload architecture, M4N delivers the extreme memory capacity, high memory bandwidth, and uncompromising I/O throughput required for the world’s most demanding mission-critical data environments." - Intel
What’s new: Scaling extreme data layers with M4N
M4N bridges two previously separate paradigms in cloud infrastructure: large memory footprints and extreme I/O performance. Engineered with custom Titanium offloads, M4N minimizes I/O bottlenecks without requiring infrastructure add-ons or compromises on memory density. Let’s take a look at how M4N fits into these environments.
1. Enabling high bandwidth data transfer
For workloads with large memory footprints, M4N provides:
-
Superior VM-to-VM bandwidth: Delivers up to 400 Gbps aggregate VM-to-VM network bandwidth and up to 50 Gbps single-flow bandwidth within the same VPC, unlocking non-blocking data exchange for distributed database clusters and real-time streaming data layers.
-
Enhanced internet and egress throughput: Enjoy up to 200 Gbps internet egress bandwidth and up to 48 MPPS packet processing performance.
-
High bandwidth out-of-the-box: Achieve full performance without needing to purchase or configure premium Tier_1 networking add-ons.
2. Dynamic storage performance with Hyperdisk
Paired with Google Cloud's next-generation storage portfolio, M4N with Hyperdisk lets you independently tune IOPS, throughput, and capacity:
-
Hyperdisk Extreme (HdX): Delivers up to 25 GiB/s aggregate block storage throughput and 1,000,000 IOPS—double the storage performance of standard M4. This is great for rapid database recovery, transactional checkpointing, and instant in-memory index reloads.
-
Hyperdisk Balanced (HdB): Scales up to 20 GiB/s throughput and 640,000 IOPS for cost-effective enterprise storage at scale.
M4N machine types and specifications
M4N instances are offered across three distinct memory-to-vCPU ratio tiers, scaling from 16 to 224 vCPUs and up to 5,952 GB of DDR5 RAM. M4N also offers predefined VM shapes across three distinct memory-to-vCPU ratios to match specific workload requirements, with support for Resource-based Committed Use Discounts (CUDs). Details here.
Get started today
The M4N instances are now available in select regions around the globe. To learn more about how the M4N family can enhance your memory- and I/O-bound applications and reduce your licensing costs, contact your account representative or explore the documentation.

-
Azure service updates
- [In preview] Public Preview: PostgreSQL skills and MCP plugin for Azure Database for PostgreSQL
[In preview] Public Preview: PostgreSQL skills and MCP plugin for Azure Database for PostgreSQL
Scaling Telco Autonomy: Leveraging GNNs with Distributed GraphFlow
The telecommunications industry is currently undergoing a paradigm shift, moving from traditional manual human-driven operations to fully Autonomous Network Operations. Modern networks have grown increasingly complex, heterogeneous, and large-scale, making handcrafted rules-based methods and traditional Machine Learning (ML) approaches alone insufficient to automate network operations. While ML methods can identify subtle patterns and make fine predictions from large amounts of structured data, they lack the ability to understand, reason about the data and the system it represents, and ultimately make the kind of decision a human operator would.
The growth of AI agents and their ability to reason is a promising solution to this shortcoming. However, in the same way a human operator is not capable of directly ingesting the statistical information spread across the billions of data points created in a large network, AI agents also lack the ability to operate at this scale. To address this challenge, telecommunications companies are adopting Graph Neural Networks (GNNs), a modern form of machine learning designed to operate natively on massive volumes of temporal and relational data. By integrating GNNs with AI agents, operators can combine advanced diagnostics such as root cause analysis, capacity planning, traffic forecasting, what-if simulations, and real-time anomaly detection with the reasoning power required to interpret these insights and execute justified actions. This powerful combination enables networks to safely move towards Level 5 Autonomy as defined by TM Forum, where the system operates autonomously.
In this post, we present the three components (Data, ML, and AI) that will power Google Cloud’s Autonomous Network Operations framework.
Google Autonomous Network Operations framework architecture
Foundation: Digital Twin on Spanner Graph
At the heart of Google Cloud’s Autonomous Network Operations framework is the network digital twin: a highly detailed, virtual replica that continuously mirrors its living telecommunications network in real time. Rather than being a static model, it is represented as a dynamic, temporal network graph that captures the evolving state and relations of its components over time. This architectural approach allows operators to "go back" in time to train and evaluate ML models on historical data, while providing AI agents with the foundational operational knowledge required to achieve Level 5 Autonomy. By simulating the impact of proposed network changes within this digital environment, the Digital Twin establishes a critical layer of trust, enabling AI agents to confidently design future states and automatically resolve network issues.
Google Cloud’s Spanner Graph is well suited to host this digital twin:
-
Scalability and Availability: Spanner Graph provides a no compromise foundation for modern applications, offering virtually unlimited scaling that grows as the network grows, along with 0-RPO/0-RTO and five 9s of availability.
-
Multi-Model Support: Supports multiple data models (Relational, Graph, Vector, and Full-Text Search) in a single platform allowing developers to build complex compositions such as graph transversals combined with nearest neighbor vector search.
-
Global Consistency: Spanner provides a globally consistent view of the network, simplifying system development.
The next figure illustrates a network topology with four node types: routers, interfaces (the physical ports), VPNs (L3VPN service instances), and flows (active traffic sessions). These are connected by directed edge types capturing the full network stack: physical containment (router-interface), physical links (interface-interface), control-plane peering (router-router via OSPF/iBGP), service membership (router-VPN), and traffic anchoring (flow-interface, flow-VPN).
High Level network topology
The ML layer: Distributed Graph Flow (DGF)
To predict how a network will behave and react, the digital twin leverages an ML layer powered by Distributed Graph Flow (DGF). By training on the vast volumes of structured historical data hosted within Spanner Graph, this layer uncovers critical predictive insights that enable human operators and AI agents to manage networks proactively rather than reactively.
DGF is a recently open-sourced Python library designed to manage the entire end-to-end lifecycle of GNN modeling. Developed by Google CoreML and Google Research, it brings a decade of internal Google-scale tools and expertise directly to Google Cloud enterprise clients. To accommodate different engineering needs, the library offers high-performance, composable, low-level primitives for advanced teams, alongside a simple API for rapid development that requires no prior GNN expertise.
For instance, training and evaluate a GNN model in GraphFlow with the high level API can be as simple as writing 5 lines of code:
- code_block
- <ListValue: [StructValue([('code', 'import dgf\r\n\r\n# Fetch the data from Spanner Graph\r\ngraph, schema = dgf.io.read_spanner_graph(...)\r\n\r\n# Train a node attribute prediction model\r\nmodel = dgf.learning.train_node_model(graph, schema, target_column="risk_score")\r\n\r\n# Evaluate the model\r\nmodel.evaluate()\r\n# Make predictions\r\nmodel.predict(graph, seed_node_idxs=[0, 1, 2])\r\n\r\n# Save the model for later\r\nmodel.save("/tmp/model")'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fce29db6350>)])]>
The DGF provides high-level concepts that map directly to Autonomous Network Operations requirements:
Use cases
By leveraging DGF and GNNs, telcos can move from reactive maintenance to proactive prevention through several advanced use cases:
-
Anomaly detection: GNNs generate node and edge embeddings that encapsulate historical patterns and current health. Any anomalous embeddings are flagged for review before they lead to service degradation.
-
Root cause analysis (RCA): DGF can output specific subgraphs containing only the relevant network instances related to an incident, such as "Attach Failures" in a specific ZIP code. This allows troubleshooting agents to perform high-speed analysis without scanning the entire global network.
-
Predictive maintenance: The system can predict the likelihood of device failures or edge breaks, such as "handover failures" for fast-moving equipment, enabling proactive load balancing or rerouting. Furthermore, by combining agents, remedial actions can be automated by adopting a ‘human-on-the-loop’/’human-in-the-loop’.
-
What-if analysis: GNNs enable Telcos to simulate scenarios like fiber cuts, or traffic surges or device configuration changes. By modeling topological dependencies, GNNs can predict how these local changes propagate across the entire network, allowing engineers to test resilience and evaluate mitigation strategies in a risk-free digital environment.
Scenario: Root cause analysis with GNNs and DGF
Once you have created a digital twin (example code), a straight-forward 5-step process can be used to implement Root Cause Analysis(RCA) detection using GNNs and DGF.
-
Connect to the Digital Twin: Use the DGF Spanner Graph connector (dgf.io.read_spanner_graph) to load the network topology directly from Spanner Graph's Digital Twin into the DGF environment.
-
Train a Supervised Node (or Edge) Prediction model: Depending on the training data and objective, you will train a supervised node prediction model to predict a target node feature or an edge prediction model to predict an edge between the root cause entity node and the affected entity node. For the given sample data you will use the high-level
dgf.learning.train_node_modelAPI to train a supervised node prediction model. -
Use the node prediction model to predict root cause node: The node prediction model can be directly used to predict the impact score on the node with the anomaly. Entity nodes affected by the anomaly with highest predicted impact score will be the top candidates for root cause.
-
Deploy to Gemini Enterprise Agent Platform (formerly Vertex AI): Export the model and host it on a Gemini Enterprise endpoint to enable scalable, low-latency predictions.
-
Real-time Inference: Make prediction calls to the inference endpoint with the anomaly date as input. The endpoint will return the predicted root cause Entity nodes.
Get started today
The integration of GNN using Distributed Graph Flow into network operations is more than just a technical upgrade; it is a critical evolution for the telco industry. By moving towards a GNN-powered autonomous framework, operators can significantly shorten outage times, optimize capacity in real-time, and ultimately deliver a superior customer experience through improved operational efficiency.
To start building your own intelligent network applications, check out the Distributed GraphFlow (DGF) library, which provides the essential primitives for scalable GNN training and inference. For a hands-on experience, follow our step-by-step code sample. You can also explore our recent award-winning Moonshot project on Business-aware GNN-healing networks, and dive deeper into our approach on self-optimizing autonomous networks by reviewing this whitepaper.
Kubernetes 1.36 restores a lost guarantee for database backups
It’s 2 a.m., and you’re restoring a PostgreSQL cluster from last night’s backup. Its data directory lives on one PersistentVolumeClaim and its write-ahead log on another, a common split for I/O isolation. Every volume snapshot reported success. The pods come back. Then Postgres refuses to start because the WAL on one volume references pages that were never captured in the data files on the other. The backup wasn’t corrupted in transit. It was inconsistent the moment it was taken.
If you run stateful workloads on Kubernetes, this failure mode has been silently lurking in your backups for years. It has nothing to do with your backup tool crashing, and everything to do with a guarantee you gave up when you moved off traditional storage.
The consistency group you lost on the way to Kubernetes
Enterprise storage arrays solved this problem decades ago with a feature called a consistency group. You told the array which LUNs belonged to the same application, and when you snapshotted the group, the array froze them all at the same instant. Every volume captured the same point in time. Restores were coherent by construction.
“The backup wasn’t corrupted in transit. It was inconsistent the moment it was taken.”
That guarantee didn’t survive the move to cloud native. The Container Storage Interface (CSI) standardized snapshots around a single object, the VolumeSnapshot, scoped to a single PersistentVolumeClaim (PVC). One PVC, one snapshot. For a stateless service with one volume, that model is fine. For anything that spreads its state across multiple volumes (a database with separate data and log disks, a sharded datastore, most real applications), the per-PVC model can’t tell which volumes belong together.
As teams migrated off proprietary SANs and virtualization stacks onto Kubernetes-native storage, they gained enormous flexibility but lost the consistency group. Most never notice since the gap only shows up at restore time after an incident, when it is far too late to do anything about it.

How individual PVC snapshots break
When protecting a multi-volume application, backup tools enumerate the PVCs and issue a VolumeSnapshot for each one, in sequence. Snapshot volume A, volume B, then volume C.
Each snapshot is individually crash-consistent, equivalent to pulling the power cord on that one volume. But they are not consistent with one another. Between snapshotting A and snapshotting B, the application keeps writing. A transaction can land in the log on volume B that references data never captured on volume A, because A was frozen a few hundred milliseconds earlier. The busier the application and the more volumes involved, the wider the inconsistency window.
“The result is a set of snapshots that each looks healthy and collectively describes a state that never existed.”
The result is a set of snapshots that each looks healthy and collectively describes a state that never existed. You can quiesce the application to close the window (freeze I/O, flush buffers, snapshot, unfreeze), but at production scale, freezing a busy database for the duration of a multi-volume snapshot is exactly the disruption backups are supposed to avoid.
VolumeGroupSnapshot: consistency groups as a Kubernetes API
VolumeGroupSnapshot is the missing primitive, and as of Kubernetes v1.36 (May 2026), it is generally available. It brings the consistency group back as a first-class, vendor-neutral Kubernetes API rather than a proprietary array feature.
The model has three objects. A VolumeGroupSnapshotClass, defined by an administrator, describes how group snapshots are created for a given CSI driver. A VolumeGroupSnapshot is the user’s request, and it carries a label selector that picks out every PVC belonging to the application. A VolumeGroupSnapshotContent tracks the provisioned result. Under the hood, the CSI driver takes one atomic, point-in-time snapshot across every selected volume: a real consistency group, with no application quiescence required, provided the underlying storage supports it.
“Under the hood, the CSI driver takes one atomic, point-in-time snapshot across every selected volume.”
The label selector is the important design choice. You don’t enumerate volumes; you describe them. A selector like `app=postgres` picks up the data and logs PVCs together, and the group boundary is expressed in Kubernetes terms that survive adding or resizing volumes over time.
Wiring it into backup: what changes
An API that produces consistent snapshots only helps if your backup workflow uses it. When I implemented VolumeGroupSnapshot support in Velero, the CNCF project that has become the de facto standard for Kubernetes backup and restore, the core change was replacing “iterate over PVCs and snapshot each” with “group the PVCs that belong together, snapshot the group as one operation, then track the per-volume members for restore.”
That last part matters. A group snapshot fans back out into individual volume snapshots, one per member, so restore still rehydrates each PVC independently, but now every member shares a single point in time. Velero was among the first backup projects to build directly on the upstream VolumeGroupSnapshot API, rather than a proprietary grouping scheme, so the consistency guarantee rides on a standard the whole ecosystem shares instead of a format locked to one tool.
Individual snapshots vs. group snapshots: which to use
This is not a wholesale replacement. Individual PVC snapshots remain the right tool for single-volume workloads and for volumes that are genuinely independent, since snapshotting those as a group buys you nothing and adds coordination overhead. Reach for VolumeGroupSnapshot when correctness depends on multiple volumes sharing a point in time.
A quick decision guide:
- One volume, or several fully independent volumes: individual VolumeSnapshots.
- Multiple volumes with cross-volume write ordering (data plus WAL, data plus index): VolumeGroupSnapshot.
- Unsure whether a skewed or partial restore would corrupt the application? Treat it as a group.
Day 2 notes
A few things to check before relying on this in production. Group snapshot support is per CSI driver: the API is standard, but the driver has to implement it, and adoption is still spreading. A growing set of CSI drivers implement it (Ceph CSI among them), so check the driver’s release notes for upstream VolumeGroupSnapshot support, and create the VolumeGroupSnapshotClass before needing it. The atomicity guarantee is only as strong as the storage backend behind the driver, so validate it by restoring, not by reading success statuses. Audit existing backups now: by protecting multi-volume applications with per-PVC snapshots today, you likely have inconsistent restore points that have never been tested under a real failure.
“With VolumeGroupSnapshot now GA and backup tooling adopting it upstream, multi-volume backups on Kubernetes are finally consistent by construction, not by luck.”
The broader arc is that Kubernetes storage is catching up to what enterprise arrays offered for years, but as an open standard rather than a capability locked to one vendor’s hardware. Consistency groups were one of the last missing pieces. With VolumeGroupSnapshot now GA and backup tooling adopting it upstream, multi-volume backups on Kubernetes are finally consistent by construction, not by luck.
The post Kubernetes 1.36 restores a lost guarantee for database backups appeared first on The New Stack.
How to migrate from Apache HBase to Cloud Bigtable with Live Migrations
Cloud Bigtable is a natural destination for Apache HBase workloads, as it is a fully managed service that is compatible with the HBase API. As a result, many customers running business-critical applications with large-scale data and low-latency needs consider migrating to Bigtable.
However, migrating from HBase to Bigtable can still be challenging since you typically have to pause your applications for migration downtime. In addition, some companies choose to write custom tools, which require extensive resources to build and test, adding months to the migration process.
Today, we’re announcing that Live Migrations from Apache HBase to Cloud Bigtable are now generally available. This enables faster and simpler migrations from HBase to Bigtable to ensure accurate data migration, reduce migration effort, and provide a better overall developer experience.
HBase to Bigtable migrations just got easier
Historically, you would need to manually create tables in Bigtable from your existing HBase tables and execute several steps to export and import data, define target tables, and validate data integrity. This process can be tedious, especially if the migration requires moving multiple tables or pre-splitting tables.
At Google Cloud, we’re always trying to find ways to make migrations from HBase to Bigtable even easier for our customers. Our latest Live Migration features aim to provide a more straightforward, more efficient, and proven way to migrate data from HBase to Bigtable with minimal downtime. All together, they provide the necessary components to complete a seamless live migration.
We have built four new features:
Schema Translation Tool automates table schema conversions.
HBase Bigtable Replication Library minimizes downtime for live migrations.
Snapshot Import Tool easily imports HBase snapshots into Cloud Bigtable.
Migration Validation Tool ensures accurate data migration.
Now, you can automate the migration process and facilitate end-to-end data pipelines. The Schema Translation Tool fully automates table conversion by connecting to HBase, copying the table schema, and creating similar tables in Bigtable. You can also import HBase snapshots and validate data migration for a more seamless migration process with our Snapshot Import and Migration Validation tools.
The HBase Bigtable Replication Library, which becomes available today, removes the need for building custom migration tools. It allows you to use HBase replication to sequence bulk imports and live writes correctly, ensuring consistent performance during migration of large workloads.
How live migrations from HBase to Bigtable works
HBase provides asynchronous replication between clusters for various use cases like disaster recovery and data aggregation workloads. The HBase Bigtable Replication Library enables Bigtable to be added as an HBase cluster replication target. HBase to Bigtable replication enables customers to sync mutations happening on their HBase cluster to Bigtable, providing near-zero downtime migrations from HBase to Cloud Bigtable.
The following diagram shows a live replication from HBase to Bigtable:
The HBase Cluster is the source database, which can be located in an on-premises network, another cloud provider, or managed data services. Once enabled, live replication allows all the writes happening on the source cluster to be replicated to the target Bigtable Instance.
Before enabling replication, you will need to create all the tables from HBase with the same column families in Bigtable. You can use the Schema Translation Tool to create target tables in Bigtable based on your existing HBase schema. To enable replication, the source cluster must be able to connect to the target Bigtable instance.
Get started with HBase to Bigtable live migrations
To learn more about HBase to Bigtable Live Migrations and how to get started, please visit our documentation page.
To learn more about Bigtable:
Create an instance or try it out with a Bigtable Qwiklab.
Check out these Youtube video tutorials for a step-by-step introduction to how Bigtable can be used for real-world applications like Personalization and Fraud detection.

Accelerate your move to the cloud with the new Database Migration Program
Today, we’re announcing the Database Migration Program, a new and stress-free approach to migrating existing open source and proprietary databases, whether on-premises or in the cloud, to Google Cloud’s industry-leading, managed database services. With the Database Migration Program, you benefit from assessments, tooling, best practices, and resources from our network of specialized database technology partners. The program also offers special incentive funding to offset migration costs, helping you to quickly and cost-effectively migrate your databases to Google Cloud. Get started today with the Database Migration Program.
Over the past decade, companies big and small have realized the benefits of the cloud for their application modernization journey, helping them become more efficient, scalable, agile, and innovative. Furthermore, managed cloud databases typically result in an overall lower cost of ownership while upskilling database administrators to focus on higher-value work like data modeling and deriving additional value from data with AI and machine learning.
Since modernizing to GKE, Istio and Cloud SQL, Auto Trader’s release cadence has improved by over 140% (year over year), enabling an impressive peak of 458 releases to production in a single day. Auto Trader’s fast-paced delivery platform managed over 36,000 releases in a year with an improved success rate of 99.87%, and it continues to growMohsin Patel
Principal Database Engineer, Auto Trader UK
Still, many companies continue to self-manage databases on cloud instances or leave databases on premises even when the application is running in the cloud. The primary reason is the complexity of database migrations. Databases are at the core of every enterprise’s day-to-day operations, making them more challenging to move without careful planning and execution. In addition, migrations can be expensive, time-consuming, and risky. Timelines can drag on and scope regularly increases, leaving customers frustrated.
Our new Database Migration Program seeks to address database migration complexity by providing comprehensive guidance and support for your migrations. Our assessments help you understand the footprint of your database fleet, its dependencies and architecture, and our specialized database partners can help with their expert knowledge of tooling and resources to move data and code without disrupting your business. Additionally, Google Cloud offers special incentive funding to offset migration costs, helping you to quickly and cost-effectively migrate your databases with the minimum amount of financial risk.
The secret to stress-free database migrations
What’s unique about this program is that you have access to a one-stop shop for all things database migrations. You can break your migrations into smaller sprints and execute one migration after another, allowing you to achieve business outcomes faster. With the Database Migration Program, you can accelerate your move from on-premises, other clouds, or self-managed databases over to Cloud SQL, Cloud Spanner, Memorystore, Firestore, and Cloud Bigtable.
Already, the Database Migration Program is transforming the way our customers and partners approach their database migrations to the cloud, allowing them to reimagine the time and resources required to deliver on their digital transformation goals without the burden of uncertain timelines and high costs.
Google Cloud’s new Database Migration Program provides a streamlined approach to seamlessly and efficiently migrate on-premise or in-cloud databases to Google’s industry-leading managed databases. This innovative program helps customers fast-track their database migration by leveraging Google Cloud’s assessments, tools, best practices, and resources.Shiwanand Pathak
Global Practice Head of Data & AI Services, Google Cloud Business, Tata Consultancy Services
Cloud and digital transformation continues to shape the strategic agenda for our clients. Data estate modernization is a key enabler for this transformation journey and clients that choose Google Cloud products typically utilize Cloud SQL for operational application databases and BigQuery for analytics. We collaborate with Google Cloud and provide strategy, implementation and operate services that enable our clients to achieve tangible business outcomes from their transformation journey using Google Cloud products.Navin Warerkar
Managing Director, US Google Cloud Data & Analytics GTM Leader, Deloitte Consulting
Three steps for a successful database migration
The Database Migration Program guides you from the initial assessment and planning phase to eventual migration with the expert assistance of qualified database partners.
Here’s how Google Cloud helps at each stage of the database migration journey:
Assess: Request a database assessment to discover and analyze your existing databases and applications. Leverage specialized tools and resources, along with assistance from Google Cloud database experts who provide guidance based on your specific needs and requirements.
Plan: Connect with specialized database partners who can help you create a migration plan, including engineering resources and cost estimates, and identify the right workload to kick off your migration.
Execute: Get special incentive funding to offset some of your migration costs by helping to pay for the specialist technology partner who performs your migration. There’s no need to move everything over at once—you can move one department or database at a time and use the program again as many times as you need.
Interested in learning more? Complete this form to get started.

-
Cloud Blog
- Modernize your Oracle workloads to PostgreSQL with Database Migration Service, now in preview
Modernize your Oracle workloads to PostgreSQL with Database Migration Service, now in preview
Many organizations have been struggling with the complexity of their legacy databases. Unfortunately, they often find themselves locked into expensive licenses and restrictive contracts, which can limit their ability to modernize and introduce new functionality. Migrating to open-source databases, especially in the cloud, can solve many of these issues and help build modern, scalable, cost-effective applications.
However, database migrations are often highly complex and may require you to convert your schema and code to the new database engine, migrate your data, and switch over your applications, all while guaranteeing minimal downtime and disruption to the business.
Last year, we announced the general availability of Database Migration Service in our mission to help migrate your databases to the cloud with a simple and secure migration path. We launched support for homogeneous migrations, where the source and target databases use the same database engine (PostgreSQL, MySQL, or SQL Server). We saw adoption by customers migrating their workloads to Cloud SQL, Google Cloud’s fully managed relational database for PostgreSQL, MySQL, and SQL Server. More than 85% of the migrations using Database Migration Service are created and started underway in less than an hour.
Announcing Oracle to PostgreSQL support
Our customers shared that they’d like a similarly simple, easy-to-use experience for Oracle to PostgreSQL migrations. Today, we’re excited to announce the preview of Database Migration Service support for Oracle to PostgreSQL schema and data migrations.
Database Migration Service can integrate with the Ora2Pg open-source tool for schema conversion so you can migrate the schema and data of your Oracle workload from on-premises or other clouds to Cloud SQL for PostgreSQL. Ora2pg allows us to map the source to the target, and then our serverless change data capture-based mechanism can move your data securely and with minimal downtime. Database Migration Service can make database migrations fast, cost-effective, and reliable, and you can now use it to modernize from legacy databases to fully managed cloud databases.
Database Migration Service has you covered
Adopting a new database technology might appear to be a challenging task at first, but we can make the migration journey easier. We understand that effective and successful modernization can require a well-rounded approach: not only differentiated tooling but also integrated support and expert services you can trust.
Database Migration Service is highly reliable and serverless, meaning you don’t need to assign resources to the migration job or predict how many resources it will need. It can move your data from Oracle databases to Cloud SQL for PostgreSQL at scale and with low latency, which can mean minimal downtime at switchover and minimal disruption to your applications and customers.
Our integration with the proven Ora2Pg tool for schema migration means you can convert your Oracle schema to PostgreSQL with this popular open-source tool. You simply feed the configuration file after configuring, converting, and applying your converted schema with Ora2pg. Database Migration Service then creates the mapping and moves the data between the source and the target. Stay tuned for enhanced, built-in schema and code conversion capabilities in DMS to create an upgraded schema, code, and data migration experience.
“At MLB, we’re on a multi-year journey to modernize our applications with PostgreSQL as the database foundation,” says Shawn O’Rourke, manager of technology at MLB. “A key step in this journey is to reliably migrate our Oracle databases to Cloud SQL for PostgreSQL securely and without any disruption to our services. We’re excited to incorporate Database Migration Service, with its simple, serverless design, into our Oracle migration toolset.”
Expert services to help accelerate your migration
By working closely with experts from Google Professional Services and experienced migration partners, we help make sure you have access to the expertise and experience you need to facilitate successful migrations across your database fleet. From guidance on migration planning to turnkey end-to-end migrations, the combination of Database Migration Service and partner services can ensure a smooth transition to the cloud.
“We see tremendous demand from our customers for migrating away from proprietary databases onto cloud database technologies”, says David Yahalom, Managing Principal, Cloud Data Solutions at EPAM Systems. “One of the key success factors in application modernization is real-time continuous data replication within a heterogeneous database environment. Real-time data replication enables near-zero and zero downtime production switchovers while maintaining data integrity. We are very excited about the addition of Oracle to PostgreSQL migration support in Database Migration Service and believe it will be of great value to our customers. It will enable us to streamline database cloud migration initiatives to Google Cloud.”
Getting started with Database Migration Service
You can start migrating your Oracle workloads today using Database Migration Service:
Navigate to the Database Migration area of your Google Cloud console, under Databases, and click Create Conversion Workspace.
Use the Conversion Workspace creation wizard to upload your Ora2PG configuration file.
Create your source and destination connection profiles. You can use this profile again later for additional migrations.
Create a migration job to connect the Cloud SQL destination Connection Profile, Oracle Connection Profile, and Conversion Workspace.
Test your migration job and make sure the test was successful as displayed below, and start it whenever you're ready.
Once the initial snapshot of data has been migrated to the new destination, Database Migration Service will keep up and replicate new changes as they happen. You can then finalize the migration job, and your new Cloud SQL instance will be ready to go. You can monitor your migration jobs on the migration jobs list, as shown in the image below:
Learn more and start your database journey
Database Migration Service schema and data migration from Oracle to Cloud SQL for PostgreSQL are available in preview in addition to the previously-announced SQL Server migration preview. If you’re interested in seeing it in action, you can request access now.
For more information to help get you started on your migration journey, head over to the documentation or start training with this Database Migration Service Qwiklab.

Boost the power of your transactional data with Cloud Spanner change streams
Data is one of the most valuable assets in today’s digital economy. One way to unlock the value of your data is to give it life after it’s first collected. A transactional database, like Cloud Spanner, captures incremental changes to your data in real time, at scale, so you can leverage it in more powerful ways. Cloud Spanner is our fully managed relational database that offers near unlimited scale, strong consistency, and industry-leading high availability of up to 99.999%.
The traditional way for downstream systems to use incremental data that’s been captured in a transactional database is through change data capture (CDC), which allows you to trigger behavior based on changes to your database, such as a deleted account or an updated inventory count.
Today, we are announcing Spanner change streams, coming soon, that lets you capture change data from Spanner databases and easily integrate it with other systems to unlock new value.
Change streams for Spanner goes above and beyond the traditional CDC capabilities of tracking inserts, updates, and deletes. Change streams are highly flexible and configurable, letting you track changes on exact tables and columns or across an entire database. You can replicate changes from Spanner to BigQuery for real-time analytics, trigger downstream application behavior using Pub/Sub, and store changes in Google Cloud Storage (GCS) for compliance. This ensures you have the freshest data to optimize business outcomes.
Change streams provides a wide range of options to integrate change data with other Google Cloud services and partner applications through turnkey connectors, including custom Dataflow processing pipelines or the change streams read API.
Spanner consistently processes over 1.2 billion requests per second. Since change streams are built right into Spanner, you not only get industry-leading availability and global scale—you also don’t have to spin up any additional resources. The same IAM permissions that already protect your Spanner databases can be used to access change streams queries.Change stream queries are protected by spanner.databases.select, and change stream DDL operations are protected by spanner.databases.updateDdl.
Change streams in action
In this section, we’ll look at how to set up a change stream that sends change data from Spanner to an analytic data warehouse in BigQuery.
Creating a change stream
As discussed above, a change stream tracks changes on an entire database, a set of tables, or a set of columns in a database. Each change stream can have a retention period of anywhere from one day to seven days, and you can set up multiple change streams to track exactly what you need for your specific business objectives.
First, we’ll create a change stream on a table called InventoryLedger. This table tracks inventory changes on two columns: InventoryLedgerProductSku and InventoryLedgerChangedUnits with a 7-day retention period.
Change records
Each change record contains a wealth of information, including primary key, the commit timestamp, transaction ID, and of course, the old and new values of the changed data, wherever applicable. This makes it easy to process change records as an entire transaction, in sequence based on their commit timestamp, or individually as they arrive, depending on your business needs.
Back to the inventory example, now that we’ve created a change stream on the InventoryLedger table, all inserts, updates, and deletes on this table will be published to the InventoryStream change stream. These changes are strongly consistent with the commits on the InventoryLedger table: When a transaction commit succeeds, the relevant changes will automatically persist in the change stream. You never have to worry about missing a change record.
Processing a change stream
There are numerous ways that you can process change streams depending on the use case:
Analytics: You can send the change records to BigQuery, either as a set of change logs or by updating the tables.
Event triggering: You can send change logs to Pub/Sub for further processing by downstream systems.
Compliance: You can retain the change log to Google Cloud Storage for archiving purposes.
The easiest way to process change stream data is to use our Spanner connector for Dataflow, where you can take advantage of Dataflow’s built-in pipelines to BigQuery, Pub/Sub, and Google Cloud Storage. The diagram below shows a Dataflow pipeline that processes this change stream and imports change data directly into BigQuery.
Alternatively, you can build a custom Dataflow pipeline to process change data with Apache Beam. In this case, we provide a Dataflow connector that outputs change data as an Apache Beam PCollection of DataChangeRecord objects.
For even more flexibility, you can use the underlying change streams query API. The query API is a powerful interface that lets you read directly from a change stream to implement your own connector and stream changes to the pipeline of your choice. On the query API side, a change stream is divided into multiple partitions, which can be used to query a change stream in parallel for higher throughput. Spanner dynamically creates these partitions based on load and size. Partitions are associated with a Spanner database split, allowing change streams to scale as effortlessly as the rest of Spanner.
Get started with change streams
With change streams, your Spanner data follows you wherever you need it, whether that’s for analytics with BigQuery, for triggering events in downstream applications, or for compliance and archiving. Change streams are highly flexible and configurable —allowing you to capture change data for the exact data you care about, and for the exact period of time that matters for your business. And because change streams are built into Spanner, there’s no software to install, and you get external consistency, high scale, and up to 99.999% availability.
There’s no extra charge for using change streams, and you’ll pay only for extra compute and storage of the change data at the regular Spanner rates.
To get started with Spanner, create an instance, or try it out with a Spanner Qwiklab.
We’re excited to see how Spanner change streams will help you unlock more value out of your data!
