❌

Vue normale

Reçu avant avant-hierInfra

Kubernetes 1.36 restores a lost guarantee for database backups

15 septembre 2026 à 15:00
Abstract digital representation of multi-volume system architecture and structural alignment for Kubernetes database storage.

It’s 2 a.m., and you’re restoring a PostgreSQL cluster from last night’s backup. Its data directory lives on one PersistentVolumeClaim and its write-ahead log on another, a common split for I/O isolation. Every volume snapshot reported success. The pods come back. Then Postgres refuses to start because the WAL on one volume references pages that were never captured in the data files on the other. The backup wasn’t corrupted in transit. It was inconsistent the moment it was taken.

If you run stateful workloads on Kubernetes, this failure mode has been silently lurking in your backups for years. It has nothing to do with your backup tool crashing, and everything to do with a guarantee you gave up when you moved off traditional storage.

The consistency group you lost on the way to Kubernetes

Enterprise storage arrays solved this problem decades ago with a feature called a consistency group. You told the array which LUNs belonged to the same application, and when you snapshotted the group, the array froze them all at the same instant. Every volume captured the same point in time. Restores were coherent by construction.

“The backup wasn’t corrupted in transit. It was inconsistent the moment it was taken.”

That guarantee didn’t survive the move to cloud native. The Container Storage Interface (CSI) standardized snapshots around a single object, the VolumeSnapshot, scoped to a single PersistentVolumeClaim (PVC). One PVC, one snapshot. For a stateless service with one volume, that model is fine. For anything that spreads its state across multiple volumes (a database with separate data and log disks, a sharded datastore, most real applications), the per-PVC model can’t tell which volumes belong together.

As teams migrated off proprietary SANs and virtualization stacks onto Kubernetes-native storage, they gained enormous flexibility but lost the consistency group. Most never notice since the gap only shows up at restore time after an incident, when it is far too late to do anything about it.

Charts showing multi-volume backup: individual snapshots vs VolumeGroupSnapshot

How individual PVC snapshots break

When protecting a multi-volume application, backup tools enumerate the PVCs and issue a VolumeSnapshot for each one, in sequence. Snapshot volume A, volume B, then volume C.

Each snapshot is individually crash-consistent, equivalent to pulling the power cord on that one volume. But they are not consistent with one another. Between snapshotting A and snapshotting B, the application keeps writing. A transaction can land in the log on volume B that references data never captured on volume A, because A was frozen a few hundred milliseconds earlier. The busier the application and the more volumes involved, the wider the inconsistency window.

“The result is a set of snapshots that each looks healthy and collectively describes a state that never existed.”

The result is a set of snapshots that each looks healthy and collectively describes a state that never existed. You can quiesce the application to close the window (freeze I/O, flush buffers, snapshot, unfreeze), but at production scale, freezing a busy database for the duration of a multi-volume snapshot is exactly the disruption backups are supposed to avoid.

VolumeGroupSnapshot: consistency groups as a Kubernetes API

VolumeGroupSnapshot is the missing primitive, and as of Kubernetes v1.36 (May 2026), it is generally available. It brings the consistency group back as a first-class, vendor-neutral Kubernetes API rather than a proprietary array feature.

The model has three objects. A VolumeGroupSnapshotClass, defined by an administrator, describes how group snapshots are created for a given CSI driver. A VolumeGroupSnapshot is the user’s request, and it carries a label selector that picks out every PVC belonging to the application. A VolumeGroupSnapshotContent tracks the provisioned result. Under the hood, the CSI driver takes one atomic, point-in-time snapshot across every selected volume: a real consistency group, with no application quiescence required, provided the underlying storage supports it.

“Under the hood, the CSI driver takes one atomic, point-in-time snapshot across every selected volume.”

The label selector is the important design choice. You don’t enumerate volumes; you describe them. A selector like `app=postgres` picks up the data and logs PVCs together, and the group boundary is expressed in Kubernetes terms that survive adding or resizing volumes over time.

Wiring it into backup: what changes

An API that produces consistent snapshots only helps if your backup workflow uses it. When I implemented VolumeGroupSnapshot support in Velero, the CNCF project that has become the de facto standard for Kubernetes backup and restore, the core change was replacing “iterate over PVCs and snapshot each” with “group the PVCs that belong together, snapshot the group as one operation, then track the per-volume members for restore.”

That last part matters. A group snapshot fans back out into individual volume snapshots, one per member, so restore still rehydrates each PVC independently, but now every member shares a single point in time. Velero was among the first backup projects to build directly on the upstream VolumeGroupSnapshot API, rather than a proprietary grouping scheme, so the consistency guarantee rides on a standard the whole ecosystem shares instead of a format locked to one tool.

Individual snapshots vs. group snapshots: which to use

This is not a wholesale replacement. Individual PVC snapshots remain the right tool for single-volume workloads and for volumes that are genuinely independent, since snapshotting those as a group buys you nothing and adds coordination overhead. Reach for VolumeGroupSnapshot when correctness depends on multiple volumes sharing a point in time.

A quick decision guide:

  • One volume, or several fully independent volumes: individual VolumeSnapshots.
  • Multiple volumes with cross-volume write ordering (data plus WAL, data plus index): VolumeGroupSnapshot.
  • Unsure whether a skewed or partial restore would corrupt the application? Treat it as a group.

Day 2 notes

A few things to check before relying on this in production. Group snapshot support is per CSI driver: the API is standard, but the driver has to implement it, and adoption is still spreading. A growing set of CSI drivers implement it (Ceph CSI among them), so check the driver’s release notes for upstream VolumeGroupSnapshot support, and create the VolumeGroupSnapshotClass before needing it. The atomicity guarantee is only as strong as the storage backend behind the driver, so validate it by restoring, not by reading success statuses. Audit existing backups now: by protecting multi-volume applications with per-PVC snapshots today, you likely have inconsistent restore points that have never been tested under a real failure.

“With VolumeGroupSnapshot now GA and backup tooling adopting it upstream, multi-volume backups on Kubernetes are finally consistent by construction, not by luck.”

The broader arc is that Kubernetes storage is catching up to what enterprise arrays offered for years, but as an open standard rather than a capability locked to one vendor’s hardware. Consistency groups were one of the last missing pieces. With VolumeGroupSnapshot now GA and backup tooling adopting it upstream, multi-volume backups on Kubernetes are finally consistent by construction, not by luck.

The post Kubernetes 1.36 restores a lost guarantee for database backups appeared first on The New Stack.

When AI agent traces become application data

26 août 2026 à 15:00
Abstract dark digital visualization with glowing red data streams representing AI agent trace telemetry.

Say a test starts failing and a developer hands it to a coding agent. It digs into the relevant files, runs the test suite, changes two files, then runs targeted validation. The task view shows the files it touched, the commands it ran, what results those commands returned, and the diff it landed on.

Before accepting the patch, the developer reviews that activity. A teammate might reopen the same run later to see why the code changed. The team building the agent can compare thousands of runs to see whether a model or prompt update improved test success or just added more tool calls and cost.

That record needs somewhere to live. Developers and reviewers want a durable version of the run whenever they need it. Engineering wants the same execution data aggregated across runs, because the agent’s behavior is nondeterministic and shifts over time.

“For many agentic products, that record turns out to be application data with a telemetry-shaped workload.”

For many agentic products, that record turns out to be application data with a telemetry-shaped workload. That’s the combination that changes the storage decision.

When a trace becomes product data

Not every agent trace counts as application data. An internal diagnostic trace that can be sampled, expired, or discarded is still telemetry. That boundary can exist within the same trace, where raw diagnostic fields remain internal and the fields needed to reconstruct the user’s task move into product state.

The boundary moves once your product has to retrieve, display, or retain a durable execution record. A developer, for example, might need to see which files an agent inspected, or a reviewer might need to check which commands ran and whether the tests passed.

This doesn’t require exposing a model’s private chain of thought. The product can instead render a projection of observable execution, showing model invocations, tool calls, file reads, command results, errors, timing, and state transitions. That record lets users verify the result and decide how much they want to trust it. In some workflows, it also becomes an audit record, which changes its access and retention requirements.

“Internal diagnostic fields don’t automatically belong in the product view just because they came from the same run.”

Once that execution record becomes product state, it must follow your application’s access model. Code, prompts, retrieved documents, tool arguments, and command output may carry tenant or user data, so the application must enforce the same authorization boundaries when storing and retrieving them. Internal diagnostic fields don’t automatically belong in the product view just because they came from the same run.

Why one agent run produces so much data

Pull request volume scales with the number of patches submitted for review, and issue volume scales with the number of development tasks. Both are just counting units of work at the workflow boundary.

Agent traces scale differently, depending on the execution graph within each task. One request to fix a failing test can trigger multiple model calls, file reads, searches, command executions, test retries, and branches before the agent ever proposes a patch. Each of those steps can produce its own span or event.

OpenTelemetry GenAI semantic conventions, still in development, define separate span types for model inference, tool execution, and retrieval. What gets captured depends on the span type and content-capture policy. Inference spans may carry model identifiers, token usage, and opt-in input and output messages, while tool-execution spans may carry opt-in arguments and results.

The span count tracks what the agent actually does inside each unit of work. If you add a new tool, a retry policy, or a branch, the data volume can increase even though the number of completed tasks hasn’t changed.

“If you add a new tool, a retry policy, or a branch, the data volume can increase even though the number of completed tasks hasn’t changed.”

This same fan-out also shows up outside coding agents. In a case study, Laminar reported more than 500,000 browser events per day. A browser-agent session could run for more than 30 minutes and generate hundreds of thousands of DOM diff events. Laminar used those events to reconstruct a video-like replay of what the agent saw.

Whether the agent writes code or navigates a browser, the same data serves two different purposes. One person loads one trace to understand a single run, and the engineering team scans many traces to find patterns. That combination of point retrieval and cohort analysis gives this data its unusual shape.

Why agent trace data behaves like telemetry

Most records in an agent trace are written once rather than updated. A model invocation or tool result describes an event that has already happened. Scores, annotations, and run status may change later, but teams can store those mutable fields separately or record the changes as new events.

The records also carry high-cardinality dimensions such as model version, prompt template, tool name, session ID, user ID, and outcome. Their value only shows up in context. On its own, an isolated tool-call span doesn’t say much, but a full trajectory can explain a failed run, while a cohort can reveal a regression.

The same dataset therefore serves several distinct readers:

ConsumerRead patternExample
Product interfacePoint lookupLoad one coding-agent run for a developer or reviewer
Evaluation pipelineCohort scanCompare test success, latency, and cost across agent versions
Platform teamTime-window aggregationGroup errors and latency by model, tool or deployment

A primary application database can serve all three patterns at modest scale, but that changes when wide scans and high-cardinality aggregations start competing with product reads and writes on the critical path.

Outgrowing the primary database

There is no universal event-count threshold for moving traces out of the primary database. The decision shows up in the workload.

The first signal is contention, where ingestion or retention work starts consuming enough I/O and CPU to affect transactional operations. Next comes analytical friction, where evaluations and debugging queries need to scan long time ranges or join large trace tables, and stop meeting your team’s latency target. Eventually, teams resort to forced sampling, discarding traces to protect the application database, even though the product or an audit process requires the complete record.

Langfuse documented both contention and analytical friction as it scaled its open-source platform for LLM observability, evaluation, and prompt management. It experienced Postgres IOPS exhaustion during ingestion and prompt API latency reaching seven seconds under heavy load. Langfuse moved its tracing data from Postgres to ClickHouse, while keeping transactional and latency-sensitive paths isolated.

The store is only half the decision

The database move solved one class of problem, but the original data model created another. Langfuse initially carried separate trace, observation, and score tables into its analytical architecture. Updates required deduplication and cross-table analysis, adding join cost.

“The store is only half the decision. The analytical storage engine addressed the workload, and a data model reduced cross-table work.”

Later, Langfuse collapsed those records into a wide, mostly immutable observations table, with one row per model call, tool execution, or agent step. Initial table loads for large datasets went from seconds to milliseconds, and dashboard load times for large projects improved by at least 10 times over longer time ranges.

Langfuse needed both changes. The analytical storage engine addressed the workload, and a data model reduced cross-table work.

How to choose a storage pattern

Start with the reads your product must support, then choose the simplest architecture that meets those requirements.

At modest volume, keeping traces in the primary database avoids another operational boundary. As analytical contention grows, the application can send trace events to a dedicated analytical store while keeping mutable business records in its transactional database.

Some applications need both systems to share data. A Postgres-backed product might keep users, permissions, and workflow state in Postgres while sending agent events to ClickHouse for analytical queries. If relevant application data already lives in Postgres, change data capture can replicate it into the analytical path.

Start with who reads the trace

Who depends on the record and how they query it matters more than whether a trace looks like a log or a pull request.

If your product needs to reconstruct a durable record of a single run and the engineering team needs to compare behavior across thousands of runs, the trace has become application data with a telemetry-like storage workload.

Map the point lookups and cross-run scans before choosing a store. If both are product requirements, design for both from the first trace you retain.

The post When AI agent traces become application data appeared first on The New Stack.

❌