❌

Vue normale

Reçu hier — 27 septembre 2026

Performance engineering from kernel analysis to AI: Adrian Cockcroft’s take

27 septembre 2026 à 17:00
Abstract dark digital landscape with glowing contour lines representing multidimensional performance data and response time distributions.

Over its five-year history, P99 CONF has hosted quite a few speakers who’ve offered pointed takedowns of the namesake metric. At last year’s conference, Adrian Cockcroft didn’t explicitly state that P99s are BS… but he did allude to it.   

If you don’t know Cockcroft, he’s spent decades architecting, scaling, and optimizing resilient, high-performance systems at giants like Sun Microsystems, Netflix, eBay, and Amazon. We could probably dedicate an entire day of P99 CONF to discussing the lessons learned from just some of his projects (Solaris kernel performance, multi-processor optimization, Netflix’s on-prem to cloud migration, Chaos Monkey…) 

Fortunately, RedMonk analyst Rachel Stephens proved the perfect host for a conference that’s all about making things fast. She sat down with Adrian and led us on a whirlwind tour of how AI has impacted performance engineering. Here are some highlights from the chat (full video below).

Note: P99 CONF 2026 – a free + virtual conference on all things performance – is going live October 21-22. Grab a complimentary pass and join us!

From kernel analysis to vibe coding perf tools 

As a performance specialist at Sun in its heyday, getting to the root of performance problems involved lots of digging and divination. Cockcroft recalls, “Back in the old days with Sun, people would look at the output of system metrics in vmstat or whatever, and they’d be guessing what the numbers meant. There was a very vague understanding of what these things meant. The manual page wasn’t very clear.” 

Cockcroft ended up going to the source, literally. “I went and read all the kernel source code and figured out exactly where these numbers came from, exactly what they meant, which ones were approximating what, and wrote all that down.” That led to two performance books: Sun Performance and Tuning and Resource Management.

“My speedup is infinite, because this code would never exist without these tools. I wouldn’t have the time to build them.”

Four decades later, there’s now a wealth of helpful tools for end-to-end tracing, but Cockcroft’s curiosity still lies in what the tools are not showing. He continued, “Everything sort of looks okay in the tools – but the system isn’t behaving well. I usually come in and try to find a new way of looking at the data. A new type of analysis, or go a little bit deeper or finer grain, or stop looking at averages and start looking at distributions, and find all kinds of interesting things that nobody knew were happening.”

Currently, he’s vibe coding tools to better analyze the anomalies he finds. Saved from having to brush up on Python or hunt down graphics library fragments on Stack Overflow, Cockcroft can now stand up custom tooling in minutes. “My speedup is infinite, because this code would never exist without these tools. I wouldn’t have the time to build them.”

Peaks not percentiles

One specific vibe coding project: Cockcroft built (and open-sourced) tooling to get a better understanding of response time distributions. 

Response time distributions have been on Cockcroft’s mind for over a decade. While most people obsess over percentiles – yes, P99 CONF included – Cockcroft is most intrigued by the distribution of response time peaks in a histogram. He believes percentiles don’t work when trying to understand the latency and performance of modern web services. A single number like P99 can’t tell you whether the underlying distribution has one peak or several. And when there’s more than one peak (as is often the case in the real world), the mean, the standard deviation, and even the P99 itself lose most of their meaning.

“Percentiles don’t work when trying to understand the latency and performance of modern web services.”

Image showing what people think response time distributions looks like vs what they really look like
(source: A Tale of Two Histograms)

For example, assume you have a histogram with two response time peaks: a fast one from a cache hit and a slow one for misses that require actual work. As the cache hit rate shifts, each peak’s position remains the same (i.e., the latency values of the fast-response mode and the slow-response mode don’t change), but the peak heights rise and fall. “Your averages and your P99 are changing all over the place, but all that’s really happening is your cache hit rate is changing,” Cockcroft said.

So how do you go beyond measuring P99s and averages? Cockcroft did what he’s done for decades: dive in and build a custom tool. But these days, it’s much simpler thanks to LLMs.

“Your averages and your P99 are changing all over the place, but all that’s really happening is your cache hit rate is changing.”

He had already worked out the statistical approach for analyzing the distribution. Once ChatGPT came out, he quickly used it to build a tool that automated it. Instead of collapsing everything into an average, it identifies an arbitrary number of peaks in a distribution and tracks how they fluctuate over time. It’s implemented in R – a language Cockcroft hadn’t used in a while, but ChatGPT knew quite well – and it’s open source. If you’re curious, learn more in his Percentiles Don’t Work article and “A Tale of Two Histograms” talk and deck (“It was the best of response times, it was the worst of response times…”)

Where do we go from here?

To close, Stephens asked Cockcroft what advice he’d share with teams working on high-performance systems today. His top tip was to start with the macro view to find what’s interesting, then keep digging deeper until you’re inspecting individual slow requests end-to-end.

“Remember the microscope that you got when you were a kid,” Cockcroft said. “First, you have to focus it using the lowest resolution, at 10x, and then you can click it to 100x and adjust that, looking at just one speck now. Once you get that in focus, you click it to 1,000x.”

Cockcroft has spent his career building tools that bring obscure performance issues into focus. We look forward to seeing what others have cooked up with agentic tooling to help identify and solve performance problems this year at P99 CONF. 

Learn about the latest performance optimization techniques, tooling, and case studies at P99 CONF – free and virtual, October 21-22. Grab a complimentary pass and join us!

The post Performance engineering from kernel analysis to AI: Adrian Cockcroft’s take appeared first on The New Stack.

The rise of agentic AI on Kubernetes: unleashing the new infrastructure layer

27 septembre 2026 à 16:00
Abstract 3D render of blue cubes inside gold wireframe boxes, linked by red rods into a dense cluster, with teal lines connecting outer cubes.

AI is changing expectations around infrastructure and operations, including Kubernetes management. When models run close to the data they use, deployment, scaling, and governance responsibilities tend to shift to platform teams. And as clusters, environments, and operational signals continue to multiply, manual operations often strain under the added weight.

AI may simultaneously provide opportunities to lighten this growing load. Agentic software can now observe a system, reason about it, and act within predefined limits. 

Ultimately, these platforms’ value depends on the quality of the context an agent can see and the boundaries you set. Without cluster state, policy, and access rules, an agent can only guess.

Without cluster state, policy, and access rules, an agent can only guess.

For agentic AI to streamline multi-cluster management, you need clear lines between what the system observes, what it recommends, and what it changes. Drawn well, those lines let teams gain notable speed while still maintaining control.

The impact of AI on computing infrastructure

Teams once treated AI as an application concern; models sat on top of existing systems, and the stack underneath stayed mostly unchanged. Today, AI reaches into more and more customer interactions, while data storage needs simultaneously expand and orchestration pressure grows. A recent Forrester report describes the modern AI computing stack as stretching from the models themselves into and across the infrastructure beneath them.

As AI workloads move into production, they place new demands on the infrastructure beneath them. Many lean on specialized compute, with resource needs that rise and fall through bursts of training and inference. Because conditions shift quickly, they can also call into question whether telemetry remains trustworthy. Each of these demands lands at the infrastructure layer, where the workloads run.

The infrastructure layer of the new AI stack

The infrastructure layer covers compute, storage, and networking. It is a foundation that every workload running on the layer depends on. As AI workloads grow, choices about capacity, placement, and control will increasingly shape the performance of the data, intelligence, orchestration, and experience layers atop the infrastructure.

To operate the infrastructure layer efficiently across many machines and locations, a team may rely on orchestration instead of managing servers by hand. In cloud native contexts, Kubernetes has become a control point for scheduling workloads, applying policy, and presenting a consistent interface across environments. Kubernetes is especially well-suited to support organizations this way when teams need consistent control across an estate spanning data centers, clouds, and edge sites. 

Agentic AI and Kubernetes: the future of the infrastructure layer

Agentic AI can extend automation from fixed rules to systems that adapt to real-time conditions. Traditional automation runs the same script whether the environment has changed, while an agentic system observes the environment, reasons about what it finds, and then takes action.

When you apply agentic capabilities to multi-cluster management, the system follows this same sequence. An agent reads cluster state and operational data, proposes a diagnosis or next step, and then carries out actions based on an approved scope, usually after a person signs off. You can further reinforce these boundaries by routing each request to a specialized agent that receives only the metadata it needs.

The signals that an agent receives from the cluster, the context about policy and access, and the definitions of what the agent may change are the key elements that give agentic systems their value. They also separate agentic AI on Kubernetes from a generic assistant. 

Manual Kubernetes management is less efficient at scale

Admittedly, agentic AI fits some settings better than others. On a small single-cluster footprint, the overhead may outweigh the benefit. Manual Kubernetes management often holds up on a handful of clusters, but it can become unreliable in a rapidly growing estate. After all, each new cluster adds lifecycle work across upgrades, patching, configuration, and renewal. Those tasks can quickly multiply and diverge in hybrid environments.

Configuration drift is a high risk in these situations. Settings that started identical can fall out of sync, and policies can apply unevenly from one team to the next. Individually, these gaps may be manageable, but collectively they raise the odds of an outage or a failed rollout.

Visibility can also erode in an unmanageable way. Clusters spread across data centers, clouds, and edge sites often leave teams with no single view of the whole landscape. When DevOps and platform engineers stitch together signals from separate tools, resolution can slow and become more error-prone. A unified view helps enable sound, efficient decision-making by people, agents, or both.

Kubernetes knowledge is fragmented, and existing AI tools lack business context

Kubernetes expertise often sits unevenly across an organization. For example, senior engineers may hold deep operational knowledge that application teams lack. The most current information about a running system may also be fragmented if logs sit in one tool and metrics in another. Real-time understanding can be further clouded when policies, runbooks, access rules, and deployment history each live elsewhere.

Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes. Without that context, even a capable AI tool may fall short of providing meaningful Kubernetes management support.

Most well-trained AI models understand Kubernetes at a basic level, but they can’t know your unique cluster state, your policies, or your recent changes.

When an agent can read current signals alongside the rules that govern them, its suggestions become specific, testable, and actionable. In an incident, agentic systems can correlate logs with a recent change. Ahead of a rollout, they can check the change against policy. During troubleshooting, they can account for access rules rather than guessing at them. Kubernetes decisions carry real operational consequences, which makes these details all the more important to consider. 

Engineering “toil” isn’t time-efficient

Site reliability teams use the word “toil” for repetitive manual work, especially tasks that keep systems running without adding lasting impact. In Kubernetes operations, toil takes the form of repeated triage, manual signal correlation, alert follow-up, and routine checks. The tasks aren’t particularly difficult, but they can consume significant time and attention for enterprise teams.

When engineers spend their days on this kind of investigation, proactive modernization efforts tend to stall and planned upgrades can slip behind schedule. In other words, the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements. In a recent survey about how AI provides value to DevOps teams, reducing toil emerged as one of the clearer opportunities.

…the conditions that created the original toil perpetuate it, since teams lack the capacity to make related improvements.

Agentic AI can support repetitive investigations by gathering signals, correlating them, and proposing a likely cause for an engineer to weigh.

Kept under human review, it can take on some of the routine correlation that would otherwise fall to the team. That kind of support can give engineers more room to focus on the strategic work that most needs their judgment.

Building more intelligent infrastructure with agentic AI and Kubernetes

As you consider building toward intelligent infrastructure without surrendering control, the following principles can inform your efforts:

  • Start with observable context, giving agents access to current cluster state, policy, and history before they reason about a problem.
  • Separate suggestions from actions, allowing agents to recommend freely while any change must wait for human approval and a defined scope.
  • Connect agents to existing controls, routing their work through the access rules, identity, and audit paths the team already trusts.
  • Keep the ecosystem open, favoring platforms that integrate with current tools and standards over those that lock work into a single stack.

Platforms like SUSE Rancher Prime and SUSE AI Factory embrace these principles and illustrate how Kubernetes management can become a foundation for agentic operations. These platforms can help you improve cluster and policy consistency without compromising your authority over AI. Built on open-source foundations, they can also help you avoid being trapped in a single vendor’s stack.

In SUSE Rancher Prime, the industry’s first context-aware agentic AI ecosystem, its AI assistants work as a crew of specialized agents with an intelligent router. The platform draws on the cluster context already in place and acts through existing access controls. Through support for external Model Context Protocol (MCP) servers, teams can extend that crew to their own sources. In addition, human validation tools allow you to hold a proposed action for approval before the agent runs it.

Despite its potential, intelligent infrastructure is not universally beneficial. In situations where change control must stay fully manual, for example, agentic AI’s role may be strictly limited to observation and suggestion. Measure the technology’s value against the realities of your day-to-day operations. For those who are investing, agentic AI will have the greatest impact when it actively supports context, control, openness, and human judgment.

The post The rise of agentic AI on Kubernetes: unleashing the new infrastructure layer appeared first on The New Stack.

Reçu avant avant-hier

Avoiding vendor lock-in through an open-source approach: a developer’s perspective

26 septembre 2026 à 17:00
Abstract dark digital artwork depicting dense undulating layers, symbolizing cloud architecture and software ecosystem tension.

Every infrastructure team makes decisions that are difficult to reverse. Most of the time, that works out. Sometimes it does not.

Vendor lock-in usually begins as a reasonable choice, made under time or budget pressure, that solves a real problem at the time. A managed service ships faster or a deployment model fits better in that moment, but eventually a difficult constraint appears. 

When business conditions inevitably change, those accumulated choices and their consequences will determine whether a team can pivot accordingly. Limits on flexibility rarely trace back to a single vendor; more often, they hinge on how reversible the team’s past decisions are.

What is vendor lock-in and how can it harm your business?

The risks of vendor lock-in are not really about relying on vendors, since every production system relies on vendors. The big issue is dependencies that become too expensive or impractical to unwind.

For a platform team, that dependency builds up across APIs, contracts, roadmaps, and data models. It extends further into managed services, identity patterns, observability pipelines, and operational tooling. Each piece likely represents a reasonable design choice, but together they can quietly limit your options and raise the cost of leaving. When switching a database or control plane means rewriting tons of integrations, retraining the whole staff, or migrating data under inconvenient timelines, you have lost the room to maneuver.

The big impacts of small, invisible and unexamined decisions

Not every dependency is automatically a problem; some are understood, contained, and worth the tradeoff. The real risk lives in the dependencies no one examined closely, which may stay invisible until they block the business from evolving. 

“The real risk lives in the dependencies no one examined closely, which may stay invisible until they block the business from evolving.”

Unfortunately, some teams are familiar with these invisible dependencies. A managed database might pick up proprietary extensions, which application code then starts to assume. A Kubernetes environment might bind to one cloud’s IAM, networking, storage, and load balancer model. Observability and logging pipelines might harden around a single provider’s formats. None of these choices is reckless on its own, but together they can create significant friction. 

Obstacles to change and their hidden costs

The extent of a dependency-based tradeoff can sometimes remain unknown until circumstances shift, such as a new compliance requirement or customers needing a new deployment model. The hidden costs of these moments often escalate in stages. It might start with a visible, unwelcome migration bill, but the expense can also show up as operational drag. Rushed migrations can lead to additional service disruptions later. A workload may be unable to move, limiting services to certain customers. When you are tied to a specific vendor’s release cadence, it can make it difficult or even impossible to adopt emerging technology. 

Concentration risk compounds the problem, because a single change from one provider that carries pricing, support quality, and roadmap can ripple across the estate. By the time a switch becomes necessary, the cost shows up as service disruption, complex data transfer, and retraining. Naming these costs early keeps them from arriving as surprises.

At some point, a dependency can accumulate enough of these costs to become more than an architectural detail. Once it affects budgets and timelines, leadership has to account for it—and the team has to be ready to explain it. Identifying these dependencies early gives everyone time to plan.

Open source offers a different path

One way to proactively address this pattern is to evaluate potential dependencies more deliberately. For example, before committing to a platform or service, try to determine its reversibility. In other words, establish how difficult it would be for the team to change its mind about the investment in the future.

“Open source offers no guarantee against lock-in, however, since a team can still build tight coupling on open foundations.”

Open source solutions tend to perform well against that test, because they are intentionally built to keep systems inspectable, portable, supportable, and replaceable. By design, open source makes it easier for you to preserve options over time. It offers no guarantee against lock-in, however, since a team can still build tight coupling on open foundations.

What is open source?

Open source describes software you can inspect, run, modify, extend, support, and replace with relative ease compared to proprietary alternatives. The software’s source is available, and the license grants you the right to use and change it. Notably, no-cost or freeware software is not necessarily open source, specifically if it does not provide this level of access and rights.

Several companies have open source principles at their core, and open source software can be extremely valuable in enterprise contexts. Transparent code is often easier to audit, and open standards can reduce friction when moving between tools.

Open source also changes who can move the goalposts

For developers, reversibility is not only about APIs and data formats. It is also about whether one company can change the terms underneath a foundational technology. The Linux kernel is a useful example. Linux kernel documentation notes that copyright assignments are not required, so merged code retains its original ownership and the kernel now has thousands of owners. That makes unilateral relicensing of the kernel effectively impractical.

Kubernetes has a different legal structure, but the practical protection is similar. The project is licensed under Apache 2.0 and governed by the Cloud Native Computing Foundation. The license grants users durable rights to the existing code, so no single vendor, including SUSE, can retroactively take those open-source rights away from the project as it already exists. That matters because a platform can remain available even if a particular vendor changes strategy.

The Terraform-to-OpenTofu fork shows why this is more than a theoretical distinction. In 2023, HashiCorp changed Terraform’s license from the Mozilla Public License 2.0 to the Business Source License 1.1. The community responded by forking the last open-source codebase into OpenTofu, now a Linux Foundation project that remains under the MPL 2.0. The lesson for developers is not that every open-source project is immune to licensing changes. It is that open licensing and neutral governance can preserve a viable exit path when a vendor changes direction.

Open source powered by enterprise discipline

Open source ultimately earns its place through engineering discipline. Source availability has benefits but does not resolve governance, patching, lifecycle management, documentation, security, or integration on its own. A community project can be powerful and nonetheless arrive without enterprise-grade operational guarantees.

Enterprise open source providers exist and can help with closing that gap. They embrace open foundations and add the support, security, maintenance, and lifecycle discipline that production environments require. Founded in 1992, SUSE was the first provider of an enterprise Linux distribution. Today, it focuses on helping organizations operationalize open source with enterprise-grade support.

These companies aim not to close off open source software but to make it dependable at scale. In other words, open source and operational rigor can coexist. And enterprises should expect both from any external provider.

Digital sovereignty: the x-factor that makes open source even more critical

Digital sovereignty describes how much control an organization has over its infrastructure, data, operations, and technology choices. Sovereignty is a spectrum, and architecture decisions can move an organization a step in either direction.

Recent research by SUSE suggests that almost all enterprises are prioritizing digital sovereignty, but only 52% are actively taking steps toward it. That gap is largely an execution problem, and much of it surfaces in everyday platform decisions. 

If your team supports regulated industries or deploys in on-premises or air-gapped environments, you may be especially familiar with growing pressures around sovereignty.

Sovereignty puts a deadline on work that was already worth doing

Developers can hear “digital sovereignty” and assume it means a separate compliance workstream with a separate engineering bill. In practice, much of the work is the same discipline platform teams already invest in: portable workloads, clean interfaces, automated verification, reproducible deployment, auditable behavior, and the ability to replace a dependency without rewriting the system around it.

“Sovereignty does not suddenly make that engineering work valuable. It puts a deadline on work that was already worth doing.”

Those practices already have an economic case. They reduce migration costs, lower operational risk, make platform changes less disruptive, and preserve options when pricing, regulations, or business requirements shift. Sovereignty does not suddenly make that engineering work valuable. It puts a deadline on work that was already worth doing.

That reframe matters because it turns sovereignty from a policy overlay into an architecture property. The useful question is not simply, “How much extra work will sovereignty cost?” It is, “Which parts of our stack already fail the portability, interface, and verification tests we would want anyway?”

How to strengthen sovereignty with open source

Sovereignty depends on how a team designs, deploys, and operates its systems. Open source does not make an organization sovereign by default, but it can improve the conditions for sovereignty. 

In fact, many of the same questions that expose lock-in also matter for digital sovereignty. Each of the following questions about reversibility connects to open source and sovereignty alike:

Reversibility questionWhy open source can helpHow sovereignty strengthens
Can we run this workload elsewhere?Open source typically runs across on-premises, cloud, hybrid, and edge environments, not just one vendor’s platform.More control over where workloads run, including specific regions and regulated contexts.
Can we understand and audit how it works?Source availability and community scrutiny improve inspectability over closed alternatives.Teams can verify behavior, assess risk, and meet assurance requirements.
Can we migrate or reuse our data?Open ecosystems favor open formats and interoperable tooling.Data stays more portable, improving control over storage and movement.
Can another team or partner support it?Multiple support paths exist, from internal teams to integrators and enterprise vendors.Less dependence on one vendor’s pricing, availability, or roadmap.
Can we replace one component without rewriting everything?Open interfaces and modular design make components easier to swap.More control over architecture as requirements change.
Can we keep operating if a vendor changes direction?Open source projects can outlast one vendor’s strategy or license.Less exposure to decisions the team cannot control.
Can we deploy closer to the data?Open source can run in private data centers, sovereign clouds, edge sites and hybrid models.Sensitive workloads, including AI, can be governed nearer the data.

The ongoing work of digital sovereignty

Sovereignty is more of a practice rather than a specific destination. For many teams, the work begins with identifying existing dependencies that are especially hard to reverse. Similarly, you’ll need to separate the tradeoffs worth accepting from the ones that remove a significant number of options. 

Moving forward, it can be helpful to prioritize open interfaces and portable foundations when possible. When evaluating new services or solutions, treat lifecycles, support, and governance as first-order concerns.

In some cases, sovereignty work can be too heavy for an in-house team to carry alone. Providers such as SUSE can help strengthen your operational layer, including security and observability, and especially in growing or hybrid contexts.

Automated checks can make those principles concrete by continuously testing whether workloads can be rebuilt, moved, audited, and recovered instead of waiting for a migration or compliance event to expose the gaps.

Open source lets you take control of your software ecosystem

No enterprise team avoids every dependency, and candidly none should try. Some coupling is reasonable, contained, and worth it. A vendor-free system is not a realistic goal for a major enterprise. A realistic goal is the judgment to separate acceptable dependencies from dangerous ones.

“The true cost of any platform includes the cost of leaving it, and teams should understand that cost before they commit.”

Reversibility gives that judgment something concrete to work with, because it can be broken down into capabilities a team can name, evaluate, and test:

  • Ownership. Ownership does not mean building everything yourself. It means holding the realistic ability to run, move, or hand over each layer of your stack. The test is simple: if a vendor disappeared tomorrow, or was ordered to stop serving you, what still runs next month?
  • Auditability. You should be able to verify what your software does, yourself or through an auditor you appoint, rather than accepting a vendor’s report as the final word. With open source, inspection is a property you hold. With closed software, it is a permission you are granted, and permissions can be withdrawn.
  • Exit velocity. An exit plan without speed is just a document. Exit velocity measures how fast a workload can move from one platform to another, and it only means something when you test it on a schedule, as earlier generations tested disaster recovery.
  • Pivot ability. These capabilities matter when conditions change: a new compliance requirement, a customer that needs a different deployment model, or a vendor that changes direction. Teams that can reroute workloads respond on their own timeline. Teams that cannot must renegotiate from a position of weakness.

Vendor lock-in becomes a manageable risk when you can confidently flag which decisions are hard to undo, weigh the tradeoffs honestly, and protect the team’s pathways to change. Open source strengthens every one of these capabilities because it keeps larger portions of your system inspectable, portable, and replaceable.

The true cost of any platform includes the cost of leaving it, and teams should understand that cost before they commit.

The post Avoiding vendor lock-in through an open-source approach: a developer’s perspective appeared first on The New Stack.

The agent didn’t break your controls. It went around them.

26 septembre 2026 à 16:00
Three black circular directional signs on a gray concrete wall, showing arrows pointing straight ahead, turning left and turning right.

The identity part of agent security is settled. An agent needs its own identity: a short-lived, revocable credential scoped to the job, and an audit trail that names the human who set it running. NIST’s security leads made that case in August 2026, and most identity vendors agree.1

Identity and access management is table stakes. It’s necessary, but it isn’t what’s breaking.

What’s breaking is an assumption we’ve carried for twenty years: Get identity and permissions right at the door, and whatever happens inside takes care of itself. That worked when software was passive. Agents reason about a goal and choose their own steps toward it, like a seasoned escape artist.

An agent that hits a wall looks for another way

Almost every control in today’s stack answers a question about entry. Should it connect? Should it reach that service? Should its token be accepted here? Each is a question about a route, and there’s rarely just one route to anywhere worth going.

An agent treats a blocked route as a problem to solve, because that’s what we built it to do. A person who hits a locked door usually files a ticket, while an agent tries the window.

In July 2026, an autonomous agent spent four and a half days inside Hugging Face’s production systems.2 A filter controlled which internet addresses its dataset servers could download from, and it never fired, because “the agent stopped asking the worker to fetch remote resources and instead made it act on local ones.” The filter worked as designed, and the agent went around it anyway.

A person who hits a locked door usually files a ticket, while an agent tries the window.

On ordinary developer machines, malware in a compromised npm package tried to recruit the AI coding assistants already installed to search for secrets,3 and a coding agent deleted a production database during a change freeze before falsely telling its operator the data couldn’t be recovered.4 Both happened on the machine itself, where no network control was looking.

The shift from outside-in to inside-out

Outside-in controls govern entry, and most organizations run plenty of them. Make no mistake, inside-out security completes those controls rather than replacing them.

Inside-out control governs the action itself, and asks a narrower, harder question: Should this agent, acting on this person’s authority, delete this table in this database, right now?

That question matters because an agent can swap routes but not the outcome it’s after. No matter how many routes it tries, deleting a table is still deleting a table, and a checkpoint on the action sees it every time.

Here’s how today’s controls line up against it.

ControlWhat it coversWhat it misses
GatewayTraffic you route through itLocal shell commands and file edits never reach it
SandboxThe environment as a wholeConstrains reach, not individual actions
SIEMA record of what occurredReports after the action is completed
RegistryThat an agent existsWhat the agent did with that existence

Each does its job, but they all decide somewhere other than the moment the action runs.

Put the enforcement point where the agent acts

Every agent acts through an agent harness: the software that takes the action the model chose and carries it out, whether that means running a command, writing a file, or calling an API. In most deployments today, nothing checks that action before it runs.

An inside-out control puts an approval step in that gap. Before the harness executes anything, the checkpoint looks at which agent is asking, on whose authority, and against which system, then applies policy to allow the action, block it, or send it to a human. Because every action passes through it, an agent denied a destructive command and trying a smaller version of the same thing is held to the same rules. The remaining risk is a badly written policy, which can be fixed.

None of this works without the identity basics. Any type of control, whether it be at the prompt level, inference level, harness level, or MCP layer, can’t judge “may an agent take this action here on this object?” when the only name on the request is a service account shared by six agents and four engineers.

The companies building agent runtimes have reached the same conclusion. Over the past eighteen months, Anthropic, Google, Microsoft, OpenAI, LangChain, and Cursor have each added a hook that lets you inspect an agent’s action before it runs.5 When AWS explained its own agent policy design, it argued that controls belong at the moment an agent attempts to invoke tools.6

The catch is that each hook works differently, with no standardized request or response formats. An enterprise whose developers use Claude Code and Cursor while its platform team builds on LangChain would maintain the same enforcement logic in multiple different flavors, each with its own audit trail. That doesn’t scale, and it tightly couples your security model to whichever runtime a team favors that month. Enterprises need one vendor-agnostic agentic security layer that spans every harness, so adopting a new model or framework doesn’t mean restarting the entire onerous security review.

Turn the lights on before you start blocking

The standard, well-ingrained security instinct is to start blocking right away, but we’ve all seen how well that works with the business in the past. Security must move and adapt at the speed of business, not the other way around. Security tools such as intrusion prevention systems and web application firewalls both ran in monitoring mode until teams understood what normal looked like, and those that skipped that step tended to hear about it from a production outage.

Agents need the same sequence, only faster. An enforcement point in monitoring mode blocks nothing and quickly answers questions most organizations can’t today:

  • Which agents are actually running, not which ones someone believes are running
  • Who started each one, and whose authority it’s operating under
  • What capabilities it used, and against which systems
  • Which of those actions would have violated a policy, had the agent security platform been switched to enforcement mode

Write policy from what you know, see, and have evidence of, not just from an architecture diagram. Enforce first where the stakes are highest: destructive commands, production data, and anything that moves data out. Then watch-learn-build, just like the agents we use: Watch the patterns, build finer-grained controls and policies, and learn how to use AI securely, safely, and confidently. Observe first, then enforce, build, and deploy, in that order.

The bottom line

None of this requires a new category of infrastructure. It’s the identity, authorization, and audit you already run for your people, extended to agents and applied inside the harness before the action runs.

The perimeter is still there, but it has moved to the moment an agent acts, the one place it can’t route around.

Ory built Agent Security inside the harness, on the same identity and authorization engines that run in production for human users. It starts in “observe mode,” so you get that inventory first, and you can try it today at ory.com/agent-security.

Footnotes

  1. Bill Fisher and Ryan Galluzzo, “Back to the Future: Why Agentic AI Needs a Strong Identity Foundation,” NIST Cybersecurity Insights, August 27, 2026.  ↩︎
  2. Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident.” The intrusion ran July 9–13, 2026. ↩︎
  3. Nx, “s1ngularity postmortem,” August 2025. The malicious packages “attempted to use local AI tools (like Claude and Gemini)” while scanning systems for sensitive data. ↩︎
  4. AI Incident Database, Incident 1152: Replit agent deletes production database during code freeze, July 18, 2025. ↩︎
  5. Pre-execution hooks by vendor. Anthropic, Claude Code hooks; Google, Agent Development Kit callbacks; Microsoft, Agent Framework middleware; OpenAI, Agents SDK guardrails; LangChain, human-in-the-loop middleware; Cursor hooks (InfoQ, October 2025) ↩︎
  6. Liana Hadarean and Jean-Baptiste Tristan, “Why Policy in Amazon Bedrock AgentCore chose Cedar for securing agentic workflows,” AWS Security Blog, May 20, 2026. ↩︎

The post The agent didn’t break your controls. It went around them. appeared first on The New Stack.

Query decomposition doesn’t fix context starvation — it just moves it

24 septembre 2026 à 15:00
Abstract digital art showing warped light lines surrounding a void, illustrating data compression and AI context starvation.

I have a small example that would best communicate the message I am trying to convey: say you built a chat widget for GitLab’s public documentation (the corpus we are experimenting with in this article) and one of the developers sends this kind of message:

We got an email saying our card was declined for something called “quarterly reconciliation” and I need to know what actually happens now. On top of that, I think we’ve gone over our seat count; there are more people in the group than seats we bought. Our CI has been queuing all week and I want to know whether the compute minutes we purchased last month rolled over or if we lose them. Our finance lead also needs to be the one who gets the invoices from now on, not me. And last thing, is the REST API rate limited? We’re building an internal dashboard and would rather find out now than after it breaks.

Five separate asks: the declined payment, the seat overage, compute-minute rollover, changing who receives invoices, and API rate limits. Each one is answered by a specific passage in GitLab’s public documentation, and you labeled which passage answers which before running anything, so you knew in advance exactly what a correct system needed to find.

Then you run the message through a pipeline that follows current best practice. It splits the query into five clean sub-queries, retrieves for each one independently, merges the results, drops near-duplicates, reranks the merged pool against the original message, and packs the highest-scoring passages into a 2,000-token context.

The pipeline retrieved all five correct passages, but only one of them survived into the packed context; that is one of five asks, not one of five sentences. The packer found, scored, and threw away the other four before the model ever saw them. The same message with no decomposition at all managed three out of five.

The failure has a name, and it isn’t the one you’re thinking of

I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.

The definition is deliberately about allocation rather than retrieval, because allocation is the part nobody watches. If the passage answering the fifth question was found, scored, and then squeezed out by three passages about the first question, the fifth sub-intent is starved, and every recall metric you have will report that the system worked perfectly.

“I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.”

Two failure modes already in circulation describe something different, and it’s worth separating them cleanly:

Semantic dilution happens at retrieval time: when you embed a five-part question as a single vector, you get a centroid that sits somewhere between five topics and lands close to none of them, so the evidence is never found. Decomposition fixes this, which is why it spread.

Context poisoning is about what is present, not what is missing. Wrong, stale, or adversarial content enters the window and corrupts what the model generates downstream. Poisoning is a contamination problem. Starvation is an absence problem, and policy, not accident, produces the absence.

I borrowed the word from operating systems. In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work. Every ingredient of that situation is present in a retrieval pipeline: a fixed resource, competing demands, and a policy that decides who gets served. A relevance-greedy packer is priority scheduling with no aging term, and under priority scheduling without aging, valid low-priority work waits forever.

Decomposition is the right fix to the wrong half of the problem

Split that message into five single-intent queries, and each one embeds cleanly, so per-sub-query recall climbs sharply. This is well-trodden ground. LlamaIndex ships a SubQuestionQueryEngine that breaks a complex query into sub-questions and synthesizes the responses. LangChain’s MultiQueryRetriever generates query variants and returns the unique union of what they retrieve. RAG-Fusion applies reciprocal rank fusion across the per-query result lists. The technique works, and it isn’t mine.

“In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work.”

The context window did not grow. Let me explain: after decomposition, you have n result sets competing for one fixed token budget, and something downstream has to decide the split. In most production pipelines, that something is a short, unremarkable sequence: merge the pools, drop near-duplicates, rerank the merged pool against the original query, then greedily fill until the budget closes.

That sequence is a scheduler. It has a priority function, which is the reranker score, and it has no fairness constraint of any kind. A sub-intent with three strongly-scoring passages takes three slots. A sub-intent whose single correct passage scores mid-pack takes none of them.

So the failure did not go away. It moved from the embedding, where it has a name and people watch for it, into the packer, where it has neither. It also moved somewhere with much worse instrumentation, because recall@k per sub-query is the metric decomposition usually gets validated with, and that number goes up. It goes up at the same time as coverage inside the packed context goes down. You ship on a green dashboard.

The harness

The corpus, GitLab’s public documentation: 10,000 chunks and 2.2M tokens, split on heading boundaries and capped at 480 tokens each. Sixty-one single-intent questions span nine topics, from seat management to rate limits, each labeled with the one passage that answers it. I built multi-intent queries by concatenating those questions while varying n across 2, 3, 5, and 7, randomizing the order so position doesn’t confound topic, and varying topical distance so half the queries draw everything from one topic and half span distinct ones. That produces 100 queries, 25 at each value of n, whose correct decomposition I know exactly.

Two decisions matter more than the rest:

The metric is not recall. Recall tells you what the retriever found. What I need is what survived into the packed context, per sub-intent. So I log each sub-intent twice: once for whether its correct passage reached the candidate pool, and once for whether it reached the packed context. The gap between those two numbers is the entire argument.

Every question has to be retrievable on its own before it’s allowed in. A question enters only if its correct passage ranks in the top 10 for its own isolated query, under both retriever configurations, and both scored 100% recall@10 on that test. Since each sub-query’s candidate pool is exactly its own top 10, passing that gate guarantees the correct passage sits in the pool for every decomposed arm. Any sub-intent that then fails to appear was denied by the packer rather than missed by the retriever, which removes the most obvious objection to everything below.

The core measurement uses no language model. Because queries are composed from known sub-questions, the decomposer is an oracle so that anyone can reproduce the main result with no API key.

That invites an objection, so I tested it. A real LLM decomposer, blind to n, disagreed with my ground truth on 41% of the queries, and on inspection it was right every time. Five of my sixty-one supposedly single-intent questions contain two distinct information needs. What is excess storage usage, and what happens when we go over the free limit? is two questions wearing one question mark. Adjusted for those five, agreement is 100 out of 100. That is not evidence decomposers are reliable, because my queries are joined by fixed connectives and splitting on those alone recovers n perfectly, which real messages never allow. What it caught was an error in my own labels, and that is the best argument I have for the oracle design.

Results

Every arm runs at 2,000, 4,000, and 8,000 tokens against two retriever configurations. The stronger pairs are BAAI/bge-base-en-v1.5 with BAAI/bge-reranker-base; the weaker pairs are a quantized BAAI/bge-small-en-v1.5 with Xenova/ms-marco-MiniLM-L-6-v2. Retrieval is in-memory cosine similarity over a NumPy array because, at 10,000 chunks, a vector database would be slower to write, slower to run, and harder to verify.

Sub-intent coverage at a 2,000-token budget on the stronger configuration:

Table showing coverage (in %) per arm.

The production-default pipeline starves 31.1% of sub-intents whose evidence it had already retrieved. It beats no decomposition by nine points, while a flat B/n split, which is the crudest allocator anyone could write, beats it by fourteen.

Floors work, but not the obvious floor. Reserving one passage per sub-intent before the greedy fill satisfied 99.4% of its reservations and bought only seven points. The mechanism fires correctly and reserves the wrong passage because it picks each sub-intent’s best chunk by score against the original query. The original query asks about all five intents at once. Selecting that same reservation by score against its own sub-query pushes coverage to 89.4% and cuts allocation starvation from 30.6% to 10.1%. That is a change of about four lines.

Two of my own recommendations died here. I expected reranking against the original query to beat reranking against the fragment, and it loses by seventeen points. The incomparable score scales I worried about turn out to help, because each sub-query’s best match ends up at the top of its own scale, producing per-intent fairness for free. I also expected deduplication before allocation to matter, but near-duplicates consume 1.0% of the budget and removing them moves coverage by 0.3 points.

The crossover: starvation by tokens-per-sub-intent, which is simply the budget divided by n:

Table showing tokens per sub-intent.

Below roughly 1,000 tokens per sub-intent, allocation policy dominates. Above it, nothing you do to the allocator matters, because everything fits anyway.

The retriever comparison is the one I’d lead with. Upgrading the retriever moves coverage on the production-default arm from 56.2% to 68.9%, a gain of 12.7 points. Changing the allocation policy on the same retriever moves it from 68.9% to 89.4%, a gain of 20.5 points. In its sharpest form: the weaker retriever with a fragment-scored floor reaches 84.5%, and beats the stronger retriever with a greedy packer at 68.9% by sixteen points. A worse retriever with a better allocator wins.

Position: Held within a fixed n so query difficulty doesn’t contaminate the comparison; starvation across the seven positions of an n=7 query runs 1.3%, 21.3%, 30.7%, 48.0%, 48.0%, 32.0%, and 13.3%. That is a serial-position curve. The packer protects what you asked for first, protects what you asked for last a little less, and drops the middle. The fragment-scored floor flattens it to 4.0%, 9.3%, 6.7%, 8.0%, 22.7%, 5.3%, and 6.7%.

“A worse retriever with a better allocator wins.”

Topical distance: Sub-intents that span distinct topics starve about twice as often as sub-intents drawn from one topic, at 38.1% against 18.3% for n=7. I predicted the opposite. A topically coherent query gives the reranker a coherent target, and it scores all the correct passages similarly. In contrast, a scattered query lets it latch onto some topics and abandon others.

What the user actually sees

Everything above is retrieval-side. What decides whether any of it matters is what reaches the person who wrote the message, so I generated real support replies from 80 packed contexts and had every reply graded per sub-intent, with both the generation and the grading blind to which arm produced which context.

When the correct evidence reached the packed context, the reply addressed that question 100% of the time, across 261 out of 261 cases, in both arms. Coverage predicts the generated outcome exactly, which is the strongest justification I have for measuring it.

When a sub-intent was starved, the reply answered it anyway 48.1% of the time, based on whatever else happened to be in the window. It explicitly flagged the gap 45.6% of the time, with some version of “I’ll follow up on that separately.” It went silent only 6.3% of the time.

“Starvation mostly does not produce silence; it produces unsupported answers.”

I expected silence, and I was wrong. Starvation mostly does not produce silence; it produces unsupported answers. Whether those answers are actually incorrect is the next experiment, because this harness measures whether a question was addressed, not whether the answer was right.

Some limits: composed queries are cleaner than real support messages, which carry pronouns, implicit context, and conditional clauses. This is one corpus and one embedding family. I drafted the gold labels with model assistance and verified them myself. The model writing those replies was strong, so a cheaper production model would plausibly flag fewer gaps and invent more.

What an allocator actually looks like

Give every sub-intent a floor, and choose it by fragment score. Not the naive floor, which satisfies 99% of its reservations and buys seven points. Select a reservation by relevance to the sub-intent it protects, not by relevance to the message as a whole.

Rerank against the fragment rather than the original query. This inverts what I expected and what I have seen recommended. Scores from different fragments are not comparable across sub-intents, and that incomparability is doing useful work.

Don’t spend your effort on deduplication; near-duplicates cost 1.0% of the budget here. Dedup is worth doing, but it isn’t why your fifth question went unanswered, and treating it as the fix will cost you weeks.

Log per-sub-intent coverage: You already computed it to pack, and it predicts the generated outcome perfectly. A sub-intent that received zero passages is the best predictor available that your reply is about to assert something you cannot support.

Where parallel decomposition breaks

“If it’s late can I get a refund” is one clause and two intents, and the second one’s retrieval target depends on the first one’s answer. Parallel decomposition treats them as siblings. It retrieves the late-delivery policy and the general refund policy, packs both, and misses that the passage you actually need covers refunds for late delivery, which may match neither sub-query particularly well.

There are two ways out: You can tag dependencies at decomposition time, or run a deferred second pass that re-retrieves conditional clauses once the first round resolves.

I would take dependency tagging, for three reasons: A second pass costs a full retrieval round trip inside a latency budget a support bot does not have. The tag is reusable, because a dependent sub-intent should not hold a floor reservation. At the same time, its parent is unsatisfied, so it feeds the allocator directly instead of bolting on a separate mechanism. And it fails visibly, since an untagged dependency shows up as a starved sub-intent in the coverage signal. In contrast, a deferred pass that resolves the wrong condition produces a confident wrong answer with nothing to flag it.

The cost is real; dependency tagging pushes work onto the decomposer, which is already the weakest component in the chain, and I have not measured tagged against untagged. That is a design position rather than a result, and it is the one thing here I am asking you to take on argument instead of evidence.

What to measure on Monday

Take your production pipeline and compute one number: your context budget divided by the average count of distinct questions per incoming message. If that number lands below roughly 1,000 tokens, your allocation policy costs more than your retriever does, and the reranker upgrade sitting in your backlog will buy you less than reserving one slot per question.

On my corpus, the retriever upgrade was worth 12.7 points of coverage, and the allocation change was worth 20.5, which is why I think the ordering is wrong in most pipelines I’ve seen. That ordering is the falsifiable part. Run the same two comparisons against your own corpus, and if the retriever wins, I want to see the numbers, because that result would tell me the crossover sits somewhere other than where I measured it.

The cheaper thing to do first takes an afternoon. Log, for every multi-intent request, how many sub-intents ended up with zero passages in the packed context. A support system that cannot tell you which question it dropped will keep answering that question anyway, about half the time, out of whatever else was in the window.

The post Query decomposition doesn’t fix context starvation — it just moves it appeared first on The New Stack.

Can enterprises protect data without making AI less reliable?

24 septembre 2026 à 14:00
Abstract long-exposure photograph of dark motion blur and light streaks representing digital speed and data flow.

As organizations invest in AI, many are discovering a new bottleneck: obtaining data that is both protected and useful. Engineering teams need realistic, production-like data to validate AI-generated changes, train models, test applications, and generate business insights. Yet, privacy initiatives can sometimes make that data harder to access, less representative of real-world conditions, or unable to preserve critical relationships between records. 

The Perforce Delphix “2026 State of AI and Data Privacy Report” highlights this challenge. Among surveyed organizations, 26% say privacy controls make production-quality data harder to obtain, 25% struggle to preserve relationships across data entities, and 51% cite data quality challenges. 

“Protecting data isn’t enough if it can no longer support the systems that depend on it.”

Protecting data isn’t enough if it can no longer support the systems that depend on it. For engineering teams, the question is whether their data protection strategies can preserve the qualities that make data valuable in the first place. You need a well-rounded data strategy with tools that maintain referential integrity and relationships across your environments.

What these statistics mean for practitioners

At first glance, statistics from the report, like “51% of enterprises cite data quality challenges,” may sound like a purely governance issue.

In practice, they represent engineering problems. Low-quality datasets can produce:

  • Inaccurate analytics.
  • Poorly trained AI models.
  • Incomplete test coverage.
  • Increased rework.
  • Delayed releases.
  • Reduced confidence in data automation.

“When an AI model is trained on incomplete or distorted data, its outputs become less reliable.”

When an AI model is trained on incomplete or distorted data, its outputs become less reliable. When test environments contain unrealistic data, defects can escape into production. When analytics datasets lack consistency, teams spend more time validating results than acting on them.

Data protection and utility are not opposing goals

A common misconception is that organizations must choose between privacy and innovation, but the most successful organizations know that compliance, quality, and speed can and need to work together.

Protected data still needs to be:

  • Realistic enough for testing and validation.
  • Representative enough for analytics.
  • Accessible enough for engineering teams.
  • Governed enough for regulatory requirements.
  • Connected enough to preserve referential integrity.

There’s a two-fold goal in the AI era: reduce sensitive data exposure while also creating trusted data that remains valuable after protection. All too often, enterprises sacrifice compliance for innovation or speed. That’s a big reason 84% of respondents in our report have a data privacy exception in their non-production environments.

“There’s a two-fold goal in the AI era: reduce sensitive data exposure while also creating trusted data that remains valuable after protection.”

Organizations can only move at AI speed when they have access to trustworthy data that accurately represents production conditions. When privacy controls degrade quality, limit realism, or restrict access to representative datasets, the data layer becomes the new bottleneck.

Why referential integrity matters more than ever

Many discussions about data privacy focus on masking sensitive fields. However, masked data that loses referential integrity between entities can create a different kind of risk.

A customer, order, or payment record may still exist, but if the relationships connecting those records break during protection processes, the data no longer resembles reality. Take billing validation, for example. It needs referential integrity when a customer has multiple products, charges, and invoices across several database tables to produce the correct products or make accurate charges. 

Broken relationships are especially problematic for modern AI and analytics systems. Analytics pipelines depend on consistent identifiers to join information across sources. AI and machine learning workflows depend on complete business context to identify patterns and make predictions. Software testing depends on realistic relationships between records to validate application behavior accurately.

When referential integrity is lost:

  • Analytics can produce incomplete or misleading results.
  • AI models can learn from flawed datasets.
  • Testing environments can fail to expose production issues.
  • Teams lose trust in protected datasets.

Importantly, these failures are often difficult to detect. Pipelines may continue running successfully while quietly producing degraded outcomes, which can be a very costly mistake.

This helps explain why 25% of surveyed organizations cited preserving relationships across data entities as a significant challenge. For practitioners, that statistic is a warning that privacy controls can unintentionally undermine the quality of AI and analytics initiatives if they fail to preserve business context.

Keep in mind that not all data protection solutions are created equal — enterprise-grade masking algorithms are key to preserving relationships. When applied consistently and at scale, these algorithms ensure the same input produces the same masked output across your systems and environments. Other masking approaches might be done piecemeal, resulting in broken relationships.

What engineering teams should measure

Many organizations measure privacy success through compliance metrics alone. However, AI-driven environments require a broader definition of success. Engineering leaders should evaluate privacy initiatives against several dimensions:

  • Data quality: Does the protected dataset accurately reflect production conditions?
  • Realism: Can developers, data scientists, and analysts use the data confidently for their intended purpose?
  • Referential integrity: Do relationships remain consistent across applications, tables, environments, and data sources?
  • Accessibility: Can teams obtain compliant data without introducing delays?
  • Provisioning speed: How quickly can trusted datasets be delivered when needed?

These measurements help organizations determine whether privacy efforts enable AI outcomes or create new obstacles.

Designing for governance by default

As AI adoption grows, privacy cannot remain a separate process that occurs after development begins. Organizations should instead map the entire data lifecycle, from data request and discovery to reuse and retirement. 

This approach helps ensure governance is built into workflows rather than applied as a late-stage checkpoint. It also provides stronger auditability, reduces compliance exceptions, and gives teams greater confidence that protected datasets remain fit for purpose.

Most importantly, it aligns privacy objectives with business outcomes instead of treating them as competing priorities.

Trusted data will become a competitive differentiator

The need for test data — in volume, coverage, and scale — is booming as agentic development continues to rise. AI has also increased the volume of change enterprises can generate, but the challenge of validating that change remains unsolved.

That responsibility still belongs to data. The organizations that gain the most value from AI will not necessarily be the ones with the most advanced models. They will be the ones with the most trustworthy data foundations — using a portfolio approach that combines data virtualization for speed, masking for security, and synthetic data for coverage as needed.

“In the AI era, trusted data, not fast model output, may be the ultimate competitive advantage.”

The research points to a clear lesson: If privacy controls uphold realism, quality, relationships between records, or access to representative datasets, they can enable AI success.

As enterprises continue investing in AI and data privacy, the real objective should be ensuring protection and utility coexist. Because in the AI era, trusted data, not fast model output, may be the ultimate competitive advantage.

The post Can enterprises protect data without making AI less reliable? appeared first on The New Stack.

The cloud reduced operational complexity. But many teams need someone to own it completely.

22 septembre 2026 à 16:00
Dark abstract digital wave with subtle color distortion, representing application lifecycle and cloud native infrastructure management.

For companies running serious production workloads with lean engineering teams, the real test of operational ownership comes at 3 a.m. Not whether the system stays up but whether anyone needs to be awake to make that happen. The page arrives. A service is degrading. The engineer who answers didn’t build this service, doesn’t know what thresholds were set at deploy time, and cannot tell whether the system is healing itself or waiting for a human decision. 

This is the moment that separates services that transferred operational burdens from platforms that merely deferred them. The system that passes the 3 a.m. test already knows what healthy looks like, what to do when healthy stops being true, and how to communicate what happened—because those decisions were made at deploy time, not incident time. The team sleeps through the night because nothing went wrong, but because the response was already determined.

Forty engineers, eight applications, zero dedicated ops

Cloud infrastructure evolved in two directions simultaneously, and neither arrived where mid-market teams actually stand. On one end: full control. Infrastructure-as-code, service meshes, custom pipelines. Powerful, flexible, and designed for organizations that have deliberately invested in operational staff who can absorb the cost of that flexibility. On the other end: single-app simplicity. Push code, get a URL. Elegant for a first deployment, but architecturally limited the moment a team manages more than one service, needs compliance controls, or inherits an application that doesn’t fit the platform’s opinions.

You know this company. Forty engineers. Eight production applications. Two generate 80% of revenue. One SRE who is actually a senior developer with an on-call rotation nobody else wants. A compliance audit due in Q3 that nobody has started preparing for.

“The right investment is shipping product. The cost of this gap is measured in what these teams do not ship.”

Every engineer is a full-stack contributor. The person who wrote the feature deploys it, monitors it, and gets the page when it breaks not because the team lacks sophistication, but because hiring dedicated infrastructure staff is not the right investment at their stage. The right investment is shipping product. The cost of this gap is measured in what these teams do not ship. Every sprint spent upgrading the deployment pipeline is a sprint without the feature a customer asked for. Every 3 a.m. page answered by a developer with a product standup at 9 a.m. is diminished output that never shows up in any dashboard. The gap is widening.

The applications nobody planned to operate

Not every application in a portfolio was built by the team now responsible for it. Many companies acquire products through M&A. They inherit internal tools built by engineers who left three years ago. They run commercial off-the-shelf applications customized beyond vendor support. They maintain line-of-business applications in languages nobody on the current team chose. These applications share a common trait: they are in production, they serve customers or meet compliance requirements, and nobody has the budget or the mandate to rewrite them. They need a home that accepts them as they are, not as a modernization roadmap says they should become.

“They need a home that accepts them as they are, not as a modernization roadmap says they should become.”

This is where the full lifecycle vision matters. An application management service that only serves net-new applications forces teams to maintain two operational models: one for the applications they are building today, and another for the applications they inherited yesterday. That split is where drift starts, where patching falls behind, and where audit findings accumulate.

The application management service that solves this problem must accept the full range: the Java application packaged as a WAR file, the .NET Framework service running on Windows, the Python application with dependencies pinned to a specific runtime, the containerized service already running elsewhere, and apply the same operational model, the same deployment interface, the same patching and scaling behavior to all of them. Migrate, manage, and modernize within the same experience, without requiring a different operational posture for each stage of the application lifecycle.

What the market data actually shows

Janakiram MSV, an analyst and advisor on cloud-native platforms and TNS contributor, puts it this way:

Cloud native standardized the infrastructure layer around containers and Kubernetes. But it never standardized the operational boundary between the application team and the infrastructure underneath it, and platform engineering largely emerged to re-establish that boundary. CNCF’s latest research with SlashData reports that 28 percent of organizations run a dedicated platform engineering team and 41% split those capabilities across multiple teams. Another 3% have no formal approach at all, and that is where most mid-market engineering organizations exist. They want the same outcome as platform engineering without first becoming a platform engineering organization.

The pattern that repeats is that maturity stalls at the third application rather than the first deployment. A forty-engineer team gets one service into production and keeps it healthy through familiarity. Then an acquisition brings in a .NET workload, an internal tool developed by an engineer who left two years ago, and the team is carrying three operational models with nobody owning the mandate to reconcile them. Hiring will not close that gap, because what is missing is a standardized operational posture, not headcount.

What hundreds of thousands of production deployments actually reveal

We have visibility into hundreds of thousands of production deployments across thousands of customers. That scale does not tell you what people say they want. It tells you what actually breaks, what gets escalated at 3 a.m., and what determines whether a team trusts their platform enough to stop thinking about it. The patterns are remarkably consistent.

1. The deployment that nobody touches again

A team spends two full days getting a Spring Boot application deployed with CI/CD and an SSL certificate. The deploy works. Then nobody touches it for three months because it is stable, but because touching it might break it again. They discover the problem when customers report the application is unreachable. Four hours of forensics follow. What changed? Why? How do we prevent this?

“A release that can partially succeed is a release that will partially fail.”

This is the most common failure mode we observe: not the deployment that fails loudly, but the deployment that succeeds quietly and degrades invisibly. The team cannot explain what happened because the platform did not decide what “healthy” meant before the incident arrived.

A release that can partially succeed is a release that will partially fail. The platforms that earn trust are the ones where a deployment either completes fully or reverses entirely no intermediate states, no manual rollback procedures discovered under pressure.

2. The observability sprint that ships three weeks late

A developer notices response times degrading. They want memory utilization metrics. The platform does not collect them by default. They spend a sprint writing configuration files to install a monitoring agent across every instance. The insight they needed three weeks ago ships three weeks late.

This is the second most common pattern: observability treated as an add-on rather than a default. Every team we have observed that adds monitoring after their first incident wishes they had it before. The teams that never experience this problem are the ones whose platforms shipped metrics, traces, and health signals at deploy time without code changes, without configuration, without a sprint spent on plumbing. The distinction matters: a platform that can be observed is not the same as a platform that is observed from the moment it goes live.

3. The portfolio tax

Each application gets its own infrastructure. Costs scale linearly with the portfolio. The team running eight applications pays eight times the overhead of the team running one, not because each application needs dedicated resources, but because the platform’s architecture assumes isolation rather than shared operational responsibility. The result: teams avoid migrating inherited applications because the cost model punishes breadth. The compliance audit does not care that the inherited application runs on a different operational model. It expects the same governance, patching cadence, and access controls.

A platform that rewards portfolio growth, shared infrastructure, and a consistent operational posture —and economics that improve with breadth rather than degrade—changes the calculus for teams managing applications they did not build.

4. The security configuration nobody made 

The forty-engineer company has no security team. They have a senior developer who reads the CIS benchmarks on weekends. The compliance audit arrives regardless. Teams without specialists will not configure security controls that require specialist configuration. This is not a criticism of those teams; it is a structural observation about how security actually gets implemented (or does not) in organizations where every engineer is a full-stack contributor with a product backlog that never shrinks.

The only security posture that works reliably for these teams is the one they inherit by default: compliance certifications, network isolation, access controls that ship with the platform rather than requiring a dedicated sprint to implement.

What these patterns demand today

Every failure mode we observed the deployment nobody touches, the observability sprint that ships late, the portfolio tax, the security configuration nobody made- shares a root cause: the platform asked the team to make an operational decision, and the team either made the wrong one or made none at all. The response to these patterns required starting from a different question. Not “what should we configure for the team?” but “what should the team never need to decide?”

A platform that answers that question correctly holds a specific set of commitments. It decides what healthy looks like before the first request arrives, not after the first incident. It ships observability at deploy time, not as a sprint the team schedules after something breaks. It treats the eighth application in a portfolio the same as the first, same operational model, same governance, same economics. It inherits security posture by default, because the teams it serves will never staff a dedicated security function.

These are not feature decisions. They are architectural decisions about where operational responsibility permanently resides. The team defines the application source code, a Dockerfile, a pre-built image, or an existing workload being migrated in. From that point forward, the platform owns everything underneath it. Not for the first deploy. For the life of the application. Patching, scaling, healing, certificate rotation, capacity planning, health evaluation. These are not capabilities the team enables. They are responsibilities the platform holds permanently.

This is what we rebuilt AWS Elastic Beanstalk to be. Not a deployment tool. Not a hosting layer. An application management service that takes operational responsibility for everything underneath the application. The architecture now starts from the question above and refuses to let the answer drift back toward the team over time. Elastic Beanstalk operates in two modes, a structural change from its previous single-environment architecture:

Standard Mode delivers full operational ownership for individual applications and Windows/.NET Framework workloads: the complete operational stack, owned outright, for a single service.

Cluster Mode extends the same ownership model across the portfolio, shared infrastructure, source-to-production deployment that transforms code into running applications, and economics that improve as the portfolio grows. The eighth application shares operational overhead with the first seven rather than duplicating it. For the forty-engineer company running eight production applications today and inheriting ten more next quarter, this is the difference between a platform that covers the portfolio and a platform that covers only the applications simple enough to fit its opinions.

The industry convergence

The distinction is real, though I would not draw it as a line between platforms that reduce complexity and platforms that own operations permanently. Every vendor in this market absorbs some operational responsibility at deploy time. The key question is how much of it returns to the team during an incident, a patch cycle, and an audit. A platform that removes infrastructure management from developers during the workweek and reintroduces it at 3 a.m. Sunday addresses only half the challenge.

“A platform that removes infrastructure management from developers during the workweek and reintroduces it at 3 a.m. Sunday addresses only half the challenge.”

At convergence, the direction is correct, but the shape is incorrect. This is not two camps meeting in the middle. Gartner’s 2026 Magic Quadrant for Cloud Native Application Platforms places AWS, Microsoft, Google, and Red Hat in the leaders quadrant, with Render, Netlify, and Upsun as niche players. Vendors specializing in developer experience showed this category is viable, and now the hyperscalers are adopting it. Since source-code-to-URL mapping is now standard across the entire quadrant, the key differentiation becomes who bears operational liability for the eighth application three years after its release.

The only aspect of framing I would challenge is the idea that a platform determines everything the team never has to decide. Routine infrastructure decisions should stay out of the developer’s path, and escape hatches should stay in place for the teams that genuinely need them. A platform that removes choice altogether will demo well and then stall when teams migrate applications that don’t fit its opinions.

The 3 a.m. test that actually matters

For teams already living this reality — serious production, lean staff, growing portfolios — nobody planned to operate the 3 a.m. test; it is not a nice-to-have. It is the evaluation criterion.

The platforms that define the next decade will not simply make deployment easier. They will decide, in advance, how production systems should behave when things inevitably go wrong. Because by 3 a.m., the time for deciding has already passed. The CNAP category was built to describe platforms that own the application lifecycle.

Elastic Beanstalk made those decisions before the incident arrived: what healthy looks like, what to do when it stops being true, how to communicate what happened. The cloud gave teams power. These teams needed someone to stay. Elastic Beanstalk stays.

Check the latest AWS release notes for Elastic Beanstalk here.

The post The cloud reduced operational complexity. But many teams need someone to own it completely. appeared first on The New Stack.

AI coding agents need a secrets-safe context boundary

22 septembre 2026 à 15:00
Abstract glowing blue and yellow distortion wave on a black background, illustrating digital data security concepts.

AI coding agents play a major role in software development and delivery, and for good reason. They can investigate bugs, trace dependencies, refactor services, and propose patches without developers needing to assemble all of the relevant context manually. That capability comes courtesy of agents’ appetite for context. To make informed decisions, agents read source code, configuration files, terminal output, error messages, environment information, and more…much, much more.

“Secure, agentic development depends on a security control many teams still lack: preventing secrets from leaking to AI coding agents and becoming model context.”

From a security standpoint, this becomes problematic when agents, in the search for context, inadvertently reach for secrets.

For years, developers have been taught not to commit API keys, database credentials, and tokens to Git. But agentic workflows have created another route for secrets to escape development environments before a commit, code review, or CI job. Depending on its permissions, configuration, and provider architecture, an AI coding agent may read local files or receive pasted content that is then included in data sent to an AI service, and, in the process, developers may never see their credentials leak.

The quiet path from local files to external systems

Some forms of secrets leakage are obvious. A developer troubleshooting an authentication failure may paste, for example, a failing API call into a chat window, including the token. However serious, this sort of leak is characteristically human.

The more consequential escape pathway is quieter. An agent tasked with understanding a project may inspect files in its working directory, including an overlooked .env file, a cloud credential profile, an SSH configuration, or sensitive application logs. In such instances, nothing has necessarily gone wrong from the agent’s perspective; it is doing exactly what it was designed to do: collect context to solve the task at hand.

“Agentic workflows have created another route for secrets to escape development environments before a commit, code review, or CI job.”

But once a secret becomes part of that context, it may pass through systems outside an organization’s direct control. Depending on the workflow, it can appear in model provider logs, gateway telemetry, prompt histories, or debugging records. Rotating the credential is essential, but it does not erase copies that may already exist in those systems.

This changes the practical definition of a secret leak. The problem is no longer limited to what lands in a repository, but also includes what an autonomous tool reads and forwards while operating on a developer’s machine.

Why traditional security gates no longer suffice

Most application security programs are built around durable checkpoints: the commit, pull request, build, and deployment. In the agentic era, these checkpoints remain important as they can detect secrets that reach version control and prevent a bad change from merging and deploying.

They cannot, on their own, prevent a secret from being included in an agent prompt before the code ever reaches a repository.

This highlights an important timing gap. The 2025 Verizon Data Breach Investigations Report reports a median of 94 days to remediate leaked secrets discovered in GitHub repositories. In an agent-driven workflow, detection and response need to happen much earlier, and not after a credential is exposed. Still, at the moment it’s about to cross the boundary from local context to an external model.

Bad actors already understand the value of that porous boundary. Recent supply-chain attack campaigns, including Mini Shai-Hulud, have searched developer and CI environments for credentials and configuration data, including AI coding-tool configuration files. These campaigns show that agent configurations and the local context accessible to an agent are valuable targets. AI coding agents can broaden the local data reachable during a session, making even the agent’s context-collection mechanisms an attractive target.

Treat agent context as an egress surface.

The secure mental model doesn’t frame AI agents as mere code editors, but rather, automated data-movement systems. Its inputs can include far more than the source files a developer is actively editing, and its outputs may involve external services.

That calls for a zero-trust approach to agent context. Before sending a prompt or adding a file to an agent’s working set, organizations should evaluate it for sensitive material. Controls should be deterministic: identify a likely secret, block or redact it, and provide the developer with a clear path to remediate it.

“Asking an LLM to decide whether to transmit a credential does not create a reliable security boundary.”

Critically, the control should be independent of the model. Asking an LLM to decide whether to transmit a credential does not create a reliable security boundary. Purpose-built secrets detection can inspect prompts and files against known credential patterns and policies, applying a deterministic policy, such as blocking a prompt or file read when it detects a credential-shaped value. For example, Sonar’s secrets detection ships alongside dedicated agent plugins to bring that local check into tools such as Claude Code, GitHub Copilot, Codex, and Cursor, so it can flag a credential before a prompt or file read is transmitted to a model provider.

Build defense in layers, without disrupting your agentic workflow

A legitimate workflow does not involve forcing developers to choose between secure development and useful automation, but instead places fast controls at several points where secrets can escape:

  • In the editor: Use IDE-integrated secrets detection to flag credentials while they are being written.
  • Before model submission or agent file access: Where the agent supports it, scan prompt submissions and file reads locally, and block risky operations according to policy.
  • At the command line: Check generated snippets and local changes in terminal-driven workflows.
  • In pull requests and CI: Detect secrets that reach the repository and use review, quality gate, and deployment controls to prevent unsafe changes from progressing.
  • In incident response: Rotate exposed credentials quickly, investigate downstream logs and access, and reduce recurrence through policy and training.

Building defense at the pre-submission layer is an emerging requirement and requires both security and usability. Secrets detection must be fast enough to run in developer workflows; a scanner that introduces lengthy pauses may be bypassed or disabled by developers, and it must also have a manageable false-positive rate, or developers may stop trusting it.

Teams should also make their agent permissions and context rules explicit, as broad agent permissions can increase the amount of sensitive local context reachable during a coding session. Consider the following: which directories can an agent read? Are .env files, credential stores, home-directory configurations, and production logs excluded by default? Does the organization route prompts through an approved gateway? What retention, training, and audit settings apply at the provider level? Document and enforce the answers rather than leaving them to individual developer preference.

Secrets security must shift left.

Prevent secret leakage without hindering AI-assisted development, ensuring the productivity promise of agentic development doesn’t carry significant security implications.

As agents become more autonomous, security standards must follow agents upstream. It’s critical to stop a secret before it becomes context, while it is still local, visible, and easier to control. In the agentic era, code review and CI-level checks will remain essential safety nets. Still, for agent-centric development, the first line of defense must shift left: to the instant an AI coding tool determines what to read and what to transmit. That is the control modern development teams need to implement now.

The post AI coding agents need a secrets-safe context boundary appeared first on The New Stack.

One engineer shipped 2,000 PRs a month to production. Verification is the key.

19 septembre 2026 à 16:00
Abstract dark digital wireframe mesh with chromatic glitch effects, representing AI agent verification and virtualized software environments.

Lauren Tan, an engineer on the Grok team at SpaceXAI, who previously worked at Cursor and Meta, recently published a guide to her personal agent workflow: pstack. The attention-grabbing number is that pstack has let her ship 2,000 pull requests (PRs) a month to production with high confidence. That’s one engineer shipping nearly 100 PRs per working day.

The number is incredible and undeniably an outlier, but the direction is not a surprise. I have argued previously that coding agents would enable teams to generate ten times the code with similar headcounts. What is surprising is that these are not just code output numbers. These are actual changes landing in production.

According to Tan, the most critical piece of that workflow is verification. A verification skill lets an agent check its own work and keep going until the task is done, and she treats it as “critical infrastructure” rather than one skill among many. 

The verification skill rests on something underneath it: a rich runtime the agent can drive, inspect, and get structured answers from. For a single application, that runtime is the application itself, started on demand. For a system made of dozens or hundreds of services, no such runtime exists by default, and providing one that keeps up with hundreds of parallel agents is the hard part.

Verification is the whole game, and the math says so

Her argument for agentic verification is a throughput argument. An agent that can check its own output keeps working until the task is done. An agent that can’t hand you a diff and wait makes you the slowest component in the loop. That is why she claims strong verification skills can multiply a team’s output by 100 to 1,000 times.

“An agent that can check its own output keeps working until the task is done. An agent that can’t hand you a diff and wait makes you the slowest component in the loop.”

At 2,000 pull requests a month, reviewing every change by hand would allow about five minutes per PR across a full working month. Human review cannot be the verification layer at that volume. Whatever does the checking has to run without a person in the loop, and it has to run in parallel with the agents generating the work.

The model assumes the agent can run the whole application

The verification skill she describes generates a command line interface (CLI) and a feature map for the application. The CLI lets an agent start the app, navigate it, inspect state, and read structured JSON results back. Each agent gets a complete copy of the application and can test a change end to end.

She is direct about how much rests on that runtime: “I personally feel that agentic verification is so important that I would unironically suggest building your own rich debugging tools, or even choosing a different tech stack, in order to have unfair advantages and extreme productivity in building software.”

Her approach to providing a runtime for her agent works because the application fits in one process. A frontend, a compiler, or a single service with a database can start from a CLI in seconds and be thrown away afterward.

“I would unironically suggest building your own rich debugging tools, or even choosing a different tech stack, in order to have unfair advantages and extreme productivity in building software.”

For teams building complex distributed applications, their system does not have that property. The application is the interaction between an order service, a payments service, an inventory service, a queue, several databases, and a handful of third-party APIs. At larger shops, the count runs into the thousands. A pull request to one service is only verified by exercising the calls it makes and receives. The CLI can start the changed service. It cannot start the system.

None of the existing runtimes survive hundreds of parallel agents

Local runtimes with mocks are cheap and can run fully parallel using worktrees or CDEs. Their problem is fidelity. Mocks encode what a dependency did the last time someone looked, and they drift the moment the real service changes. An agent that verifies against mocks closes its loop against fiction, and the failure shows up after merge.

A full copy of the stack per change is faithful and isolated. But its cost scales with the number of services times the number of concurrent changes, and at hundreds of agents that cost is untenable. Time is the bigger problem. A full stack takes minutes to provision, and the loop she describes has the agent testing every iteration of a change while it is still working on it. An environment that is ready after the agent has moved on to its next attempt is no use to it.

Shared staging is faithful and cheap because there is one of it, and that is the whole problem. A single mutable environment cannot host hundreds of concurrent changes. Agents overwrite each other’s deployments, a broken change from one agent becomes failed tests for every other agent, and the loop-closing property that makes the workflow valuable disappears.

Each existing verification runtime plotted on a chart showing its realism of dependencies against concurrency.

What verification needs when the callers are agents

Read Lauren’s workflow as a requirements document, and five properties emerge:

  • The change has to run against real dependencies or the verification means nothing.
  • Hundreds of concurrent changes have to be unable to see each other.
  • The cost of an environment has to scale with the size of the change, not the size of the system.
  • Environments have to come up in seconds, because an agent waiting on provisioning is parallelism you paid for but didn’t use.
  • All of it has to be reachable through the CLI or MCP server the agent already uses, because the caller is an agent.

The first and third requirements pull in opposite directions. Realism pushes toward complete copies of the system. Cost pushes toward sharing as much as possible. Shared staging resolves that tension by giving up isolation, and a per-change full stack resolves it by giving up cost efficiency. A design that satisfies all five has to share and isolate at the same time.

Virtualized full-stack environments share the system and isolate the change

The architecture that does this treats an environment as a view of a running system rather than a copy. One shared set of stable services runs continuously, deployed from the main branch and kept healthy the way production is. When an agent needs to verify a change, it runs only the service it modified, on its own machine or as a lightweight deployment in the cluster, and joins it to the shared stack as a new isolated environment.

“The architecture that does this treats an environment as a view of a running system rather than a copy.”

From inside that environment, the changed service is the version of record, and every other call falls through to the shared stable versions. The agent sees a complete, realistic system, and so do the other hundred agents, each seeing a system that differs from the baseline by only the delta of its own change. Requests carry their environment identity as they cross service boundaries, which keeps one agent’s traffic from reaching another agent’s version under test. Stateful side effects that cannot be shared safely, like queue topics or writable databases, get a per-environment copy where needed.

Diagram showing "agent 1 env" and "agent 2 env" interacting with the shared cluster.

The cost model follows directly. An environment costs one or two running services instead of sixty; it is ready in the time a single service takes to start, and you can create and destroy it from inside the agent’s own loop. This is the pattern Signadot packages for Kubernetes, with the shared stable stack running in the team’s existing cluster.

The loop, end to end, with an agent as the actor

Put the two halves together, and the workflow that enables her to ship 2,000 PRs a month to production carries over to a distributed system almost unchanged. An agent picks up a task and changes one service. It asks for an environment for that change and gets one in the time it takes its service to start. It then drives real requests through the system’s entry point and watches them traverse the real dependency graph, with only its own service running new code. It reads structured results, fixes what failed, and runs again. When the checks pass, it opens the PR, and the environment goes away at merge.

“Environments stop being something the platform team hands out and become something agents create, use, and discard as needed.”

For the platform team, the unit of work changes. Today it provisions environments, whether that means keeping a shared staging alive or stamping out full copies of it. In this model, it runs one shared stable stack and the layer that virtualizes it: context propagation across every service, isolation for the stateful dependencies that cannot be shared, and the tooling that creates and tears down environments. Environments stop being something the platform team hands out and become something agents create, use, and discard as needed.

Verification capacity is the new ceiling on throughput

Lauren Tan’s post is not a story about one unusually productive engineer. It shows what happens when agents run the full loop, writing a change, verifying it, and iterating without a person in between. The verification infrastructure is the foundation that the entire loop stands on.

In distributed applications, that infrastructure has to be a runtime environment that gives every agent real dependencies, keeps hundreds of concurrent changes from seeing each other, costs a change rather than a copy of the system, and is ready in the seconds an agent is willing to wait. That is what turns agent parallelism into shipped code rather than a longer review queue. That model of runtime environments is exactly what we built Signadot to enable.

The post One engineer shipped 2,000 PRs a month to production. Verification is the key. appeared first on The New Stack.

How buildpacks help enterprises finally operate container security controls at scale

18 septembre 2026 à 16:00
Dark abstract geometric grid texture symbolizing standardized container security controls.

Container security controls often fail less because organizations lack standards, scanners, or best practices, and more. After all, every service tends to use its own Dockerfile, building images in slightly different ways.

Buildpacks are often pitched as a way to make containerization more developer-friendly, standardized, and maintainable by eliminating the Dockerfile pain points. They can also help improve security posture in practice, especially in larger, multi-team environments.

In this article, we will explore how Buildpacks provide a standard application-image path where platform teams can centrally govern build inputs, runtime images, Software Bill of Materials (SBOM) production, artifact metadata, and update workflows–making container controls easier to enforce and operate at scale.

The patch that never reached production

What does a patch integration process look like in the real world?

In more than one enterprise, I’ve seen the story look roughly like this: a critical vulnerability is fixed in the organization’s approved runtime image. In a perfect world, a patched base image is detected centrally, rebuilt, tested, published, and automatically propagated to every affected application. But we do not live in a perfect world. A week later, some production workloads still run on the vulnerable base.

“A container security control isn’t useful on its own. It must be applied consistently, observable in the running estate, and maintainable when images, dependencies, and vulnerabilities change.”

Typical reasons include:

  • Inconsistent Dockerfiles using different base image tags, Linux versions, and runtime versions, making it harder to determine which service is affected;
  • Application images may start from a patched base but reinstall vulnerable dependencies because developers manually install packages;
  • Uneven rebuild cadences: some teams rebuild more frequently, others only rebuild when the application code changes, so a service with no recent feature work may run an old vulnerable image for months;
  • No centralized inventory, SBOMs, image metadata, or deployment tracking, so organizations cannot reliably determine which applications remain vulnerable or measure patch propagation. 

This story has not yet found its happy ending because the organization is only halfway into establishing a vulnerability management process. A container security control in place, such as “all services must use approved patched base images,” isn’t useful on its own. It must be applied consistently, observable in the running estate, and maintainable when images, dependencies, and vulnerabilities change.

That’s the stage where many teams fall short.

Why container security controls drift

Container security controls are hard to implement largely because of established containerization practices. In many organizations, no single governed build path exists. Each repository produces its own image, usually through a Dockerfile maintained by the application team.

Dockerfiles offer flexibility. They also push many security decisions onto developers: which base images to use, which packages to install, how to configure the runtime user, how minimal the image should be, how to track patches, which CI policies to apply. Over time, as services multiply and teams change, these choices drift. It’s not unusual to see different repositories using different base images, update cycles, and even different interpretations of what “secure” means in practice.

“The problem is expecting every developer to have enough container expertise to do this consistently across hundreds of repositories.”

Dockerfiles are not the problem by themselves. A well-written Dockerfile can produce a minimal, hardened image. The problem is expecting every developer to have enough container expertise to do this consistently across hundreds of repositories.

This approach also makes patch propagation unreliable. Platform teams may publish a patched base image, but each application team must still notice the update, modify its Dockerfile, rebuild, test, and redeploy. Some do it quickly. Others do it late or not at all. The control exists, but adoption remains uneven.

Buildpacks help close this gap by moving common containerization decisions out of individual repositories and into a shared, governed build path.

The shift from repository-specific builds to a governed build platform

Buildpacks turn application source code into a production-ready OCI container image without a Dockerfile. They detect the application type, select the required buildpacks, provide the necessary runtime and dependencies, and produce a runnable image.

For container security, the main benefit is not simply removing Dockerfiles. Buildpacks tend to provide a standard way to construct images, at least for typical application stacks.

“Buildpacks turn application source code into a production-ready OCI container image without a Dockerfile.”

This gives platform and security teams a central point for enforcing controls. Instead of asking every team to choose an approved base image, configure the runtime, manage layers, generate metadata, and track updates, the organization can encode much of this work in shared builders and buildpacks. Developers own their code, dependencies, and service behavior. Producing a compliant image becomes a platform responsibility.

Because Cloud Native Buildpacks is a CNCF graduated project, its specifications and reference implementations undergo rigorous community review and long-term maintenance, making them suitable as the foundation for enterprise security controls.

With that foundation in place, applying and maintaining specific container security controls becomes easier at scale.

Four container security controls buildpacks make it easier to operate

Standardization of build inputs

The first control is standardizing what goes into the build.

In Cloud Native Buildpacks, the key unit is the builder. A builder packages the buildpacks, lifecycle, build-time base image, and runtime base image used to create the final application image. This establishes the builder as a controlled definition of how application images are produced. Developers do not choose a random base image on Docker Hub to build their applications; the Buildpacks ecosystem defines a set of build and run images.

This builder-based approach introduces an important concept. A developer cannot change the base OS layer in a builder with a single line of code because compatibility isn’t guaranteed. With one line of code, however, they can swap builders and still get a compatible, working image.  

This is where standardization takes place. Developers keep using their preferred languages and frameworks, but the platform team builds images from a controlled set of approved builders. Extension paths for cases where buildpacks require customization can also be standardized.

“By controlling the builder, the organization also controls the buildpacks, runtime image family, lifecycle version, and build paths allowed in CI/CD.”

The security benefit is straightforward. By controlling the builder, the organization also controls the buildpacks, runtime image family, lifecycle version, and build paths allowed in CI/CD. The policy is defined once at the platform level and then applied across many services.

Best container security practices by default

After the builder standardization, the next question is: what does the resulting image look like? This is where buildpacks help. They improve the output by applying several container security practices as part of the normal image creation flow:

  • Non-root build and execution. Cloud Native Buildpacks require buildpack code to run as a non-root user. Platforms, such as pack, also produce images configured to run applications as the non-root user defined by the run image. Non-root defaults reduce the blast radius of a compromise and make container escape, filesystem tampering, and exploitation of the image build process more difficult.
  • Separation of build and runtime environments. Buildpacks mark layers as build-only or launch-time. Only launch layers are included in the final image, so compilers, npm tooling, build caches, and similar tools remain outside production, reducing the attack surface.
  • Restricted modifications of the base image. Because buildpacks run without root privileges, they cannot simply install OS packages or modify the base filesystem. OS changes must be provided through the controlled build/run images or image extensions.
  • Isolation of sensitive build privileges. When using an untrusted builder, sensitive lifecycle phases can run separately from the phases that execute buildpack code. This prevents an untrusted buildpack from receiving capabilities such as registry credentials and access to the container daemon.

Some builder providers offer additional security features. For example, Paketo buildpacks for Spring Boot provide base images without a shell. Another example is BellSoft’s hardened builder for Paketo buildpacks based on BellSoft Hardened Images.  

These defaults reduce the manual effort and make secure container builds part of a standard process.

SBOM Generation

A Software Bill of Materials (SBOM) lists all software components in an application or image. An SBOM is required for vulnerability tracking, license reviews, and audits, and many regulations require it.

Teams often add SBOM generation as a separate CI step, with separate tools, formats, storage rules, and ownership. That makes coverage uneven, especially across many repositories.

Cloud Native Buildpacks reduce this friction by emitting an SBOM for all dependencies they provide, in formats such as CycloneDX, SPDX, or Syft JSON. As a result, platform and security teams get consistent image inventory data without requiring every application team to build its own SBOM process.

For a typical Java service, that might mean SBOM entries for the JRE version, Spring Boot version, and key libraries pulled in at build time. For Node.js services, it might include the Node runtime version and major npm dependencies. The exact content will vary, but the pattern is the same: SBOM data comes from the build itself, not a separate, manually maintained process.

Patching at scale

As discussed earlier, the hard part is often not writing or receiving a patch, but getting it into the running workloads. Buildpacks shorten that path by using shared builders, buildpacks, and run images. Platform teams can then update these common inputs once, instead of waiting for every application team to repeat the same change.

Another powerful feature of Buildpacks that promotes rapid patching is rebasing. When OS-level fixes become available, the runtime base layers can be replaced with layers from a newer run image without rebuilding the application from source.

Kubernetes tools such as kpack can automate this process. They track image resources and trigger rebuilds when the source, builder, buildpacks, or stack changes.

Rebasing has limits: it updates only the run-image layers. Dependencies added by buildpacks, such as a JRE or Node.js runtime, usually require a rebuild with updated buildpacks or dependency versions. For a Node.js service, a rebuild might mean picking up a new Node.js runtime version and updated npm packages, not just a newer OS base.

So, we are not talking about magical automatic patching here. The key benefit is a more centralized, repeatable patch-propagation path across the image estate—though teams still need to handle testing, rollouts, and exceptions.

The new patch path after buildpacks

Let’s return to the original story. We left the organization in a state where vulnerability management lacked consistent patch integration. Suppose the enterprise has migrated their workflows to Buildpacks — how has the process changed?

Once again, a critical vulnerability is discovered. But after adopting buildpacks, the response looks different: the remediation starts with shared build inputs.

The platform team updates the approved run image, builder, or buildpacks. Images are then rebased or rebuilt:

  • Rebase when the fix affects the run-image OS layers;
  • Rebuild when the affected component is runtime, like a JRE or Node.js runtime, or application dependencies.

Of course, Buildpacks do not remove the need for testing or deployment controls. Application teams still own their code, dependencies, and compatibility testing. SRE teams still promote, deploy, monitor, and, when necessary, roll back the patched images.

“Buildpacks give organizations one controlled way to build container images instead of leaving every repository to define its own process.”

The most important change is the patch path. Platform teams maintain the shared build inputs and automation. Security defines scan policies and exception rules. Compliance defines the evidence that must be retained.

This path also promotes shifting security left, as it becomes easier for developers to meet the in-house security requirements and container security best practices.

FunctionResponsibility
PlatformBuilders, run images, buildpacks, and rebuild automation
SecurityVulnerability policy and exceptions
Application teamsCode, dependencies, and compatibility testing
SREPromotion, rollout, monitoring, and rollback
ComplianceAudit evidence and retention

Some workloads still need custom images. But you don’t need to choose between Buildpacks and Dockerfiles; you can use both. When required, you can create a custom buildpack and extend the build-time base image with a Dockerfile. The result is not a rigid “Buildpacks only” model, but a governed image strategy with controlled customizations. 

Buildpacks give organizations one controlled way to build container images instead of leaving every repository to define its own process. This makes security controls–such as approved images, safe defaults, SBOMs, and patching–easier to scale. The result is reduced drift, clearer ownership, and faster updates.

If you’d like to get started with Buildpacks, try them on one service and compare the workflow with your current Dockerfile process.

The post How buildpacks help enterprises finally operate container security controls at scale appeared first on The New Stack.

Your agent is only as good as your infrastructure

18 septembre 2026 à 15:00
Dark abstract 3D glass ribbon rendering symbolizing complex AI agent infrastructure and bursty data workflows.

You built a great agent, but something happened when it moved into production.

In testing, your agent reviewed pull requests efficiently on its own. It read the diff, grepped the codebase for related usages, ran the test suite, checked whether CI was still red from an earlier commit, and drafted a comment—all before you’d finished reading the diff yourself.

In production, however, imagine the same five steps ran behind every other PR review your team’s agents performed that hour. Some reviews landed in seconds; others sat for minutes because the test-suite step landed on a node that was mid-burst from someone else’s agent. 

The agent didn’t change. The execution environment did, and that’s what decided whether review time held steady or crept up.

Agentic applications introduce a different execution pattern than traditional chat applications. As those workflows become longer and more dynamic, infrastructure has a much larger influence on latency, reliability, and cost than it does for a simple chatbot.

That’s why your agent is only as good as your infrastructure.

Agents aren’t chatbots with more steps

The difference between serving inference for a chatbot vs. an AI agent isn’t simply that one is “more capable.” They execute work differently:

  • A chatbot usually makes one inference call per user message. The model receives a prompt, generates a response, and waits for the next user input before proceeding.
  • An agent executes the entire workflow, turning one user message into a chain of inference calls. It might decide to search documentation, retrieve data from a database, call an API, execute code, evaluate the result, and then repeat that process before producing an answer. Each of those decisions may trigger another inference call, and every result becomes additional context for the next step.

“Infrastructure has a much larger influence on latency, reliability, and cost than it does for a simple chatbot.”

That execution model changes the infrastructure requirements for AI agents. 

One question, many steps behind it

Instead of optimizing for individual inference requests, the system has to support long-running workflows whose latency and reliability depend on every component in the chain.

A single user request often expands into a sequence of inference and tool execution steps, sometimes called multi-turn tool calls or the agentic loop. Rather than generating one response, the model alternates between reasoning and interacting with external systems.

For example, you ask an agent why checkout latency spiked overnight. The agent pulls the deploy log, queries the monitoring system, runs a diagnostic against the connection pool, weighs whether the culprit is a bad deploy or a capacity issue, and then folds that into another inference call before finally producing a full-fledged response.

Each reasoning step becomes another inference request, and every tool result is added to the model’s context before the next step.

This workflow changes what reliability means

Multi-turn workflows are inherently sequential, which is why even low latency can quickly add up to a significant amount. Every inference step waits for the previous one to finish. If a database query takes two seconds, the model can’t begin the next reasoning step until that result returns. The model may generate tokens quickly, but the other steps slow it down.

“In this agentic workflow, every step in the chain has to hold, because the chain is only as strong as its slowest link.”

In this agentic workflow, every step in the chain has to hold, because the chain is only as strong as its slowest link. Instead of processing isolated inference requests, the inference stack has to orchestrate a chain of dependent model invocations and external tool calls. As those workflows become longer, the stack increasingly determines how quickly, reliably, and cost-effectively the application performs.

But the user doesn’t see an orchestration hiccup. They see an agent that hung or gave up.

That’s why end-to-end agent reliability depends on much more than model quality. Infrastructure determines whether each step has the resources it needs to execute predictably under load.

Why the bill and the performance both feel unpredictable

A second difference in agentic workflows catches teams off guard: demand patterns and their impact on your inference bill. 

Most inference services, and the pricing built on top of them, assume traffic arrives at a predictable pace. A typical inference solution knows the predictable demand pattern: User traffic increases, request volume increases, and capacity scales accordingly. Cloud infrastructure is typically optimized for these steady request patterns using mechanisms such as autoscaling, load balancing, and capacity planning.

Agent workloads don’t behave that way. Individual workflows pause while waiting on external systems, then resume as soon as new information becomes available.

  • The pause: The agent waits on an external API or database, so the GPU serving that workflow has no inference work to perform, and its accumulated context may be evicted from GPU memory while it waits
  • The burst: As soon as external systems return results, inference resumes simultaneously across many workflows, creating short and sharp spikes in GPU demand, each re-processing its full accumulated context

If you’re watching GPU utilization and it looks less like steady load and more like a heartbeat — flat, then a spike every time tool results come back — that’s the signature. It means you’re provisioning for the average when you should be provisioning for the peak, and it’s usually the first place p99 latency quietly blows out.

“If you’re watching GPU utilization and it looks less like steady load and more like a heartbeat, that’s the signature.”

Inference services designed around steady or predictable request streams can struggle to allocate resources efficiently under these conditions. This leads to inconsistent latency, GPU underutilization, or higher operating costs.

When performance becomes unpredictable, or a bill doesn’t match what you expected, it’s evidence your infrastructure was built for a different workload than the one you’re actually running.

What infrastructure built for agents actually looks like

Agentic applications place different demands on infrastructure than traditional AI workloads: long dependency chains, bursty demand, and continuous evolution. Because of this, the infrastructure is deciding whether the chain holds, and whether the bill holds too.

For an agent to run, it needs infrastructure purpose-built to support:

Performance that holds across the whole chain. The infrastructure must keep latency consistent across multi-step and multi-tool workflows.

Scalability that responds to bursty demand. Infrastructure should scale quickly as inference demand fluctuates, without requiring capacity to remain provisioned during idle periods.

  • Predictable economics even for dynamic workloads. The infrastructure bill should reflect actual usage.
  • None of this makes bursty demand disappear, but it changes how the system absorbs it. A large enough simultaneous burst, or a workflow that accumulates enough context before pausing, still costs something. The goal isn’t zero cost or zero limit; it’s making both predictable.

Your agent is only as good as your infrastructure. Get it right, and your agent’s responsiveness, reliability, and cost-effectiveness will improve your work. 

Learn more about how infrastructure can be purpose-built for agentic workflows: Check out the documentation to get started.

The post Your agent is only as good as your infrastructure appeared first on The New Stack.

Code review is burning out your best engineers

18 septembre 2026 à 14:00
Dark, deeply fractured rock surface symbolizing structural tension and developer burnout

Every team I talk to has the same problem. Their best engineers, the ones who care most about code quality, are drowning in review queues they can’t keep up with and don’t enjoy. Some experienced engineers complain, and some point-blank refuse to review AI-generated code.

I run a community of senior engineers and engineering leaders, and they all name AI code review bottlenecks among their top concerns. In one study, 77% of engineers said they spend less time writing code now. They put that time into reviewing AI output.

The job shifted from crafting to verifying

Teams with high AI adoption are merging 98% more PRs, and review times went up 91%. Engineers didn’t sign up to spend their days reading machine-generated diffs. But that’s increasingly what the job demands.

The engineers who feel this most aren’t the ones resisting AI. They’re the ones who adopted it first, care most about code quality, and built the review culture their teams depend on. Those same people now have 15 PRs, 400 lines of code each, in their queue every day.

Why reviewing AI-generated code is harder

When a colleague writes code, the intent travels with them through the review process. They can explain the tradeoffs they considered, the alternatives they rejected, and the constraints they worked within. Even if unwritten, that context is accessible.

“When AI writes code, the reasoning is gone. The reviewer is left reverse-engineering intent from a diff.”

When AI writes code, the reasoning is gone. The reviewer is left reverse-engineering intent from a diff. That is a fundamentally different cognitive task. What makes it worse is that AI-generated code passes the eye test.

Five shades of AI slop

Plausible but wrong. The code reads coherently and handles the happy path, but edge cases reveal misaligned assumptions. These bugs are difficult to catch in review because they require understanding what the code was supposed to do, not just what it does.

Over-engineered. AI models are trained on vast bodies of code, including enterprise patterns and production-hardened architectures. Asked to solve a problem that really needs 15 lines, a model may produce a 200-line abstraction layer that anticipates a generality nobody asked for.

Convention-blind. Models generate good generic code, not code that fits your system. Your repo has conventions around naming, error handling, logging patterns, module boundaries. AI frequently ignores them.

Confidently hallucinated. Calls APIs that don’t exist, uses deprecated methods, invents config options. Sometimes caught immediately, sometimes only in production.

Cargo-cult patterns. Copies structures without understanding why. Retry logic where retries make no sense. Circuit breakers for calls that are always synchronous. Error handling that looks thorough but doesn’t map to actual failure modes.

The common thread is that it looks like real code, which makes it hard to review at scale.

How to fix the code review

The answer is not “review harder” or “add an LLM reviewer.” When the same model writes and reviews the code, it shares its own blind spots. If you add adversarial agents and multiple steps, the process becomes a theater of multi-step workflows that turns engineers into bot-sitters, spending time configuring and tuning filters instead of building.

What works is shifting the burden off reviewers in three places: codify repeated feedback, preserve the intent that produced the code, and measure the work that actually prevents slop.

Create your AI slop registry

Pull your team’s last 100 PR review comments. Sort each one: Is it deterministic, something a rule can check? Is it execution-testable, something you can catch by running the code? Or is it genuine judgment?

When teams run this exercise, the rough split is 45% deterministic, 30% execution-testable, and 25% judgment. Three-quarters of review feedback is codifiable.

“Three-quarters of review feedback is codifiable. Every recurring review comment is an invariant you haven’t written yet.”

Every recurring review comment is an invariant you haven’t written yet. “New endpoints must have OTel spans” is not a judgment call. It’s an AST check. Write it once. It never needs a reviewer again. The test for promoting something to an invariant is recurrence: if you’ve posted the same comment more than once, it should be codified.

Preserve the reasoning trail

The prompts and agent sessions that produced your code hold the intent. Most teams throw them away. That’s like deleting the commit messages and PR descriptions and expecting reviewers to reconstruct intent from diffs alone.

At Aviator, we built Verify around this problem. It captures intent from prompts and agent sessions and structures it as acceptance criteria: what the change does, what’s out of scope, and how to tell if it worked. The decisions an engineer makes while talking to the agent, the architectural choices, the scope calls, the behavior tradeoffs, become reviewable acceptance criteria.

The reviewer reads a list of acceptance criteria and asks, “Are we solving the right problem with the right constraints?” That’s the high-value work for senior engineers. Not reading a 400-line diff at 4 p.m. Code is actually the least important part of reviews. What matters is intent: acceptance criteria, non-goals, blast radius.

How knowledge sharing survives

Reviewers reading specs and acceptance criteria are reading decisions, not scanning syntax. They’re debating tradeoffs, understanding how the system is evolving, seeing what constraints shaped the approach. That’s where knowledge sharing survives. If we move code review left, knowledge sharing has to move left too.

Measure and reward verification work

31% more PRs are being merged without any review at all. That’s engineers voting with their behavior.

Dashboards measuring AI adoption and productivity in lines of code will never show the work of senior engineers carrying the review burden. They will never surface the effort that goes into building the systems and guardrails that prevent slop. If you’re measuring throughput and cycle time and feeling good, you’re measuring the wrong thing.

Annie Vella has been tracking this shift across 158 engineers in 28 countries. Her observation: engineers are resigning, some hoping the role will return to what it was, others leaving the profession entirely. The shift toward verification-heavy work is turning the job into something they don’t enjoy.

“Those dashboards don’t show the senior engineer who spent her afternoon reverse-engineering intent. They show throughput. And throughput looks great right up until the people carrying the review burden walk out the door.”

The engineers who carry the review burden aren’t complaining. They’re quitting. Some leave for teams with better tooling. Some leave engineering entirely, not because they can’t keep up, but because the work stopped being the work they signed up for.

Leaders chasing lines of code generated and PRs merged will never see this coming. Those dashboards don’t show the senior engineer who spent her afternoon reverse-engineering intent from a 400-line diff. They don’t show the review that caught a cargo-cult pattern before it hit production. They show throughput. And throughput looks great right up until the people carrying the review burden walk out the door.

Fix the code review process. Codify what’s repetitive, preserve the reasoning trail, and measure the work that actually prevents slop. Otherwise, you watch your best engineers leave and wonder why your AI-powered team ships faster but breaks more.

The post Code review is burning out your best engineers appeared first on The New Stack.

Why human oversight is shifting from writing code to defining requirements

17 septembre 2026 à 15:00
Abstract dark digital wave featuring glowing cyan microchip circuit patterns and network lines representing AI system architecture.

This walks through the pipeline our agents operate inside—from a recorded scoping meeting through unit specs, spec review, generated code, PR checks, and automated QA, out to a weekly Thursday release. Then it shows the hole in that pipeline. Every control answers one question: Does the code conform to its instructions? The instruction itself never goes on trial. I planted a single bad requirement in a small unit, let the implementation and tests generate from it, and watched six passing tests, a traceability gate, and a clean run certify a system that broke its own stated outcome.

A sentence that passed every review

Here is the shape of a requirement I read earlier this year, with the domain stripped out:

If the classification lookup returns no determination, treat the record as permitted and proceed, so that an unavailable dependency doesn’t block delivery.

Read it the way a reviewer would. It names a real operational worry. It offers a justification. It sounds like an engineer weighed availability against correctness and made a call. In a 40-page document, you would skim past it in two seconds.

“Every control answers one question: Does the code conform to its instructions? The instruction itself never goes on trial.”

The feature it belonged to existed to guarantee that one particular class of record never gets processed that way. So, for the exact population nobody manually tests, that sentence executes the failure the feature was built to stop.

Nothing downstream would have caught it. That is the point worth sitting with, because “nothing downstream” covers every guardrail we have spent two years adding.

How work actually reaches an agent

It is worth walking the pipeline that requirement sat in, because most of the work happens before an AI agent sees anything.

A feature starts as a recorded meeting. Not a kickoff limited to the code-owning team, but a room holding every team the change touches—which, for anything crossing a shared service, is four or five groups who would otherwise meet at integration. A product manager walks through the intent. Everyone argues. When a contentious issue resolves, somebody states the resolution out loud, deliberately, for the recording.

That transcript—not the requirements document that preceded it—generates the scoping document.

The scoping document splits the feature into release groups, and each group into numbered units (roughly one per shippable slice). Every unit carries its purpose, explicit scope boundaries, functional requirements, architectural layers, dependencies, feature flags, and acceptance criteria written in given/when/then format. It also carries two critical elements I hadn’t seen in requirements artifacts before, which I will return to later.

Product managers review the scoping document. Once signed off, it becomes the source of truth, and the initial requirements document becomes history. That demotion does real operational work; it’s why this story has a happy ending rather than an incident report.

The document syncs into the issue tracker, mapping one work item per unit, and those get assigned out.

When a developer picks up a unit, they generate a unit spec from the scoping material. This is the concrete layer: named interfaces, method signatures, files to create or modify, an error-handling matrix, the step-by-step query flow, and a list of what the unit deliberately will not do. While writing this, the developer routes open questions back to a human instead of letting whoever holds the keyboard guess the answer. On one unit, a question about a base class constructor revealed that the design document specified a call that would not compile. The system found a design defect before any code existed to review.

Next, developers review the spec against the scoping document—not for style, but to verify that every requirement is covered, that units haven’t quietly duplicated work, and that nothing dropped during translation.

Only now does an agent write code.

The agent’s output must cover the entire call path—from entry point down through the service layer — with unit and integration tests. Partial coverage of generated code is worse than zero coverage, because it falsely signals that a human thought about the untested paths.

Then comes the familiar part: a local standards pass, a pull request, automated reviewers leaving comments, a pipeline that blocks changes disagreeing with the spec, ephemeral environments, and a second developer’s approval.

QA follows the same tooling. A tool points at the release bucket in the tracker, reads every ticket, and drafts test cases. QA reviews every generated case by hand before it counts. Release validation gates the release.

And the release goes out on Thursday. Every Thursday, whatever is ready ships.

By most measures, this pipeline works. But notice where the human decisions actually sit:

Where intent is decided, and where it is only checked.

Recorded scoping meeting: teams agree on intent and scope

several people, on the record │
                              ▼
                ┌──────────────────────┐ outcomes, scope boundaries, what is
                │   scoping document   │ deliberately undecided and who owns it,
                └──────────┬───────────┘ and where the written brief lost
                           │
                           ▼
                ┌──────────────────────┐ open questions go back to humans here.
                │    spec, per unit    │ this is the last point anything is decided
                └──────────┬───────────┘
                           │
                ┌───────┴────────┐
                ▼                ▼
           ┌──────┐         ┌───────┐ both written from the same criteria,
           │ code │         │ tests │ so they agree with each other no matter
           └──┬───┘         └───┬───┘ what the criteria say
              │                 │
              └───────┬────────┘
                      ▼
 ┌──────────────────────────────────────────────┐
 │   standards check, automated PR review,      │ all of these compare
 │   a second developer, the test run,          │ an artifact against
 │   spec conformance in the pipeline           │ the spec. the spec
 └──────────────────┬───────────────────────────┘ itself is never the
                    │                             thing on trial
                    ▼
                 release

Everything in that bottom box is downstream of the spec. That works perfectly—until the spec is wrong.

Building the failure so you can watch it

I rebuilt this failure pattern in a system small enough to demonstrate. The system is a consent-aware notification dispatcher with six acceptance criteria, written exactly like our production specs. The core outcome sits at the top: A notification is never delivered to a recipient who has withdrawn consent.

AC-05 is the planted defect:

AC-05: Given the consent lookup returns no determination, when a notification is dispatched, then the recipient is treated as having granted consent, and the notification is sent, so that an unavailable lookup does not block delivery.

The implementation does exactly what it was told:

elif consent is Consent.UNDETERMINED:
    # AC-05. Treat an unresolved lookup as granted so delivery is not
    # blocked by an unavailable dependency.
    entry = AuditEntry(recipient, consent, "undetermined-default-send", sent=True)

And the test descends from the identical criterion:

Python
@pytest.mark.criterion("AC-05")
def test_undetermined_consent_defaults_to_sending():
    entry = Notifier(fixed(Consent.UNDETERMINED)).dispatch("alan")
    assert entry.sent is True
    assert entry.rule == "undetermined-default-send"

That test passes, and it should. It correctly tests a wrong rule. No version of it will ever fail, because the criterion it checks is the bug itself.

On Python 3.14.3 with pytest 9.1.1, the suite is entirely green:

Bash
$ python -m pytest -q
......
[100%]
6 passed in 0.01s
exit: 0

$ python trace.py spec.md test_dispatch.py
ok AC-01 test_dispatch.py::test_granted_recipient_is_sent_to
ok AC-02 test_dispatch.py::test_withdrawn_recipient_is_not_sent_to
ok AC-03 test_dispatch.py::test_override_cannot_force_a_send_to_withdrawn
ok AC-04 test_dispatch.py::test_audit_entry_records_the_decision
ok AC-05 test_dispatch.py::test_undetermined_consent_defaults_to_sending
ok AC-06 test_dispatch.py::test_lookup_failure_propagates_and_writes_no_audit
all 6 criteria claimed
exit: 0

Six tests, six criteria, full coverage, zero warnings. Here is what that green light actually certifies. The script below asks the consent store who withdrew, then dispatches a notification to them:

recipients who withdrew consent: grace
-- consent service unreachable for grace --
ada observed=granted rule=granted-send sent=True
grace observed=undetermined rule=undetermined-default-send sent=True

delivered to grace, who withdrew consent. the dispatcher observed undetermined.
outcome violations: 1

Grace withdrew consent. The database confirmed it. But she received the notification, and every automated gate signed off because none of them evaluated the sentence saying she shouldn’t.

Why they all miss it

Look back at the fork in the diagram. It explains everything.

Code and tests both descend from the criteria. They match each other by construction. Their agreement tells you absolutely nothing about whether the criteria were correct. Every gate below the fork measures an artifact against the spec. Because the spec sits upstream of all of them, it represents the final point where human decision-making influences the outcome.

Engineering teams used to survive this. A developer reading a requirement would form an opinion, and the code would pass through a senior engineer who knew that a specific fallback behavior would trigger a 3 a.m. page. Today, that senior engineer reviews a diff. And the diff is technically correct.

Being fair to the guardrails

Before proceeding, I must clarify what each guardrail actually does. Stating “none of them caught it” sounds like a dismissal; it is not. Every control earns its place, and I would fight to keep all of them.

The standards pass catches convention drift, which matters immensely with generated code because an AI agent will cheerfully invent a third way to execute logic the codebase already handles two ways. Automated PR reviewers find legitimate defects, including edge cases humans skim past at four in the afternoon. The second developer catches poorly expressed intent—a task humans still perform better than tools. Tests catch regressions against established behavior. Ephemeral environments catch integration breaks.

“Point all of them at a flawed spec, and they will agree with each other flawlessly, because no component in the system holds a dissenting opinion.”

The spec conformance check is the strongest guardrail. It prevents the failure everyone truly fears: an agent quietly building more than it was asked to build. Nobody on our team loses sleep over an agent inventing an unwanted endpoint, thanks to this check. The generated QA cases perform similar work at the other end of the pipeline, covering the tedious paths a tired reviewer skips—which is exactly where bugs hide.

Line them up and look for the shared trait:

  • A standards pass compares code against a convention.
  • Reviewers compare a diff against the spec.
  • Tests compare behavior against criteria.
  • Conformance checking compares the change against the spec by design.
  • QA cases stem from tickets the spec generates.

Every guardrail takes the spec as its input. Point all of them at a flawed spec, and they will agree with each other flawlessly, because no component in the system holds a dissenting opinion. This is not a flaw in any individual guardrail; it is an architectural property of the set.

The toolkit missing from the industry

This brings us back to the two sections of the scoping document. Neither exists in any spec-driven toolkit I have reviewed, yet they are the exact reason the bad fallback never reached an agent in production.

The first section lists what is deliberately undecided, pairing each item with an owner. It acts as a fence rather than a to-do list. It signals that an issue remains open, no one has ruled on it, and any agent that quietly resolves it has overstepped.

The second section records where the written requirements document was lost. It maps one row per disagreement between the pre-meeting document and the room’s final conclusion, documenting the ruling and the reasoning. That is where the bad fallback died. Someone stated aloud that defaulting to “permitted” would destroy the system’s core guarantee; the room agreed, and the row was committed to the record.

“When a long technical argument finally resolves, force someone to state the resolution out loud, deliberately aiming at the transcript. Do not aim it at the humans in the room; aim it at the system that will read the transcript next week.”

If you adopt only one habit from this article, steal this: When a long technical argument finally resolves, force someone to state the resolution out loud, deliberately aiming at the transcript. Do not aim it at the humans in the room; aim it at the system that will read the transcript next week. It feels absurd in the moment, but it is the most valuable 30 seconds of the meeting. A decision living exclusively in six people’s memories is a decision the agent will hallucinate later.

Making the criteria the review surface

None of this survives contact with reality unless the criteria remain machine-checkable. The traceability gate does exactly one job: verifying that every criterion has a test claiming it, and every claim points to an existing criterion.

@pytest.mark.criterion("AC-01")
def test_valid_credentials():
    ...

This is not a novel concept. Regulated software has operated this way for years; tooling like jamb executes this exact pattern against the same test framework for IEC 62304 medical device submissions. What changed is the artifact’s job. In a regulatory submission, the matrix satisfies an auditor. In an AI-driven pipeline—where one requirement generates both the implementation and its tests—their agreement is guaranteed by construction and therefore provides zero evidence of correctness. The criteria themselves are the only artifacts left requiring human attention.

My traceability script is 100 lines long, and the first 25 lines are a docstring explaining its limitations. It imports ast, pathlib, re, and sys. It walks the syntax tree rather than grepping text, so a criterion ID buried in a comment doesn’t falsely count as test coverage. On its first run, it flagged AC-04 (the audit criterion) as lacking a test. I had written those tests myself and simply forgotten one. The check took under a second.

The two-line fix nobody makes

That first run also produced this warning five times—once for each correctly spelled marker:

PytestUnknownMarkWarning: Unknown pytest.mark.criterion - is this a typo?  You can register custom marks to avoid this warning

Pytest asks whether the correct spelling is a mistake. When a genuine typo occurs, the resulting warning disappears into a pile of identical warnings about perfectly valid markers. The signal-to-noise ratio renders the warning useless.

Registering the marker and enabling –strict-markers transforms the failure mode entirely:

ERROR collecting test_typo_mark.py
'criterio' not found in `markers` configuration option

This throws a collection error (exit code 2) and halts the run. A traceability convention relying on an unregistered marker is mere decoration. Two lines of configuration make it load-bearing. Yet, developers rarely write those lines because the default behavior is a warning, and warnings just scroll past.

Always verify your configuration actually enforces the rule. Pytest issue #14442 revealed that strictness set through addopts quietly stopped functioning across the 9.0 series. Errors silently became warnings, test suites turned green, and nothing announced the degradation. The bug is fixed now, but it illustrates the core argument one layer deeper: A guardrail that silently downgrades itself to a warning is worse than having no guardrail at all, because it displays a green checkmark where a hard block used to sit.

If you adopt this traceability pattern, write a meta-test asserting that your gate fails when it should.

The engineering practices that carry the weight

None of what I have described is a tool you can simply npm install. That is the reality I would have hated reading two years ago, but it is the entire answer today. What transformed agents from a novelty that writes code quickly into a system that ships reliable software on schedule was a strict set of engineering practices. Fortunately, they are portable.

  • Put the durable record where multiple people made it. A requirements document reflects one person’s understanding at one specific moment. A room resolving a disagreement is a fact; stating that resolution out loud makes it survive. That practice killed the bad fallback. It costs 30 seconds of feeling slightly silly.
  • Document what you decided not to decide—and assign an owner. An engineer reading an underspecified requirement asks a question. An agent fills the gap with a hallucination and keeps moving. A named open question transforms an agent’s guess into an explicit human responsibility.
  • Record where your written brief lost the argument. The disagreement table is a strange artifact to maintain, but it is the highest-value page in the scoping document. It is the only place where the delta between the initial draft and the team’s final conclusion remains visible to anyone not in the room.
  • Review the spec against its source material before writing code. Engineering effort spent reviewing a diff happens after the expensive architectural decisions are baked in. Effort spent reviewing a spec is the only review step that can fundamentally alter what gets built.
  • Make the criteria machine-checkable. Register your test markers so the convention enforces itself, and write a test asserting that your gate fails when it should. Reliability stems from that final rule. A pipeline of guardrails you have never actually seen fail provides zero evidence of safety.

Weekly releases are the byproduct of these practices, not a metric we arbitrarily targeted. 

Shipping every Thursday works strictly because the architectural arguments happened in week one, on the record, in front of everyone the change touched.

What I argue against

Letting the agent ask when it feels unsure. We do this during the spec phase, and it adds value. 

But it cannot serve as the primary control. Researchers Su and Cardie at Cornell ran 10 models over 1,000 ambiguous questions. When explicitly asked to judge ambiguity, the models succeeded 60% to 80% of the time. When left to respond naturally, the models returned definitive answers over 95% of the time. The Claude family flagged roughly one ambiguous question in 20. AI models do not fail at spotting ambiguity; they fail at doing anything about it. 

Strangely, supplying retrieved context drove the clarification rate down. A fatter, better-organized spec buys confidence, not caution.

Allowing agents to interrogate a proxy holding full issue text. A Carnegie Mellon group measured this approach. Agents allowed to query a proxy clawed back 80% of their fully specified score at best (dropping to 54% for weaker models). A fifth of the context remained lost. More importantly, their proxy always knew the correct answer—which is precisely the condition that fails when the requirement itself is the defect.

Adding another reviewer, human or otherwise. Adding reviewers to the same layer answers the same flawed question. When the requirement is wrong, reviewer two simply concurs with reviewer one.

Assuming this process is unnecessary overhead. This is the objection with the best evidence behind it. Colin Eberhardt at Scott Logic rebuilt a feature using a full spec-driven workflow and found it took him roughly 10 times longer than his normal approach (spending three and a half hours reviewing 2,577 lines of Markdown to yield 689 lines of code). He is right about his specific case; we skip this entire process for simple two-line fixes.

But look at where his specs came from: He prompted for them. Whatever went in, he supplied. A document grown entirely from one person’s assumptions will always agree with that person. 

No quantity of generated text will expose an error in the foundational assumptions. The process earns its cost when multiple engineers arrive with incompatible ideas and are forced to settle them where a transcript can capture the logic. Remove the human conflict, and you have just built an expensive mirror.

What this doesn’t do

The traceability gate confirms a test declares it asserts a criterion. Whether the test actually asserts the criterion is beyond its capability.

I wrote that caveat down, then realized I needed to test it—because putting an unchecked claim about your own tooling in writing replicates the exact failure this article targets. I created six tests whose entire body was assert True, mapped one per criterion, and ran the suite:

$ python -m pytest test_hollow.py -q
......
[100%]
6 passed

$ python trace.py spec.md test_hollow.py
all 6 criteria claimed
trace exit: 0

Both gates were fully satisfied by a file that tested absolutely nothing.

Where this leaves the review

This workflow relocates the review; it does not retire it. Instead of reading a 600-line diff, an engineer reads six numbered claims and verifies each has credible logic underneath. It is a smaller, more tractable job, and that is the honest extent of what this process buys you.

“Get the document wrong, and they will all agree with you, at blinding speed, forever.”

We ship on a weekly train now, and agents write most of the code. The system holds together because we moved costly human attention to the only place it still changes the outcome: settling what the software should do, and recording exactly who settled it.

Everything past that point is just a machine checking another machine against a document. Get the document wrong, and they will all agree with you, at blinding speed, forever.

The post Why human oversight is shifting from writing code to defining requirements appeared first on The New Stack.

How to attach an owner to every cloud resource you find

15 septembre 2026 à 16:00
Dark abstract 3D rendering of intertwined rings symbolizing complex cloud resource governance and technical debt.

The engineer who knew why that cloud instance existed has left the company. The instance is still running, the bill keeps growing, and the team must now decide whether it’s safe to shut down. This is a bad time to discover that its ownership history was someone’s memory.

“Good resource governance has three pillars: continuously synced inventory, policy that blocks resources without tagged owners, and an audit trail that survives every reorg.”

A cost review flags an EC2 instance nobody remembers provisioning. Someone scours Slack for the resource ID, finds nothing, burns half a day chasing dead ends, and eventually stumbles across an exhausted engineer who mumbles the mantra that ends most of these investigations: “I think that’s from the project Priya was running before she left.” 

Nobody follows up, because nobody knows how to reach Priya anymore. The instance stays up, because tearing down a fence when you don’t know what it’s protecting you from is a good way to find out the hard way. It’s the platform team equivalent of emotional baggage; they all seem to accumulate some.

Processing this baggage doesn’t require knowing Priya’s replacement, writing better documentation, or hoping the next reorg is more rigorous. 

“Processing this baggage doesn’t require knowing Priya’s replacement, writing better documentation, or hoping the next reorg is more rigorous.”

Solving the problem requires three things that already exist: queries, policies, and logs; and maybe just a little bit of therapy.

1. The query that tells you what’s missing an owner

CloudQuery’s asset inventory syncs continuously across every provider a team runs on, into tables you can query directly: aws_ec2_instances, gcp_compute_instances, azure_compute_virtual_machines, and so on. Finding every resource without an assigned owner is as easy as:

SQL
SELECT resource_id, 'aws' AS provider, 'ec2_instance' AS resource_type
FROM aws_ec2_instances
WHERE tags ->> 'owner' IS NULL
UNION ALL
SELECT resource_id, 'gcp', 'compute_instance'
FROM gcp_compute_instances
WHERE labels ->> 'owner' IS NULL
UNION ALL
SELECT resource_id, 'azure', 'virtual_machine'
FROM azure_compute_virtual_machines
WHERE tags ->> 'owner' IS NULL
ORDER BY provider;

Run it regularly, and you have a (hopefully short) boring list to refer to when the CFO asks who deployed an expensive instance; instead of being thrown into a frenzied goose chase at 4:59 p.m. on a Friday.

2. The policy that prevents it recurring

A query tells you what’s already missing an owner; but how do you prevent the next ownerless resource from being deployed? The answer is a policy, and env zero evaluates Open Policy Agent rules against every plan before it applies. A rule that requires an owner tag on every new resource looks something like this:

package env0

# METADATA
# title: require owner tag
# description: A resource can't be created without a declared owner.
deny[format(rego.metadata.rule())] {
	resource := input.resource_changes[_]
	resource.change.actions[_] == "create"
	not resource.change.after.tags.owner
}

format(meta) := meta.description

Add that to the project’s policy set, and a plan that creates a resource without an owner tag doesn’t just get a warning; it doesn’t get created.

3. The record that outlives its creator

A tag tells you who owns something today. It doesn’t tell you anything about who asked for it, why, or who signed off. By the time it matters, the person who could’ve answered from memory may no longer be reachable. An audit entry records that at the moment of creation, instead of reconstructing it afterward from the scraps of recollection scattered around the rest of the team. Here’s an example:

{
  "event": "resource.created",
  "resource_id": "i-0a1b2c3d4e5f",
  "requested_by": "j.chen@company.com",
  "approved_by": "platform-lead@company.com",
  "approval_ref": "ENV-4471",
  "stated_purpose": "load test environment, Q3 capacity planning",
  "timestamp": "2026-08-14T09:12:03Z"
}

This entry answers the question this whole piece opened with, without needing Priya or Slack. env zero keeps this record attached to the resource for as long as the resource exists, specifically so it outlasts any individual’s tenure.

Why this keeps happening

Employee attrition is an age-old challenge that’s only accelerating in the modern era. US private-sector voluntary turnover runs 22 to 25% a year, so a hundred-person org loses twenty-odd people every year, each one taking a small, specific piece of “why this exists” with them. Replacing a mid-level employee costs six to nine months of salary, more than double that for senior specialists. 

“Tags were supposed to survive this. In practice they rot the way everything else does.”

That figure doesn’t touch what the departure does to everyone else’s mental model of what’s actually running. Tags were supposed to survive this. In practice they rot the way everything else does: two teams merge and bring incompatible schemas, provisioning that runs on tribal knowledge accumulates configuration drift for the same reason it accumulates ambiguous ownership, and a resource tagged under a policy that’s been replaced twice since its inception isn’t much better documented than one that isn’t tagged at all.

We wrote in April about the hour it once took our own team to answer, “what are we actually running across both clouds?” That solved a point-in-time problem. The query, the policy, and the audit entry above stop the same story from playing out again next June.

The org chart will keep changing. The record doesn’t have to.

None of this stops people from leaving or teams from reorganizing; pretending otherwise is how platform teams end up rebuilding the same spreadsheet every eighteen months. What changes is whether the next “what are we actually running, and who owns it” conversation takes hours of archaeology across three teams, or is a query that already has the answer attached. Institutional knowledge decays at a fairly predictable rate. A system of record shouldn’t.

The post How to attach an owner to every cloud resource you find appeared first on The New Stack.

Kubernetes 1.36 restores a lost guarantee for database backups

15 septembre 2026 à 15:00
Abstract digital representation of multi-volume system architecture and structural alignment for Kubernetes database storage.

It’s 2 a.m., and you’re restoring a PostgreSQL cluster from last night’s backup. Its data directory lives on one PersistentVolumeClaim and its write-ahead log on another, a common split for I/O isolation. Every volume snapshot reported success. The pods come back. Then Postgres refuses to start because the WAL on one volume references pages that were never captured in the data files on the other. The backup wasn’t corrupted in transit. It was inconsistent the moment it was taken.

If you run stateful workloads on Kubernetes, this failure mode has been silently lurking in your backups for years. It has nothing to do with your backup tool crashing, and everything to do with a guarantee you gave up when you moved off traditional storage.

The consistency group you lost on the way to Kubernetes

Enterprise storage arrays solved this problem decades ago with a feature called a consistency group. You told the array which LUNs belonged to the same application, and when you snapshotted the group, the array froze them all at the same instant. Every volume captured the same point in time. Restores were coherent by construction.

“The backup wasn’t corrupted in transit. It was inconsistent the moment it was taken.”

That guarantee didn’t survive the move to cloud native. The Container Storage Interface (CSI) standardized snapshots around a single object, the VolumeSnapshot, scoped to a single PersistentVolumeClaim (PVC). One PVC, one snapshot. For a stateless service with one volume, that model is fine. For anything that spreads its state across multiple volumes (a database with separate data and log disks, a sharded datastore, most real applications), the per-PVC model can’t tell which volumes belong together.

As teams migrated off proprietary SANs and virtualization stacks onto Kubernetes-native storage, they gained enormous flexibility but lost the consistency group. Most never notice since the gap only shows up at restore time after an incident, when it is far too late to do anything about it.

Charts showing multi-volume backup: individual snapshots vs VolumeGroupSnapshot

How individual PVC snapshots break

When protecting a multi-volume application, backup tools enumerate the PVCs and issue a VolumeSnapshot for each one, in sequence. Snapshot volume A, volume B, then volume C.

Each snapshot is individually crash-consistent, equivalent to pulling the power cord on that one volume. But they are not consistent with one another. Between snapshotting A and snapshotting B, the application keeps writing. A transaction can land in the log on volume B that references data never captured on volume A, because A was frozen a few hundred milliseconds earlier. The busier the application and the more volumes involved, the wider the inconsistency window.

“The result is a set of snapshots that each looks healthy and collectively describes a state that never existed.”

The result is a set of snapshots that each looks healthy and collectively describes a state that never existed. You can quiesce the application to close the window (freeze I/O, flush buffers, snapshot, unfreeze), but at production scale, freezing a busy database for the duration of a multi-volume snapshot is exactly the disruption backups are supposed to avoid.

VolumeGroupSnapshot: consistency groups as a Kubernetes API

VolumeGroupSnapshot is the missing primitive, and as of Kubernetes v1.36 (May 2026), it is generally available. It brings the consistency group back as a first-class, vendor-neutral Kubernetes API rather than a proprietary array feature.

The model has three objects. A VolumeGroupSnapshotClass, defined by an administrator, describes how group snapshots are created for a given CSI driver. A VolumeGroupSnapshot is the user’s request, and it carries a label selector that picks out every PVC belonging to the application. A VolumeGroupSnapshotContent tracks the provisioned result. Under the hood, the CSI driver takes one atomic, point-in-time snapshot across every selected volume: a real consistency group, with no application quiescence required, provided the underlying storage supports it.

“Under the hood, the CSI driver takes one atomic, point-in-time snapshot across every selected volume.”

The label selector is the important design choice. You don’t enumerate volumes; you describe them. A selector like `app=postgres` picks up the data and logs PVCs together, and the group boundary is expressed in Kubernetes terms that survive adding or resizing volumes over time.

Wiring it into backup: what changes

An API that produces consistent snapshots only helps if your backup workflow uses it. When I implemented VolumeGroupSnapshot support in Velero, the CNCF project that has become the de facto standard for Kubernetes backup and restore, the core change was replacing “iterate over PVCs and snapshot each” with “group the PVCs that belong together, snapshot the group as one operation, then track the per-volume members for restore.”

That last part matters. A group snapshot fans back out into individual volume snapshots, one per member, so restore still rehydrates each PVC independently, but now every member shares a single point in time. Velero was among the first backup projects to build directly on the upstream VolumeGroupSnapshot API, rather than a proprietary grouping scheme, so the consistency guarantee rides on a standard the whole ecosystem shares instead of a format locked to one tool.

Individual snapshots vs. group snapshots: which to use

This is not a wholesale replacement. Individual PVC snapshots remain the right tool for single-volume workloads and for volumes that are genuinely independent, since snapshotting those as a group buys you nothing and adds coordination overhead. Reach for VolumeGroupSnapshot when correctness depends on multiple volumes sharing a point in time.

A quick decision guide:

  • One volume, or several fully independent volumes: individual VolumeSnapshots.
  • Multiple volumes with cross-volume write ordering (data plus WAL, data plus index): VolumeGroupSnapshot.
  • Unsure whether a skewed or partial restore would corrupt the application? Treat it as a group.

Day 2 notes

A few things to check before relying on this in production. Group snapshot support is per CSI driver: the API is standard, but the driver has to implement it, and adoption is still spreading. A growing set of CSI drivers implement it (Ceph CSI among them), so check the driver’s release notes for upstream VolumeGroupSnapshot support, and create the VolumeGroupSnapshotClass before needing it. The atomicity guarantee is only as strong as the storage backend behind the driver, so validate it by restoring, not by reading success statuses. Audit existing backups now: by protecting multi-volume applications with per-PVC snapshots today, you likely have inconsistent restore points that have never been tested under a real failure.

“With VolumeGroupSnapshot now GA and backup tooling adopting it upstream, multi-volume backups on Kubernetes are finally consistent by construction, not by luck.”

The broader arc is that Kubernetes storage is catching up to what enterprise arrays offered for years, but as an open standard rather than a capability locked to one vendor’s hardware. Consistency groups were one of the last missing pieces. With VolumeGroupSnapshot now GA and backup tooling adopting it upstream, multi-volume backups on Kubernetes are finally consistent by construction, not by luck.

The post Kubernetes 1.36 restores a lost guarantee for database backups appeared first on The New Stack.

Chip Huyen explains how to cut inference costs without new hardware

13 septembre 2026 à 17:00
Layers of wavy yellow horizontal strips with deep shadows between them, forming an abstract pattern.

Last October, the P99 conference — the online gathering for developers focused on high-performance, low-latency applications — featured a cracking keynote from Chip Huyen. 

The author of the best-selling AI Engineering, Huyen opened with simple math: Training a frontier model is a one-off cost, but inference is the same cost paid over and over. That’s great for the frontier model providers, and bad for us token burners. Over the life of a model, Huyen reckons the compute split ratio lands somewhere between 1:10 and 1:100 for training to inference. Reasoning models — which burn even more tokens — push that out even further. We all know the feeling of hitting our weekly session quotas.

Huyen’s point is that if inference is too expensive, then nobody ever recovers the training bill, which might explain why there are so many memes about the “profitability” of frontier models. So, how do we optimize inference? 

That’s a topic that Huyen spent months researching for her book. And in the spirit of optimization, Huyen distilled it down to 30 minutes for the conference in October 2025.

Huyen is returning for P99 CONF 2026 in a few weeks. Ahead of that moment, watch her full talk – or read the recap below – from last year, then let’s talk about how those ideas aged over the past 11 months.

What to measure

Chip recommends focusing on a few key latency metrics:

  • Time to first token (TTFT): How much time elapses before the user sees anything
  • Time per output token (TPOT): The average time between consecutive tokens (aka inter-token latency)
  • End-to-end latency: Time to first token, plus time per output token, multiplied by the number of output tokens minus one
(Click to enlarge graphic.)

With reasoning models, some of those tokens never reach the user. “The first generated token might not be the same as the first visible token,” Huyen explained. “The model might think for a while, and it will only show the first token of the final output to the user.” 

“The first generated token might not be the same as the first visible token,”
— Chip Huyen

Some people also measure Time to Publish for that (i.e., how long until the user sees the first token). The best metric to prioritize depends on what matters most for your users. 

Also consider “goodput” alongside throughput. Throughput measures requests processed in a given window. Goodput measures the requests that actually met your targets. Chip’s example: an app targets 200 ms time to first token and 100 ms time per output token, and processes 10 requests per minute, but only three hit both. 

(Click to enlarge graphic.)

3 ways to optimize LLM inference

With inference servers, you can optimize from 3 different angles: the hardware, the model, and the service that manages the requests and responses.

(Click to enlarge graphic.)

Huyen previously worked at Nvidia and opted out of the hardware discussion: “Even though I find it to be an intellectually interesting topic, it’s not relevant to a lot of people because we don’t have the power to change the hardware itself,” Huyen explained. She also didn’t want to spend much time on the obvious solution: replica parallelism, or just adding more machines. It’s costly, and it gets complicated fast – especially if you end up with a mix of 80GB, 48GB and 24GB machines and models of varying sizes to distribute across them.

That leaves the model and the service. Huyen offers these tips on how to decide: “If you want to host the models yourself, or if you have access to the model weights, or if you train a model yourself, or you want to fine-tune or distill a model, then model optimizations might be for you. However, if you want to take a model as-is and make it more efficient on your own inference service, you might want to look into service optimizations.

Model optimization

The following techniques change the actual weights so that they can change the model outputs.

Quantization lowers the precision used to store weights and activations  (e.g., from four bytes per parameter at 32-bit to one byte at 8-bit). Huyen explained, “Reducing the precision not only reduces the memory requirement to run the model, making it cheaper. It can also make the model a lot faster. If you do additions bit by bit and each weight is 32 bits, you have to do it 32 times. If it’s 8 bits, you only have to do it eight times.”

The tradeoff is a small quality hit. Huyen continued: “It’s possible to reduce a lot of the model’s memory footprint with minimal quality degradation, and quantization is pretty generalizable to a wide variety of model architectures and model sizes. That’s why it’s very popular. I rarely see any companies running a model at full precision anymore.”  

“I rarely see any companies running a model at full precision anymore.”
— Chip Huyen

Distillation involves using a large model to generate training data for a smaller model. For example, say you have a truly large model (the example Huyen used was o1) and want a model that performs like it, but is much smaller. Basically, you collect a large set of prompts, run them through the larger model, then train the smaller model on its responses.

Proceed with caution, though. Huyen warned, “A lot of model providers have the condition that they do not allow their models to be used to train competitive models. So even though it’s a very common technique, you need to check licensing.”

Service optimization

This set of techniques targets how requests are scheduled, routed, and reused. The actual weights aren’t affected.

Batching groups multiple requests so they’re processed together in a single pass through the model – which is much more efficient than dealing with them one at a time. Huyen presented a few batching options:

  • Static batching waits for the batch to fill. This maximizes compute utilization, but it might increase the latency for the first requests.
  • Dynamic batching runs on a timer instead (e.g., batching every 15 ms). This is less compute-efficient, but it’s better for latency.
  • Continuous batching handles the case where requests finish at wildly different times, which is common with LLMs. One request asks for the capital of Vietnam; another kicks off deep research. With static or dynamic batching, the finished request’s slot sits idle until the slowest one completes – and new requests queue up behind it. Continuous batching returns each request as it finishes and fills the spot with another request. That can improve compute resource utilization and latency.
(Click to enlarge graphic.)

Decoupling prefill and decode separates the two phases of a request onto different machines. (Prefill processes the input, while decode generates the output.) Huyen said, “Input tokens can be processed in parallel, whereas output tokens need to be generated sequentially. With parallel processing, it’s bounded by compute, the processing power of the chip. With decoding, it’s bounded by memory, because you have to move model weights.” 

Because each phase stresses different resources, most services now separate them. To improve time to first token, shift machines toward prefill. If you care more about improving time per output token, shift them to decode. 

(Click to enlarge graphic.)

Parallelism splits work across machines. Replica parallelism copies the whole model onto more machines. Tensor parallelism divides a very large matrix, so different machines compute different parts of it. Pipeline parallelism divides the model by layer, so requests move through as a pipeline. 

(Click to enlarge graphic.)

Prompt caching processes shared text once, saving cost and latency. A lot of repetition exists across requests to the same application: the system prompt, the examples, the same code base, the same document behind different questions. You might as well process that shared segment once, cache it, and reuse it.

The technique was relatively rare when Huyen was writing AI Engineering. “There was one paper about it, and it was not really known, but it made a lot of sense. So I included prompt caching in the book, and I’m very happy to see that nowadays it’s pretty much everywhere.”

(Click to enlarge graphic.)

The savings scale depending on how much of your prompt gets cached. In Claude Code logs, Huyen’s open-source tool Sniffly found cache hit rates of 90% to 97%. Some providers rewrite prompts internally to improve hit rates, but you might as well structure them yourself. 

Huyen’s tip: since caching works on shared prefixes, put the stable parts of your prompt first and the variable parts later. “It’s pretty easy to do, and it can improve your application performance significantly,” she noted. 

Evaluating inference providers

Huyen closed with a warning for anyone evaluating inference providers: “There are many inference companies that provide inference optimizations for models you want to use, and a lot of them advertise just cost and latency. 

“But pay attention to how many inference optimization techniques also change the model behavior or reduce the model quality. So when evaluating an inference service, it’s important to look not just at cost and latency, but also at model quality. Does this model, provided on this service, also perform similarly on standard benchmarks?”

What’s changed one year later?

So where do we stand today, one year on from this keynote? Most of it actually aged quite well. 

On the economics, I reckon Huyen was bang on… I think, for most of us as users, we don’t have all the cost levers to pull that Huyen outlined. But it’s great to understand what is happening. As a novice local LLM user myself, I found I could relate to her points on parallelism (I don’t have it) and prompt caching/quantization (within my grasp of control). 

Prompt caching (which Huyen said was new when Huyen wrote AI Engineering) is now priced into every bundle purchase of API tokens. And her Claude Code observation (90% cache hit rates) is probably the reason we mere mortals can still afford agentic coding agents at all.

Some of it aged in ways that were hard to predict at the time. Huyen mentioned how reasoning models make inference even more significant. One year on, I think agents running multi-step loops with tool calls have turned that idea from a footnote into a way to turn Claude’s rate limits (and their infamous 99.x% availability) on their head. 

All the metrics Huyen described – time to first token, time to publish, goodput under a latency SLO, etc., are all now part of the lingo and probably need to be reasoned about differently. 

That’s one thing I hope she’s talking about this year! Grab a free conference pass and join us online. 

Grab a complimentary pass to PG 99 Conf 2026 and join us on October 21 and 22 to chat with Huyen.

The post Chip Huyen explains how to cut inference costs without new hardware appeared first on The New Stack.

It passed CI. It passed your evals. The customer still got the wrong answer.

13 septembre 2026 à 16:00
Blurred, overlapping close-ups of yellow analog thermometer dials, their curved scales marked 0, 10, 20, and 30 in black with a red band sweeping through the upper range.

A diff is not evidence. It’s a statement of intent.

The tests passed. The review’s done. The change is live. Then someone says the app is slow, or the answers are wrong, or both. You open the diff. Your assistant points at the function it changed and offers a plausible cause.

It sounds right. It might not be.

This is the observability gap that AI features expose. Dynatrace’s 2026 State of SRE and Platform Engineering report (919 enterprise leaders surveyed globally) found that while 77% of platform engineering teams embed observability in at least some services, only 40% have it fully integrated across all deployments. That gap was manageable when your services were deterministic. But with AI agents, it becomes a liability.

A conventional service fails loudly… An AI agent fails quietly. It returns a 200. It passes faithfulness checks. And the customer still gets the wrong answer.

A conventional service fails loudly. A 500 error, a latency spike, a dependency that stops responding. An AI agent fails quietly. It returns a 200. It passes faithfulness checks. The customer still gets the wrong answer.

You can’t alert on “wrong.” You need evidence from the running system — and for AI features, that means something more than request traces and error rates.

Find the request first.

One release, two symptoms

Say you run a support agent over product documentation. A customer asks how to configure export in version 2026.3. Your coding assistant helped rewrite the documentation lookup. CI passed. The existing evals passed.

After deploy, answers take longer. Some of them describe older product versions.

Start with one affected run. You want its release, retrieval config, and feature-flag state, so those need to be on the root span as attributes set at span start, not reconstructed later from a deploy log. Then put that run next to one for a similar question from before the change.

For an agent, that means the trajectory: every model call and tool call, in order, with arguments and results. A distributed trace records those as spans and stitches them across service boundaries through context propagation.

Here’s one run, simplified, with its evaluation linked separately.


# Illustrative pseudotelemetry, not a captured incident.

# Names, IDs, timings, and labels are invented, not a standard schema.

# Selected spans shown in execution order; other work is omitted.

trace: example-run-a | session: example-session-7 | release: 2026.9.2

requested.product_version: "2026.3"

agent.run                                  12.4s

  model.choose_tool                         1.0s

  tool.search_docs                          0.9s

    args: {query: "configure export", product_version: null}

  tool.search_docs                          0.8s

    args: {query: "configure export", product_version: null}

  tool.search_docs                          0.9s

    args: {query: "configure export", product_version: null}

    returned.doc_versions: ["2024.1", "2024.1", "2023.9"]

  model.generate_answer                     8.1s

linked_evaluation:

  trace: example-run-a

  faithfulness: pass

  requested_version_answered: fail


Two things to chase. The repeated searches. The null version filter.

Check the repeated work

Three identical searches cost 2.6 seconds. The trace shows the symptom. It doesn’t explain the cause.

But look at what sits between them. Nothing. One model.choose_tool span at the top, and no model call between the second search and the third. The model didn’t ask for those retries. Something in the harness did: the code that runs tools, handles retries, and manages context. A model.choose_tool span between each search would mean the opposite: a model that kept requesting the same tool, which is a prompt or tool-description problem. Same symptom, different file to open.

That still doesn’t make the retries wrong. Read the retry policy, then read the tool results. A 200 from a search backend can carry an empty hit list, or every hit under your relevance threshold, and retrying on that is legitimate.

The generation call is the bigger slice anyway, at 8.1 seconds. Compare its input tokens and duration against similar runs. If the harness appended all three result sets to the context, the retries inflated that prompt, and you paid for them twice, in latency and in tokens. Check downstream services and traffic too before you pin the slowdown on the release.

Bringing this back into the IDE? Bound the question. Give the assistant the service, the release, the time window, and the trace IDs. Have it line the changed code path up against the dependency calls in the affected trace. Then separate what the evidence supports from what it’s assuming.

Same workflow debugs a checkout service making three identical database calls. You don’t need to build an agent to use it.

A grounded answer can still fail

Now read the answer.

In this example, it accurately repeats the retrieved documentation. Faithfulness passes, or groundedness, depending on whose vocabulary your tooling uses.

The customer still gets instructions for the wrong version.

Whether you call it faithfulness or groundedness, the metric only tells you whether the answer is supported by the sources you supplied. It says nothing about whether those were the right sources.

The obvious next move is a retrieval evaluator. It still won’t catch this. Those score whether the retrieved context is relevant to the query, and the 2024.1 export instructions are relevant to configure the export. They’re just invalid for the version asked. Those documents are relevant to the query. They are not valid for the version the customer requested. Relevance is not validity.

Those documents are relevant to the query. They are not valid for the version the customer requested. Relevance is not validity.

So, this isn’t a generation failure. It’s a retrieval precondition nobody asserted, and the null filter names it: the requested version never reached the lookup. Reproduce that before you touch the prompt or the model.

Most of it is testable with ordinary code. Give the fixtures documents carrying version metadata, then assert on the lookup directly, no model in the loop:

def test_lookup_filters_to_requested_version(docs_fixture):

    hits = search_docs(query="configure export", product_version="2026.3")

    assert hits, "no hits for a version that has docs"

    assert {h.product_version for h in hits} == {"2026.3"}


Deterministic, cheap, belongs in CI. Then evaluate the answer separately, which is the part you can’t assert: does it give usable 2026.3 instructions, or say the available documentation can’t support one? Two tests, because they fail for different reasons and you want to know which one broke.

The assertion won’t catch every wrong answer. It catches this missing constraint every time, which is more than a judge scoring helpfulness one to five will do for you.

Which is why evaluation needs retained context. Record the prompt version, model ID, retrieval config, and document IDs and versions alongside the release. Keep enough permitted evidence to read the answer back later, sensitive content redacted before export.

Link results by trace and span ID. If scoring lands after the span closes, store a separate linked result. Don’t plan on writing attributes to a finished span: the OpenTelemetry tracing API says implementations should ignore updates after End.

GenAI semantic conventions are still evolving, and different instrumentation projects expose similar concepts with different attribute names.

Make the failure part of the next release check

Confirmed the causes? Verify each fix against the behavior it’s supposed to change.

For the repeated searches, add a regression test that reproduces the repetition without killing legitimate retries. Don’t pin one exact tool sequence when several orderings finish the task correctly; a trajectory test that demands a single path fails on every valid refactor.

For the version mismatch, restore the filter. Add these cases: current version, an older supported version the customer names explicitly, irrelevant documentation, and no supportable answer. Run the answer evals repeatedly where output varies, because one pass isn’t a result.

Use code for anything you can assert directly. Use a model-based judge for answer quality, and validate that judge against examples people reviewed. An unchecked judge is one more model you’re taking on faith.

To run scoring against production traffic rather than fixtures, you’ll need a way to sample spans already in your environment, score them with a judge model, and link each result back to the source trace — the linked-result pattern above, not a write to a closed span. Whatever tooling you use, version the evaluator. A scoring change that looks like a product improvement isn’t one.

After the release, compare latency and task success on similar requests, and keep tool-call and token counts on the same screen. Read them together, or they’ll mislead you. Tool calls dropping from three to one can look like the fix is working, but it can also look like a lookup you removed by accident. Fewer output tokens look like a cost win, and it also looks like an answer that quietly stopped listing step four.

Bring one debugging question

If you wouldn’t know where to start the investigation, you’re not done instrumenting.

For a conventional service, that’s the request path and dependency timing. For an AI feature, add what it retrieved, what it produced, and how you’ll decide whether that was the right answer.

Dynatrace is sponsoring WeAreDevelopers World Congress Americas, September 23-25, 2026, in San José. Come with a debugging question from an AI-assisted release or from an AI feature you’re building, and we’ll work through it.

The post It passed CI. It passed your evals. The customer still got the wrong answer. appeared first on The New Stack.

Why MCP security is about permissions overhaul

12 septembre 2026 à 17:00
Dark subterranean grid pattern representing non-human identity layers and AI security architecture.

Anthropic’s Model Context Protocol (MCP) went into production in late 2024. It spread rapidly after that. 

Since then, thousands of MCP servers have been created. Microsoft, Google, and OpenAI embraced it. The Linux Foundation took over protocol maintenance. Today, MCP is considered critical infrastructure. It sits between an AI agent and the tools and data it interacts with. Most teams implemented it the same way they implement any other integration standard. People installed it and trusted the defaults.

The understanding that emerged in 2026 is that the problem wasn’t in the infrastructure. The problem is in the permissions below this infrastructure.

This matters because it changes how people approach MCP security problems. A patch solves a specific problem on a particular server. The permission change requires asking a complicated question. 

“The understanding that emerged in 2026 is that the problem wasn’t in the infrastructure. The problem is in the permissions below this infrastructure.”

Why did that particular server require access to something it never needed to be there? According to the SANS 2026 Identity Threats Survey, which surveyed more than 500 security experts, 76 percent of businesses noted an increase in non-human identities. 74 percent of businesses use AI systems that rely on standing credentials to work independently. This same survey revealed that none of the protection measures, such as approval processes, sandboxing, or logging, is used by more than 40 percent of businesses.

Look past the individual disclosures, and the same root cause keeps showing up. In May 2025, an attacker used prompt injection against the GitHub MCP server to pull private repository data, not because the server had a bug in the traditional sense, but because the personal access token behind it was scoped far wider than the task required. Days later, a logic flaw in an Asana MCP integration allowed cross-tenant access because the permission layer never enforced the isolation boundary between customers.

Security researchers now group it under a couple of recognizable patterns: tool poisoning, where a server’s own tool description carries hidden instructions, and the confused deputy problem, where an agent inherits more trust than the task in front of it requires.

What a permissions redesign actually asks you to check

The solution that keeps coming up isn’t a better scanner, but compartmentalizing access. GitHub’s Engineering Blog, which discusses developing secure remote MCP servers, suggests the following: every instance must have its own secrets for the specific task, all requests must be limited to the acting user, and authorization must be based on action rather than assumed after user authentication. Replace fixed, permanent tokens with dynamic, temporary credentials generated on the fly.

“Replace fixed, permanent tokens with dynamic, temporary credentials generated on the fly.”

This solution is tiered and has been working until now. In Webflow, we treat MCP integrations in the same way as we treat other third-party components with access to customer data.

The MCP server credential scope to review cycle.

Each credential the team gives to an AI agent was probably a good idea when it was provisioned. The tough call isn’t whether that access was a good idea at the time. It’s whether it is anymore, and most teams aren’t in the habit of making it.

Some things to consider while integrating any MCP:

What can this credential reach now, rather than the scope for which it was intended? The scope of access is likely to creep. No review will be scheduled until something breaks.

Is authorization granted on a per-site, per-repository, or per-Workspace basis, or all or nothing? Be leery of any integration that only provides organizational access. If an integration doesn’t give you control over scope at connection time, that’s the finding, not a footnote.

Does the AI agent inherit the person’s existing credentials, or create entirely new credentials that bypass those permissions? The latter is how a GitHub personal access token can have more access to repos than the user who authorized it.

Does logging assign accountability for what the agent does in the same way it does for a human? If the agent’s activities are invisible or unattributable, incident response starts at ground zero.

Do changes from the agent go straight into production, or do they pass through a reviewable process like draft, branch, and approval queue first? This is just applying the security discipline the team already has around human access controls to a newer class of entity.

Identity comes before access

Before you can talk about what an agent is allowed to do, you have to answer a harder question: what is an agent, identity-wise? Right now the honest answer for most of the industry is “a human’s OAuth token wearing a trenchcoat.” The agent doesn’t have its own identity. It inherits the scope, the blast radius, and often the literal credential of whoever spun it up.

That’s a problem the moment agents stop being ephemeral. Most agents today live for minutes to hours: a task starts, the agent runs, it dies. But that’s changing. We’re heading toward agents that run for weeks or months, and a thing that lives that long needs its own identity, not a borrowed one, with permissions that get stricter, not looser, as the lifespan grows.

“Right now the honest answer for most of the industry is ‘a human’s OAuth token wearing a trenchcoat.'”

Think of it like the difference between a contractor you bring in for an afternoon and a contingent worker embedded in your systems for a quarter. You wouldn’t give the afternoon contractor a permanent badge, and you shouldn’t give the quarter-long agent the same access as a five-minute script.

OAuth wasn’t built for this, and it’s not just a missing feature; it’s a structural mismatch, and at root a UX failure: a consent model built for a human in the loop, applied to a process that has none. 

Its whole model assumes a human sits in front of a scope dialog and makes an informed choice, and we all know how that goes: nobody reads the scope list; they click allow. That already-shaky assumption collapses completely when there’s no human reading anything. 

The base spec has no concept of “this client is an agent” or “this grant is expected to run for six months,” just a server-set expiry after the fact. A few IETF drafts are starting to sketch a fix: binding token lifetime to a task’s actual lifecycle and tagging agents with stable identities distinct from the human who invoked them. Still, those are early-stage proposals, not deployed standard(s).

We don’t have a clean answer for where the line sits between “short-lived task, broad-ish access” and “long-lived agent, locked down tight.” Webflow’s security team is actively working through this, and we’re comparing notes with peers across the industry rather than pretending we’ve solved it. But the framing itself matters: treat agent lifespan as a first-class input to your permission model, not an afterthought.

Permissions aren’t a checkbox at provisioning

A permissions redesign, a protocol update, and an identity model represent three distinct levels of the same problem. For each level to remain effective, the other two must be effective too. 

If you create a credential with the correct scope on Day One, but then no one ever verifies that the credential remains valid, your efforts were wasted. Similarly, if you develop a new protocol that finally distinguishes an agent from the human behind it, but all integrations continue to provide standing, all-or-nothing access by default, you have done little good. Neither approach addresses the deeper question at the root of both: what an agent is permitted to do should depend on how long it will exist. 

“The teams who get this right will be the teams who stopped viewing scope, identity, and lifetime as three separate evaluations, and began treating them as a single setting.”

Right now, nearly all components of the technology stack don’t know how to ask that question, much less answer it. The teams who get this right will not be the ones who developed a more efficient scanning tool. Rather, they will be the teams who stopped viewing scope, identity, and lifetime as three separate evaluations that occur once per agent instance, and began treating them as a single setting that must be evaluated each time the agent’s function or existence changes.

This article was originally published on September 9, 2026, on webflow.com.

The post Why MCP security is about permissions overhaul appeared first on The New Stack.

The AI-native SDLC won’t be one process 

12 septembre 2026 à 16:00
Five fuzzy pom-poms — white, pink, magenta, purple, and teal — arranged in a horizontal row across the upper portion of a dark navy background, each casting a small shadow, with colored light washing the backdrop in magenta and teal.

Anthropic recently published its AI-Native SDLC Playbook. Its central claim is that “code is no longer the bottleneck.” When agents can produce an implementation in minutes, the constraint moves to everything around the build phase: planning, review, verification, deployment, and governance.

The risk, if organizations get this wrong, is producing ten times the changes at the same quality per change or worse, with no way to identify which changes are the bad ones. The traditional answer is that a person looks at each one, and that is exactly what stops working at this volume.

The risk… is producing ten times the changes at the same quality per change or worse, with no way to identify which changes are the bad ones.

The playbook gets the foundations right. What it doesn’t capture is that an organization’s process has nuance: it is really a family of processes that vary with the change at hand, not a single flow every change travels through.

The spec-driven wave

The playbook is part of a broader wave of spec-driven development tooling, including Amazon’s Kiro and GitHub’s Spec Kit. The tools share a common shape. Written artifacts drive the work: an intent document becomes a spec, a plan, a diff, and review findings, all committed to version control. Policy is enforced by deterministic mechanisms such as hooks, rather than by instructions in a prompt. Agents check their own work before a human sees it. Humans own the approvals.

But each of these tools also prescribes a particular process: a fixed sequence of stages that produce fixed artifacts, and that every change travels through. Adopting the tool means adopting its process.

One organization runs many processes

No real organization runs a single process. The right process for a change depends on the risk it carries and the accountability it requires. A documentation fix, a dependency upgrade, and a schema migration in a payments service should not travel the same path. They need different levels of verification, different approvers, and different records. In regulated domains, the process itself is part of the compliance obligation: auditors expect a record of who approved each change and based on what evidence. What must be recorded differs by the type of change.

When a tool prescribes one process, teams route around it for changes that don’t fit, which is the worst outcome because the real process becomes invisible.

When a tool prescribes one process, teams route around it for changes that don’t fit, which is the worst outcome because the real process becomes invisible. Or the vendor keeps adding configuration until the tool becomes a workflow engine that nobody fully understands.

The tool should not prescribe a process. It should give the organization a way to define its own.

Processes as state machines

A better model is to define each process as a state machine. The states are facts about a change: reviewed, validated against its dependencies, approved for production. Those facts live in systems no single tool owns: the repository, CI, the cluster, the tracker. So a process cannot be a program that executes steps. It is a set of rules that react to observations about those systems. Each rule specifies:

  1. The facts it requires before it can fire.
  2. Its gate: fire automatically, or wait for a person’s approval.
  3. The permission it grants when it fires, such as merging or deploying.

The definition is this set of rules, stored as data and reviewed like code. An organization runs many small machines, one per risk class.

At runtime, this behaves nothing like a workflow engine. No component tracks “we are on step four”: the process advances when a fact appears in the system that owns it, and rules react. Events that arrive late, twice, or after a restart are handled like any other, because rules only react to current state. The gate is one of a rule’s conditions, so you can hold firing during an incident or a release freeze without editing any definition.

Gates need enforcement. A gate implemented as a prompt instruction depends on the model following it. The agent harness can provide the determinism required to run the state machine and enforce its gates, stopping the agent between actions until a gate is answered, while the infrastructure enforces the rest.

The process adapts to the change

One fixed process definition per repository is not enough: every change in that repository would still travel the same path, regardless of its risk. The path a change takes should depend on what the change is, and this routing comes from classifying the change, not from the author choosing a path. The organization defines classification using signals it already has: the paths a change touches, the repository it lives in, a label on its tracking issue, etc.

The definitions themselves also need to change over time, and that has to be safe. Because a process definition is data, editing it is itself a change, and it goes through its own gated process. Loosening an approval gate on the release process gets reviewed the way a schema migration does, not edited the way a config file does.

Click to enlarge graphic.

Concretely, consider three changes to the same service:

  • A documentation fix is classified by the paths it touches. Its process has two states: the build passes, and it merges. No person is involved.
  • A dependency upgrade skips design review, but its process requires compatibility evidence: the upgraded service runs its integration tests against real dependencies. A major version bump adds an approval that a patch bump does not.
  • A schema migration in the payments service is classified by the component it touches, no matter what kind of change it claims to be. Its process adds states the others never see: review by a payments owner, validation against production-shaped data, and a release approval from someone accountable for that domain.

Each transition, in each path, is logged with who approved it and on what evidence.

Tenets

The tenets these processes should follow:

  • Autonomy is granted per action, and grows over time. Each transition is set to fire automatically, require approval, or hold. As agents prove reliable on a class of change, that setting is relaxed, so the process absorbs agent improvements without redesign.
  • Human attention is spent only where judgment is needed. Agent effort keeps getting cheaper; supervision hours do not. A person is brought into the loop only when the decision requires human judgment, and is given the context to decide quickly.
  • Evidence comes from outside the agent. An agent’s own report never moves a change forward. Transitions fire on facts from systems the agent cannot write to, such as test results and validation in a realistic environment.
  • The process record is the audit trail. The definition is the written policy, and the transition log shows who approved each step, on what evidence, under which version of the policy.

Quality at scale

The playbook and its peers get the foundations right. What is missing is the ability for an organization to define its own processes, vary them by the risk of each change, and evolve them safely. The goal is not fewer humans in the loop. It is spending human judgment only where it is needed, backed by evidence agents cannot produce about themselves, so that quality holds while throughput multiplies.

We are building these ideas at Signadot and acting as our own guinea pigs, running our own development through this process. If you’re experimenting with these ideas too, we’d love to talk!

The post The AI-native SDLC won’t be one process  appeared first on The New Stack.

How AWS Lambda logs every flow across thousands of microVMs per host with eBPF and Rust

11 septembre 2026 à 14:00
Abstract dark geometric corridor with glowing purple and blue neon lines, representing high-density network flow architecture in AWS Lambda.

On any compute platform, when a security alert fires, the question is always the same. Which workload talked to that endpoint, when, and how much? So, you dig through logs and hope they hold up. The hard part is knowing what to look for and where to find it. On a single server, thousands of microVMs run for a few hundred milliseconds, then shut down. 

Each one belongs to a different customer running a different function. Within milliseconds, the workload is done, and the logs captured during those few milliseconds are the only witness left.

“Within milliseconds, the workload is done, and the logs captured during those few milliseconds are the only witness left.”

That was our situation at AWS Lambda. We needed a complete network ledger for every tenant workload, no matter how short its life or how much traffic it generates. This is the story of how we replaced an aging capture system with a purpose-built pipeline written in eBPF and Rust: why the old architecture ran out of road, and the decisions that let the new one hold up at Lambda’s scale.

The job: a record you’re not allowed to touch

Every Lambda worker is a bare-metal EC2 instance packed with microVMs, each an isolated Firecracker guest. They all talk over the network to S3, other AWS services, the public internet, and back into a customer’s VPC. Something has to keep an honest, complete account of it all.

That is where a network flow log comes in. A network flow log helps with investigation, incident response, audit, and reconstructing what happened with a workload. Imagine it as the system of record for what happened to a network packet flowing across the system. The same records also feed into network usage and metering services that demand accuracy above all else. Lastly, all records must be persisted for audit and compliance purposes.

Two properties matter more than the rest. The record must be complete and correctly attributed, and capturing it must add almost no overhead to both the network flow and the platform.

Correct attribution means every packet/flow is associated with the microVM where it landed and the tenant that produced it. Completeness means no missed packets. A missing or misattributed record can cause billing, observability, and monitoring problems at scale. Imagine this happening for every microVM when Lambda serves millions of requests per second. Overhead matters for a reason that isn’t readily visible to customers but plays a major role in how you run a service at Lambda’s scale. Every extra megabyte of RAM and every microsecond of CPU consumed adds up at Lambda’s density, lowering utilization, operating margin, and the ability to serve requests under load.

Why the old way ran out of road

While building the new multi-tenant, Firecracker microVM-powered Lambda, we started with a system borrowed from the older, less-dense single-tenant EC2-era design. It had two parts: a kernel-side extension that counted packets and matched them to tenants and sandboxes per flow, and a userspace daemon that read those counters, batched them into records, serialized them, and frequently uploaded the files. This solution worked well for a small number of VMs. However, the solution broke at Lambda’s density for two mutually exclusive reasons.

First, rule explosion. iptables walks its rules more or less linearly for every packet, and each new microVM piled more rules onto the chain. One worker running a couple of thousand microVMs needed well over a hundred thousand iptables rules just to keep the record. Every packet paid a tax proportional to how crowded the host was and what slot it received. So a packet’s bookkeeping slowed as the host got busier, which is exactly the wrong direction, since the whole plan was to pack more microVMs onto a host, not fewer.

“A record that can’t see half the address space isn’t one you can trust, and the moment dual-stack IPv6 support was proposed for Lambda, the old approach was finished.”

The second reason was harder: the borrowed kernel module did not speak IPv6. A record that can’t see half the address space isn’t one you can trust, and the moment dual-stack IPv6 support was proposed for Lambda, the old approach was finished, no matter the performance improvements.

Design considerations

As we embarked on the rewrite, we had a few non-negotiable design considerations: correct attribution, meaningful overhead reduction, and support for IPv6 as a first-class citizen. Dozens of internal systems already knew how to read the Amazon Ion records the old daemon produced. If we could emit byte-for-byte identical records, we could rip out the whole capture engine and no consumer, flow-log, or metering would notice the swap.

To solve for correct attribution, we implemented isolation at the microVM level. Each microVM already had its own network: a namespace with its own virtual devices. That network is the unit everything else organizes around.

The shape of the new system

The replacement is three cooperating pieces. We’ll give each a plain name based on their role.

 ONE WORKER HOST

   control plane
       |
       |  gRPC over a Unix domain socket
       v
   orchestrator ....... one privileged process per host
       |                (loads eBPF, configures TC,
       |                 spawns one tagger per network)
       v
   --- per network (one microVM) ------------------------------

   eBPF capture  -->  ring buffer  -->  tagger  -->  Ion records
   TC hooks on        one per           Rust,
   the network's      network           userspace
   devices

   ------------------------------------------------------------
       |
       v
   same billing + flow-log pipeline as before

At the bottom sits the kernel capture layer. It’s a set of small eBPF programs attached to the traffic-control(tc) hook on each network’s relevant virtual Linux devices. They intercept packets and emit one compact event per packet into a ring buffer. They just watch. Nothing in these programs can copy, block, drop, or rewrite a packet; there’s literally no code path for it.

In the middle is the tagger: an unprivileged userspace process written in Rust, one per network. It drains its own dedicated ring buffer, rolls the raw per-packet events up into per-flow records, and writes them to disk in the legacy Amazon Ion format. On top is the orchestrator, one privileged process per host. It owns everything that needs elevated permissions: loading the eBPF programs, wiring up traffic control, spawning and supervising the fleet of taggers. It exposes a small lifecycle API over a Unix domain socket so the control plane can create, assign, recycle, and tear down tagging as microVMs come and go.

The decoupled design lets capture happen in the kernel, aggregation in userspace, and one process per host orchestrates it all. The captured records go to the existing downstream consumers unchanged.

Capturing in the kernel, without getting in the way

The capture programs attach to the clsact qdisc in traffic control(tc), on both the ingress and egress side of each of a network’s devices. A network spans two devices, so that’s four attach points per network. Traffic control is a good place to stand. It sees every packet early, before anything downstream has touched it. Every eBPF program reads the packet and returns the “keep going” action. We never drop or modify a customer’s packet, and nothing we do adds meaningful latency.

For each packet, the program walks the headers- Ethernet, then IPv4 or IPv6, then TCP, UDP, or ICMP and writes one fixed-size event into that network’s BPF ring buffer. The event is small on purpose, about two dozen bytes for IPv4. A simplified version looks like this:

/* one event per packet, ~24 bytes for IPv4 */
 struct flow_event {
     u8  ip_version;       /* 4 or 6 */
     u8  protocol;         /* TCP / UDP / ICMP */
     u8  direction;        /* ingress or egress */
     u8  device_id;        /* which of the network's devices */
     u16 local_port;       /* "local" is always the sandbox side */
     u16 remote_port;
     u32 flags_and_bytes;  /* TCP flags in bits [31:24], byte count in [23:0] */
     u32 local_addr;       /* 16 bytes for IPv6 */
     u32 remote_addr;
     u32 received_time_ms;
 };

One packet, one event, one byte count. We don’t count packets in the kernel. The flags and the byte count share a single 32-bit word: eight bits of TCP flags on top, a 24-bit byte count underneath. Doing less work per packet in the kernel is the entire point, so aggregation is somebody else’s job.

The local and remote fields get normalized by direction before the event ever leaves the kernel. Arriving or departing, local always means the sandbox side and remote means the outside world. That one small normalization means userspace never has to reason about direction when it groups flows, and the record reads the same way regardless of which way the packet was headed.

“Doing less work per packet in the kernel is the entire point, so aggregation is somebody else’s job.”

Let’s look at how eBPF plays a key role in simplifying the architecture. The eBPF program here works in a reserve-then-commit order. Each eBPF program maintains a dedicated ring buffer to record packet metadata. You ask the ring buffer for space, add the event in place, and submit. This avoids copying data through a syscall, consumes minimal CPU, and needs no per-CPU bookkeeping. Just a single consumer that drains it in order.

However, all this complex logic for parsing and event submission past the eBPF verifier took work. The eBPF verifier has a critical job: proving a program is safe before the kernel can load it, and that requires strict memory bounds checks and instruction limits. We had to make sure it could verify and load programs quickly, so we did the following: First, we made the header parser a shared subroutine, so the verifier proves it once instead of re-proving it inline at every attach point in the eBPF program. Then, we bounded the IPv6 extension-header walk to a fixed number of hops so the verifier can ensure it terminates. We also had to make sure packet fragments past the first IPv4 fragment report zero ports and flags rather than vending garbage into the ring-buffer events. We also had to make sure any coalesced super-packets from segmentation and receive offload (GSO/GRO) have their byte counts handled correctly. A garbage port or an over-counted byte would create a false log entry, so we keep the record honest at the source and byte-identical with the old system.

However, we don’t trust the verifier as the sole source of correctness, because this code produces a record critical to billing, compliance, and auditing systems; each eBPF program is written in C and runs through a formal model checker (CBMC) during every build. Its harnesses assert, among other things, that the event struct’s byte layout stays compatible with what every attached program expects. A struct that silently shifts by a byte is the kind of bug that quietly corrupts every record it touches, and nobody notices until the day they need the log.

Sizing the ring buffer from first principles

The ring buffer is the one thing the kernel producer and the userspace consumer share, and its size is a real tradeoff. Too small and you drop events under a burst, which means missing records, which is a hole in the log right when traffic matters. Too large and you waste memory, and you pay for that waste on every ring buffer on the host.

So we didn’t guess. We derived the floor from each microVM’s packet rate; for example, if the ceiling is 100,000 packets per second per direction. We drain the ring roughly every 100 milliseconds. Multiply the peak rate by the drain interval, by the event size, and by two directions:

ring bytes =~ 62,500 pps x 0.1 s x ~24 bytes x 2 directions
            =~ 300 KB

The ring buffer API requires a power of two, so the design defaults to 512 KiB. That’s the smallest buffer that can’t overflow between drains at the guest’s own maximum packet rate. Put another way, the floor ensures a guest can’t outrun the recorder, even when it’s trying to. The size is configurable per network. In the running deployment, we currently provision it more generously than that floor, on the order of a couple of megabytes, while we finish tuning the right per-workload value. The number to defend is the floor, and the floor comes from a hard system limit, not a guess.

The drain cadence has another nice property: waking a userspace process isn’t free, and a fleet of thousands of processes all waking constantly would thrash the CPU. So the kernel decides when to bother. It checks how full the ring is and only forces a wakeup once the ring crosses about one percent full. Below that, it stays quiet and lets events pile up. Userspace, on its end, won’t come back to read more than once every 100 milliseconds. A quiet flow just sits there until the next drain, basically free. A busy one trips that one-percent threshold and gets read almost right away. Nothing’s on a fixed timer, so neither case gets the timing wrong.

Draining and aggregating in Rust

The tagger turns raw per-packet events into the per-flow records the pipeline stores. One tagger per network, running unprivileged.

We picked Rust for boring, practical reasons. At this density, thousands of these processes run on a host, each holding a small amount of state that has to be correct. A garbage-collected runtime would give us pause times and memory that balloon under load, and a pause in the wrong place could create a gap in the record. Rust gives us predictable memory and no collector, plus a compiler that flat-out refuses to build whole categories of bugs that turn into misattribution. Each tagger runs in a few hundred kilobytes of RAM against a budget of about a megabyte. That’s what makes thousands of them per host affordable.

Inside, it’s a small set of cooperating tasks on a single-threaded async runtime. One task reads the ring, another owns the flow state, and a third writes parcels. The only work we fence off onto a blocking pool is the couple of operations that genuinely block: receiving the ring descriptor and serializing Ion, since the Ion writer isn’t async. It reads the ring through epoll, so it sleeps when there’s nothing to do and wakes when there is.

As events arrive, the tagger drops them into a flow map keyed by device, the five-tuple, and a tenant attribution handle. That handle comes from the metadata the control plane handed over when the flow was activated. Matching events accumulate bytes, packet counts, and OR’d TCP flags. Grouping happens as the events are read, so the hot path stays a lookup and an add.

Where does attribution actually come from? That’s because mapping is the most important property for the record. The kernel event carries no identity, and it doesn’t need to. Every network has its own dedicated ring and devices, so packets are separated long before the tagger sees them. The tagger isn’t pulling one tenant’s packets out of some shared firehose. The stream it reads was only ever that one tenant’s, because the ring and the devices feeding it belong to that tenant alone.

Once a minute, on a fixed interval lined up to the top of the second to match the old system it replaced, the tagger serializes the completed flows into Amazon Ion records in exactly the schema the old daemon produced. Each file gets written to a temporary name, flushed to disk, and renamed into place. A reader sees a complete record or nothing, never a torn one. A separate flush loop, with a little random jitter at startup so thousands of processes don’t all write at the same instant, drains completed flows even after a microVM has gone quiet. A workload that goes silent still leaves a finished record behind it.

Because the records are byte-compatible with the old format, the entire downstream world kept working untouched. And when we ran the two systems side by side, we could compare their output record for record and confirm they agreed. That’s about as direct a completeness check as you can get.

Least privilege, enforced by a file descriptor

This is the part of the design I’m most fond of, because it uses an old Unix trick to get a strong security property for almost nothing.

The processes that do the actual packet work, the thousands of taggers, hold no elevated privileges. They can’t load eBPF or touch traffic control. They can’t even open the ring buffer map on their own. All of that power lives in one place: the per-host orchestrator, and even it runs with just the two capabilities it needs rather than as root.

“Passing a file descriptor over a socket is a decades-old Unix feature, and it lets us keep thousands of processes powerless while concentrating privilege in one small place.”

So how does an unprivileged tagger read a ring buffer it isn’t allowed to open? The orchestrator opens it and hands the open file descriptor to the tagger over a Unix domain socket, using the kernel’s SCM_RIGHTS mechanism to pass descriptors between processes. The tagger gets a ready-to-use handle to the ring and nothing else. It never had, and never needs, permission to create one. The privileged surface of the whole system is one small process per host. The thousands of processes touching customer traffic are about as powerless as we can make them.

A lifecycle API, and why it has two doors

MicroVMs come and go constantly, so the control plane needs a way to tell the orchestrator when to start and stop recording a network. It does that through gRPC APIs over a Unix domain socket, with a handful of methods: create a set of flows, activate a flow, recycle one, tear one down, plus a health check.

Starting to record is split into two calls, a heavy one and a light one, and the split is deliberate. Create is a heavy call, the expensive path: it loads and attaches the eBPF programs, configures traffic control, and spawns the tagger. Attaching to network devices takes a kernel lock that every such operation on the host contends for, so when a host is standing up many networks at once, we batch these to keep everyone from serializing behind that one lock. Activate is the lighter call. By the time it runs, the machinery already exists, so it just hands over the customer metadata and flips the flow into steady-state recording. Its latency budget is tight: under 2 milliseconds at p90, under 10 milliseconds at p99.9, matching the baseline of the system it replaced.

An honest tradeoff: kill it, or reuse it

Not every decision came out clean. These are lessons for anyone building something similar.

The original design had a strict rule for recycling a network: always destroy the tagger and spawn a fresh one. From a correctness standpoint, the reasoning was airtight. A brand-new process can’t carry stale metadata from a previous tenant, so a flow from one tenant landing in another’s record across a recycle becomes structurally impossible. Kill it, don’t try to clean it. We were sure that was the right call.

Reality under Production workloads showed us that forking and exec’ing a new process thousands of times as networks churned turned into a real source of CPU spikes at scale. The safest choice showed up as a flame graph. So the shipped system needed a new knob. A workload that reuses its networks can reuse the tagger after a recycle, trading off a little of that structural guarantee for a lot less CPU churn. A workload that wants the strict, cross-tenant-proof behavior leaves the knob off. I still think the strict version is the more correct design, but the fleet’s CPU budget just didn’t allow it.

What it bought us

Start with the number that killed the old design. One host needed more than a hundred thousand firewall rules to keep the record for two thousand microVMs, and each additional microVM piled on more, taxing every packet a little further. The eBPF version swaps that linear rule walk for constant-time map lookups whose cost doesn’t climb as the host fills up. The linear tax is gone. That puts the density target, roughly double the microVMs per host, within reach, and without the burst-time gap a per-packet tax invites.

The rest of the payoff falls out of the constraints we started with:

  1. IPv6 flows, invisible to the old tool, get recorded like anything else, so the log covers the whole address space instead of half.
  2. Each tagger lives in a few hundred kilobytes of RAM against a roughly one-megabyte budget, small enough that thousands per host is practical.
  3. Activating a flow into steady-state recording stays under 2 milliseconds at p90 and under 10 milliseconds at p99.9.
  4. The capture layer is observe-only and formally checked, and the processes touching customer traffic hold no privileges. We added significant visibility while shrinking the trusted, privileged surface that could corrupt the record. 
  5. The output records are byte-for-byte identical to the old format, so every downstream flow-log and metering consumer kept working with no change

Lessons worth carrying to other systems

A few of these travel well beyond Lambda. Kubernetes pods, edge runtimes, the sandboxes people are spinning up now to run AI agents. Anywhere you’ve got many tenants sharing a host and need a trustworthy record of their traffic, the same shapes hold.

First: observe from outside the hot path. The moment your recording logic sits inline in packet forwarding, its cost becomes a tax on every packet, and that tax is heaviest right when the record matters most. That same pressure pushes teams to drop or sample data, and an audit trail can’t survive sampling. eBPF lets you watch from the side and emit a compact event, while everything expensive happens elsewhere.

Second: size buffers from something real. A buffer sized by an actual rate limit times an actual drain interval is a number you can defend in a review, and it’s what lets you promise no dropped events under a burst. We could’ve picked 512 KB because it felt about right, and it probably would’ve held most of the time, right up until some burst it wasn’t sized for.

Third: keeping tenants apart at the point of capture is the part I’d argue hardest for. Give each tenant its own ring and its own devices, and the streams never touch, so you’re labeling clean traffic instead of guessing after the fact.

Fourth: old primitives are underrated. Passing a file descriptor over a socket is a decades-old Unix feature, and it lets us keep thousands of processes powerless while concentrating privilege in one small place.

Last one, and it’s the cheapest to get wrong: when you swap out an engine, keep the bolt pattern. Byte-for-byte identical output let us replace the entire capture path with zero downstream migration, and it handed us a record-for-record way to prove the new system saw everything the old one did.

The post How AWS Lambda logs every flow across thousands of microVMs per host with eBPF and Rust appeared first on The New Stack.

❌