❌

Vue normale

Reçu avant avant-hier

One engineer shipped 2,000 PRs a month to production. Verification is the key.

19 septembre 2026 à 16:00
Abstract dark digital wireframe mesh with chromatic glitch effects, representing AI agent verification and virtualized software environments.

Lauren Tan, an engineer on the Grok team at SpaceXAI, who previously worked at Cursor and Meta, recently published a guide to her personal agent workflow: pstack. The attention-grabbing number is that pstack has let her ship 2,000 pull requests (PRs) a month to production with high confidence. That’s one engineer shipping nearly 100 PRs per working day.

The number is incredible and undeniably an outlier, but the direction is not a surprise. I have argued previously that coding agents would enable teams to generate ten times the code with similar headcounts. What is surprising is that these are not just code output numbers. These are actual changes landing in production.

According to Tan, the most critical piece of that workflow is verification. A verification skill lets an agent check its own work and keep going until the task is done, and she treats it as “critical infrastructure” rather than one skill among many. 

The verification skill rests on something underneath it: a rich runtime the agent can drive, inspect, and get structured answers from. For a single application, that runtime is the application itself, started on demand. For a system made of dozens or hundreds of services, no such runtime exists by default, and providing one that keeps up with hundreds of parallel agents is the hard part.

Verification is the whole game, and the math says so

Her argument for agentic verification is a throughput argument. An agent that can check its own output keeps working until the task is done. An agent that can’t hand you a diff and wait makes you the slowest component in the loop. That is why she claims strong verification skills can multiply a team’s output by 100 to 1,000 times.

“An agent that can check its own output keeps working until the task is done. An agent that can’t hand you a diff and wait makes you the slowest component in the loop.”

At 2,000 pull requests a month, reviewing every change by hand would allow about five minutes per PR across a full working month. Human review cannot be the verification layer at that volume. Whatever does the checking has to run without a person in the loop, and it has to run in parallel with the agents generating the work.

The model assumes the agent can run the whole application

The verification skill she describes generates a command line interface (CLI) and a feature map for the application. The CLI lets an agent start the app, navigate it, inspect state, and read structured JSON results back. Each agent gets a complete copy of the application and can test a change end to end.

She is direct about how much rests on that runtime: “I personally feel that agentic verification is so important that I would unironically suggest building your own rich debugging tools, or even choosing a different tech stack, in order to have unfair advantages and extreme productivity in building software.”

Her approach to providing a runtime for her agent works because the application fits in one process. A frontend, a compiler, or a single service with a database can start from a CLI in seconds and be thrown away afterward.

“I would unironically suggest building your own rich debugging tools, or even choosing a different tech stack, in order to have unfair advantages and extreme productivity in building software.”

For teams building complex distributed applications, their system does not have that property. The application is the interaction between an order service, a payments service, an inventory service, a queue, several databases, and a handful of third-party APIs. At larger shops, the count runs into the thousands. A pull request to one service is only verified by exercising the calls it makes and receives. The CLI can start the changed service. It cannot start the system.

None of the existing runtimes survive hundreds of parallel agents

Local runtimes with mocks are cheap and can run fully parallel using worktrees or CDEs. Their problem is fidelity. Mocks encode what a dependency did the last time someone looked, and they drift the moment the real service changes. An agent that verifies against mocks closes its loop against fiction, and the failure shows up after merge.

A full copy of the stack per change is faithful and isolated. But its cost scales with the number of services times the number of concurrent changes, and at hundreds of agents that cost is untenable. Time is the bigger problem. A full stack takes minutes to provision, and the loop she describes has the agent testing every iteration of a change while it is still working on it. An environment that is ready after the agent has moved on to its next attempt is no use to it.

Shared staging is faithful and cheap because there is one of it, and that is the whole problem. A single mutable environment cannot host hundreds of concurrent changes. Agents overwrite each other’s deployments, a broken change from one agent becomes failed tests for every other agent, and the loop-closing property that makes the workflow valuable disappears.

Each existing verification runtime plotted on a chart showing its realism of dependencies against concurrency.

What verification needs when the callers are agents

Read Lauren’s workflow as a requirements document, and five properties emerge:

  • The change has to run against real dependencies or the verification means nothing.
  • Hundreds of concurrent changes have to be unable to see each other.
  • The cost of an environment has to scale with the size of the change, not the size of the system.
  • Environments have to come up in seconds, because an agent waiting on provisioning is parallelism you paid for but didn’t use.
  • All of it has to be reachable through the CLI or MCP server the agent already uses, because the caller is an agent.

The first and third requirements pull in opposite directions. Realism pushes toward complete copies of the system. Cost pushes toward sharing as much as possible. Shared staging resolves that tension by giving up isolation, and a per-change full stack resolves it by giving up cost efficiency. A design that satisfies all five has to share and isolate at the same time.

Virtualized full-stack environments share the system and isolate the change

The architecture that does this treats an environment as a view of a running system rather than a copy. One shared set of stable services runs continuously, deployed from the main branch and kept healthy the way production is. When an agent needs to verify a change, it runs only the service it modified, on its own machine or as a lightweight deployment in the cluster, and joins it to the shared stack as a new isolated environment.

“The architecture that does this treats an environment as a view of a running system rather than a copy.”

From inside that environment, the changed service is the version of record, and every other call falls through to the shared stable versions. The agent sees a complete, realistic system, and so do the other hundred agents, each seeing a system that differs from the baseline by only the delta of its own change. Requests carry their environment identity as they cross service boundaries, which keeps one agent’s traffic from reaching another agent’s version under test. Stateful side effects that cannot be shared safely, like queue topics or writable databases, get a per-environment copy where needed.

Diagram showing "agent 1 env" and "agent 2 env" interacting with the shared cluster.

The cost model follows directly. An environment costs one or two running services instead of sixty; it is ready in the time a single service takes to start, and you can create and destroy it from inside the agent’s own loop. This is the pattern Signadot packages for Kubernetes, with the shared stable stack running in the team’s existing cluster.

The loop, end to end, with an agent as the actor

Put the two halves together, and the workflow that enables her to ship 2,000 PRs a month to production carries over to a distributed system almost unchanged. An agent picks up a task and changes one service. It asks for an environment for that change and gets one in the time it takes its service to start. It then drives real requests through the system’s entry point and watches them traverse the real dependency graph, with only its own service running new code. It reads structured results, fixes what failed, and runs again. When the checks pass, it opens the PR, and the environment goes away at merge.

“Environments stop being something the platform team hands out and become something agents create, use, and discard as needed.”

For the platform team, the unit of work changes. Today it provisions environments, whether that means keeping a shared staging alive or stamping out full copies of it. In this model, it runs one shared stable stack and the layer that virtualizes it: context propagation across every service, isolation for the stateful dependencies that cannot be shared, and the tooling that creates and tears down environments. Environments stop being something the platform team hands out and become something agents create, use, and discard as needed.

Verification capacity is the new ceiling on throughput

Lauren Tan’s post is not a story about one unusually productive engineer. It shows what happens when agents run the full loop, writing a change, verifying it, and iterating without a person in between. The verification infrastructure is the foundation that the entire loop stands on.

In distributed applications, that infrastructure has to be a runtime environment that gives every agent real dependencies, keeps hundreds of concurrent changes from seeing each other, costs a change rather than a copy of the system, and is ready in the seconds an agent is willing to wait. That is what turns agent parallelism into shipped code rather than a longer review queue. That model of runtime environments is exactly what we built Signadot to enable.

The post One engineer shipped 2,000 PRs a month to production. Verification is the key. appeared first on The New Stack.

The AI-native SDLC won’t be one process 

12 septembre 2026 à 16:00
Five fuzzy pom-poms — white, pink, magenta, purple, and teal — arranged in a horizontal row across the upper portion of a dark navy background, each casting a small shadow, with colored light washing the backdrop in magenta and teal.

Anthropic recently published its AI-Native SDLC Playbook. Its central claim is that “code is no longer the bottleneck.” When agents can produce an implementation in minutes, the constraint moves to everything around the build phase: planning, review, verification, deployment, and governance.

The risk, if organizations get this wrong, is producing ten times the changes at the same quality per change or worse, with no way to identify which changes are the bad ones. The traditional answer is that a person looks at each one, and that is exactly what stops working at this volume.

The risk… is producing ten times the changes at the same quality per change or worse, with no way to identify which changes are the bad ones.

The playbook gets the foundations right. What it doesn’t capture is that an organization’s process has nuance: it is really a family of processes that vary with the change at hand, not a single flow every change travels through.

The spec-driven wave

The playbook is part of a broader wave of spec-driven development tooling, including Amazon’s Kiro and GitHub’s Spec Kit. The tools share a common shape. Written artifacts drive the work: an intent document becomes a spec, a plan, a diff, and review findings, all committed to version control. Policy is enforced by deterministic mechanisms such as hooks, rather than by instructions in a prompt. Agents check their own work before a human sees it. Humans own the approvals.

But each of these tools also prescribes a particular process: a fixed sequence of stages that produce fixed artifacts, and that every change travels through. Adopting the tool means adopting its process.

One organization runs many processes

No real organization runs a single process. The right process for a change depends on the risk it carries and the accountability it requires. A documentation fix, a dependency upgrade, and a schema migration in a payments service should not travel the same path. They need different levels of verification, different approvers, and different records. In regulated domains, the process itself is part of the compliance obligation: auditors expect a record of who approved each change and based on what evidence. What must be recorded differs by the type of change.

When a tool prescribes one process, teams route around it for changes that don’t fit, which is the worst outcome because the real process becomes invisible.

When a tool prescribes one process, teams route around it for changes that don’t fit, which is the worst outcome because the real process becomes invisible. Or the vendor keeps adding configuration until the tool becomes a workflow engine that nobody fully understands.

The tool should not prescribe a process. It should give the organization a way to define its own.

Processes as state machines

A better model is to define each process as a state machine. The states are facts about a change: reviewed, validated against its dependencies, approved for production. Those facts live in systems no single tool owns: the repository, CI, the cluster, the tracker. So a process cannot be a program that executes steps. It is a set of rules that react to observations about those systems. Each rule specifies:

  1. The facts it requires before it can fire.
  2. Its gate: fire automatically, or wait for a person’s approval.
  3. The permission it grants when it fires, such as merging or deploying.

The definition is this set of rules, stored as data and reviewed like code. An organization runs many small machines, one per risk class.

At runtime, this behaves nothing like a workflow engine. No component tracks “we are on step four”: the process advances when a fact appears in the system that owns it, and rules react. Events that arrive late, twice, or after a restart are handled like any other, because rules only react to current state. The gate is one of a rule’s conditions, so you can hold firing during an incident or a release freeze without editing any definition.

Gates need enforcement. A gate implemented as a prompt instruction depends on the model following it. The agent harness can provide the determinism required to run the state machine and enforce its gates, stopping the agent between actions until a gate is answered, while the infrastructure enforces the rest.

The process adapts to the change

One fixed process definition per repository is not enough: every change in that repository would still travel the same path, regardless of its risk. The path a change takes should depend on what the change is, and this routing comes from classifying the change, not from the author choosing a path. The organization defines classification using signals it already has: the paths a change touches, the repository it lives in, a label on its tracking issue, etc.

The definitions themselves also need to change over time, and that has to be safe. Because a process definition is data, editing it is itself a change, and it goes through its own gated process. Loosening an approval gate on the release process gets reviewed the way a schema migration does, not edited the way a config file does.

Click to enlarge graphic.

Concretely, consider three changes to the same service:

  • A documentation fix is classified by the paths it touches. Its process has two states: the build passes, and it merges. No person is involved.
  • A dependency upgrade skips design review, but its process requires compatibility evidence: the upgraded service runs its integration tests against real dependencies. A major version bump adds an approval that a patch bump does not.
  • A schema migration in the payments service is classified by the component it touches, no matter what kind of change it claims to be. Its process adds states the others never see: review by a payments owner, validation against production-shaped data, and a release approval from someone accountable for that domain.

Each transition, in each path, is logged with who approved it and on what evidence.

Tenets

The tenets these processes should follow:

  • Autonomy is granted per action, and grows over time. Each transition is set to fire automatically, require approval, or hold. As agents prove reliable on a class of change, that setting is relaxed, so the process absorbs agent improvements without redesign.
  • Human attention is spent only where judgment is needed. Agent effort keeps getting cheaper; supervision hours do not. A person is brought into the loop only when the decision requires human judgment, and is given the context to decide quickly.
  • Evidence comes from outside the agent. An agent’s own report never moves a change forward. Transitions fire on facts from systems the agent cannot write to, such as test results and validation in a realistic environment.
  • The process record is the audit trail. The definition is the written policy, and the transition log shows who approved each step, on what evidence, under which version of the policy.

Quality at scale

The playbook and its peers get the foundations right. What is missing is the ability for an organization to define its own processes, vary them by the risk of each change, and evolve them safely. The goal is not fewer humans in the loop. It is spending human judgment only where it is needed, backed by evidence agents cannot produce about themselves, so that quality holds while throughput multiplies.

We are building these ideas at Signadot and acting as our own guinea pigs, running our own development through this process. If you’re experimenting with these ideas too, we’d love to talk!

The post The AI-native SDLC won’t be one process  appeared first on The New Stack.

Commits on GitHub have doubled in four months. Verification capacity has not.

29 août 2026 à 16:00
Abstract dark digital artwork of diverging red and blue wave lines, symbolizing AI code generation and software verification bottlenecks.

I have been waiting for agent-generated code to show up in public infrastructure data rather than in vendor benchmarks. In August, it did. GitHub now handles 2.9 billion commits a month and says it cannot keep up. Monthly commit volume more than doubled in four months, from 1.4 billion in April to 2.9 billion in August.

The growth broke the platform. On August 17, GitHub went down for 7 hours and 47 minutes after a core infrastructure component in its Central US data center failed to scale with traffic. The postmortem from CTO Vladimir Fedorov did not hedge: If you were trying to ship software that day, GitHub let you down. The remediation is what a hyperscaler does when demand outruns supply. More than 3 million new CPU cores, 120 petabytes of high-speed storage, and an accelerated migration to Azure, which now serves 58% of platform load.

Most coverage treated this as a capacity story, and for GitHub it is one. The number that should worry you is one the postmortem never touches. Every one of those 2.9 billion commits carried an implicit claim that the change works. Almost nothing in the system that produced, transported, and merged them checked that claim against a running system.

“Code generation has become machine-paced, and its volume curve is exponential. Verification is still human-paced, and its capacity curve is close to flat.”

That is the warning inside GitHub’s data. Code generation has become machine-paced, and its volume curve is exponential. Verification, the work of proving a change does what was intended without breaking what already worked, is still human-paced, and its capacity curve is close to flat. The distance between those two curves is the defining infrastructure problem of the next three years.

The commit curve is the first public trace of machine-paced development

GitHub’s telemetry is a proxy for your own organization: it aggregates what thousands of engineering teams are doing at once. Alongside the commit number, the postmortem reports about 130 million merged pull requests and 24 million new repositories a month. Engadget’s reporting notes that GitHub attributes the surge to AI-generated code.

The shape of the curve matters more than its height. Commit volume grew for years at roughly the same rate as the developer population, because commits tracked people. Then it doubled in four months, because it stopped tracking people. A developer who runs three coding agent sessions in parallel produces commits at a rate no hiring plan ever predicted.

“A developer who runs three coding agent sessions in parallel produces commits at a rate no hiring plan ever predicted.”

Your internal dashboards almost certainly show the same shape in miniature: pull request counts climbing quarter over quarter, more commits per engineer, more branches open at once. The public number matters because it proves your curve is not a local anomaly. This is what development looks like when generation is no longer the scarce step.

GitHub’s bottleneck is capacity. Yours is confidence

GitHub’s problem, for all its severity, has a known fix: When traffic outgrows infrastructure, you add infrastructure. Cores, disks, and data centers scale with money, and Microsoft has plenty.

The problem on your side of the platform does not respond to money the same way. A commit is not traffic. It is a claim about behavior: This change does what its description says and breaks nothing downstream. In a distributed, cloud-native system, checking that claim means running the change against the services, data, and traffic it will meet after merge.

The pipeline in front of that check keeps getting faster. AI code review tools triage diffs before a human looks at them, CI has learned test selection and caching, and static analysis catches more than it used to. Those are real gains, and none of them runs the change. The step that verifies behavior, integration, and end-to-end tests against a live system still funnels through a shared staging environment or waits on a full copy of the stack, which takes too long and costs too much to stand up for each change.

That step has a hard ceiling. Staging is one environment per organization, so it functions as a queue. Full-stack duplicates are expensive enough that teams ration them. Neither doubles in four months because you approved a budget. Generation now scales like GitHub. Verification still moves one change at a time.

Workflow diagrams comparing a single shared queue vs one environment per change.

Verification was sized for human pace, and agents broke the sizing

None of this is new. Large engineering organizations were complaining about staging contention and review backlogs years before coding agents existed. The apparatus was built when code arrived at the pace humans type, and at that pace its costs were a tax teams could manage: an occasional staging conflict, a review queue that cleared by Friday. Deferring expensive full-fidelity testing to the end of the pipeline was a reasonable trade while changes were scarce.

Agents do not introduce the bottleneck. They multiply it past the point where the old coping strategies work, and they bring it to smaller teams. Throughput that used to strain a 500-engineer platform now appears on a 50-engineer team running agents in parallel. The assumption under all of those design decisions, that code is the scarce input, is gone, and the checking machinery built on it has not moved. The two curves that used to track each other roughly have come apart. GitHub’s chart is the aggregate picture of that separation.

Chart showing monthly commits vs verification capacity

The gap between the curves fills with unverified merges

Teams respond to the widening gap in a few ways. The first is to make review faster. AI code review tools sit in front of human reviewers, catch real defects in the diff, and keep improving. They also have a limit: A reviewer, human or model, is reasoning about code they have not run, and a diff cannot tell you whether the change holds up against its real dependencies.

The second is to throttle the agents, capping the amount of generated work that enters the pipeline. That protects the verification queue by returning most of the throughput the agents were designed to deliver.

The third is to merge anyway and absorb downstream failures. Changes that were never run against a real system land in main at the rate commits arrive, and the cost surfaces later as broken staging environments and lengthy debugging sessions. Or worse, production incidents and rollbacks.

The August 17 outage is a preview of how that ends. The root cause was not a bad change. It was a component that everyone depended on, and nobody had scaled, failing on the day traffic finally exceeded its design. In most delivery systems, verification is that component.

Verification has to run on the same curve as generation

The structural fix is to make verification match the shape of generation: parallel and per-change. For a single application, this is nearly solved: an agent runs the app on a laptop or in a CI container and checks whether the change works. The hard case is the cloud-native one, where the change’s behavior only exists when interacting with other services, databases, and queues. Both familiar options fail at agent volume: A shared staging environment serializes everything into a queue, and duplicating the full stack for each change is too costly and too time-consuming.

There is a third shape. Keep one shared environment running the stable version of every service, and for each change, deploy only the services that the change touched. Test traffic carries the change’s routing key, so at each hop, a request for that change reaches the changed version while every other request flows through the stable one. Each change gets isolation where it matters, at the services it modified, and shares everything else: the same cluster, the same data, the same downstream dependencies.

Workflow diagram showing request routing in the shared environment

That shape changes the economics and the actor. Verifying one more change costs one extra deployment, not another copy of the stack, so hundreds of changes can be checked in parallel on the cluster you already run.

Because an environment appears in seconds, an agent can use one inside its loop: open a change, run functional checks against real upstream and downstream services, read the failures, and iterate until they pass. The pull request that reaches a human arrives already exercised against the real system, and review time goes to intent and design.

Match the curves or lose the generation gains

GitHub’s 2.9 billion commits quantify a shift that every engineering organization is living through on a smaller scale. Generation is now effectively free and unlimited, so it is no longer the source of advantage.

“The teams pulling ahead are not the ones producing the most commits. They are the ones whose verification capacity rises with their generation capacity.”

The teams pulling ahead are not the ones producing the most commits. They are the ones whose verification capacity rises with their generation capacity, so more generated code becomes more shipped code, rather than a longer review queue, a deeper staging backlog, and a bigger incident bill. Ask what happens to your pipeline when commit volume doubles in four months, because that is no longer a hypothetical. It is the gap we work on at Signadot, and it is worth closing before the curve doubles again.

The post Commits on GitHub have doubled in four months. Verification capacity has not. appeared first on The New Stack.

❌