❌

Vue normale

Reçu avant avant-hierThe New Stack

One engineer shipped 2,000 PRs a month to production. Verification is the key.

19 septembre 2026 à 16:00
Abstract dark digital wireframe mesh with chromatic glitch effects, representing AI agent verification and virtualized software environments.

Lauren Tan, an engineer on the Grok team at SpaceXAI, who previously worked at Cursor and Meta, recently published a guide to her personal agent workflow: pstack. The attention-grabbing number is that pstack has let her ship 2,000 pull requests (PRs) a month to production with high confidence. That’s one engineer shipping nearly 100 PRs per working day.

The number is incredible and undeniably an outlier, but the direction is not a surprise. I have argued previously that coding agents would enable teams to generate ten times the code with similar headcounts. What is surprising is that these are not just code output numbers. These are actual changes landing in production.

According to Tan, the most critical piece of that workflow is verification. A verification skill lets an agent check its own work and keep going until the task is done, and she treats it as “critical infrastructure” rather than one skill among many. 

The verification skill rests on something underneath it: a rich runtime the agent can drive, inspect, and get structured answers from. For a single application, that runtime is the application itself, started on demand. For a system made of dozens or hundreds of services, no such runtime exists by default, and providing one that keeps up with hundreds of parallel agents is the hard part.

Verification is the whole game, and the math says so

Her argument for agentic verification is a throughput argument. An agent that can check its own output keeps working until the task is done. An agent that can’t hand you a diff and wait makes you the slowest component in the loop. That is why she claims strong verification skills can multiply a team’s output by 100 to 1,000 times.

“An agent that can check its own output keeps working until the task is done. An agent that can’t hand you a diff and wait makes you the slowest component in the loop.”

At 2,000 pull requests a month, reviewing every change by hand would allow about five minutes per PR across a full working month. Human review cannot be the verification layer at that volume. Whatever does the checking has to run without a person in the loop, and it has to run in parallel with the agents generating the work.

The model assumes the agent can run the whole application

The verification skill she describes generates a command line interface (CLI) and a feature map for the application. The CLI lets an agent start the app, navigate it, inspect state, and read structured JSON results back. Each agent gets a complete copy of the application and can test a change end to end.

She is direct about how much rests on that runtime: “I personally feel that agentic verification is so important that I would unironically suggest building your own rich debugging tools, or even choosing a different tech stack, in order to have unfair advantages and extreme productivity in building software.”

Her approach to providing a runtime for her agent works because the application fits in one process. A frontend, a compiler, or a single service with a database can start from a CLI in seconds and be thrown away afterward.

“I would unironically suggest building your own rich debugging tools, or even choosing a different tech stack, in order to have unfair advantages and extreme productivity in building software.”

For teams building complex distributed applications, their system does not have that property. The application is the interaction between an order service, a payments service, an inventory service, a queue, several databases, and a handful of third-party APIs. At larger shops, the count runs into the thousands. A pull request to one service is only verified by exercising the calls it makes and receives. The CLI can start the changed service. It cannot start the system.

None of the existing runtimes survive hundreds of parallel agents

Local runtimes with mocks are cheap and can run fully parallel using worktrees or CDEs. Their problem is fidelity. Mocks encode what a dependency did the last time someone looked, and they drift the moment the real service changes. An agent that verifies against mocks closes its loop against fiction, and the failure shows up after merge.

A full copy of the stack per change is faithful and isolated. But its cost scales with the number of services times the number of concurrent changes, and at hundreds of agents that cost is untenable. Time is the bigger problem. A full stack takes minutes to provision, and the loop she describes has the agent testing every iteration of a change while it is still working on it. An environment that is ready after the agent has moved on to its next attempt is no use to it.

Shared staging is faithful and cheap because there is one of it, and that is the whole problem. A single mutable environment cannot host hundreds of concurrent changes. Agents overwrite each other’s deployments, a broken change from one agent becomes failed tests for every other agent, and the loop-closing property that makes the workflow valuable disappears.

Each existing verification runtime plotted on a chart showing its realism of dependencies against concurrency.

What verification needs when the callers are agents

Read Lauren’s workflow as a requirements document, and five properties emerge:

  • The change has to run against real dependencies or the verification means nothing.
  • Hundreds of concurrent changes have to be unable to see each other.
  • The cost of an environment has to scale with the size of the change, not the size of the system.
  • Environments have to come up in seconds, because an agent waiting on provisioning is parallelism you paid for but didn’t use.
  • All of it has to be reachable through the CLI or MCP server the agent already uses, because the caller is an agent.

The first and third requirements pull in opposite directions. Realism pushes toward complete copies of the system. Cost pushes toward sharing as much as possible. Shared staging resolves that tension by giving up isolation, and a per-change full stack resolves it by giving up cost efficiency. A design that satisfies all five has to share and isolate at the same time.

Virtualized full-stack environments share the system and isolate the change

The architecture that does this treats an environment as a view of a running system rather than a copy. One shared set of stable services runs continuously, deployed from the main branch and kept healthy the way production is. When an agent needs to verify a change, it runs only the service it modified, on its own machine or as a lightweight deployment in the cluster, and joins it to the shared stack as a new isolated environment.

“The architecture that does this treats an environment as a view of a running system rather than a copy.”

From inside that environment, the changed service is the version of record, and every other call falls through to the shared stable versions. The agent sees a complete, realistic system, and so do the other hundred agents, each seeing a system that differs from the baseline by only the delta of its own change. Requests carry their environment identity as they cross service boundaries, which keeps one agent’s traffic from reaching another agent’s version under test. Stateful side effects that cannot be shared safely, like queue topics or writable databases, get a per-environment copy where needed.

Diagram showing "agent 1 env" and "agent 2 env" interacting with the shared cluster.

The cost model follows directly. An environment costs one or two running services instead of sixty; it is ready in the time a single service takes to start, and you can create and destroy it from inside the agent’s own loop. This is the pattern Signadot packages for Kubernetes, with the shared stable stack running in the team’s existing cluster.

The loop, end to end, with an agent as the actor

Put the two halves together, and the workflow that enables her to ship 2,000 PRs a month to production carries over to a distributed system almost unchanged. An agent picks up a task and changes one service. It asks for an environment for that change and gets one in the time it takes its service to start. It then drives real requests through the system’s entry point and watches them traverse the real dependency graph, with only its own service running new code. It reads structured results, fixes what failed, and runs again. When the checks pass, it opens the PR, and the environment goes away at merge.

“Environments stop being something the platform team hands out and become something agents create, use, and discard as needed.”

For the platform team, the unit of work changes. Today it provisions environments, whether that means keeping a shared staging alive or stamping out full copies of it. In this model, it runs one shared stable stack and the layer that virtualizes it: context propagation across every service, isolation for the stateful dependencies that cannot be shared, and the tooling that creates and tears down environments. Environments stop being something the platform team hands out and become something agents create, use, and discard as needed.

Verification capacity is the new ceiling on throughput

Lauren Tan’s post is not a story about one unusually productive engineer. It shows what happens when agents run the full loop, writing a change, verifying it, and iterating without a person in between. The verification infrastructure is the foundation that the entire loop stands on.

In distributed applications, that infrastructure has to be a runtime environment that gives every agent real dependencies, keeps hundreds of concurrent changes from seeing each other, costs a change rather than a copy of the system, and is ready in the seconds an agent is willing to wait. That is what turns agent parallelism into shipped code rather than a longer review queue. That model of runtime environments is exactly what we built Signadot to enable.

The post One engineer shipped 2,000 PRs a month to production. Verification is the key. appeared first on The New Stack.

Commits on GitHub have doubled in four months. Verification capacity has not.

29 août 2026 à 16:00
Abstract dark digital artwork of diverging red and blue wave lines, symbolizing AI code generation and software verification bottlenecks.

I have been waiting for agent-generated code to show up in public infrastructure data rather than in vendor benchmarks. In August, it did. GitHub now handles 2.9 billion commits a month and says it cannot keep up. Monthly commit volume more than doubled in four months, from 1.4 billion in April to 2.9 billion in August.

The growth broke the platform. On August 17, GitHub went down for 7 hours and 47 minutes after a core infrastructure component in its Central US data center failed to scale with traffic. The postmortem from CTO Vladimir Fedorov did not hedge: If you were trying to ship software that day, GitHub let you down. The remediation is what a hyperscaler does when demand outruns supply. More than 3 million new CPU cores, 120 petabytes of high-speed storage, and an accelerated migration to Azure, which now serves 58% of platform load.

Most coverage treated this as a capacity story, and for GitHub it is one. The number that should worry you is one the postmortem never touches. Every one of those 2.9 billion commits carried an implicit claim that the change works. Almost nothing in the system that produced, transported, and merged them checked that claim against a running system.

“Code generation has become machine-paced, and its volume curve is exponential. Verification is still human-paced, and its capacity curve is close to flat.”

That is the warning inside GitHub’s data. Code generation has become machine-paced, and its volume curve is exponential. Verification, the work of proving a change does what was intended without breaking what already worked, is still human-paced, and its capacity curve is close to flat. The distance between those two curves is the defining infrastructure problem of the next three years.

The commit curve is the first public trace of machine-paced development

GitHub’s telemetry is a proxy for your own organization: it aggregates what thousands of engineering teams are doing at once. Alongside the commit number, the postmortem reports about 130 million merged pull requests and 24 million new repositories a month. Engadget’s reporting notes that GitHub attributes the surge to AI-generated code.

The shape of the curve matters more than its height. Commit volume grew for years at roughly the same rate as the developer population, because commits tracked people. Then it doubled in four months, because it stopped tracking people. A developer who runs three coding agent sessions in parallel produces commits at a rate no hiring plan ever predicted.

“A developer who runs three coding agent sessions in parallel produces commits at a rate no hiring plan ever predicted.”

Your internal dashboards almost certainly show the same shape in miniature: pull request counts climbing quarter over quarter, more commits per engineer, more branches open at once. The public number matters because it proves your curve is not a local anomaly. This is what development looks like when generation is no longer the scarce step.

GitHub’s bottleneck is capacity. Yours is confidence

GitHub’s problem, for all its severity, has a known fix: When traffic outgrows infrastructure, you add infrastructure. Cores, disks, and data centers scale with money, and Microsoft has plenty.

The problem on your side of the platform does not respond to money the same way. A commit is not traffic. It is a claim about behavior: This change does what its description says and breaks nothing downstream. In a distributed, cloud-native system, checking that claim means running the change against the services, data, and traffic it will meet after merge.

The pipeline in front of that check keeps getting faster. AI code review tools triage diffs before a human looks at them, CI has learned test selection and caching, and static analysis catches more than it used to. Those are real gains, and none of them runs the change. The step that verifies behavior, integration, and end-to-end tests against a live system still funnels through a shared staging environment or waits on a full copy of the stack, which takes too long and costs too much to stand up for each change.

That step has a hard ceiling. Staging is one environment per organization, so it functions as a queue. Full-stack duplicates are expensive enough that teams ration them. Neither doubles in four months because you approved a budget. Generation now scales like GitHub. Verification still moves one change at a time.

Workflow diagrams comparing a single shared queue vs one environment per change.

Verification was sized for human pace, and agents broke the sizing

None of this is new. Large engineering organizations were complaining about staging contention and review backlogs years before coding agents existed. The apparatus was built when code arrived at the pace humans type, and at that pace its costs were a tax teams could manage: an occasional staging conflict, a review queue that cleared by Friday. Deferring expensive full-fidelity testing to the end of the pipeline was a reasonable trade while changes were scarce.

Agents do not introduce the bottleneck. They multiply it past the point where the old coping strategies work, and they bring it to smaller teams. Throughput that used to strain a 500-engineer platform now appears on a 50-engineer team running agents in parallel. The assumption under all of those design decisions, that code is the scarce input, is gone, and the checking machinery built on it has not moved. The two curves that used to track each other roughly have come apart. GitHub’s chart is the aggregate picture of that separation.

Chart showing monthly commits vs verification capacity

The gap between the curves fills with unverified merges

Teams respond to the widening gap in a few ways. The first is to make review faster. AI code review tools sit in front of human reviewers, catch real defects in the diff, and keep improving. They also have a limit: A reviewer, human or model, is reasoning about code they have not run, and a diff cannot tell you whether the change holds up against its real dependencies.

The second is to throttle the agents, capping the amount of generated work that enters the pipeline. That protects the verification queue by returning most of the throughput the agents were designed to deliver.

The third is to merge anyway and absorb downstream failures. Changes that were never run against a real system land in main at the rate commits arrive, and the cost surfaces later as broken staging environments and lengthy debugging sessions. Or worse, production incidents and rollbacks.

The August 17 outage is a preview of how that ends. The root cause was not a bad change. It was a component that everyone depended on, and nobody had scaled, failing on the day traffic finally exceeded its design. In most delivery systems, verification is that component.

Verification has to run on the same curve as generation

The structural fix is to make verification match the shape of generation: parallel and per-change. For a single application, this is nearly solved: an agent runs the app on a laptop or in a CI container and checks whether the change works. The hard case is the cloud-native one, where the change’s behavior only exists when interacting with other services, databases, and queues. Both familiar options fail at agent volume: A shared staging environment serializes everything into a queue, and duplicating the full stack for each change is too costly and too time-consuming.

There is a third shape. Keep one shared environment running the stable version of every service, and for each change, deploy only the services that the change touched. Test traffic carries the change’s routing key, so at each hop, a request for that change reaches the changed version while every other request flows through the stable one. Each change gets isolation where it matters, at the services it modified, and shares everything else: the same cluster, the same data, the same downstream dependencies.

Workflow diagram showing request routing in the shared environment

That shape changes the economics and the actor. Verifying one more change costs one extra deployment, not another copy of the stack, so hundreds of changes can be checked in parallel on the cluster you already run.

Because an environment appears in seconds, an agent can use one inside its loop: open a change, run functional checks against real upstream and downstream services, read the failures, and iterate until they pass. The pull request that reaches a human arrives already exercised against the real system, and review time goes to intent and design.

Match the curves or lose the generation gains

GitHub’s 2.9 billion commits quantify a shift that every engineering organization is living through on a smaller scale. Generation is now effectively free and unlimited, so it is no longer the source of advantage.

“The teams pulling ahead are not the ones producing the most commits. They are the ones whose verification capacity rises with their generation capacity.”

The teams pulling ahead are not the ones producing the most commits. They are the ones whose verification capacity rises with their generation capacity, so more generated code becomes more shipped code, rather than a longer review queue, a deeper staging backlog, and a bigger incident bill. Ask what happens to your pipeline when commit volume doubles in four months, because that is no longer a hypothetical. It is the gap we work on at Signadot, and it is worth closing before the curve doubles again.

The post Commits on GitHub have doubled in four months. Verification capacity has not. appeared first on The New Stack.

❌