❌

Vue normale

Reçu avant avant-hierThe New Stack

One engineer shipped 2,000 PRs a month to production. Verification is the key.

19 septembre 2026 à 16:00
Abstract dark digital wireframe mesh with chromatic glitch effects, representing AI agent verification and virtualized software environments.

Lauren Tan, an engineer on the Grok team at SpaceXAI, who previously worked at Cursor and Meta, recently published a guide to her personal agent workflow: pstack. The attention-grabbing number is that pstack has let her ship 2,000 pull requests (PRs) a month to production with high confidence. That’s one engineer shipping nearly 100 PRs per working day.

The number is incredible and undeniably an outlier, but the direction is not a surprise. I have argued previously that coding agents would enable teams to generate ten times the code with similar headcounts. What is surprising is that these are not just code output numbers. These are actual changes landing in production.

According to Tan, the most critical piece of that workflow is verification. A verification skill lets an agent check its own work and keep going until the task is done, and she treats it as “critical infrastructure” rather than one skill among many. 

The verification skill rests on something underneath it: a rich runtime the agent can drive, inspect, and get structured answers from. For a single application, that runtime is the application itself, started on demand. For a system made of dozens or hundreds of services, no such runtime exists by default, and providing one that keeps up with hundreds of parallel agents is the hard part.

Verification is the whole game, and the math says so

Her argument for agentic verification is a throughput argument. An agent that can check its own output keeps working until the task is done. An agent that can’t hand you a diff and wait makes you the slowest component in the loop. That is why she claims strong verification skills can multiply a team’s output by 100 to 1,000 times.

“An agent that can check its own output keeps working until the task is done. An agent that can’t hand you a diff and wait makes you the slowest component in the loop.”

At 2,000 pull requests a month, reviewing every change by hand would allow about five minutes per PR across a full working month. Human review cannot be the verification layer at that volume. Whatever does the checking has to run without a person in the loop, and it has to run in parallel with the agents generating the work.

The model assumes the agent can run the whole application

The verification skill she describes generates a command line interface (CLI) and a feature map for the application. The CLI lets an agent start the app, navigate it, inspect state, and read structured JSON results back. Each agent gets a complete copy of the application and can test a change end to end.

She is direct about how much rests on that runtime: “I personally feel that agentic verification is so important that I would unironically suggest building your own rich debugging tools, or even choosing a different tech stack, in order to have unfair advantages and extreme productivity in building software.”

Her approach to providing a runtime for her agent works because the application fits in one process. A frontend, a compiler, or a single service with a database can start from a CLI in seconds and be thrown away afterward.

“I would unironically suggest building your own rich debugging tools, or even choosing a different tech stack, in order to have unfair advantages and extreme productivity in building software.”

For teams building complex distributed applications, their system does not have that property. The application is the interaction between an order service, a payments service, an inventory service, a queue, several databases, and a handful of third-party APIs. At larger shops, the count runs into the thousands. A pull request to one service is only verified by exercising the calls it makes and receives. The CLI can start the changed service. It cannot start the system.

None of the existing runtimes survive hundreds of parallel agents

Local runtimes with mocks are cheap and can run fully parallel using worktrees or CDEs. Their problem is fidelity. Mocks encode what a dependency did the last time someone looked, and they drift the moment the real service changes. An agent that verifies against mocks closes its loop against fiction, and the failure shows up after merge.

A full copy of the stack per change is faithful and isolated. But its cost scales with the number of services times the number of concurrent changes, and at hundreds of agents that cost is untenable. Time is the bigger problem. A full stack takes minutes to provision, and the loop she describes has the agent testing every iteration of a change while it is still working on it. An environment that is ready after the agent has moved on to its next attempt is no use to it.

Shared staging is faithful and cheap because there is one of it, and that is the whole problem. A single mutable environment cannot host hundreds of concurrent changes. Agents overwrite each other’s deployments, a broken change from one agent becomes failed tests for every other agent, and the loop-closing property that makes the workflow valuable disappears.

Each existing verification runtime plotted on a chart showing its realism of dependencies against concurrency.

What verification needs when the callers are agents

Read Lauren’s workflow as a requirements document, and five properties emerge:

  • The change has to run against real dependencies or the verification means nothing.
  • Hundreds of concurrent changes have to be unable to see each other.
  • The cost of an environment has to scale with the size of the change, not the size of the system.
  • Environments have to come up in seconds, because an agent waiting on provisioning is parallelism you paid for but didn’t use.
  • All of it has to be reachable through the CLI or MCP server the agent already uses, because the caller is an agent.

The first and third requirements pull in opposite directions. Realism pushes toward complete copies of the system. Cost pushes toward sharing as much as possible. Shared staging resolves that tension by giving up isolation, and a per-change full stack resolves it by giving up cost efficiency. A design that satisfies all five has to share and isolate at the same time.

Virtualized full-stack environments share the system and isolate the change

The architecture that does this treats an environment as a view of a running system rather than a copy. One shared set of stable services runs continuously, deployed from the main branch and kept healthy the way production is. When an agent needs to verify a change, it runs only the service it modified, on its own machine or as a lightweight deployment in the cluster, and joins it to the shared stack as a new isolated environment.

“The architecture that does this treats an environment as a view of a running system rather than a copy.”

From inside that environment, the changed service is the version of record, and every other call falls through to the shared stable versions. The agent sees a complete, realistic system, and so do the other hundred agents, each seeing a system that differs from the baseline by only the delta of its own change. Requests carry their environment identity as they cross service boundaries, which keeps one agent’s traffic from reaching another agent’s version under test. Stateful side effects that cannot be shared safely, like queue topics or writable databases, get a per-environment copy where needed.

Diagram showing "agent 1 env" and "agent 2 env" interacting with the shared cluster.

The cost model follows directly. An environment costs one or two running services instead of sixty; it is ready in the time a single service takes to start, and you can create and destroy it from inside the agent’s own loop. This is the pattern Signadot packages for Kubernetes, with the shared stable stack running in the team’s existing cluster.

The loop, end to end, with an agent as the actor

Put the two halves together, and the workflow that enables her to ship 2,000 PRs a month to production carries over to a distributed system almost unchanged. An agent picks up a task and changes one service. It asks for an environment for that change and gets one in the time it takes its service to start. It then drives real requests through the system’s entry point and watches them traverse the real dependency graph, with only its own service running new code. It reads structured results, fixes what failed, and runs again. When the checks pass, it opens the PR, and the environment goes away at merge.

“Environments stop being something the platform team hands out and become something agents create, use, and discard as needed.”

For the platform team, the unit of work changes. Today it provisions environments, whether that means keeping a shared staging alive or stamping out full copies of it. In this model, it runs one shared stable stack and the layer that virtualizes it: context propagation across every service, isolation for the stateful dependencies that cannot be shared, and the tooling that creates and tears down environments. Environments stop being something the platform team hands out and become something agents create, use, and discard as needed.

Verification capacity is the new ceiling on throughput

Lauren Tan’s post is not a story about one unusually productive engineer. It shows what happens when agents run the full loop, writing a change, verifying it, and iterating without a person in between. The verification infrastructure is the foundation that the entire loop stands on.

In distributed applications, that infrastructure has to be a runtime environment that gives every agent real dependencies, keeps hundreds of concurrent changes from seeing each other, costs a change rather than a copy of the system, and is ready in the seconds an agent is willing to wait. That is what turns agent parallelism into shipped code rather than a longer review queue. That model of runtime environments is exactly what we built Signadot to enable.

The post One engineer shipped 2,000 PRs a month to production. Verification is the key. appeared first on The New Stack.

AI’s best coding agent fails 60% of the time — and the data backs it up

15 septembre 2026 à 00:22
abstract screen

Claude Fable 5.1 just won a new coding benchmark despite failing more than six out of 10 times. Its 38.8% score comes from Real-SWE, a benchmark from Y Combinator-backed Specific Labs that takes a different approach to testing coding agents. Instead of giving them problems pulled from public repositories, it drops them into private codebases from real companies and asks them to tackle problems similar to those engineers deal with every day.

The code and its solutions aren’t publicly available, which makes it even less likely that they showed up in a model’s training data.

Specific Labs can’t guarantee that a model has never encountered any of the code, but the company says using private code makes that much less likely. It also estimates that 99% of tokens in real-world enterprises are hidden from frontier models. Once the agents were dropped into unfamiliar territory, the scores fell fast.

The code and its solutions aren’t publicly available, which makes it even less likely that they showed up in a model’s training data.

Private code changes the test

Fable 5.1, running via Claude Code, led the pack at 38.8%. GPT-6 Astra on Codex CLI followed at 33.8%, with Gemini 3.8 Flash on Gemini CLI at 31.2%.

After that, the scores dropped significantly. GLM 5.3 scored 28.8%, Grok 4.6 and Muse Spark 1.3 tied at 23.8%, Kimi K3 hit 18.8%, and GPT-5.6 Sol finished at 16.2%.

Each model got eight tries at every task. Real-SWE also tested each model with its own coding tool — Fable 5.1 with Claude Code, Astra with Codex CLI, and Gemini 3.8 Flash with Gemini CLI — so the scores reflect the full setup (not just the model).

As GPT-6 Astra’s ARC-AGI score showed, changing the scaffolding around a model can change its performance. Fable 5.1 in Cursor, for example, could produce a very different result.

Six tasks stumped everyone

Real-SWE uses proprietary code licensed from real businesses, including a consumer product with more than 200,000 users and a fintech platform that has processed more than 100,000 bank statements. The work also spreads across the codebase, with Real-SWE solutions touching a median of 11 files, nearly double the six-file median in benchmarks like FrontierCode and DeepSWE.

On individual tasks, the scores fell even further, with six of the 10 posting success rates below 15%.

On individual tasks, the scores fell even further, with six of the 10 posting success rates below 15%. A billing schedule migration had a 14.1% fix rate, API token metering landed at 12.5%, S3 storage tracking hit 10.9% and a linearizable scan came in at 4.7%, while a tax jurisdiction bug was patched just 3.1% of the time.

Not a single model solved the analytics stream reducer across 64 attempts. Astra and Gemini, meanwhile, went eight for eight on a multi-region sweep and Fable solved seven of eight, yet all three failed every attempt at the linearizable scan. No agent was consistently reliable across the benchmark.

Not a single model solved the analytics stream reducer across 64 attempts.

Where the agents broke down

Fable 5.1 most often missed requirements (36.7%) or ran into integration errors (34.7%). Astra’s failures were split between integration errors and unverified assumptions, both at 34%.

Integration errors appeared in nearly half of Gemini 3.8 Flash’s failed runs, while GPT-5.6 Sol made unverified assumptions in 43.3% of its failures.

What 38.8% really means

Real-SWE doesn’t prove that public coding benchmarks are inflated by data contamination, and 10 tasks is still a small sample.

But the top-performing agent still failed more than 60% of the time on private code it likely hadn’t seen before, suggesting that solving a coding problem is very different from finding your way through an unfamiliar production codebase.

The post AI’s best coding agent fails 60% of the time — and the data backs it up appeared first on The New Stack.

“Valuable warning shots”: How Anthropic now views Claude’s cyber incidents

10 septembre 2026 à 21:54

This week, Anthropic acknowledged that the three cyber incidents it disclosed this summer weren’t just the result of a misconfigured test environment. It turns out that Claude’s own behavior was part of the problem. 

Recall in July when the AI company released a report on three cases where Claude models reached the open internet from misconfigured test environments and compromised real third-party systems — a telling example of the limits of AI safety tests and why AI evaluation infrastructure needs production-grade security.

Upon deeper review, Anthropic now says that explanation was incomplete. Beyond the operational failure, the models themselves displayed two recurring alignment failures: biased reasoning and recklessness. And, actually, there was a fourth incident that the AI company only found out about after further investigation. 

It’s awkward timing, given one of Anthropic’s pretraining researchers, Jacob Coxon, resigned this week due to concerns about superintelligence that “could kill us all,” taking to X to write: 

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.

— Jacob Coxon (@hilbertspaess) September 9, 2026

“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately.”

The incidents definitely don’t prove this kind of risk, but they do illuminate how models can behave unexpectedly — and how existing evaluations can fail to catch it.

The problem was more than an operational failure

This summer, when Anthropic shared its initial findings, it said the three cyber incidents were “closer to a harness and operational failure than a model alignment failure,” with Claude seeming to believe it was operating in a simulation without internet access. 

But a closer look at the transcripts revealed that the problem was not simply the model reaching the internet by mistake; Claude also showed signs of misaligned reasoning. 

“When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm.”

After widening the initial search to about 481 million transcripts, including those from the company’s Frontier Red Team, as well as non-cyber evaluations, reinforcement learning (RL) environments, subagent logs, and others, Anthropic conducted two scans. 

First, it scanned these transcripts for signs of internet access, whittling down the pool to 9.2 million flagged for escalation. It then used Claude to review this smaller batch, re-identify the known incidents, and look for others of similar or worse severity. 

What did it find? It found more than an operational failure. Anthropic says its investigation identified “two recurring alignment issues, present at varying levels of severity across the incidents.”

It names these issues as biased reasoning, where Claude disregarded or misinterpreted evidence that was, in fact, on the internet, and recklessness, where Claude was willing to take harmful actions to pursue a task.

Anthropic was already concerned that Claude Mythos 5 uploaded a malicious package to PyPI, but the model repeatedly said it believed it was operating in a simulation. Anthropic now says the simulation explanation does not fully account for the model’s behavior and that its reasoning appears to have been biased.

“When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm.”

And there was a fourth incident Anthropic missed the first time

Anthropic’s new alignment assessment also contains another less than reassuring revelation: It missed one of the four incidents the first time around. 

As the AI company explains, its first review relied on an agentic search due to “the volume of transcripts and [its] desire to disclose incidents quickly.” But that scan missed a set of transcripts that had also had internet access. 

How did the cat get let out of the bag? It happened in August when Anthropic was assembling transcripts to share with METR (Model Evaluation and Threat Research). This research nonprofit measures whether and when AI systems might threaten catastrophic harm to society. The fourth incident dates from January 2026 and involves an early version of Claude Opus 4.6.

Discovering the fourth incident pushed Anthropic to broaden its search to the roughly 481 million transcripts described above, giving it more information to analyze and ultimately surfacing new insights into Claude’s “biased reasoning” and “recklessness.”

While this broader search didn’t turn up any more incidents of what Anthropic describes as “similar or worse severity,” the investigation isn’t over. Per the AI company, it’s signed an agreement with METR, granting the organization wide-ranging access to transcripts and Anthropic employees for eight weeks to investigate the incidents independently. The agreement can be extended if both parties agree.

Anthropic reproduced the failure to try to understand if newer models would do the same

Beyond reviewing the original incidents, Anthropic is taking other steps to try to understand how things went pear-shaped in the first place. It put Claude Opus 5 and Claude Mythos 5.1 in a simulated replication of the Mythos 5 scenario — the one where the model uploaded a malicious package to PyPI. Anthropic reports:

“Both take harmful actions substantially less often than Claude Mythos 5 did, but they still engage in the same behaviors at concerning rates.”

It notes that both models only took harmful actions after stating that they were, indeed, acting in a simulated environment. Anthropic also says that Mythos 5 appears to be “an outlier in terms of poor performance” on evaluations for biased reasoning, recklessness, and willingness to probe boundaries.

“Both take harmful actions substantially less often than Claude Mythos 5 did, but they still engage in the same behaviors at concerning rates.”

Still, it’s not such a clear indication of how the newer models would behave in the real world. As the AI company adds, “because the reproduction is, in fact, a simulation, we are cautious about how far these results generalize to the real world.”

Additionally, Anthropic looked back into its training to search for a root cause of Claude Mythos 5’s biased reasoning, but came up short. While it didn’t identify a single root cause, the company says biased reasoning decreased across its production models over time, potentially due to updated RL and alignment training.

Anthropic admits it still has work to do. Since its pre-release auditing didn’t catch misalignment of this severity, it’s since added evaluations to specifically target these behaviors, though reliably doing so is still a challenge.

Anthropic admits the failures were serious, but says they’re “narrow”

While Anthropic clearly calls out the four cyber incidents as “valuable warning shots,” it also points out that their scope was limited: 

“While Claude’s actions were misaligned, they remained within a narrow scope,” the AI company says, pointing to the fact that all incidents included a single Claude instance and the model never attempted to coordinate with other agents or conceal evidence of its actions. 

It says these behaviors playing out in ordinary use are “unlikely” and that, if they did, the safeguards shipped with production models would add more defenses that didn’t exist in these evaluations.

But following Coxon’s remarks about the risks of superintelligence, and Anthropic’s own alignment science lead, Evan Hubinger, responding that Anthropic “really do[es] earnestly believe AI could kill all humans,” seeing Claude go off the rails isn’t comforting.

The post “Valuable warning shots”: How Anthropic now views Claude’s cyber incidents appeared first on The New Stack.

AI broke code review. Two experts disagree on what replaces it.

8 septembre 2026 à 17:35
Detective holding up a magnifying lens in front of eye

Ask two experienced engineers about how to handle the flood of AI-generated code in their review queues, and you’ll get two different answers.

The debate remains very much unsettled. And on Tuesday, September 29, two industry leaders will join a live event to hash out what to do.

John Bristowe, Principal Developer Advocate at Octopus Deploy, will join Viktor Farcic, the platform engineering voice behind DevOps Toolkit, for the live conversation we’re calling “Human Review vs. Verified Pipelines: What Catches Bugs in the Age of AI Code.”

REGISTER NOW FOR THIS WEBINAR
By registering, you consent to The New Stack’s Privacy Policy, Terms of Use and to receiving email communication from The New Stack and our event partner. You may opt out at any time.

Here are the facts: Developers have adopted AI en masse. According to the 2026 DORA report, 90% of developers now use AI at work. The result? Developers are merging 98% more pull requests than they managed in the pre-AI era. 

But all that AI-generated code is leaving a mess. Bugs per developer are up 54%, and one analysis of 10,000 developers found that incidents per pull request have climbed a staggering 243%. Octopus Deploy’s own AI Pulse report found that while AI usage enables faster code creation, it can “degrade overall performance” because coding agents write large code updates that humans struggle to fully understand.

Part of the problem is that developers have adopted automated code generation faster than they have adopted automated code review, effectively moving the human bottleneck further down the software creation chain without removing it entirely. And AI code review may have the same shortcomings as the coding agents.

Bristowe argues that code review has quietly become little more than theater. No human reviewer can quickly audit a 40,000-line, agent-created pull request, since they were not part of the reasoning that produced it and cannot realistically understand everything it may change. 

What does Bristowe recommend? Moving the quality gate off the humans’ desks and into the delivery pipeline itself. Does that mean more AI? Not necessarily, with the developer advocate arguing that building robust “policy-as-code” rules into deployment standards can flag only what goes against those policies. Humans can handle those exceptions, without pretending they are “reviewing” the entire package.

Expect Farcic to press Bristowe on how well a policy-as-code setup can truly absorb judgment, and whether we’re simply creating another accountability sink in software development. The conversation will also explore the plight of the junior engineer, who can no longer expect to join a team of humans writing code that other humans review and discuss.

The debate kicks off at 2:30 p.m. Eastern/11:30 a.m. Pacific on Tuesday, September 29. It’s free to attend, and attendees will receive a companion resource built from Octopus Deploy’s AI Pulse data, available immediately for participants who show up live. Register today.

What you’ll take away:

  • Why AI-generated code broke the assumptions code review was built on, and why more review isn’t the fix
  • Why using AI to review AI’s own code doesn’t close the gap (same training data, same blind spots)
  • How to build a pipeline that verifies every deployment against a defined set of rules, no matter who or what wrote the code
  • Where code review still earns its keep, and where it needs to step aside for the pipeline

The post AI broke code review. Two experts disagree on what replaces it. appeared first on The New Stack.

AI agent evaluations are part of the product

4 septembre 2026 à 16:00
Abstract dark blue digital geometric mesh representing AI agent evaluation frameworks and execution paths

A team builds an agent, gives it a few representative questions in a test chat, and watches it produce useful answers. Someone tries a slightly harder prompt, and that works too. The team records a demo, approves the change, and ships it.

Then the retrieval configuration changes. A model upgrade follows a few weeks later. The agent still answers the original questions, but now it skips a required citation on one task and selects an unintended customer lookup tool on another. The issue may not become visible until user feedback or monitoring surfaces it. 

A good demonstration tells you that an agent worked once, under the conditions you happened to give it. It doesn’t fully establish whether the next version will consistently preserve the behavior your users and operators need. For that, evaluation has to become part of the delivery process.

“If it can’t reproduce a run or a material regression in a high-risk workflow, the product isn’t ready to pass the release gate.”

A repeatable evaluation system runs fixed scenarios along the product’s execution path and records sufficient evidence to determine whether a release should proceed. It should exercise the code that assembles context, the tools the agent can call, and the permissions the runtime enforces. If it can’t reproduce a run or a material regression in a high-risk workflow, the product isn’t ready to pass the release gate.

Define correct behavior before writing tests

“The answer was good” isn’t a requirement anyone can test twice. Before choosing an evaluation tool, write down the jobs the agent performs, the limits around each job, and the outcomes that fall outside the product’s accepted operating boundaries.

For a support agent, a useful job might be to answer a billing question using records from the correct account and cite the policy currently in force. Its limits may forbid changing a plan or exposing another customer’s data. When a policy can’t be found, the agent should acknowledge the gap and avoid presenting an unsupported answer as fact. The agent may also need to request an account number before continuing, or route an exception to someone with the appropriate access.

Separate the result from the process that produced it. An agent can give the right answer after retrieving the wrong document, or complete a task after calling an unnecessary tool or searching outside the customer’s scope. It can even escalate a routine request it should have handled on its own. Those runs may look successful in a transcript while masking weaknesses that may appear under different requests.

Start with a few observable requirements for each job. Required facts must be supported by named sources, and writes must wait for confirmation. When data is missing, the agent should ask rather than guess. High-risk rules get exact assertions; the wording around them can tolerate some variation.

Build scenarios from real user work

The first test set should be small enough for someone to maintain. Ten real tasks are more valuable than a large benchmark filled with prompts your users never send.

“Ten real tasks are more valuable than a large benchmark filled with prompts your users never send.”

Support tickets and workflow logs are good raw material. So are incident reports and conversations with users. Include the ordinary requests that make up most of the workload, then add cases with unclear instructions or missing account data. Test what happens when a document is outdated or a tool times out. Some scenarios should require approval before the agent can act, and a few should cover unusual but still perfectly valid requests.

Agents operate across turns, so some scenarios should too. Ask for an account change, provide the missing identifier in the next message, and confirm the proposed change in a third. The test should verify that the agent carries the account identifier and proposed change across turns without dragging unrelated details into the final action.

Each scenario also needs fixtures. Freeze the documents and tool responses used during the run, and pin the account state to a known snapshot. The policy version and the agent’s permissions matter just as much. A failure you cannot reproduce becomes a debate about what the agent may have seen. Fixed fixtures turn it into an engineering problem.

Production failures should be treated as permanent regression cases. Over time, the suite records the mistakes the team has already paid for and learned from.

Test the entire execution path

Final-answer scoring misses much of what distinguishes an agent from a chatbot. An agent retrieves data and chooses which tools to call. It supplies the arguments, reads the results, and then decides whether to continue. Any step in that loop can diverge from the intended path even when the response appears convincing.

Capture the request and system instructions. Record the exact model and application build. Version the prompt and retrieval configuration, including the tool schemas. Then record every retrieved source with its version, as well as every tool call, its arguments, and its result. Permission checks and the final response belong in the trace too, along with latency, token usage, and cost. The trace should answer practical questions without requiring someone to reconstruct the run from unrelated logs.

“Final-answer scoring misses much of what distinguishes an agent from a chatbot. Any step in that loop can diverge from the intended path even when the response appears convincing.”

When combined with server-side enforcement and audit records, the trace should show that a search remained within the correct tenant and customer account. It should also identify the approved policy source and the records cited in the answer. For a write, it should show that the user confirmed the change and that the server-side permission check passed. Those are deterministic checks: they either happened or they didn’t.

Clarity and usefulness are less deterministic. A human reviewer or model-based evaluator can score whether the response answered the request, explained a limitation, or asked a sensible follow-up question. Keep those judgments attached to the trace. When a score drops, the team should be able to find the step that changed.

This also makes evaluator failures easier to spot. A model-based evaluator may produce different judgments after an upgrade or respond differently to a revised rubric. Save the evaluator’s version and instructions with its result. Regularly compare a sample of those scores with human reviews.

Side-by-side workflows of final-answer scoring and execution-path trace
Both columns end in a convincing answer. Only one of them can tell you whether the agent was correctly grounded.

Keep fixed rules separate from variable scores

Agent quality doesn’t fit into one unexplained number. Track task completion and factual support separately from retrieval quality. Keep policy compliance distinct from user experience, latency, and cost.

Some of those signals have tolerances: a response that takes 200 milliseconds longer may still be acceptable, and a slightly longer answer may even be clearer. Others allow no failures. An unapproved update, cross-tenant retrieval, or missing approval should remain a release-blocking condition regardless of other scores.

Compare a candidate against a known baseline on the same scenarios and fixtures. Show the reviewer the changed answers and the records behind them, then let the tool paths and individual scores explain why. If the new version completes more tasks but doubles latency, that may be a reasonable product decision. If it improves the average score while bypassing one permission gate, it is not.

Repeat scenarios when behavior is variable. A task that succeeds inconsistently—for example, once in ten attempts—does not yet meet a reliable release threshold. Set thresholds based on risk, and reserve absolute gates for rules the system must obey every time.

None of this requires building your own tooling from scratch. Tools such as Promptfoo, DeepEval, LangSmith, and Braintrust provide capabilities for building evaluation workflows. Depending on the tool, that support can include running scenarios and capturing traces. Some also use models to grade the result. 

The metrics vocabulary is worth learning too. Conceptually, pass@k asks whether at least one of k attempts succeeds, while pass^k asks whether all k attempts do. Pass^k is useful when consistent behavior matters, but it doesn’t replace exact gates for rules an agent must obey.

Evaluation also has a cost. Every live, end-to-end run that calls a model spends tokens. Judge models cost more than string checks, and a large suite on every commit adds up quickly. Save the expensive judgments for the scenarios that carry real risk.

Make evaluation a release gate

Run the suite whenever the team changes a model or a prompt. A new retrieval configuration counts, as does any change to a memory policy or a tool interface. Use a fast set for ordinary changes and a broader set before a major release or model migration. When a behavior change is intentional, require a reviewer to approve the new expectation rather than rewrite the test.

The records behind this process need the same controls as the agent itself because evaluation inputs may contain customer data. Traces can include retrieved text and internal instructions. They may also capture tool arguments containing credentials or personal data. Version the records and scope access carefully. Consider redacting sensitive values before persistence and applying appropriate encryption, access controls, and retention policies based on the data involved. 

Keeping more of that work near the operational data can shorten the path. Oracle AI Vector Search stores vector embeddings alongside business data, and SQL queries can combine similarity search with relational filters and lexical search. A team using Oracle AI Database can keep operational records and their vectors in a data platform it already controls.

The same platform can hold evaluation traces and enforce access rules. Database-enforced access controls can apply row- and column-level policies within the database, providing another layer for enforcing data-access boundaries. The exact platform matters less than the invariant: the team needs durable evaluation cases, the inputs used for each run, and a record of why each build passed.

A release gate should prevent releases with known high-risk conditions while recognizing that some judgments require context. Exact checks block permission and policy regressions. Thresholds catch measurable quality drops, and human review handles the ambiguous changes scores cannot settle.

Diagram showing the workflow of a release gate that runs on every model, prompt, retrieval, or tool change.
Hard rules block on any failure; everything else gets a tolerance or a reviewer.

Start with ten cases and keep every important failure

Pick ten real tasks this week. Record the expected result and the evidence the agent should use. Note the actions it must not take. Freeze the fixtures, capture the trace, and run the cases before the next release.

“Trust in an agent grows when the team can replay what happened and show that the next release still respects the boundaries users depend on.”

When an incident happens, add it to the suite. When a user finds a failure that nobody predicted, keep it. The suite will grow with the product.

Trust in an agent grows when the team can replay what happened and show that the next release still respects the boundaries users depend on. The evaluation system is part of the product that ships.

Need help with agent evaluations? Working examples of these patterns in Oracle AI Database, including agentic RAG patterns with hybrid search, are available in Oracle’s AI Developer Hub.

The post AI agent evaluations are part of the product appeared first on The New Stack.

AI Agents built a 3D city for $33 in two hours —and exposed a major flaw

3 septembre 2026 à 19:34
abstract city

PhiloLabs wanted to see how far a group of AI coding agents could get building something where working code wasn’t enough. In a recently published experiment, the company set Claude Fable 5.1 agents loose on a 3D reconstruction of San Francisco’s Union Square, built from real-world geographic data and reference images.

Two hours later, the agents had built a working Three.js version of Union Square in the browser. The experiment included 453 building footprints, 75 custom façades and 129 named storefronts, along with 220 pedestrians and 109 vehicles moving through the scene, including Powell Street’s cable cars.

But getting the application to run was only one part of the experiment. PhiloLabs also wanted the agents to catch visual problems that conventional tests would miss, so it put Playwright into the development loop.

The entire run used roughly 8 million tokens and cost about $33 in API calls.

The entire run used roughly 8 million tokens and cost about $33 in API calls.

Playwright as agent vision

PhiloLabs divided the reconstruction among subagents that handled geographic research, building geometry, textures, storefronts, and other parts of the scene. Once their work was running in the browser, Playwright moved through 34 predetermined camera positions and captured screenshots that could be checked against photographs of the real Union Square.

In all, the agents produced 147 comparison sheets, making it easier to spot things that were technically correct but still looked wrong. For instance, a building might be in the right place but have the wrong proportions, or a storefront might end up on the wrong side of the street. Using the same camera positions each time also made it easier to see what changed from one pass to the next.

In all, the agents produced 147 comparison sheets, making it easier to spot things that were technically correct but still looked wrong.

Agents reviewing agents

PhiloLabs then had specialist agents examine the material, with some focused on architecture and geography and others on technical art and interactions. Together, they produced nine reports on the Union Square build.

Those reports became a punch list for the next pass. The development agents could address the reviewers’ findings and rerun the scene.

That division is useful because not every mistake translates neatly into a test. You can check whether a building was placed at the right coordinates. It’s much harder to write a test that tells you whether the street actually looks like Union Square.

Filling gaps in source data

The agents weren’t starting with a finished 3D model they could simply recreate. They had to pull together open geographic data (primarily OpenStreetMap and USGS elevation data) along with information about the real location, then turn all of it into geometry, façades, and objects that would run in a browser.

There were still plenty of gaps to fill because geographic data could tell the agents where a building belonged without showing what its façade looked like, while photographs only captured the parts of the building visible from a particular angle, leaving the agents to make their own calls when neither source provided an answer.

Another agent checking the work doesn’t guarantee that those decisions are right, especially when the source material is incomplete to begin with, because the reviewer can miss the same thing the first agent did.

What $33 buys

Eight million tokens is a lot of model activity for a single application. Yet the reported API cost for the Union Square run was about $33.

PhiloLabs split the job among subagents, ran tasks in parallel, and reused cached context rather than having a single agent repeatedly work through the entire project from scratch.

Still, Union Square was a fairly contained experiment. Screenshots are not as useful once agents move into complicated applications. Spline recently rebuilt its 3D editor using Claude Code agents, but the finished interface only shows part of what those agents built. Problems buried in the code or triggered by the way people actually use the editor may never show up in a screenshot.

Eight million tokens is a lot of model activity for a single application. Yet the reported API cost for the Union Square run was about $33.

The post AI Agents built a 3D city for $33 in two hours —and exposed a major flaw appeared first on The New Stack.

Vercel built a feedback loop that treats agent instructions like software

2 septembre 2026 à 17:29
Red and pink dots sweep in looping waves across a black background, forming a flowing abstract pattern.

Vercel ran more than 200 agent runs to build design.md, a new public prompt file designed to help agents create web pages that look and feel like Vercel, even when they don’t have access to the company’s internal codebase.

In a post published Monday, Vercel shared the behind-the-scenes process of how it built and evaluated the file to determine whether the corrections it encoded to prevent failures worked — the answer is yes, but not perfectly. 

The release offers a broader lesson for developers: Encoding human judgment into reusable agent guidance can help reduce recurring failures, but it’s no silver bullet. 

In three desktop scenarios, when Codex with GPT-5.5 generated the page once with design.md loaded and once without, Vercel’s deterministic checks counted 39 instances of known failure modes with design.md, compared to 91 without it — a 57% reduction in this six-page test.

Encoding human judgment into reusable agent guidance can help reduce recurring failures, but it’s not a silver bullet. 

While Vercel acknowledged that every one of the six pages (an admittedly small sample size) had a failure large enough to prevent shipping, the experiment suggests that failures explicitly named and encoded are less likely to recur. 

The problem with keeping design knowledge in the codebase

In June, Vercel shared the thinking behind product design, a skill that teaches coding agents working in its codebase to design pages with the brand’s look and feel. Vercel says it’s proven valuable — but only for agents working inside the codebase. Once it’s time to get tools that live elsewhere to produce the same on-brand content, they can’t reach the same design context. 

So the company decided to build a public file that any agent or tool, even outside Vercel’s development environment, can load to access the relevant design knowledge and produce on-brand pages. 

First, Vercel tried to port product design to a public prompt simply, but that proved a bust. Because the prompt included subjective design language, each model interpreted it differently. Plus, the prompt only tells half the story; important information about design and implementation lives in the codebase, which works for product design, not for a public prompt. 

Making that design knowledge accessible, then, Vercel decided, would take a new file, built from the ground up. 

How Vercel built and tested design.md

In building that file, Vercel tested every iteration against a repeatable set of seven evaluation prompts, designed to expose how changes in guidance ultimately affected the generated output. With every round, these prompts allowed Vercel to measure two things: 1) what changes the file caused; 2) how different agents interpreted it. 

As design.md cycled through different tests and iterations, ultimately covering more than 200 agent runs, Vercel settled on a three-part system to make the guidance both reusable and testable. 

First, the prompt file itself gives agents guidance on how to make design decisions, including writing copy, composing hierarchy, typography, and color, as well as publishing. Importantly, it also spells out what design patterns are not allowed so agents can avoid them. 

A public stylesheet, meanwhile, defines reusable implementation details so agents don’t go rogue on things like spacing and layout. Finally, an evaluation loop turns human feedback into updated guidance and deterministic checks for mechanical failures. 

Vercel then built a local app to serve as an eval harness for reviewing the generated pages from each round and for storing the prompt, inputs, model configuration, file version, screenshots, and reviewer feedback. Corrections were routed to each layer accordingly: judgment changes were encoded as prose directly in the file; the stylesheet captured reusable mechanics; mechanical failures became deterministic checks in code. 

With each round of feedback, Vercel reviewed the results, encoded the accepted corrections, and reran the same eval prompts to ensure no changes inadvertently broke anything else. 

How ongoing feedback becomes updates

To keep the shipped file up to date, Vercel then built design-agent. With just a mention in a Slack thread, the agent loads the current design.md, builds the requested website page using the published stylesheet, and then posts a full screenshot and URL back in Slack, giving Vercel a compact record that connects the request, output, and subsequent feedback. 

At the end of the week, all that feedback is consolidated with comments from GitHub reviews and Figma, and each repeated complaint automatically becomes a proposed change for human review and approval. 

Better guidance still doesn’t guarantee agent reliability.

Over time, Vercel says it tracks how often each complaint recurs. Ideally, once a fix is implemented, that count should start to decline; if it doesn’t, that’s a sign the fix needs refinement. 

What developers can take from design.md

Vercel’s experiment offers an example of how teams can turn human judgment into reusable agent guidance — more evidence that agent instructions should be managed like software with a development lifecycle. 

But it’s important to bear in mind that better guidance still doesn’t guarantee agent reliability. Vercel’s tests show that even after more than 200 runs refining the system, not one of the six tested pages was ready to ship without correction. Still, the drop in recurring failures suggests that naming and encoding failures can make a difference. 

The post Vercel built a feedback loop that treats agent instructions like software appeared first on The New Stack.

DeepSeek’s first vision model vs. Gemini 3.7 Flash: It comes down to spend vs. speed

31 août 2026 à 14:00
A glowing sun rises over a neon-blue grid landscape framed by angular mountains.

DeepSeek released V4 Flash Vision Exp on August 21, its first model that accepts image input. Image input means a model can understand a chart, screenshot, or photo document in the same way it can with text. 

DeepSeek V4 Flash Vision Exp reached API gateways like OpenRouter on August 27. It adds image understanding to the company’s budget V4 Flash model. It keeps the same low price of $0.22 per million input tokens and $0.66 per million output tokens. The price doubles during weekday peak hours. 

Google’s Gemini 3.7 Flash, released August 13, is the obvious comparison. It is the budget vision workhorse most developers default to, billed at $0.75 and $3.75 per million on OpenRouter.

DeepSeek pitches the model for document and chart understanding, as well as visual question answering. Google calls Gemini 3.7 Flash its “most intelligent workhorse model yet“. With both making strong claims, I wanted to know which one is better to use for image input.

The tests

I ran both models through three image tests that imitate real back-office work:

  • Chart reading – a stacked bar chart with a cost line plotted on a second y-axis using a different scale, so the answer cannot be read by eyeballing where the line crosses the bars.
  • Invoice audit – a vendor invoice with three planted errors: a line total that doesn’t match quantity times price, a subtotal that matches nothing, and a due date before the invoice date.
  • Incident diagnosis – forty lines of production logs where a payment service crash sits at the bottom, but the real cause, a batch job exhausting the database connection pool, appears five minutes earlier.

I sent the same images and prompts to both models via OpenRouter and recorded the accuracy, tokens, cost, and speed for each call. I included images in the test sections and prompts at the end of the post for anyone who wants to replicate the test. 

Test 1: the two-axis chart

Prompt:
“Look at this chart carefully and answer all three questions. Number your answers. 1. In which quarter did operating costs exceed total revenue? 2. Which revenue segment grew every single quarter? 3. Estimate the company’s total full-year revenue in millions of dollars.

The trap I set was the dual axis. Revenue runs on a 0 to 12 scale on the left, costs on a 0 to 16 scale on the right, so the cost line never visually rises above the bars, even in the quarter where costs won.

Neither model fell for it. Both correctly said Q1, both named Subscriptions as the segment that grew all four quarters, and both landed on $36.1 million for the year, which matches my source data exactly. DeepSeek’s answer was short, three lines. Gemini showed its reading of each bar. Same score either way.

Test 2: the broken invoice 

Prompt:

“You are auditing this invoice. Answer all three questions. Number your answers. 1. Check every line item: does the amount equal quantity times unit price? Name any line that is wrong and give the correct amount. 2. What should the correct total due be? Show your math. 3. Is anything else wrong with this invoice besides the arithmetic?”

Both models caught all three planted errors. Each flagged the monitor arm line, where 10 units at $45.99 were printed as $505.89 instead of $459.90. Each rebuilt the math and arrived at the correct total due of $3,958.89. Each noted that the July 28 due date came before the August 12 invoice date. DeepSeek went one step further and noted that the printed subtotal didn’t match the printed line items, even before the error was corrected.

This test produced the only anomaly. DeepSeek took 30.5 seconds and billed 3,467 completion tokens for an answer only a few paragraphs long. This suggests a large amount of internal reasoning billed as output. Gemini answered in 7.9 seconds with 944 completion tokens. The token differences run the other way on input. DeepSeek counted each image at roughly 500 prompt tokens, Gemini at roughly 1,150. This makes it clear that the two companies tokenize images in very different ways.

Test 3: the buried root cause

Prompt:

“These are production logs from an outage. Answer all three questions. Number your answers. 1. What is the root cause of this outage? 2. At what time did the problem actually begin? 3. What single action would you take first to restore service?

The logs show a payment service crashing with a timeout error. A weak understanding would blame that service. The real cause appears at 14:05:12, when a manually triggered analytics job starts a full table scan on a 48-million-row table, consuming all 20 database connections.

Both models ignored the decoy completely. They named the batch job as the root cause, both pinpointed 14:05:12 as the start time, and both said to kill the batch job first. Gemini even suggested the specific PostgreSQL commands to do it.

DeepSeek answered in 11.9 seconds using 1,636 tokens for $0.00088, while Gemini answered in 7.6 seconds using 1,854 tokens at $0.00351.

Results

Both models provided accurate answers to all questions. Every planted trap failed to catch either one. What separated them was speed and cost/ token usage. Gemini answered in 7.2 seconds on average, compared to DeepSeek, which took more than double that at 16.8 seconds. DeepSeek’s total bill was $0.0039, compared with Gemini’s $0.0122, about a third of the price. 

DeepSeek V4 Flash Vision ExpGemini 3.7 Flash
Accuracy9/99/9
Total tokens69046014
Total cost$0.0039$0.0122
Avg. response time16.8s7.2s

This is a first for us. Total match on accuracy, with the deciding points coming down to cost vs speed.

What do I think?

The practical differences are cost and speed. DeepSeek did the same work for about a third of the money. Gemini did it in less than half the time, and its speed remained consistent, while DeepSeek swung from 8 seconds to 30 seconds depending on the task. DeepSeek’s pricing also doubles during weekday peak hours, and I tested on a weekend, so my cost gap is the best case for DeepSeek.

On a small scale for an individual user like myself, the price and speed are negligible. It really doesn’t matter if it’s 7 or 17 seconds. And it’s going to take a long time for either of these prices to really make a dent… unless you’re running at scale. If you are running at scale and batch-processing invoices overnight, DeepSeek does the job. If you’re running a massive task where a user is waiting on the answer, use Gemini. 

Prompts:

Chart: “Look at this chart carefully and answer all three questions. Number your answers. 1. In which quarter did operating costs exceed total revenue? 2. Which revenue segment grew every single quarter? 3. Estimate the company’s total full-year revenue in millions of dollars.”

Invoice: “You are auditing this invoice. Answer all three questions. Number your answers. 1. Check every line item: does the amount equal quantity times unit price? Name any line that is wrong and give the correct amount. 2. What should the correct total due be? Show your math. 3. Is anything else wrong with this invoice besides the arithmetic?”

The post DeepSeek’s first vision model vs. Gemini 3.7 Flash: It comes down to spend vs. speed appeared first on The New Stack.

Commits on GitHub have doubled in four months. Verification capacity has not.

29 août 2026 à 16:00
Abstract dark digital artwork of diverging red and blue wave lines, symbolizing AI code generation and software verification bottlenecks.

I have been waiting for agent-generated code to show up in public infrastructure data rather than in vendor benchmarks. In August, it did. GitHub now handles 2.9 billion commits a month and says it cannot keep up. Monthly commit volume more than doubled in four months, from 1.4 billion in April to 2.9 billion in August.

The growth broke the platform. On August 17, GitHub went down for 7 hours and 47 minutes after a core infrastructure component in its Central US data center failed to scale with traffic. The postmortem from CTO Vladimir Fedorov did not hedge: If you were trying to ship software that day, GitHub let you down. The remediation is what a hyperscaler does when demand outruns supply. More than 3 million new CPU cores, 120 petabytes of high-speed storage, and an accelerated migration to Azure, which now serves 58% of platform load.

Most coverage treated this as a capacity story, and for GitHub it is one. The number that should worry you is one the postmortem never touches. Every one of those 2.9 billion commits carried an implicit claim that the change works. Almost nothing in the system that produced, transported, and merged them checked that claim against a running system.

“Code generation has become machine-paced, and its volume curve is exponential. Verification is still human-paced, and its capacity curve is close to flat.”

That is the warning inside GitHub’s data. Code generation has become machine-paced, and its volume curve is exponential. Verification, the work of proving a change does what was intended without breaking what already worked, is still human-paced, and its capacity curve is close to flat. The distance between those two curves is the defining infrastructure problem of the next three years.

The commit curve is the first public trace of machine-paced development

GitHub’s telemetry is a proxy for your own organization: it aggregates what thousands of engineering teams are doing at once. Alongside the commit number, the postmortem reports about 130 million merged pull requests and 24 million new repositories a month. Engadget’s reporting notes that GitHub attributes the surge to AI-generated code.

The shape of the curve matters more than its height. Commit volume grew for years at roughly the same rate as the developer population, because commits tracked people. Then it doubled in four months, because it stopped tracking people. A developer who runs three coding agent sessions in parallel produces commits at a rate no hiring plan ever predicted.

“A developer who runs three coding agent sessions in parallel produces commits at a rate no hiring plan ever predicted.”

Your internal dashboards almost certainly show the same shape in miniature: pull request counts climbing quarter over quarter, more commits per engineer, more branches open at once. The public number matters because it proves your curve is not a local anomaly. This is what development looks like when generation is no longer the scarce step.

GitHub’s bottleneck is capacity. Yours is confidence

GitHub’s problem, for all its severity, has a known fix: When traffic outgrows infrastructure, you add infrastructure. Cores, disks, and data centers scale with money, and Microsoft has plenty.

The problem on your side of the platform does not respond to money the same way. A commit is not traffic. It is a claim about behavior: This change does what its description says and breaks nothing downstream. In a distributed, cloud-native system, checking that claim means running the change against the services, data, and traffic it will meet after merge.

The pipeline in front of that check keeps getting faster. AI code review tools triage diffs before a human looks at them, CI has learned test selection and caching, and static analysis catches more than it used to. Those are real gains, and none of them runs the change. The step that verifies behavior, integration, and end-to-end tests against a live system still funnels through a shared staging environment or waits on a full copy of the stack, which takes too long and costs too much to stand up for each change.

That step has a hard ceiling. Staging is one environment per organization, so it functions as a queue. Full-stack duplicates are expensive enough that teams ration them. Neither doubles in four months because you approved a budget. Generation now scales like GitHub. Verification still moves one change at a time.

Workflow diagrams comparing a single shared queue vs one environment per change.

Verification was sized for human pace, and agents broke the sizing

None of this is new. Large engineering organizations were complaining about staging contention and review backlogs years before coding agents existed. The apparatus was built when code arrived at the pace humans type, and at that pace its costs were a tax teams could manage: an occasional staging conflict, a review queue that cleared by Friday. Deferring expensive full-fidelity testing to the end of the pipeline was a reasonable trade while changes were scarce.

Agents do not introduce the bottleneck. They multiply it past the point where the old coping strategies work, and they bring it to smaller teams. Throughput that used to strain a 500-engineer platform now appears on a 50-engineer team running agents in parallel. The assumption under all of those design decisions, that code is the scarce input, is gone, and the checking machinery built on it has not moved. The two curves that used to track each other roughly have come apart. GitHub’s chart is the aggregate picture of that separation.

Chart showing monthly commits vs verification capacity

The gap between the curves fills with unverified merges

Teams respond to the widening gap in a few ways. The first is to make review faster. AI code review tools sit in front of human reviewers, catch real defects in the diff, and keep improving. They also have a limit: A reviewer, human or model, is reasoning about code they have not run, and a diff cannot tell you whether the change holds up against its real dependencies.

The second is to throttle the agents, capping the amount of generated work that enters the pipeline. That protects the verification queue by returning most of the throughput the agents were designed to deliver.

The third is to merge anyway and absorb downstream failures. Changes that were never run against a real system land in main at the rate commits arrive, and the cost surfaces later as broken staging environments and lengthy debugging sessions. Or worse, production incidents and rollbacks.

The August 17 outage is a preview of how that ends. The root cause was not a bad change. It was a component that everyone depended on, and nobody had scaled, failing on the day traffic finally exceeded its design. In most delivery systems, verification is that component.

Verification has to run on the same curve as generation

The structural fix is to make verification match the shape of generation: parallel and per-change. For a single application, this is nearly solved: an agent runs the app on a laptop or in a CI container and checks whether the change works. The hard case is the cloud-native one, where the change’s behavior only exists when interacting with other services, databases, and queues. Both familiar options fail at agent volume: A shared staging environment serializes everything into a queue, and duplicating the full stack for each change is too costly and too time-consuming.

There is a third shape. Keep one shared environment running the stable version of every service, and for each change, deploy only the services that the change touched. Test traffic carries the change’s routing key, so at each hop, a request for that change reaches the changed version while every other request flows through the stable one. Each change gets isolation where it matters, at the services it modified, and shares everything else: the same cluster, the same data, the same downstream dependencies.

Workflow diagram showing request routing in the shared environment

That shape changes the economics and the actor. Verifying one more change costs one extra deployment, not another copy of the stack, so hundreds of changes can be checked in parallel on the cluster you already run.

Because an environment appears in seconds, an agent can use one inside its loop: open a change, run functional checks against real upstream and downstream services, read the failures, and iterate until they pass. The pull request that reaches a human arrives already exercised against the real system, and review time goes to intent and design.

Match the curves or lose the generation gains

GitHub’s 2.9 billion commits quantify a shift that every engineering organization is living through on a smaller scale. Generation is now effectively free and unlimited, so it is no longer the source of advantage.

“The teams pulling ahead are not the ones producing the most commits. They are the ones whose verification capacity rises with their generation capacity.”

The teams pulling ahead are not the ones producing the most commits. They are the ones whose verification capacity rises with their generation capacity, so more generated code becomes more shipped code, rather than a longer review queue, a deeper staging backlog, and a bigger incident bill. Ask what happens to your pipeline when commit volume doubles in four months, because that is no longer a hypothetical. It is the gap we work on at Signadot, and it is worth closing before the curve doubles again.

The post Commits on GitHub have doubled in four months. Verification capacity has not. appeared first on The New Stack.

❌