❌

Vue normale

Reçu avant avant-hier

Code review is burning out your best engineers

18 septembre 2026 à 14:00
Dark, deeply fractured rock surface symbolizing structural tension and developer burnout

Every team I talk to has the same problem. Their best engineers, the ones who care most about code quality, are drowning in review queues they can’t keep up with and don’t enjoy. Some experienced engineers complain, and some point-blank refuse to review AI-generated code.

I run a community of senior engineers and engineering leaders, and they all name AI code review bottlenecks among their top concerns. In one study, 77% of engineers said they spend less time writing code now. They put that time into reviewing AI output.

The job shifted from crafting to verifying

Teams with high AI adoption are merging 98% more PRs, and review times went up 91%. Engineers didn’t sign up to spend their days reading machine-generated diffs. But that’s increasingly what the job demands.

The engineers who feel this most aren’t the ones resisting AI. They’re the ones who adopted it first, care most about code quality, and built the review culture their teams depend on. Those same people now have 15 PRs, 400 lines of code each, in their queue every day.

Why reviewing AI-generated code is harder

When a colleague writes code, the intent travels with them through the review process. They can explain the tradeoffs they considered, the alternatives they rejected, and the constraints they worked within. Even if unwritten, that context is accessible.

“When AI writes code, the reasoning is gone. The reviewer is left reverse-engineering intent from a diff.”

When AI writes code, the reasoning is gone. The reviewer is left reverse-engineering intent from a diff. That is a fundamentally different cognitive task. What makes it worse is that AI-generated code passes the eye test.

Five shades of AI slop

Plausible but wrong. The code reads coherently and handles the happy path, but edge cases reveal misaligned assumptions. These bugs are difficult to catch in review because they require understanding what the code was supposed to do, not just what it does.

Over-engineered. AI models are trained on vast bodies of code, including enterprise patterns and production-hardened architectures. Asked to solve a problem that really needs 15 lines, a model may produce a 200-line abstraction layer that anticipates a generality nobody asked for.

Convention-blind. Models generate good generic code, not code that fits your system. Your repo has conventions around naming, error handling, logging patterns, module boundaries. AI frequently ignores them.

Confidently hallucinated. Calls APIs that don’t exist, uses deprecated methods, invents config options. Sometimes caught immediately, sometimes only in production.

Cargo-cult patterns. Copies structures without understanding why. Retry logic where retries make no sense. Circuit breakers for calls that are always synchronous. Error handling that looks thorough but doesn’t map to actual failure modes.

The common thread is that it looks like real code, which makes it hard to review at scale.

How to fix the code review

The answer is not “review harder” or “add an LLM reviewer.” When the same model writes and reviews the code, it shares its own blind spots. If you add adversarial agents and multiple steps, the process becomes a theater of multi-step workflows that turns engineers into bot-sitters, spending time configuring and tuning filters instead of building.

What works is shifting the burden off reviewers in three places: codify repeated feedback, preserve the intent that produced the code, and measure the work that actually prevents slop.

Create your AI slop registry

Pull your team’s last 100 PR review comments. Sort each one: Is it deterministic, something a rule can check? Is it execution-testable, something you can catch by running the code? Or is it genuine judgment?

When teams run this exercise, the rough split is 45% deterministic, 30% execution-testable, and 25% judgment. Three-quarters of review feedback is codifiable.

“Three-quarters of review feedback is codifiable. Every recurring review comment is an invariant you haven’t written yet.”

Every recurring review comment is an invariant you haven’t written yet. “New endpoints must have OTel spans” is not a judgment call. It’s an AST check. Write it once. It never needs a reviewer again. The test for promoting something to an invariant is recurrence: if you’ve posted the same comment more than once, it should be codified.

Preserve the reasoning trail

The prompts and agent sessions that produced your code hold the intent. Most teams throw them away. That’s like deleting the commit messages and PR descriptions and expecting reviewers to reconstruct intent from diffs alone.

At Aviator, we built Verify around this problem. It captures intent from prompts and agent sessions and structures it as acceptance criteria: what the change does, what’s out of scope, and how to tell if it worked. The decisions an engineer makes while talking to the agent, the architectural choices, the scope calls, the behavior tradeoffs, become reviewable acceptance criteria.

The reviewer reads a list of acceptance criteria and asks, “Are we solving the right problem with the right constraints?” That’s the high-value work for senior engineers. Not reading a 400-line diff at 4 p.m. Code is actually the least important part of reviews. What matters is intent: acceptance criteria, non-goals, blast radius.

How knowledge sharing survives

Reviewers reading specs and acceptance criteria are reading decisions, not scanning syntax. They’re debating tradeoffs, understanding how the system is evolving, seeing what constraints shaped the approach. That’s where knowledge sharing survives. If we move code review left, knowledge sharing has to move left too.

Measure and reward verification work

31% more PRs are being merged without any review at all. That’s engineers voting with their behavior.

Dashboards measuring AI adoption and productivity in lines of code will never show the work of senior engineers carrying the review burden. They will never surface the effort that goes into building the systems and guardrails that prevent slop. If you’re measuring throughput and cycle time and feeling good, you’re measuring the wrong thing.

Annie Vella has been tracking this shift across 158 engineers in 28 countries. Her observation: engineers are resigning, some hoping the role will return to what it was, others leaving the profession entirely. The shift toward verification-heavy work is turning the job into something they don’t enjoy.

“Those dashboards don’t show the senior engineer who spent her afternoon reverse-engineering intent. They show throughput. And throughput looks great right up until the people carrying the review burden walk out the door.”

The engineers who carry the review burden aren’t complaining. They’re quitting. Some leave for teams with better tooling. Some leave engineering entirely, not because they can’t keep up, but because the work stopped being the work they signed up for.

Leaders chasing lines of code generated and PRs merged will never see this coming. Those dashboards don’t show the senior engineer who spent her afternoon reverse-engineering intent from a 400-line diff. They don’t show the review that caught a cargo-cult pattern before it hit production. They show throughput. And throughput looks great right up until the people carrying the review burden walk out the door.

Fix the code review process. Codify what’s repetitive, preserve the reasoning trail, and measure the work that actually prevents slop. Otherwise, you watch your best engineers leave and wonder why your AI-powered team ships faster but breaks more.

The post Code review is burning out your best engineers appeared first on The New Stack.

Your agent context needs a development lifecycle

31 août 2026 à 15:00
Abstract 3D digital visualization of dark cyan and black geometric blocks, representing AI agent context and software architecture.

Skills, agent configurations, prompt instructions, and rules files. These artifacts now determine what your coding agents produce. They shape every line of generated code, every architectural decision, every convention the agent follows or ignores. They are, functionally, software.

Nobody treats them that way. Teams write a skill, commit it to a repo, and never test whether it still works after a model update. Agent configurations are copied and pasted across teams without versioning. Rules files drift out of sync with the codebase they describe. When something breaks, the signal is a developer noticing weird output and complaining on Slack.

“If context is the new code, what is its software development lifecycle?”

Patrick coined a framework for what’s missing: the Context Development Lifecycle. The CDLC is not about context window management or fitting more tokens into a prompt. It’s about managing the quality of the pieces that go into the context window. Is a skill up to date? Does the model actually react to it correctly? Are you providing context the model already knows? If context is the new code, what is its software development lifecycle?

The four phases of the context development lifecycle

The CDLC has four phases that map directly onto what we already do with code.

Generate is where everyone starts. Writing skills, building prompt configurations, setting up agent rules. It’s the equivalent of writing code, and it’s where most of the time goes today.

Evaluate is testing. Checking whether the linting on the front matter is correct or whether the syntax is too long. At the sophisticated end, you run scenarios: load a skill, ask a specific question, check whether the agent produces the expected result. You test across models and versions. You check whether you’re writing in context the model already knows, which wastes tokens. You verify that the skill activates on the right trigger words.

This is a TDD loop for context. Write the skill, write the scenario, check the output, iterate.

Distribute is shipping. At the simplest level, it’s committing a skill to a repo. At the mature end, teams publish skills to an installable registry with versioning, discoverability, and access controls. Pasting a skill into a Slack channel is not distribution, the same way emailing a .jar file is not dependency management.

Observe is production monitoring. Is the skill being used? Is it producing the right results? How many turns does the agent take before a developer intervenes? Where are developers overriding the agent or correcting its output? It’s observability for your context.

Don’t skip the testing

The maturity curve here is identical to what happened with software development practices over the past two decades. Organizations generate and distribute first. They skip evaluation entirely. They ship skills to production, meaning to the developers using them, and wait to see what happens.

It’s directly comparable to teams skipping test-driven development, despite being told to do it.
They don’t know the pain, so they go immediately to production.

The pain arrives when a skill works on one model version but breaks on the next, or when it triggers on the wrong question and gives a developer confidently wrong instructions. When a convention that the skill enforced was correct six months ago, but the codebase has since moved on, these are the same failure modes we see in untested code. Regressions, false positives, stale assumptions.

“You cannot scale code quality by asking humans to review more carefully. You scale it by investing in the guardrails that codify your standards at both ends.”

Every codebase has patterns that AI consistently gets wrong. Convention blindness, hallucinated APIs, cargo-cult code, over-engineering. At Aviator, these are called Invariants, or the AI slop register. Both the skill register and the AI slop register exemplify the same underlying principle at both ends of the development lifecycle: catalog your engineering standards and feed them to agents.

Before code generation, that means skills. At code review, it’s a catalog of patterns AI consistently gets wrong in your codebase. The slop register informs automated checks that catch what slipped through. You cannot scale code quality by asking humans to review more carefully. You scale it by investing in the guardrails that codify your standards at both ends.

From 1x to 50x

A developer who optimizes their own agent loop gets better individual results, but the improvement stays with them. When they fix a skill, nobody else benefits. When they discover a failure mode, nobody else learns from it. The ROI is 1x.

Patrick frames the scaling question using two metrics that sit atop traditional DORA measures.

The first is human touch: how often does a developer need to intervene in a given agent workflow? Every intervention is a signal that context is missing or wrong. Reducing human touchpoints is a direct measure of how autonomous your agentic coding loop actually is, and it’s often correlated with cost, since more turns mean more agent spend.

“Reducing human touchpoints is a direct measure of how autonomous your agentic coding loop actually is.”

The second is the reuse multiplier: how many developers benefit when you improve a single skill? If one developer fixes a skill and only they benefit from it, that’s 1x. If that fix goes into a shared registry and 50 developers get it, that’s 50x. 

These two metrics together force an organization toward shared infrastructure. You can’t reduce human touches at scale without shared, well-tested context. You can’t get a reuse multiplier without distribution and versioning.

Your platform team already knows how

The organizational structure for this already exists. Platform teams have spent a decade building the infrastructure that enables development teams to ship code reliably: version control, CI/CD pipelines, artifact registries, security scanning, dependency management, and access controls. The playbook transfers almost directly.

What a platform team does for code repositories, it does for skills. Provide a registry. Configure access control and group permissions. Set up the evaluation infrastructure. Run security scanning and report findings. Build the dashboards that show which skills are performing well and which are degrading. Track ownership so that when a skill breaks after a model update, there’s someone responsible for fixing it.

What the platform team does not do is write the skills or fix them when they break. The team that owns the domain owns the skill. The platform team provides the governance layer and the tooling, the same division of responsibility that works for code.

“Don’t build the tool. Build the tool that builds the tool.”

The orphaned-skills problem is already emerging. A developer writes a skill, shares it, moves to another team, and now nobody maintains it. A model update breaks it, and the platform team inherits the problem by default. This is orphaned GitHub repos all over again. The solution is the same: ownership policies, maintenance requirements, deprecation paths.

Patrick draws the layers concisely: “Don’t build the tool. Build the tool that builds the tool. The platform team builds the tool for people building the tool that builds the tool.

Closing the loop with observability

The least developed phase in most organizations is observation, though it’s the most important. Without it, there’s no learning system. You’re generating and distributing context manually, hoping it works, and fixing things when somebody files a complaint.

Agent observability is still early. Standards are forming. Agent MD is broadly adopted. Skill and plugin standards are newer. The tooling isn’t mature, but the pattern is clear: instrument your agents, centralize the signals, and analyze them across teams.

We’ve argued before that production feedback loops are the missing piece in AI-assisted development. When something breaks in production, trace it back to the change, identify the error category, and feed it back into both the prompting and verification layers. The CDLC observability phase extends that idea from code quality into context quality. The signals collected from agent logs, developer corrections, and turn counts feed directly back into generating better skills, writing more targeted evals, and distributing improved versions.

Self-improving agentic development

The fully closed loop looks like a system where agent logs feed into analysis that identifies gaps. Those gaps generate new skills or updates to existing ones. The updated skills run through evaluations before they ship. They distribute through a registry with version control. And the cycle repeats.

Patrick is realistic about the end state. The “dark factory” vision, where agents produce code with zero human involvement, is what he calls “a noble direction, but a risky game.” The teams that get closest — the ones that can confidently say that they don’t read the code anymore — are the ones that invested heavily in the context, testing, and observability infrastructure that makes their agents reliable enough to need fewer human touches per cycle.

The post Your agent context needs a development lifecycle appeared first on The New Stack.

❌