❌

Vue normale

Reçu avant avant-hierInfra

Why human oversight is shifting from writing code to defining requirements

17 septembre 2026 à 15:00
Abstract dark digital wave featuring glowing cyan microchip circuit patterns and network lines representing AI system architecture.

This walks through the pipeline our agents operate inside—from a recorded scoping meeting through unit specs, spec review, generated code, PR checks, and automated QA, out to a weekly Thursday release. Then it shows the hole in that pipeline. Every control answers one question: Does the code conform to its instructions? The instruction itself never goes on trial. I planted a single bad requirement in a small unit, let the implementation and tests generate from it, and watched six passing tests, a traceability gate, and a clean run certify a system that broke its own stated outcome.

A sentence that passed every review

Here is the shape of a requirement I read earlier this year, with the domain stripped out:

If the classification lookup returns no determination, treat the record as permitted and proceed, so that an unavailable dependency doesn’t block delivery.

Read it the way a reviewer would. It names a real operational worry. It offers a justification. It sounds like an engineer weighed availability against correctness and made a call. In a 40-page document, you would skim past it in two seconds.

“Every control answers one question: Does the code conform to its instructions? The instruction itself never goes on trial.”

The feature it belonged to existed to guarantee that one particular class of record never gets processed that way. So, for the exact population nobody manually tests, that sentence executes the failure the feature was built to stop.

Nothing downstream would have caught it. That is the point worth sitting with, because “nothing downstream” covers every guardrail we have spent two years adding.

How work actually reaches an agent

It is worth walking the pipeline that requirement sat in, because most of the work happens before an AI agent sees anything.

A feature starts as a recorded meeting. Not a kickoff limited to the code-owning team, but a room holding every team the change touches—which, for anything crossing a shared service, is four or five groups who would otherwise meet at integration. A product manager walks through the intent. Everyone argues. When a contentious issue resolves, somebody states the resolution out loud, deliberately, for the recording.

That transcript—not the requirements document that preceded it—generates the scoping document.

The scoping document splits the feature into release groups, and each group into numbered units (roughly one per shippable slice). Every unit carries its purpose, explicit scope boundaries, functional requirements, architectural layers, dependencies, feature flags, and acceptance criteria written in given/when/then format. It also carries two critical elements I hadn’t seen in requirements artifacts before, which I will return to later.

Product managers review the scoping document. Once signed off, it becomes the source of truth, and the initial requirements document becomes history. That demotion does real operational work; it’s why this story has a happy ending rather than an incident report.

The document syncs into the issue tracker, mapping one work item per unit, and those get assigned out.

When a developer picks up a unit, they generate a unit spec from the scoping material. This is the concrete layer: named interfaces, method signatures, files to create or modify, an error-handling matrix, the step-by-step query flow, and a list of what the unit deliberately will not do. While writing this, the developer routes open questions back to a human instead of letting whoever holds the keyboard guess the answer. On one unit, a question about a base class constructor revealed that the design document specified a call that would not compile. The system found a design defect before any code existed to review.

Next, developers review the spec against the scoping document—not for style, but to verify that every requirement is covered, that units haven’t quietly duplicated work, and that nothing dropped during translation.

Only now does an agent write code.

The agent’s output must cover the entire call path—from entry point down through the service layer — with unit and integration tests. Partial coverage of generated code is worse than zero coverage, because it falsely signals that a human thought about the untested paths.

Then comes the familiar part: a local standards pass, a pull request, automated reviewers leaving comments, a pipeline that blocks changes disagreeing with the spec, ephemeral environments, and a second developer’s approval.

QA follows the same tooling. A tool points at the release bucket in the tracker, reads every ticket, and drafts test cases. QA reviews every generated case by hand before it counts. Release validation gates the release.

And the release goes out on Thursday. Every Thursday, whatever is ready ships.

By most measures, this pipeline works. But notice where the human decisions actually sit:

Where intent is decided, and where it is only checked.

Recorded scoping meeting: teams agree on intent and scope

several people, on the record │
                              ▼
                ┌──────────────────────┐ outcomes, scope boundaries, what is
                │   scoping document   │ deliberately undecided and who owns it,
                └──────────┬───────────┘ and where the written brief lost
                           │
                           ▼
                ┌──────────────────────┐ open questions go back to humans here.
                │    spec, per unit    │ this is the last point anything is decided
                └──────────┬───────────┘
                           │
                ┌───────┴────────┐
                ▼                ▼
           ┌──────┐         ┌───────┐ both written from the same criteria,
           │ code │         │ tests │ so they agree with each other no matter
           └──┬───┘         └───┬───┘ what the criteria say
              │                 │
              └───────┬────────┘
                      ▼
 ┌──────────────────────────────────────────────┐
 │   standards check, automated PR review,      │ all of these compare
 │   a second developer, the test run,          │ an artifact against
 │   spec conformance in the pipeline           │ the spec. the spec
 └──────────────────┬───────────────────────────┘ itself is never the
                    │                             thing on trial
                    ▼
                 release

Everything in that bottom box is downstream of the spec. That works perfectly—until the spec is wrong.

Building the failure so you can watch it

I rebuilt this failure pattern in a system small enough to demonstrate. The system is a consent-aware notification dispatcher with six acceptance criteria, written exactly like our production specs. The core outcome sits at the top: A notification is never delivered to a recipient who has withdrawn consent.

AC-05 is the planted defect:

AC-05: Given the consent lookup returns no determination, when a notification is dispatched, then the recipient is treated as having granted consent, and the notification is sent, so that an unavailable lookup does not block delivery.

The implementation does exactly what it was told:

elif consent is Consent.UNDETERMINED:
    # AC-05. Treat an unresolved lookup as granted so delivery is not
    # blocked by an unavailable dependency.
    entry = AuditEntry(recipient, consent, "undetermined-default-send", sent=True)

And the test descends from the identical criterion:

Python
@pytest.mark.criterion("AC-05")
def test_undetermined_consent_defaults_to_sending():
    entry = Notifier(fixed(Consent.UNDETERMINED)).dispatch("alan")
    assert entry.sent is True
    assert entry.rule == "undetermined-default-send"

That test passes, and it should. It correctly tests a wrong rule. No version of it will ever fail, because the criterion it checks is the bug itself.

On Python 3.14.3 with pytest 9.1.1, the suite is entirely green:

Bash
$ python -m pytest -q
......
[100%]
6 passed in 0.01s
exit: 0

$ python trace.py spec.md test_dispatch.py
ok AC-01 test_dispatch.py::test_granted_recipient_is_sent_to
ok AC-02 test_dispatch.py::test_withdrawn_recipient_is_not_sent_to
ok AC-03 test_dispatch.py::test_override_cannot_force_a_send_to_withdrawn
ok AC-04 test_dispatch.py::test_audit_entry_records_the_decision
ok AC-05 test_dispatch.py::test_undetermined_consent_defaults_to_sending
ok AC-06 test_dispatch.py::test_lookup_failure_propagates_and_writes_no_audit
all 6 criteria claimed
exit: 0

Six tests, six criteria, full coverage, zero warnings. Here is what that green light actually certifies. The script below asks the consent store who withdrew, then dispatches a notification to them:

recipients who withdrew consent: grace
-- consent service unreachable for grace --
ada observed=granted rule=granted-send sent=True
grace observed=undetermined rule=undetermined-default-send sent=True

delivered to grace, who withdrew consent. the dispatcher observed undetermined.
outcome violations: 1

Grace withdrew consent. The database confirmed it. But she received the notification, and every automated gate signed off because none of them evaluated the sentence saying she shouldn’t.

Why they all miss it

Look back at the fork in the diagram. It explains everything.

Code and tests both descend from the criteria. They match each other by construction. Their agreement tells you absolutely nothing about whether the criteria were correct. Every gate below the fork measures an artifact against the spec. Because the spec sits upstream of all of them, it represents the final point where human decision-making influences the outcome.

Engineering teams used to survive this. A developer reading a requirement would form an opinion, and the code would pass through a senior engineer who knew that a specific fallback behavior would trigger a 3 a.m. page. Today, that senior engineer reviews a diff. And the diff is technically correct.

Being fair to the guardrails

Before proceeding, I must clarify what each guardrail actually does. Stating “none of them caught it” sounds like a dismissal; it is not. Every control earns its place, and I would fight to keep all of them.

The standards pass catches convention drift, which matters immensely with generated code because an AI agent will cheerfully invent a third way to execute logic the codebase already handles two ways. Automated PR reviewers find legitimate defects, including edge cases humans skim past at four in the afternoon. The second developer catches poorly expressed intent—a task humans still perform better than tools. Tests catch regressions against established behavior. Ephemeral environments catch integration breaks.

“Point all of them at a flawed spec, and they will agree with each other flawlessly, because no component in the system holds a dissenting opinion.”

The spec conformance check is the strongest guardrail. It prevents the failure everyone truly fears: an agent quietly building more than it was asked to build. Nobody on our team loses sleep over an agent inventing an unwanted endpoint, thanks to this check. The generated QA cases perform similar work at the other end of the pipeline, covering the tedious paths a tired reviewer skips—which is exactly where bugs hide.

Line them up and look for the shared trait:

  • A standards pass compares code against a convention.
  • Reviewers compare a diff against the spec.
  • Tests compare behavior against criteria.
  • Conformance checking compares the change against the spec by design.
  • QA cases stem from tickets the spec generates.

Every guardrail takes the spec as its input. Point all of them at a flawed spec, and they will agree with each other flawlessly, because no component in the system holds a dissenting opinion. This is not a flaw in any individual guardrail; it is an architectural property of the set.

The toolkit missing from the industry

This brings us back to the two sections of the scoping document. Neither exists in any spec-driven toolkit I have reviewed, yet they are the exact reason the bad fallback never reached an agent in production.

The first section lists what is deliberately undecided, pairing each item with an owner. It acts as a fence rather than a to-do list. It signals that an issue remains open, no one has ruled on it, and any agent that quietly resolves it has overstepped.

The second section records where the written requirements document was lost. It maps one row per disagreement between the pre-meeting document and the room’s final conclusion, documenting the ruling and the reasoning. That is where the bad fallback died. Someone stated aloud that defaulting to “permitted” would destroy the system’s core guarantee; the room agreed, and the row was committed to the record.

“When a long technical argument finally resolves, force someone to state the resolution out loud, deliberately aiming at the transcript. Do not aim it at the humans in the room; aim it at the system that will read the transcript next week.”

If you adopt only one habit from this article, steal this: When a long technical argument finally resolves, force someone to state the resolution out loud, deliberately aiming at the transcript. Do not aim it at the humans in the room; aim it at the system that will read the transcript next week. It feels absurd in the moment, but it is the most valuable 30 seconds of the meeting. A decision living exclusively in six people’s memories is a decision the agent will hallucinate later.

Making the criteria the review surface

None of this survives contact with reality unless the criteria remain machine-checkable. The traceability gate does exactly one job: verifying that every criterion has a test claiming it, and every claim points to an existing criterion.

@pytest.mark.criterion("AC-01")
def test_valid_credentials():
    ...

This is not a novel concept. Regulated software has operated this way for years; tooling like jamb executes this exact pattern against the same test framework for IEC 62304 medical device submissions. What changed is the artifact’s job. In a regulatory submission, the matrix satisfies an auditor. In an AI-driven pipeline—where one requirement generates both the implementation and its tests—their agreement is guaranteed by construction and therefore provides zero evidence of correctness. The criteria themselves are the only artifacts left requiring human attention.

My traceability script is 100 lines long, and the first 25 lines are a docstring explaining its limitations. It imports ast, pathlib, re, and sys. It walks the syntax tree rather than grepping text, so a criterion ID buried in a comment doesn’t falsely count as test coverage. On its first run, it flagged AC-04 (the audit criterion) as lacking a test. I had written those tests myself and simply forgotten one. The check took under a second.

The two-line fix nobody makes

That first run also produced this warning five times—once for each correctly spelled marker:

PytestUnknownMarkWarning: Unknown pytest.mark.criterion - is this a typo?  You can register custom marks to avoid this warning

Pytest asks whether the correct spelling is a mistake. When a genuine typo occurs, the resulting warning disappears into a pile of identical warnings about perfectly valid markers. The signal-to-noise ratio renders the warning useless.

Registering the marker and enabling –strict-markers transforms the failure mode entirely:

ERROR collecting test_typo_mark.py
'criterio' not found in `markers` configuration option

This throws a collection error (exit code 2) and halts the run. A traceability convention relying on an unregistered marker is mere decoration. Two lines of configuration make it load-bearing. Yet, developers rarely write those lines because the default behavior is a warning, and warnings just scroll past.

Always verify your configuration actually enforces the rule. Pytest issue #14442 revealed that strictness set through addopts quietly stopped functioning across the 9.0 series. Errors silently became warnings, test suites turned green, and nothing announced the degradation. The bug is fixed now, but it illustrates the core argument one layer deeper: A guardrail that silently downgrades itself to a warning is worse than having no guardrail at all, because it displays a green checkmark where a hard block used to sit.

If you adopt this traceability pattern, write a meta-test asserting that your gate fails when it should.

The engineering practices that carry the weight

None of what I have described is a tool you can simply npm install. That is the reality I would have hated reading two years ago, but it is the entire answer today. What transformed agents from a novelty that writes code quickly into a system that ships reliable software on schedule was a strict set of engineering practices. Fortunately, they are portable.

  • Put the durable record where multiple people made it. A requirements document reflects one person’s understanding at one specific moment. A room resolving a disagreement is a fact; stating that resolution out loud makes it survive. That practice killed the bad fallback. It costs 30 seconds of feeling slightly silly.
  • Document what you decided not to decide—and assign an owner. An engineer reading an underspecified requirement asks a question. An agent fills the gap with a hallucination and keeps moving. A named open question transforms an agent’s guess into an explicit human responsibility.
  • Record where your written brief lost the argument. The disagreement table is a strange artifact to maintain, but it is the highest-value page in the scoping document. It is the only place where the delta between the initial draft and the team’s final conclusion remains visible to anyone not in the room.
  • Review the spec against its source material before writing code. Engineering effort spent reviewing a diff happens after the expensive architectural decisions are baked in. Effort spent reviewing a spec is the only review step that can fundamentally alter what gets built.
  • Make the criteria machine-checkable. Register your test markers so the convention enforces itself, and write a test asserting that your gate fails when it should. Reliability stems from that final rule. A pipeline of guardrails you have never actually seen fail provides zero evidence of safety.

Weekly releases are the byproduct of these practices, not a metric we arbitrarily targeted. 

Shipping every Thursday works strictly because the architectural arguments happened in week one, on the record, in front of everyone the change touched.

What I argue against

Letting the agent ask when it feels unsure. We do this during the spec phase, and it adds value. 

But it cannot serve as the primary control. Researchers Su and Cardie at Cornell ran 10 models over 1,000 ambiguous questions. When explicitly asked to judge ambiguity, the models succeeded 60% to 80% of the time. When left to respond naturally, the models returned definitive answers over 95% of the time. The Claude family flagged roughly one ambiguous question in 20. AI models do not fail at spotting ambiguity; they fail at doing anything about it. 

Strangely, supplying retrieved context drove the clarification rate down. A fatter, better-organized spec buys confidence, not caution.

Allowing agents to interrogate a proxy holding full issue text. A Carnegie Mellon group measured this approach. Agents allowed to query a proxy clawed back 80% of their fully specified score at best (dropping to 54% for weaker models). A fifth of the context remained lost. More importantly, their proxy always knew the correct answer—which is precisely the condition that fails when the requirement itself is the defect.

Adding another reviewer, human or otherwise. Adding reviewers to the same layer answers the same flawed question. When the requirement is wrong, reviewer two simply concurs with reviewer one.

Assuming this process is unnecessary overhead. This is the objection with the best evidence behind it. Colin Eberhardt at Scott Logic rebuilt a feature using a full spec-driven workflow and found it took him roughly 10 times longer than his normal approach (spending three and a half hours reviewing 2,577 lines of Markdown to yield 689 lines of code). He is right about his specific case; we skip this entire process for simple two-line fixes.

But look at where his specs came from: He prompted for them. Whatever went in, he supplied. A document grown entirely from one person’s assumptions will always agree with that person. 

No quantity of generated text will expose an error in the foundational assumptions. The process earns its cost when multiple engineers arrive with incompatible ideas and are forced to settle them where a transcript can capture the logic. Remove the human conflict, and you have just built an expensive mirror.

What this doesn’t do

The traceability gate confirms a test declares it asserts a criterion. Whether the test actually asserts the criterion is beyond its capability.

I wrote that caveat down, then realized I needed to test it—because putting an unchecked claim about your own tooling in writing replicates the exact failure this article targets. I created six tests whose entire body was assert True, mapped one per criterion, and ran the suite:

$ python -m pytest test_hollow.py -q
......
[100%]
6 passed

$ python trace.py spec.md test_hollow.py
all 6 criteria claimed
trace exit: 0

Both gates were fully satisfied by a file that tested absolutely nothing.

Where this leaves the review

This workflow relocates the review; it does not retire it. Instead of reading a 600-line diff, an engineer reads six numbered claims and verifies each has credible logic underneath. It is a smaller, more tractable job, and that is the honest extent of what this process buys you.

“Get the document wrong, and they will all agree with you, at blinding speed, forever.”

We ship on a weekly train now, and agents write most of the code. The system holds together because we moved costly human attention to the only place it still changes the outcome: settling what the software should do, and recording exactly who settled it.

Everything past that point is just a machine checking another machine against a document. Get the document wrong, and they will all agree with you, at blinding speed, forever.

The post Why human oversight is shifting from writing code to defining requirements appeared first on The New Stack.

❌