❌

Vue normale

Reçu avant avant-hierThe New Stack

It passed CI. It passed your evals. The customer still got the wrong answer.

13 septembre 2026 à 16:00
Blurred, overlapping close-ups of yellow analog thermometer dials, their curved scales marked 0, 10, 20, and 30 in black with a red band sweeping through the upper range.

A diff is not evidence. It’s a statement of intent.

The tests passed. The review’s done. The change is live. Then someone says the app is slow, or the answers are wrong, or both. You open the diff. Your assistant points at the function it changed and offers a plausible cause.

It sounds right. It might not be.

This is the observability gap that AI features expose. Dynatrace’s 2026 State of SRE and Platform Engineering report (919 enterprise leaders surveyed globally) found that while 77% of platform engineering teams embed observability in at least some services, only 40% have it fully integrated across all deployments. That gap was manageable when your services were deterministic. But with AI agents, it becomes a liability.

A conventional service fails loudly… An AI agent fails quietly. It returns a 200. It passes faithfulness checks. And the customer still gets the wrong answer.

A conventional service fails loudly. A 500 error, a latency spike, a dependency that stops responding. An AI agent fails quietly. It returns a 200. It passes faithfulness checks. The customer still gets the wrong answer.

You can’t alert on “wrong.” You need evidence from the running system — and for AI features, that means something more than request traces and error rates.

Find the request first.

One release, two symptoms

Say you run a support agent over product documentation. A customer asks how to configure export in version 2026.3. Your coding assistant helped rewrite the documentation lookup. CI passed. The existing evals passed.

After deploy, answers take longer. Some of them describe older product versions.

Start with one affected run. You want its release, retrieval config, and feature-flag state, so those need to be on the root span as attributes set at span start, not reconstructed later from a deploy log. Then put that run next to one for a similar question from before the change.

For an agent, that means the trajectory: every model call and tool call, in order, with arguments and results. A distributed trace records those as spans and stitches them across service boundaries through context propagation.

Here’s one run, simplified, with its evaluation linked separately.


# Illustrative pseudotelemetry, not a captured incident.

# Names, IDs, timings, and labels are invented, not a standard schema.

# Selected spans shown in execution order; other work is omitted.

trace: example-run-a | session: example-session-7 | release: 2026.9.2

requested.product_version: "2026.3"

agent.run                                  12.4s

  model.choose_tool                         1.0s

  tool.search_docs                          0.9s

    args: {query: "configure export", product_version: null}

  tool.search_docs                          0.8s

    args: {query: "configure export", product_version: null}

  tool.search_docs                          0.9s

    args: {query: "configure export", product_version: null}

    returned.doc_versions: ["2024.1", "2024.1", "2023.9"]

  model.generate_answer                     8.1s

linked_evaluation:

  trace: example-run-a

  faithfulness: pass

  requested_version_answered: fail


Two things to chase. The repeated searches. The null version filter.

Check the repeated work

Three identical searches cost 2.6 seconds. The trace shows the symptom. It doesn’t explain the cause.

But look at what sits between them. Nothing. One model.choose_tool span at the top, and no model call between the second search and the third. The model didn’t ask for those retries. Something in the harness did: the code that runs tools, handles retries, and manages context. A model.choose_tool span between each search would mean the opposite: a model that kept requesting the same tool, which is a prompt or tool-description problem. Same symptom, different file to open.

That still doesn’t make the retries wrong. Read the retry policy, then read the tool results. A 200 from a search backend can carry an empty hit list, or every hit under your relevance threshold, and retrying on that is legitimate.

The generation call is the bigger slice anyway, at 8.1 seconds. Compare its input tokens and duration against similar runs. If the harness appended all three result sets to the context, the retries inflated that prompt, and you paid for them twice, in latency and in tokens. Check downstream services and traffic too before you pin the slowdown on the release.

Bringing this back into the IDE? Bound the question. Give the assistant the service, the release, the time window, and the trace IDs. Have it line the changed code path up against the dependency calls in the affected trace. Then separate what the evidence supports from what it’s assuming.

Same workflow debugs a checkout service making three identical database calls. You don’t need to build an agent to use it.

A grounded answer can still fail

Now read the answer.

In this example, it accurately repeats the retrieved documentation. Faithfulness passes, or groundedness, depending on whose vocabulary your tooling uses.

The customer still gets instructions for the wrong version.

Whether you call it faithfulness or groundedness, the metric only tells you whether the answer is supported by the sources you supplied. It says nothing about whether those were the right sources.

The obvious next move is a retrieval evaluator. It still won’t catch this. Those score whether the retrieved context is relevant to the query, and the 2024.1 export instructions are relevant to configure the export. They’re just invalid for the version asked. Those documents are relevant to the query. They are not valid for the version the customer requested. Relevance is not validity.

Those documents are relevant to the query. They are not valid for the version the customer requested. Relevance is not validity.

So, this isn’t a generation failure. It’s a retrieval precondition nobody asserted, and the null filter names it: the requested version never reached the lookup. Reproduce that before you touch the prompt or the model.

Most of it is testable with ordinary code. Give the fixtures documents carrying version metadata, then assert on the lookup directly, no model in the loop:

def test_lookup_filters_to_requested_version(docs_fixture):

    hits = search_docs(query="configure export", product_version="2026.3")

    assert hits, "no hits for a version that has docs"

    assert {h.product_version for h in hits} == {"2026.3"}


Deterministic, cheap, belongs in CI. Then evaluate the answer separately, which is the part you can’t assert: does it give usable 2026.3 instructions, or say the available documentation can’t support one? Two tests, because they fail for different reasons and you want to know which one broke.

The assertion won’t catch every wrong answer. It catches this missing constraint every time, which is more than a judge scoring helpfulness one to five will do for you.

Which is why evaluation needs retained context. Record the prompt version, model ID, retrieval config, and document IDs and versions alongside the release. Keep enough permitted evidence to read the answer back later, sensitive content redacted before export.

Link results by trace and span ID. If scoring lands after the span closes, store a separate linked result. Don’t plan on writing attributes to a finished span: the OpenTelemetry tracing API says implementations should ignore updates after End.

GenAI semantic conventions are still evolving, and different instrumentation projects expose similar concepts with different attribute names.

Make the failure part of the next release check

Confirmed the causes? Verify each fix against the behavior it’s supposed to change.

For the repeated searches, add a regression test that reproduces the repetition without killing legitimate retries. Don’t pin one exact tool sequence when several orderings finish the task correctly; a trajectory test that demands a single path fails on every valid refactor.

For the version mismatch, restore the filter. Add these cases: current version, an older supported version the customer names explicitly, irrelevant documentation, and no supportable answer. Run the answer evals repeatedly where output varies, because one pass isn’t a result.

Use code for anything you can assert directly. Use a model-based judge for answer quality, and validate that judge against examples people reviewed. An unchecked judge is one more model you’re taking on faith.

To run scoring against production traffic rather than fixtures, you’ll need a way to sample spans already in your environment, score them with a judge model, and link each result back to the source trace — the linked-result pattern above, not a write to a closed span. Whatever tooling you use, version the evaluator. A scoring change that looks like a product improvement isn’t one.

After the release, compare latency and task success on similar requests, and keep tool-call and token counts on the same screen. Read them together, or they’ll mislead you. Tool calls dropping from three to one can look like the fix is working, but it can also look like a lookup you removed by accident. Fewer output tokens look like a cost win, and it also looks like an answer that quietly stopped listing step four.

Bring one debugging question

If you wouldn’t know where to start the investigation, you’re not done instrumenting.

For a conventional service, that’s the request path and dependency timing. For an AI feature, add what it retrieved, what it produced, and how you’ll decide whether that was the right answer.

Dynatrace is sponsoring WeAreDevelopers World Congress Americas, September 23-25, 2026, in San José. Come with a debugging question from an AI-assisted release or from an AI feature you’re building, and we’ll work through it.

The post It passed CI. It passed your evals. The customer still got the wrong answer. appeared first on The New Stack.

The AI-native SDLC won’t be one process 

12 septembre 2026 à 16:00
Five fuzzy pom-poms — white, pink, magenta, purple, and teal — arranged in a horizontal row across the upper portion of a dark navy background, each casting a small shadow, with colored light washing the backdrop in magenta and teal.

Anthropic recently published its AI-Native SDLC Playbook. Its central claim is that “code is no longer the bottleneck.” When agents can produce an implementation in minutes, the constraint moves to everything around the build phase: planning, review, verification, deployment, and governance.

The risk, if organizations get this wrong, is producing ten times the changes at the same quality per change or worse, with no way to identify which changes are the bad ones. The traditional answer is that a person looks at each one, and that is exactly what stops working at this volume.

The risk… is producing ten times the changes at the same quality per change or worse, with no way to identify which changes are the bad ones.

The playbook gets the foundations right. What it doesn’t capture is that an organization’s process has nuance: it is really a family of processes that vary with the change at hand, not a single flow every change travels through.

The spec-driven wave

The playbook is part of a broader wave of spec-driven development tooling, including Amazon’s Kiro and GitHub’s Spec Kit. The tools share a common shape. Written artifacts drive the work: an intent document becomes a spec, a plan, a diff, and review findings, all committed to version control. Policy is enforced by deterministic mechanisms such as hooks, rather than by instructions in a prompt. Agents check their own work before a human sees it. Humans own the approvals.

But each of these tools also prescribes a particular process: a fixed sequence of stages that produce fixed artifacts, and that every change travels through. Adopting the tool means adopting its process.

One organization runs many processes

No real organization runs a single process. The right process for a change depends on the risk it carries and the accountability it requires. A documentation fix, a dependency upgrade, and a schema migration in a payments service should not travel the same path. They need different levels of verification, different approvers, and different records. In regulated domains, the process itself is part of the compliance obligation: auditors expect a record of who approved each change and based on what evidence. What must be recorded differs by the type of change.

When a tool prescribes one process, teams route around it for changes that don’t fit, which is the worst outcome because the real process becomes invisible.

When a tool prescribes one process, teams route around it for changes that don’t fit, which is the worst outcome because the real process becomes invisible. Or the vendor keeps adding configuration until the tool becomes a workflow engine that nobody fully understands.

The tool should not prescribe a process. It should give the organization a way to define its own.

Processes as state machines

A better model is to define each process as a state machine. The states are facts about a change: reviewed, validated against its dependencies, approved for production. Those facts live in systems no single tool owns: the repository, CI, the cluster, the tracker. So a process cannot be a program that executes steps. It is a set of rules that react to observations about those systems. Each rule specifies:

  1. The facts it requires before it can fire.
  2. Its gate: fire automatically, or wait for a person’s approval.
  3. The permission it grants when it fires, such as merging or deploying.

The definition is this set of rules, stored as data and reviewed like code. An organization runs many small machines, one per risk class.

At runtime, this behaves nothing like a workflow engine. No component tracks “we are on step four”: the process advances when a fact appears in the system that owns it, and rules react. Events that arrive late, twice, or after a restart are handled like any other, because rules only react to current state. The gate is one of a rule’s conditions, so you can hold firing during an incident or a release freeze without editing any definition.

Gates need enforcement. A gate implemented as a prompt instruction depends on the model following it. The agent harness can provide the determinism required to run the state machine and enforce its gates, stopping the agent between actions until a gate is answered, while the infrastructure enforces the rest.

The process adapts to the change

One fixed process definition per repository is not enough: every change in that repository would still travel the same path, regardless of its risk. The path a change takes should depend on what the change is, and this routing comes from classifying the change, not from the author choosing a path. The organization defines classification using signals it already has: the paths a change touches, the repository it lives in, a label on its tracking issue, etc.

The definitions themselves also need to change over time, and that has to be safe. Because a process definition is data, editing it is itself a change, and it goes through its own gated process. Loosening an approval gate on the release process gets reviewed the way a schema migration does, not edited the way a config file does.

Click to enlarge graphic.

Concretely, consider three changes to the same service:

  • A documentation fix is classified by the paths it touches. Its process has two states: the build passes, and it merges. No person is involved.
  • A dependency upgrade skips design review, but its process requires compatibility evidence: the upgraded service runs its integration tests against real dependencies. A major version bump adds an approval that a patch bump does not.
  • A schema migration in the payments service is classified by the component it touches, no matter what kind of change it claims to be. Its process adds states the others never see: review by a payments owner, validation against production-shaped data, and a release approval from someone accountable for that domain.

Each transition, in each path, is logged with who approved it and on what evidence.

Tenets

The tenets these processes should follow:

  • Autonomy is granted per action, and grows over time. Each transition is set to fire automatically, require approval, or hold. As agents prove reliable on a class of change, that setting is relaxed, so the process absorbs agent improvements without redesign.
  • Human attention is spent only where judgment is needed. Agent effort keeps getting cheaper; supervision hours do not. A person is brought into the loop only when the decision requires human judgment, and is given the context to decide quickly.
  • Evidence comes from outside the agent. An agent’s own report never moves a change forward. Transitions fire on facts from systems the agent cannot write to, such as test results and validation in a realistic environment.
  • The process record is the audit trail. The definition is the written policy, and the transition log shows who approved each step, on what evidence, under which version of the policy.

Quality at scale

The playbook and its peers get the foundations right. What is missing is the ability for an organization to define its own processes, vary them by the risk of each change, and evolve them safely. The goal is not fewer humans in the loop. It is spending human judgment only where it is needed, backed by evidence agents cannot produce about themselves, so that quality holds while throughput multiplies.

We are building these ideas at Signadot and acting as our own guinea pigs, running our own development through this process. If you’re experimenting with these ideas too, we’d love to talk!

The post The AI-native SDLC won’t be one process  appeared first on The New Stack.

❌