❌

Vue lecture

Query decomposition doesn’t fix context starvation — it just moves it

Abstract digital art showing warped light lines surrounding a void, illustrating data compression and AI context starvation.

I have a small example that would best communicate the message I am trying to convey: say you built a chat widget for GitLab’s public documentation (the corpus we are experimenting with in this article) and one of the developers sends this kind of message:

We got an email saying our card was declined for something called “quarterly reconciliation” and I need to know what actually happens now. On top of that, I think we’ve gone over our seat count; there are more people in the group than seats we bought. Our CI has been queuing all week and I want to know whether the compute minutes we purchased last month rolled over or if we lose them. Our finance lead also needs to be the one who gets the invoices from now on, not me. And last thing, is the REST API rate limited? We’re building an internal dashboard and would rather find out now than after it breaks.

Five separate asks: the declined payment, the seat overage, compute-minute rollover, changing who receives invoices, and API rate limits. Each one is answered by a specific passage in GitLab’s public documentation, and you labeled which passage answers which before running anything, so you knew in advance exactly what a correct system needed to find.

Then you run the message through a pipeline that follows current best practice. It splits the query into five clean sub-queries, retrieves for each one independently, merges the results, drops near-duplicates, reranks the merged pool against the original message, and packs the highest-scoring passages into a 2,000-token context.

The pipeline retrieved all five correct passages, but only one of them survived into the packed context; that is one of five asks, not one of five sentences. The packer found, scored, and threw away the other four before the model ever saw them. The same message with no decomposition at all managed three out of five.

The failure has a name, and it isn’t the one you’re thinking of

I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.

The definition is deliberately about allocation rather than retrieval, because allocation is the part nobody watches. If the passage answering the fifth question was found, scored, and then squeezed out by three passages about the first question, the fifth sub-intent is starved, and every recall metric you have will report that the system worked perfectly.

“I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.”

Two failure modes already in circulation describe something different, and it’s worth separating them cleanly:

Semantic dilution happens at retrieval time: when you embed a five-part question as a single vector, you get a centroid that sits somewhere between five topics and lands close to none of them, so the evidence is never found. Decomposition fixes this, which is why it spread.

Context poisoning is about what is present, not what is missing. Wrong, stale, or adversarial content enters the window and corrupts what the model generates downstream. Poisoning is a contamination problem. Starvation is an absence problem, and policy, not accident, produces the absence.

I borrowed the word from operating systems. In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work. Every ingredient of that situation is present in a retrieval pipeline: a fixed resource, competing demands, and a policy that decides who gets served. A relevance-greedy packer is priority scheduling with no aging term, and under priority scheduling without aging, valid low-priority work waits forever.

Decomposition is the right fix to the wrong half of the problem

Split that message into five single-intent queries, and each one embeds cleanly, so per-sub-query recall climbs sharply. This is well-trodden ground. LlamaIndex ships a SubQuestionQueryEngine that breaks a complex query into sub-questions and synthesizes the responses. LangChain’s MultiQueryRetriever generates query variants and returns the unique union of what they retrieve. RAG-Fusion applies reciprocal rank fusion across the per-query result lists. The technique works, and it isn’t mine.

“In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work.”

The context window did not grow. Let me explain: after decomposition, you have n result sets competing for one fixed token budget, and something downstream has to decide the split. In most production pipelines, that something is a short, unremarkable sequence: merge the pools, drop near-duplicates, rerank the merged pool against the original query, then greedily fill until the budget closes.

That sequence is a scheduler. It has a priority function, which is the reranker score, and it has no fairness constraint of any kind. A sub-intent with three strongly-scoring passages takes three slots. A sub-intent whose single correct passage scores mid-pack takes none of them.

So the failure did not go away. It moved from the embedding, where it has a name and people watch for it, into the packer, where it has neither. It also moved somewhere with much worse instrumentation, because recall@k per sub-query is the metric decomposition usually gets validated with, and that number goes up. It goes up at the same time as coverage inside the packed context goes down. You ship on a green dashboard.

The harness

The corpus, GitLab’s public documentation: 10,000 chunks and 2.2M tokens, split on heading boundaries and capped at 480 tokens each. Sixty-one single-intent questions span nine topics, from seat management to rate limits, each labeled with the one passage that answers it. I built multi-intent queries by concatenating those questions while varying n across 2, 3, 5, and 7, randomizing the order so position doesn’t confound topic, and varying topical distance so half the queries draw everything from one topic and half span distinct ones. That produces 100 queries, 25 at each value of n, whose correct decomposition I know exactly.

Two decisions matter more than the rest:

The metric is not recall. Recall tells you what the retriever found. What I need is what survived into the packed context, per sub-intent. So I log each sub-intent twice: once for whether its correct passage reached the candidate pool, and once for whether it reached the packed context. The gap between those two numbers is the entire argument.

Every question has to be retrievable on its own before it’s allowed in. A question enters only if its correct passage ranks in the top 10 for its own isolated query, under both retriever configurations, and both scored 100% recall@10 on that test. Since each sub-query’s candidate pool is exactly its own top 10, passing that gate guarantees the correct passage sits in the pool for every decomposed arm. Any sub-intent that then fails to appear was denied by the packer rather than missed by the retriever, which removes the most obvious objection to everything below.

The core measurement uses no language model. Because queries are composed from known sub-questions, the decomposer is an oracle so that anyone can reproduce the main result with no API key.

That invites an objection, so I tested it. A real LLM decomposer, blind to n, disagreed with my ground truth on 41% of the queries, and on inspection it was right every time. Five of my sixty-one supposedly single-intent questions contain two distinct information needs. What is excess storage usage, and what happens when we go over the free limit? is two questions wearing one question mark. Adjusted for those five, agreement is 100 out of 100. That is not evidence decomposers are reliable, because my queries are joined by fixed connectives and splitting on those alone recovers n perfectly, which real messages never allow. What it caught was an error in my own labels, and that is the best argument I have for the oracle design.

Results

Every arm runs at 2,000, 4,000, and 8,000 tokens against two retriever configurations. The stronger pairs are BAAI/bge-base-en-v1.5 with BAAI/bge-reranker-base; the weaker pairs are a quantized BAAI/bge-small-en-v1.5 with Xenova/ms-marco-MiniLM-L-6-v2. Retrieval is in-memory cosine similarity over a NumPy array because, at 10,000 chunks, a vector database would be slower to write, slower to run, and harder to verify.

Sub-intent coverage at a 2,000-token budget on the stronger configuration:

Table showing coverage (in %) per arm.

The production-default pipeline starves 31.1% of sub-intents whose evidence it had already retrieved. It beats no decomposition by nine points, while a flat B/n split, which is the crudest allocator anyone could write, beats it by fourteen.

Floors work, but not the obvious floor. Reserving one passage per sub-intent before the greedy fill satisfied 99.4% of its reservations and bought only seven points. The mechanism fires correctly and reserves the wrong passage because it picks each sub-intent’s best chunk by score against the original query. The original query asks about all five intents at once. Selecting that same reservation by score against its own sub-query pushes coverage to 89.4% and cuts allocation starvation from 30.6% to 10.1%. That is a change of about four lines.

Two of my own recommendations died here. I expected reranking against the original query to beat reranking against the fragment, and it loses by seventeen points. The incomparable score scales I worried about turn out to help, because each sub-query’s best match ends up at the top of its own scale, producing per-intent fairness for free. I also expected deduplication before allocation to matter, but near-duplicates consume 1.0% of the budget and removing them moves coverage by 0.3 points.

The crossover: starvation by tokens-per-sub-intent, which is simply the budget divided by n:

Table showing tokens per sub-intent.

Below roughly 1,000 tokens per sub-intent, allocation policy dominates. Above it, nothing you do to the allocator matters, because everything fits anyway.

The retriever comparison is the one I’d lead with. Upgrading the retriever moves coverage on the production-default arm from 56.2% to 68.9%, a gain of 12.7 points. Changing the allocation policy on the same retriever moves it from 68.9% to 89.4%, a gain of 20.5 points. In its sharpest form: the weaker retriever with a fragment-scored floor reaches 84.5%, and beats the stronger retriever with a greedy packer at 68.9% by sixteen points. A worse retriever with a better allocator wins.

Position: Held within a fixed n so query difficulty doesn’t contaminate the comparison; starvation across the seven positions of an n=7 query runs 1.3%, 21.3%, 30.7%, 48.0%, 48.0%, 32.0%, and 13.3%. That is a serial-position curve. The packer protects what you asked for first, protects what you asked for last a little less, and drops the middle. The fragment-scored floor flattens it to 4.0%, 9.3%, 6.7%, 8.0%, 22.7%, 5.3%, and 6.7%.

“A worse retriever with a better allocator wins.”

Topical distance: Sub-intents that span distinct topics starve about twice as often as sub-intents drawn from one topic, at 38.1% against 18.3% for n=7. I predicted the opposite. A topically coherent query gives the reranker a coherent target, and it scores all the correct passages similarly. In contrast, a scattered query lets it latch onto some topics and abandon others.

What the user actually sees

Everything above is retrieval-side. What decides whether any of it matters is what reaches the person who wrote the message, so I generated real support replies from 80 packed contexts and had every reply graded per sub-intent, with both the generation and the grading blind to which arm produced which context.

When the correct evidence reached the packed context, the reply addressed that question 100% of the time, across 261 out of 261 cases, in both arms. Coverage predicts the generated outcome exactly, which is the strongest justification I have for measuring it.

When a sub-intent was starved, the reply answered it anyway 48.1% of the time, based on whatever else happened to be in the window. It explicitly flagged the gap 45.6% of the time, with some version of “I’ll follow up on that separately.” It went silent only 6.3% of the time.

“Starvation mostly does not produce silence; it produces unsupported answers.”

I expected silence, and I was wrong. Starvation mostly does not produce silence; it produces unsupported answers. Whether those answers are actually incorrect is the next experiment, because this harness measures whether a question was addressed, not whether the answer was right.

Some limits: composed queries are cleaner than real support messages, which carry pronouns, implicit context, and conditional clauses. This is one corpus and one embedding family. I drafted the gold labels with model assistance and verified them myself. The model writing those replies was strong, so a cheaper production model would plausibly flag fewer gaps and invent more.

What an allocator actually looks like

Give every sub-intent a floor, and choose it by fragment score. Not the naive floor, which satisfies 99% of its reservations and buys seven points. Select a reservation by relevance to the sub-intent it protects, not by relevance to the message as a whole.

Rerank against the fragment rather than the original query. This inverts what I expected and what I have seen recommended. Scores from different fragments are not comparable across sub-intents, and that incomparability is doing useful work.

Don’t spend your effort on deduplication; near-duplicates cost 1.0% of the budget here. Dedup is worth doing, but it isn’t why your fifth question went unanswered, and treating it as the fix will cost you weeks.

Log per-sub-intent coverage: You already computed it to pack, and it predicts the generated outcome perfectly. A sub-intent that received zero passages is the best predictor available that your reply is about to assert something you cannot support.

Where parallel decomposition breaks

“If it’s late can I get a refund” is one clause and two intents, and the second one’s retrieval target depends on the first one’s answer. Parallel decomposition treats them as siblings. It retrieves the late-delivery policy and the general refund policy, packs both, and misses that the passage you actually need covers refunds for late delivery, which may match neither sub-query particularly well.

There are two ways out: You can tag dependencies at decomposition time, or run a deferred second pass that re-retrieves conditional clauses once the first round resolves.

I would take dependency tagging, for three reasons: A second pass costs a full retrieval round trip inside a latency budget a support bot does not have. The tag is reusable, because a dependent sub-intent should not hold a floor reservation. At the same time, its parent is unsatisfied, so it feeds the allocator directly instead of bolting on a separate mechanism. And it fails visibly, since an untagged dependency shows up as a starved sub-intent in the coverage signal. In contrast, a deferred pass that resolves the wrong condition produces a confident wrong answer with nothing to flag it.

The cost is real; dependency tagging pushes work onto the decomposer, which is already the weakest component in the chain, and I have not measured tagged against untagged. That is a design position rather than a result, and it is the one thing here I am asking you to take on argument instead of evidence.

What to measure on Monday

Take your production pipeline and compute one number: your context budget divided by the average count of distinct questions per incoming message. If that number lands below roughly 1,000 tokens, your allocation policy costs more than your retriever does, and the reranker upgrade sitting in your backlog will buy you less than reserving one slot per question.

On my corpus, the retriever upgrade was worth 12.7 points of coverage, and the allocation change was worth 20.5, which is why I think the ordering is wrong in most pipelines I’ve seen. That ordering is the falsifiable part. Run the same two comparisons against your own corpus, and if the retriever wins, I want to see the numbers, because that result would tell me the crossover sits somewhere other than where I measured it.

The cheaper thing to do first takes an afternoon. Log, for every multi-intent request, how many sub-intents ended up with zero passages in the packed context. A support system that cannot tell you which question it dropped will keep answering that question anyway, about half the time, out of whatever else was in the window.

The post Query decomposition doesn’t fix context starvation — it just moves it appeared first on The New Stack.

  •  

Why human oversight is shifting from writing code to defining requirements

Abstract dark digital wave featuring glowing cyan microchip circuit patterns and network lines representing AI system architecture.

This walks through the pipeline our agents operate inside—from a recorded scoping meeting through unit specs, spec review, generated code, PR checks, and automated QA, out to a weekly Thursday release. Then it shows the hole in that pipeline. Every control answers one question: Does the code conform to its instructions? The instruction itself never goes on trial. I planted a single bad requirement in a small unit, let the implementation and tests generate from it, and watched six passing tests, a traceability gate, and a clean run certify a system that broke its own stated outcome.

A sentence that passed every review

Here is the shape of a requirement I read earlier this year, with the domain stripped out:

If the classification lookup returns no determination, treat the record as permitted and proceed, so that an unavailable dependency doesn’t block delivery.

Read it the way a reviewer would. It names a real operational worry. It offers a justification. It sounds like an engineer weighed availability against correctness and made a call. In a 40-page document, you would skim past it in two seconds.

“Every control answers one question: Does the code conform to its instructions? The instruction itself never goes on trial.”

The feature it belonged to existed to guarantee that one particular class of record never gets processed that way. So, for the exact population nobody manually tests, that sentence executes the failure the feature was built to stop.

Nothing downstream would have caught it. That is the point worth sitting with, because “nothing downstream” covers every guardrail we have spent two years adding.

How work actually reaches an agent

It is worth walking the pipeline that requirement sat in, because most of the work happens before an AI agent sees anything.

A feature starts as a recorded meeting. Not a kickoff limited to the code-owning team, but a room holding every team the change touches—which, for anything crossing a shared service, is four or five groups who would otherwise meet at integration. A product manager walks through the intent. Everyone argues. When a contentious issue resolves, somebody states the resolution out loud, deliberately, for the recording.

That transcript—not the requirements document that preceded it—generates the scoping document.

The scoping document splits the feature into release groups, and each group into numbered units (roughly one per shippable slice). Every unit carries its purpose, explicit scope boundaries, functional requirements, architectural layers, dependencies, feature flags, and acceptance criteria written in given/when/then format. It also carries two critical elements I hadn’t seen in requirements artifacts before, which I will return to later.

Product managers review the scoping document. Once signed off, it becomes the source of truth, and the initial requirements document becomes history. That demotion does real operational work; it’s why this story has a happy ending rather than an incident report.

The document syncs into the issue tracker, mapping one work item per unit, and those get assigned out.

When a developer picks up a unit, they generate a unit spec from the scoping material. This is the concrete layer: named interfaces, method signatures, files to create or modify, an error-handling matrix, the step-by-step query flow, and a list of what the unit deliberately will not do. While writing this, the developer routes open questions back to a human instead of letting whoever holds the keyboard guess the answer. On one unit, a question about a base class constructor revealed that the design document specified a call that would not compile. The system found a design defect before any code existed to review.

Next, developers review the spec against the scoping document—not for style, but to verify that every requirement is covered, that units haven’t quietly duplicated work, and that nothing dropped during translation.

Only now does an agent write code.

The agent’s output must cover the entire call path—from entry point down through the service layer — with unit and integration tests. Partial coverage of generated code is worse than zero coverage, because it falsely signals that a human thought about the untested paths.

Then comes the familiar part: a local standards pass, a pull request, automated reviewers leaving comments, a pipeline that blocks changes disagreeing with the spec, ephemeral environments, and a second developer’s approval.

QA follows the same tooling. A tool points at the release bucket in the tracker, reads every ticket, and drafts test cases. QA reviews every generated case by hand before it counts. Release validation gates the release.

And the release goes out on Thursday. Every Thursday, whatever is ready ships.

By most measures, this pipeline works. But notice where the human decisions actually sit:

Where intent is decided, and where it is only checked.

Recorded scoping meeting: teams agree on intent and scope

several people, on the record │
                              ▼
                ┌──────────────────────┐ outcomes, scope boundaries, what is
                │   scoping document   │ deliberately undecided and who owns it,
                └──────────┬───────────┘ and where the written brief lost
                           │
                           ▼
                ┌──────────────────────┐ open questions go back to humans here.
                │    spec, per unit    │ this is the last point anything is decided
                └──────────┬───────────┘
                           │
                ┌───────┴────────┐
                ▼                ▼
           ┌──────┐         ┌───────┐ both written from the same criteria,
           │ code │         │ tests │ so they agree with each other no matter
           └──┬───┘         └───┬───┘ what the criteria say
              │                 │
              └───────┬────────┘
                      ▼
 ┌──────────────────────────────────────────────┐
 │   standards check, automated PR review,      │ all of these compare
 │   a second developer, the test run,          │ an artifact against
 │   spec conformance in the pipeline           │ the spec. the spec
 └──────────────────┬───────────────────────────┘ itself is never the
                    │                             thing on trial
                    ▼
                 release

Everything in that bottom box is downstream of the spec. That works perfectly—until the spec is wrong.

Building the failure so you can watch it

I rebuilt this failure pattern in a system small enough to demonstrate. The system is a consent-aware notification dispatcher with six acceptance criteria, written exactly like our production specs. The core outcome sits at the top: A notification is never delivered to a recipient who has withdrawn consent.

AC-05 is the planted defect:

AC-05: Given the consent lookup returns no determination, when a notification is dispatched, then the recipient is treated as having granted consent, and the notification is sent, so that an unavailable lookup does not block delivery.

The implementation does exactly what it was told:

elif consent is Consent.UNDETERMINED:
    # AC-05. Treat an unresolved lookup as granted so delivery is not
    # blocked by an unavailable dependency.
    entry = AuditEntry(recipient, consent, "undetermined-default-send", sent=True)

And the test descends from the identical criterion:

Python
@pytest.mark.criterion("AC-05")
def test_undetermined_consent_defaults_to_sending():
    entry = Notifier(fixed(Consent.UNDETERMINED)).dispatch("alan")
    assert entry.sent is True
    assert entry.rule == "undetermined-default-send"

That test passes, and it should. It correctly tests a wrong rule. No version of it will ever fail, because the criterion it checks is the bug itself.

On Python 3.14.3 with pytest 9.1.1, the suite is entirely green:

Bash
$ python -m pytest -q
......
[100%]
6 passed in 0.01s
exit: 0

$ python trace.py spec.md test_dispatch.py
ok AC-01 test_dispatch.py::test_granted_recipient_is_sent_to
ok AC-02 test_dispatch.py::test_withdrawn_recipient_is_not_sent_to
ok AC-03 test_dispatch.py::test_override_cannot_force_a_send_to_withdrawn
ok AC-04 test_dispatch.py::test_audit_entry_records_the_decision
ok AC-05 test_dispatch.py::test_undetermined_consent_defaults_to_sending
ok AC-06 test_dispatch.py::test_lookup_failure_propagates_and_writes_no_audit
all 6 criteria claimed
exit: 0

Six tests, six criteria, full coverage, zero warnings. Here is what that green light actually certifies. The script below asks the consent store who withdrew, then dispatches a notification to them:

recipients who withdrew consent: grace
-- consent service unreachable for grace --
ada observed=granted rule=granted-send sent=True
grace observed=undetermined rule=undetermined-default-send sent=True

delivered to grace, who withdrew consent. the dispatcher observed undetermined.
outcome violations: 1

Grace withdrew consent. The database confirmed it. But she received the notification, and every automated gate signed off because none of them evaluated the sentence saying she shouldn’t.

Why they all miss it

Look back at the fork in the diagram. It explains everything.

Code and tests both descend from the criteria. They match each other by construction. Their agreement tells you absolutely nothing about whether the criteria were correct. Every gate below the fork measures an artifact against the spec. Because the spec sits upstream of all of them, it represents the final point where human decision-making influences the outcome.

Engineering teams used to survive this. A developer reading a requirement would form an opinion, and the code would pass through a senior engineer who knew that a specific fallback behavior would trigger a 3 a.m. page. Today, that senior engineer reviews a diff. And the diff is technically correct.

Being fair to the guardrails

Before proceeding, I must clarify what each guardrail actually does. Stating “none of them caught it” sounds like a dismissal; it is not. Every control earns its place, and I would fight to keep all of them.

The standards pass catches convention drift, which matters immensely with generated code because an AI agent will cheerfully invent a third way to execute logic the codebase already handles two ways. Automated PR reviewers find legitimate defects, including edge cases humans skim past at four in the afternoon. The second developer catches poorly expressed intent—a task humans still perform better than tools. Tests catch regressions against established behavior. Ephemeral environments catch integration breaks.

“Point all of them at a flawed spec, and they will agree with each other flawlessly, because no component in the system holds a dissenting opinion.”

The spec conformance check is the strongest guardrail. It prevents the failure everyone truly fears: an agent quietly building more than it was asked to build. Nobody on our team loses sleep over an agent inventing an unwanted endpoint, thanks to this check. The generated QA cases perform similar work at the other end of the pipeline, covering the tedious paths a tired reviewer skips—which is exactly where bugs hide.

Line them up and look for the shared trait:

  • A standards pass compares code against a convention.
  • Reviewers compare a diff against the spec.
  • Tests compare behavior against criteria.
  • Conformance checking compares the change against the spec by design.
  • QA cases stem from tickets the spec generates.

Every guardrail takes the spec as its input. Point all of them at a flawed spec, and they will agree with each other flawlessly, because no component in the system holds a dissenting opinion. This is not a flaw in any individual guardrail; it is an architectural property of the set.

The toolkit missing from the industry

This brings us back to the two sections of the scoping document. Neither exists in any spec-driven toolkit I have reviewed, yet they are the exact reason the bad fallback never reached an agent in production.

The first section lists what is deliberately undecided, pairing each item with an owner. It acts as a fence rather than a to-do list. It signals that an issue remains open, no one has ruled on it, and any agent that quietly resolves it has overstepped.

The second section records where the written requirements document was lost. It maps one row per disagreement between the pre-meeting document and the room’s final conclusion, documenting the ruling and the reasoning. That is where the bad fallback died. Someone stated aloud that defaulting to “permitted” would destroy the system’s core guarantee; the room agreed, and the row was committed to the record.

“When a long technical argument finally resolves, force someone to state the resolution out loud, deliberately aiming at the transcript. Do not aim it at the humans in the room; aim it at the system that will read the transcript next week.”

If you adopt only one habit from this article, steal this: When a long technical argument finally resolves, force someone to state the resolution out loud, deliberately aiming at the transcript. Do not aim it at the humans in the room; aim it at the system that will read the transcript next week. It feels absurd in the moment, but it is the most valuable 30 seconds of the meeting. A decision living exclusively in six people’s memories is a decision the agent will hallucinate later.

Making the criteria the review surface

None of this survives contact with reality unless the criteria remain machine-checkable. The traceability gate does exactly one job: verifying that every criterion has a test claiming it, and every claim points to an existing criterion.

@pytest.mark.criterion("AC-01")
def test_valid_credentials():
    ...

This is not a novel concept. Regulated software has operated this way for years; tooling like jamb executes this exact pattern against the same test framework for IEC 62304 medical device submissions. What changed is the artifact’s job. In a regulatory submission, the matrix satisfies an auditor. In an AI-driven pipeline—where one requirement generates both the implementation and its tests—their agreement is guaranteed by construction and therefore provides zero evidence of correctness. The criteria themselves are the only artifacts left requiring human attention.

My traceability script is 100 lines long, and the first 25 lines are a docstring explaining its limitations. It imports ast, pathlib, re, and sys. It walks the syntax tree rather than grepping text, so a criterion ID buried in a comment doesn’t falsely count as test coverage. On its first run, it flagged AC-04 (the audit criterion) as lacking a test. I had written those tests myself and simply forgotten one. The check took under a second.

The two-line fix nobody makes

That first run also produced this warning five times—once for each correctly spelled marker:

PytestUnknownMarkWarning: Unknown pytest.mark.criterion - is this a typo?  You can register custom marks to avoid this warning

Pytest asks whether the correct spelling is a mistake. When a genuine typo occurs, the resulting warning disappears into a pile of identical warnings about perfectly valid markers. The signal-to-noise ratio renders the warning useless.

Registering the marker and enabling –strict-markers transforms the failure mode entirely:

ERROR collecting test_typo_mark.py
'criterio' not found in `markers` configuration option

This throws a collection error (exit code 2) and halts the run. A traceability convention relying on an unregistered marker is mere decoration. Two lines of configuration make it load-bearing. Yet, developers rarely write those lines because the default behavior is a warning, and warnings just scroll past.

Always verify your configuration actually enforces the rule. Pytest issue #14442 revealed that strictness set through addopts quietly stopped functioning across the 9.0 series. Errors silently became warnings, test suites turned green, and nothing announced the degradation. The bug is fixed now, but it illustrates the core argument one layer deeper: A guardrail that silently downgrades itself to a warning is worse than having no guardrail at all, because it displays a green checkmark where a hard block used to sit.

If you adopt this traceability pattern, write a meta-test asserting that your gate fails when it should.

The engineering practices that carry the weight

None of what I have described is a tool you can simply npm install. That is the reality I would have hated reading two years ago, but it is the entire answer today. What transformed agents from a novelty that writes code quickly into a system that ships reliable software on schedule was a strict set of engineering practices. Fortunately, they are portable.

  • Put the durable record where multiple people made it. A requirements document reflects one person’s understanding at one specific moment. A room resolving a disagreement is a fact; stating that resolution out loud makes it survive. That practice killed the bad fallback. It costs 30 seconds of feeling slightly silly.
  • Document what you decided not to decide—and assign an owner. An engineer reading an underspecified requirement asks a question. An agent fills the gap with a hallucination and keeps moving. A named open question transforms an agent’s guess into an explicit human responsibility.
  • Record where your written brief lost the argument. The disagreement table is a strange artifact to maintain, but it is the highest-value page in the scoping document. It is the only place where the delta between the initial draft and the team’s final conclusion remains visible to anyone not in the room.
  • Review the spec against its source material before writing code. Engineering effort spent reviewing a diff happens after the expensive architectural decisions are baked in. Effort spent reviewing a spec is the only review step that can fundamentally alter what gets built.
  • Make the criteria machine-checkable. Register your test markers so the convention enforces itself, and write a test asserting that your gate fails when it should. Reliability stems from that final rule. A pipeline of guardrails you have never actually seen fail provides zero evidence of safety.

Weekly releases are the byproduct of these practices, not a metric we arbitrarily targeted. 

Shipping every Thursday works strictly because the architectural arguments happened in week one, on the record, in front of everyone the change touched.

What I argue against

Letting the agent ask when it feels unsure. We do this during the spec phase, and it adds value. 

But it cannot serve as the primary control. Researchers Su and Cardie at Cornell ran 10 models over 1,000 ambiguous questions. When explicitly asked to judge ambiguity, the models succeeded 60% to 80% of the time. When left to respond naturally, the models returned definitive answers over 95% of the time. The Claude family flagged roughly one ambiguous question in 20. AI models do not fail at spotting ambiguity; they fail at doing anything about it. 

Strangely, supplying retrieved context drove the clarification rate down. A fatter, better-organized spec buys confidence, not caution.

Allowing agents to interrogate a proxy holding full issue text. A Carnegie Mellon group measured this approach. Agents allowed to query a proxy clawed back 80% of their fully specified score at best (dropping to 54% for weaker models). A fifth of the context remained lost. More importantly, their proxy always knew the correct answer—which is precisely the condition that fails when the requirement itself is the defect.

Adding another reviewer, human or otherwise. Adding reviewers to the same layer answers the same flawed question. When the requirement is wrong, reviewer two simply concurs with reviewer one.

Assuming this process is unnecessary overhead. This is the objection with the best evidence behind it. Colin Eberhardt at Scott Logic rebuilt a feature using a full spec-driven workflow and found it took him roughly 10 times longer than his normal approach (spending three and a half hours reviewing 2,577 lines of Markdown to yield 689 lines of code). He is right about his specific case; we skip this entire process for simple two-line fixes.

But look at where his specs came from: He prompted for them. Whatever went in, he supplied. A document grown entirely from one person’s assumptions will always agree with that person. 

No quantity of generated text will expose an error in the foundational assumptions. The process earns its cost when multiple engineers arrive with incompatible ideas and are forced to settle them where a transcript can capture the logic. Remove the human conflict, and you have just built an expensive mirror.

What this doesn’t do

The traceability gate confirms a test declares it asserts a criterion. Whether the test actually asserts the criterion is beyond its capability.

I wrote that caveat down, then realized I needed to test it—because putting an unchecked claim about your own tooling in writing replicates the exact failure this article targets. I created six tests whose entire body was assert True, mapped one per criterion, and ran the suite:

$ python -m pytest test_hollow.py -q
......
[100%]
6 passed

$ python trace.py spec.md test_hollow.py
all 6 criteria claimed
trace exit: 0

Both gates were fully satisfied by a file that tested absolutely nothing.

Where this leaves the review

This workflow relocates the review; it does not retire it. Instead of reading a 600-line diff, an engineer reads six numbered claims and verifies each has credible logic underneath. It is a smaller, more tractable job, and that is the honest extent of what this process buys you.

“Get the document wrong, and they will all agree with you, at blinding speed, forever.”

We ship on a weekly train now, and agents write most of the code. The system holds together because we moved costly human attention to the only place it still changes the outcome: settling what the software should do, and recording exactly who settled it.

Everything past that point is just a machine checking another machine against a document. Get the document wrong, and they will all agree with you, at blinding speed, forever.

The post Why human oversight is shifting from writing code to defining requirements appeared first on The New Stack.

  •  

47,000 job listings reveal the engineering roles that AI is creating

Abstract overlapping circles in black, green, orange, and pale yellow on a cream background.

Every major transformation in tech has led to roles merging, then new ones emerging. Friction between developers and operations drove the creation of the DevOps engineer. Then, when security needed to be considered throughout the delivery pipeline, DevSecOps emerged.

The team beyond the AI-native talent and services platform Andela analyzed 47,000 recent engineering job postings from Fortune 500 companies. This research, released on Thursday, uncovered more than 2,000 skills that pour into 23 emerging job titles. None of these are coming out of nowhere; they strategically merge existing skill sets to create new roles. 

Among 1,832 postings titled primarily for AI or ML engineers, 53% contained at least two skills drawn from different established roles, Andela finds.

In today’s tighter economy and amid AI, companies seem to be going one of three ways. They are lumping too much work and required experience into now-nebulous AI engineer or machine learning (ML) engineer job titles. They might be looking to replace tech workers with AI. But more forward-thinking organizations are reworking job titles and descriptions to reflect the demands of getting AI safely and efficiently through the software delivery lifecycle. 

Cory Hymel, head of research at Andela, tells The New Stack, “If you’re going to look to deploy AI within your organization, the way to look at it is that an AI has a certain set of skills, and then a human has a certain set of skills.

“If you Venn diagram those and see where they cross over, an AI should do the skills it can. But the human circle is still exponentially larger than that of AI.”

“When you’re looking to deploy AI, it’s not about trying to replace that human circle with an AI one. It’s about what certain skills you need to carve out and delegate to it.”

Read on for the top engineering jobs that are emerging because of AI, how to attract tech talent for them, and what you need to focus on to get a tech job in this tough market.

Click image to enlarge.

AI is not serving the generalist. Specialization is still key.

Citing the leading AI CEOs, Hymel remarks, “You’ve heard from the AI salespeople of the world that AI is going to push people to be more generalist, and the data that we found here doesn’t necessarily support it.”

“You’ve heard from the AI salespeople of the world that AI is going to push people to be more generalist, and the data that we found here doesn’t necessarily support it.”

Overall, they found that these emerging job titles aren’t generalist at all. These emerging roles bridge skill sets from several existing ones, but each addresses a specific operational or product need, some tied to AI adoption. 

The top five new engineering job roles discovered are:

  1. MLOps pipeline engineer, who builds and runs the automated infrastructure to deploy, version, and monitor machine-learning models in production, with 46% ML engineer, 23% DevOps engineer, 15% data engineer skills, and 8% each AI engineer and data scientist roles.
  2. LLM application engineer, who builds on and evaluates foundational models via large language model application and conversation systems, bringing 48% AI engineer and 34% ML engineer, with a touch of product designer, software architect, and embedded software engineer roles.
  3. FinOps reliability engineer runs cloud infrastructure for both reliability and cost, bridging 36% DevOps engineer, 27% site reliability engineer (SRE), 18% cloud engineer, and 9% each DevSecOps engineer and cloud solutions architect.
  4. Docs-as-Code engineer applies program management and DevOps engineering skills to the traditional technical writer’s role, pivoting from stagnant docs to specification-as-code.
  5. Product frontend engineer is about a third traditional frontend engineer and a third product manager, with a touch of full-stack engineer, UX researcher, and product designer.

“If you’re a DevOps engineer, historically, your skill bundle might have allocated 30 to 40% of pure DevOps-required skills that are rich and specific to that role, and you have a remaining bundle that is cross-role habitable, meaning that those skills would translate between DevOps or to an engineer or to a technical product manager,” Hymel explains. “Some of those skills can now be replaced with AI, which means that those skills that are more directly focused on your role become more important than ever.” 

So-called “soft” business skills are also increasingly crucial, he contends. However, he seriously doubts anyone will ever be able to slide between finance, marketing, engineering, and sales roles. 

Where enterprise engineering job descriptions falter

“Job descriptions and resumes right now are the best worst thing that we have. When you’re talking about large enterprises, and you’re having to deal with scale, your hiring process gets farther away from the work,” Hymel explains. 

Especially when the hiring process starts in HR, not engineering, “you’re needing to put language in place that will survive the chain of custody, with the naming of the job [coming from] the engineer that’s closest to the work.”

It’s not uncommon for an enterprise to have 50 different front-end developer job listings, each with very different skill requirements. It’s better for candidates and for fit to be as specific as possible, including embracing new job titles.

This habit of generic job titles used to be positive because it brought in more applicants, but nowadays, with so many engineers on the market, it further dilutes your hiring pool, leaving you with the 100 fastest applicants—who are often AI-generated anyway.

“Any company that has not taken a hard look at revising their job postings and job titles is at an extreme disadvantage because there’s a very high probability that you’re going to end up hiring the wrong person simply because you didn’t take the time to describe the role well enough,” Hymel remarks, which leads to dire consequences. 

“There’s potential churn, so you just spend all this time and cost to go headhunt and find someone. Two, if they do get in there, you have to pay for their ramp time to get up to speed because they were sold a different bill of goods than what was in the description. And then three, it impacts overall roadmaps and timelines because now you might have to replace, and, again, you have to wait for people to get up to speed.”

On top of this, HR and engineering hiring managers alike are using AI to generate job descriptions. It still isn’t recommended to have AI generate something so human and essential to your core success.

Especially in this time of flux, when no one may have the required experience, companies should start job descriptions with what they want the future hire to achieve.

“The cost of code is going nearer to zero.”

“The cost of code is going nearer to zero.” Hymel explains organizations should think more like, “Here are the outcomes that we’re looking for. If you have the soft skills and additional skills around it to get there, whether that is backlog prioritization, being able to be collaborative, having worked on project deployments before, and we don’t necessarily care that you can score a 10 out of 10 on Python anymore.”

Which emerging roles engineers should pursue

The familiar claim that women apply only when they meet every qualification is not well supported; recent research finds that application behavior is more complicated. Still, clearly separating essential qualifications from preferences can reduce ambiguity and unnecessary barriers.

Focusing on outcomes and clearly distinguishing required from preferred skills may broaden the applicant pool, although it does not guarantee greater diversity.

For example, if you’re an engineer who enjoys having a product focus, collaboration, and strategy, Hymel recommends looking toward the new product front-end engineer role, which owns the full user-facing feature lifecycle, from definition to shipping.

“You are required to have more mindshare towards prioritization of features,” he says, shifting away from a ticket person, because “now you have more control because AI allows you to span out a little bit deeper.”

Similarly, AI has the back-end engineer thinking beyond the back-end stack to deployments, scalability, and the reliability of underlying infrastructure systems, giving rise to roles like the polyglot back-end integration engineer. 

Technical writers — reasonably worried about their jobs in the face of AI-generated documentation — should look toward new docs-as-code engineer positions, which add technical program management and DevOps engineering skills.

“If you’re writing the docs, you’re essentially writing the specs that enable spec-driven development. You now have the capability to actually contribute software,” Hymel observes. “And it starts all the way at the top too. If you’re a product manager, you can now start building and contributing code, like a product experience designer.”

Read the full Emergent Role Research. If any of these AI engineering job descriptions ring truer than what you were hired for, we hope it empowers your next conversation with HR or for you to apply for a different job title. 

The post 47,000 job listings reveal the engineering roles that AI is creating appeared first on The New Stack.

  •  

Stop AI code sprawl before it destroys your software design

Dark abstract digital render of curved metallic lines spiraling into a void, representing software architecture boundaries and AI code sprawl.

While AI code generators help teams ship faster than ever, that speed brings a hidden killer: Comprehension Debt. As soon as an AI produces functionally correct code that violates your domain boundaries, the team loses its mental model of the system. Here, I’ll show how to switch from passive documentation to Executable Architecture using Python-based testing tools like pytest-archon and CI/CD pipelines.

The most dangerous thing an AI coding agent can do is generate code that works.

If a junior developer writes poor code, it breaks the build or staging environment. The team catches it, reverts it, and discusses it. But if an AI coding agent produces 500 lines of functionally correct and bug-free code that subtly violates your system’s boundaries, it merges without issues.

“The most dangerous thing an AI coding agent can do is generate code that works.”

Gradually, the AI connects your billing service to the user authentication component. It gives your presentation layer database access. It wires dependencies in a way that works but violates the design assumptions of human developers who maintain the system. 

Technical debt has given way to something far more pressing — Comprehension Debt: the growing gap between how fast code gets written and how well the human team understands its architecture. The problem isn’t messy logic; it’s a lost mental model. That happens when the team no longer knows why the codebase exists.

If you view AI as a mere machine for faster typing, the architecture has started to degrade. To endure in the era of AI-driven coding, architecture enforcement must shift — from documentation to Executable Architecture.

The illusion of documentation

The accepted guidance for AI-assisted development is: “Make better documentation so the AI understands the rules.”

This is a fallacy. Documentation will become obsolete. If your AI agent finds an easier way to reach its objectives by skipping a service layer, it will take it. And since human reviewers increasingly struggle to review thousands of AI-generated pull requests, these detours slip through code review undetected.

“You cannot depend on human beings to detect architectural drift. You have to trust the CI/CD pipeline.”

You cannot depend on human beings to detect architectural drift. You have to trust the CI/CD pipeline. 

If your architectural boundaries matter, check them the same way you’d check any business requirement. We need fitness functions that fail the build when an AI agent violates a boundary condition.

Introducing executable architecture in Python 

In the Java ecosystem, tools such as ArchUnit have traditionally enforced architectural boundaries. In Python, tools like pytest-archon do the same job.

Consider a concrete example. You’ve built a modular monolith for an e-commerce application and established strict boundaries:

  • The Billing domain should never import from the Shipping domain.
  • Domain model code should not import from infrastructure (AWS SDK, SQLAlchemy, etc.).

You task the AI agent with adding shipping cost calculations based on the user’s billing tier. Without thinking about the architecture, the AI imports the Shipping Calculator directly into the billing service. Test passes. The application works. But the architecture fails.

Here’s how pytest-archon prevents the agent from doing that.

Step 1: Install the dependency

First, install the architectural testing dependency.

Python
pip install pytest-archon

Step 2: Define the architectural rules as tests

Instead of finding the rules on the Wiki page, we define them as pytest features. We create a test_architecture.py file in the test folder.

Python
from pytest_archon import archrule

def test_billing_is_isolated_from_shipping():
    """
    Ensure the billing module never imports shipping logic.
    This prevents the AI from creating tight coupling between distinct domains.
    """
    (
        archrule("billing_isolation", comment="Billing must not know about shipping")
        .match("ecommerce.billing*")
        .should_not_import("ecommerce.shipping*")
        .check("ecommerce")
    )

def test_domain_models_are_pure():
    """
    Ensure domain models only depend on standard libraries or pydantic.
    Prevents the AI from leaking infrastructure (DBs, APIs) into the core logic.
    """
    (
        archrule("pure_domain", comment="Domain models must not import infrastructure")
        .match("ecommerce.*.models")
        .should_not_import("sqlalchemy*")
        .should_not_import("boto3*")
        .check("ecommerce")
    )

Step 3: Close the agent feedback loop

Then, once the AI agent pushes its pull request, pytest runs automatically as part of the CI workflow. Regardless of how well the AI agent generates code that calculates the Shipping fee, the build will immediately fail with something similar to this:

text
FAILED tests/test_architecture.py::test_billing_is_isolated_from_shipping -
AssertionError: Rule 'billing_isolation' violated:
ecommerce.billing.invoice imports ecommerce.shipping.calculator

A human reviewer doesn’t have to track down the entire import tree manually. Most importantly, the best engineering teams never rely on humans for this.

Once again, we feed the output of these failing pytest tests directly back into the AI agent’s context window using Aider or custom CI/CD scripts, and the AI can fix architectural problems without human help.

Strategies for avoiding Comprehension Debt

Running architectural tests alone is not enough. Here’s how to shield your team from Comprehension Debt:

1. Hard boundaries vs. soft conventions

AI agent obeys hard constraints but not soft suggestions. Get rid of sloppy folder-based architecture and establish clear module boundaries instead. Use tools like import-linter or pytest-archon to block forbidden imports with physical barriers. The path of least resistance must be the most architecturally sound.

2. Limit automated complexity

Well-defined APIs and boundaries are good, but not enough to let you off the hook for messy, complex implementation. If AI creates spaghetti code in your billing module, causing downtime from race conditions at 3 AM, a human engineer will still need to maintain and understand that codebase.

For this purpose, run architectural tests alongside cyclomatic complexity gatekeepers such as Ruff, Radon, or SonarQube as part of your CI pipeline. Set hard limits on complexity to force AI to decompose huge functions into smaller ones.

3. Examine the interfaces, not just the implementation

In code reviews of AI-generated PRs, the developer’s mind is a precious resource. Stop looking at each line, trying to decipher loops and variable assignments. Look at what changes the system from the outside. Instead, are there new dependencies? Did the PR expose new API endpoints? Did it change the data schema? If not, your mental model remains intact.

Conclusion

AI coders are very strong, but they have one big flaw — they are very pragmatic. The maintainability of your code doesn’t interest them — they care only about completing the task you assign them.

“AI coders are very strong, but they have one big flaw. They care only about completing the task you assign them.”

If you try to control your system design by relying on the human factor only, you will drown in Comprehension Debt sooner or later. It’s not a question of slowing down your AI implementation process— it’s a question of making your environment more resistant.

You don’t need to study every line of AI-generated code. You just need to create a cage for this AI.

The post Stop AI code sprawl before it destroys your software design appeared first on The New Stack.

  •  

The systems guide to production token optimization

Collage of a woman climbing progressively taller stacks of coins, with arrows tracing her upward path.

When enterprise AI applications scale, they inevitably hit a wall. For many engineering teams, this wall is initially diagnosed as a billing issue, a monthly API invoice that has grown out of control. However, viewing token consumption purely as a financial metric fundamentally misunderstands how LLMs operate in production. Token optimization, unlike its other optimization cousins, is not an accounting exercise; it’s a distributed systems and hardware utilization challenge.

“Token optimization, unlike its other optimization cousins, is not an accounting exercise; it’s a distributed systems and hardware utilization challenge.”

In this guide, we explore, through the lens of Concierge (a latency-sensitive, synchronous customer support agent) and Pathfinder (an asynchronous, multi-step autonomous CI debugging agent), how these systems fell victim to autoregressive bottlenecks as they grew, and how we fixed these issues.

What you’re actually paying for

A token is not a word: treating it like one will break your budgeting “models.” Every major LLM provider tokenizes text using byte-pair encoding (BPE), breaking words into subword units. While common words stay intact, rarer words or punctuation split into fragments. As a rule of thumb, 1 token = 4 characters, or 0.75 words in standard English prose.

When budgeting for production, you must account for the structural pricing spread; providers bill input tokens and output tokens at different rates. Output tokens are typically 4-5X more expensive than input tokens.

As a baseline, assume a mid-tier frontier model runs roughly $3 per million input tokens and $15 per million output tokens.

The quadratic history tax

LLM provider APIs are completely stateless; to make an LLM behave as if it remembers past events, you must resend the entire history of the session and input with every single API call.

This means that a model’s own previous outputs are continuously re-billed to you as inputs on subsequent steps. This triggers a compounding cost that impacts both Concierge and Pathfinder, though their curves scale differently.

Let S be the static system context (instructions and schemas), u be the incoming data per step, and r be the model’s response payload. The input cost for every given turn k is

Formula defining the input cost for every given turn.

When you sum this across a complete execution run of N steps, the total input token volume compounds quadratically. 

Formula for the total input cost, i.e. the sum of input costs of every given turn in an execution run of N steps.

This O(N^2) accumulation of history is the exact mechanism that causes the explosion in cost and latency.

VariableConcierge (Chat system)Pathfinder (Autonomous agent)
Static context (S)3,100 tokens (Full returns/shipping policies and brand guidelines)1,200 tokens (Tool definitions, system constraints, CI environment data)
Incoming data (u)80 tokens (Short customer chat replies)900 tokens (Massive raw text payloads: log excerpts, file reads, shell outputs)
Response payload (r)220 tokens (Polite customer-facing answers)300 tokens (Internal monologue + JSON Tool Arguments)
Step multiplier (N)10 turns (Average support thread length)15 steps (Average agent troubleshooting loop length)

When we calculate the total input tokens consumed by a single session using the quadratic formula

  • Concierge: Consumed 45,300 tokens per 10-turn ticket
  • Pathfinder: Consumed 150,000 tokens per 15-step turn

Because Pathfinder’s step increment was 4X larger than Concierge, its token cost curve was drastically steeper. If Pathfinder were to get stuck in an infinite tool-use loop and hit 30 steps, a single run could consume 570,000 tokens.

The solution

Fixing the individual call

Prompt hygiene: Hardcoding static reference documentation into the system prompt means you pay to parse identical text on every turn. So we stripped static text from the prompt and switched to dynamic injection. 

For Concierge, we implemented a RAG step to fetch only the 2-3 policy snippets relevant to the ticket. The prompt dropped from 3,100 tokens to 380. A 60% reduction for a 10-turn thread.

For Pathfinder, we applied automated prompt compression using LLMLingua-2 to compress verbose CI log files before sending them to the model. By filtering out non-essential log lines, we reduced the size of incoming tool observations by 3X without sacrificing debugging accuracy.

from llmlingua import PromptCompressor

compressor = PromptCompressor(
model_name="microsoft/llmlingua-2-xlm-roberta-large-meetingbank",
use_llmlingua2=True
)

try:
compressed_result = compressor.compress_prompt(
raw_ci_log_text,
rate=0.33,
force_tokens=["Error", "Exception", "Failed", "Traceback", "FATAL"]
)
# Pass high-density payload to the frontier model
compact_prompt = compressed_result["compressed_prompt"]
except Exception as e:
print(f"Compression failed, falling back to raw log text: {e}")
# Graceful degradation: pass the raw (or truncated) log if compression fails
compact_prompt = raw_ci_log_text

Eliminating the retry: Relying on open-ended prose instructions “return JSON” caused malformation. When parsing failed, the system initiated a synchronous retry, sending the entire accumulated context as if it were a new attempt. We replaced the entire natural language formatting request with strict structural contracts via forced schema validation.

In both Concierge and Pathfinder, we converted the output format to a strict pydantic schema for tool-calling mode and tool-execution payloads. Malformed outputs across both systems dropped to under 0.5%, eliminating tail latency spikes caused by cascading queues.

# Unified Schema Enforcement for Concierge Responses & Pathfinder Tool Execution
from pydantic import BaseModel
from typing import Literal

class TicketResponse(BaseModel):
    reply: str
    category: Literal["shipping", "returns", "billing", "product", "other"]
    escalate: bool
    confidence: float

# The API is structurally locked into emitting validated JSON matching the schema
response = client.messages.create(
    model="claude-opus-4",
    system=SYSTEM_PROMPT,
    messages=messages,
    tools=[
    {
    "name": "respond_to_ticket",
    "description": "Formulate a response and classify the support ticket.",
    "input_schema": TicketResponse.model_json_schema()
    }
    ],
    tool_choice={"type": "tool", "name": "respond_to_ticket"},
)

Output token bounding: Models naturally generate verbose reasoning chains and conversational filler, inflating expensive output tokens. Where the LLM provider exposes logit bias, you can directly suppress every token outside the valid set at decode time; where it doesn’t, constrained decoding libraries (Outlines, Guidance) or a forced tool call with an enum-typed schema will get you the same guarantee.

class ClassifyOnly(BaseModel):
    category: Literal["shipping", "returns", "billing", "product", "other"]
    priority: Literal["low", "medium", "high", "urgent"]

State management 

The stateless nature of the models meant we had to parse the static prompt prefix and historical steps on every turn. We introduced explicit cache breakpoints to allow the inference engine to reuse the states of static blocks. We altered both Concierge and Pathfinder to flag stable, historical segments for caching. Under standard vendor pricing, cache reads are discounted by 90%. It is important to check with your vendor on whether caching is enabled. 

# Caching the stable history prefix for a multi-turn session
response = client.messages.create(
    model="claude-sonnet-4",
    max_tokens=4096,
    system=[{
        "type": "text",
        "text": SYSTEM_PROMPT,
        "cache_control": {"type": "ephemeral"} # Cache hits drop prefix costs by 90%
    }],
    tools=TOOL_SCHEMAS,
    messages=session_history + [{"role": "user", "content": current_step_input}],
)

For a 10-turn Concierge chat, this dropped input costs by ~70%. For a 15-step Pathfinder trajectory, it resulted in a 76% cost reduction.

Semantic caching 

Duplicate queries across separate sessions were triggering redundant frontier model invocations. We implemented a vector similarity cache layer upstream of the LLM using Redis. Our Concierge service analysis showed that 34% of customer support tickets were semantic duplicates of common FAQs.

Intercepting these requests reduced latency to sub-50ms for hits. Because of the nature of CI pipeline logs, we have not yet found a suitable cache for Pathfinder’s inputs. 

import os
import json
import redis
from redis.commands.search.query import Query

# Configure connection via environment variable for environment portability
redis_url = os.environ.get("REDIS_URL", "redis://localhost:6379")
r = redis.Redis.from_url(redis_url)

def get_cached_response(tenant_id, query_text, threshold=0.92):
    try:
results = r.ft(f"cache_idx:{tenant_id}").search( # scoped by tenant -- see below
Query("*=>[KNN 1 @vector $vec AS score]").sort_by("score").dialect(2),
query_params={"vec": query_vec.tobytes()},
)
except redis.RedisError as e:
print(f"Redis cache error: {e}")
return None # Fail-open: gracefully fall back to a cache miss)
  query_vec = embed(query_text) # small, fast bi-encoder -- not the frontier model

try:
results = r.ft(f"cache_idx:{tenant_id}").search( # scoped by tenant -- see below
Query("*=>[KNN 1 @vector $vec AS score]").sort_by("score").dialect(2),
query_params={"vec": query_vec.tobytes()},
)
except redis.RedisError as e:
print(f"Redis cache error: {e}")
return None # Fail-open: gracefully fall back to a cache miss

Semantic caching could be a security problem, because if you choose a global cache, Customer A’s account-specific answer could get served to Customer B because their phrasing embeddings are close enough. To mitigate this, we split the cache into two tiers: a global cache for tenant-agnostic content, and a per-tenant, per-user namespace keyed with the tenant ID baked into the prefix itself for anything touching account state.

Cache poisoning is another risk: we only write to the cache from responses that passed schema validation and the injection-pattern classifier, we stamp every cache entry with its source traceId, and we encourage routine purging of unknown caches.

Context compaction

Uncapped conversation or agent trajectories allowed N to grow continuously, expanding the cost curve and causing latency degradation. We capped N by implementing a sliding window that summarizes historical context via a small, ultra-cheap model. For Concierge, we kept the last 3 turns verbatim while condensing older turns into a rolling metadata block.

For Pathfinder, when the debugging steps exceeded 4 runs, we trimmed and summarized the oldest tool execution outputs into a compact chronological timeline, transforming the open-ended quadratic cost explosion into a predictable, bounded window.

def compact_session_history(history_steps: List[Dict[str, Any]], keep_recent: int = 3) ->     List[Dict[str, Any]];
    """Flattens older history into a cheap summary block, preserving recent context."""
    if len(history_steps) <= keep_recent:
        return history_steps
    old_steps = history_steps[:-keep_recent]
    recent_steps = history_steps[-keep_recent:]
    
# Compress the old history using a fast, low-cost utility model
try:
historical_summary = summarize_with_utility_model(old_steps)
# Note: Anthropic prohibits 'system' roles in the messages array.
# Using 'assistant' ensures cross-provider compatibility.
return [{"role": "assistant", "content": f"[System Context: Summary of prior steps: {historical_summary}]"}] + recent_steps
except Exception as e:
print(f"History compression failed: {e}")
# Fallback: Return the uncompressed history to gracefully degrade
return history_steps

Model cascading

Directing every single operation to an expensive frontier model represents massive overprovisioning for mundane tasks. We integrated LiteLLM as an internal routing gateway to implement model cascading, routing every request to the lowest-cost model capable of completing the task.

“Directing every single operation to an expensive frontier model represents massive overprovisioning for mundane tasks.”

# litellm_config.yaml
model_list:
  - model_name: fast-path
    litellm_params:
      model: openai/mistral-support-ft
      api_base: http://vllm-internal:8000/v1
  - model_name: frontier-path
    litellm_params:
      model: anthropic/claude-opus-4

Simple, repetitive tasks are routed to a lower model, which offloads 70% of Concierge chats from the frontier model. For Pathfinder, we broke the agent loop down into separate sub-tasks: high-level planning, tool selection, and code-patch synthesis remained with the frontier model, while mechanical, text-heavy operations, such as log parsing, regex extraction, and error-string formatting, were offloaded to the lower models. This hybrid orchestration reduced Pathfinder’s token costs by more than 50%.

What’s next

The transformation of Concierge and Pathfinder proves a fundamental truth about production AI. You cannot achieve scale by simply relying on the natural language capabilities of a frontier model. You must engineer the system around it. By shifting your focus from naive token reduction to maximizing system resource utilization, we reclaimed absolute control over the infrastructure.

“Efficiency in the era of gen AI is not defined by how cheaply you can operate but by how densely you can pack information.”

Efficiency in the era of gen AI is not defined by how cheaply you can operate but by how densely you can pack information, how quickly you can serve it, and how reliably you can parse the output. The architectural decisions detailed here represent more than just a token optimization strategy; they are a required foundation for building high-throughput, battle-tested, and resilient AI systems at scale.

The post The systems guide to production token optimization appeared first on The New Stack.

  •  

Why basic RAG fails at multi-hop reasoning (and how GraphRAG fixes it)

Abstract dark metallic fluid curves reflecting light, representing complex multi-hop data pathways in GraphRAG AI architecture.

The current approach to designing LLMs within AI engineering is oversimplified. According to the echo chamber’s view, solving LLM hallucinations is easy: simply design a standard Retrieval-Augmented Generation (RAG) system in which you break your PDFs into 1,000-token chunks, embed them, insert them into a vector database, and perform cosine similarity searches.

It works perfectly well…until you actually deploy it.

After deployment, companies quickly discover that using only chunked text for retrieval does not work well for complex questions. The standard RAG assumes that semantic similarity implies relevance, which is not necessarily the case. When users ask a “multi-hop” question that requires making connections between Concept A and Concept B through Concept C, standard RAG will not work because these concepts usually don’t coexist in the same chunk of text. RAG also falls apart at global summarization (“what are the major risk factors discussed in all of our compliance reports?”)

“The standard RAG assumes that semantic similarity implies relevance, which is not necessarily the case.”

If you are designing AI for enterprise systems, you need structured reasoning. Chunking text is fine, but you should stop doing it randomly. What you need is GraphRAG.

GraphRAG offers a powerful combination of the structural knowledge of knowledge graphs along with the semantic capabilities of vector search. This tutorial will explain how your current pipeline fails and help you implement a GraphRAG workflow using Python.

The limits of naive vector search 

We will now take apart a simple misconception that vector embeddings will save the day!

Assume you have a dataset of corporate contracts stored in a vector database.

  • Chunk 1: “Acme Corp acquired BetaTech in 2022.”
  • Chunk 2: “Sarah Connor was appointed CEO of BetaTech in 2023.”

The Query: “Who leads the company that Acme Corp acquired?”

Naive vector search will embed the query and find the closest chunks. Inevitably, chunks with high similarity to “Acme Corp” and “leadership” will be found, but “Chunk 2” will be overlooked because “Sarah Connor” and “CEO of BetaTech” have nothing to do with “Acme Corp” semantically. 

The regular RAG approach lacks a structured relationship with different entities. To answer multi-hop queries, you need to understand that there is an (“Acme Corp”) - [ACQUIRED] ->  (BetaTech) relationship, and (“Sarah Connor”) - [LEADS] -> (BetaTech)

This is where GraphRAG comes in.

Enter GraphRAG 

GraphRAG requires the LLM to process the documents during the ingestion step, extracting nodes (entities) and edges (relationships) to create a knowledge graph from your chunks.

Unlike traditional approaches, where queries consist of isolated text strings, in GraphRAG we first identify the starting node via vector search and then traverse the graph’s relationships to obtain a highly relevant subgraph for reasoning with the LLM.

Recently, frameworks such as Neo4j’s neo4j-graphrag have made this architecture accessible to everyone. We’ll use this framework to build a robust pipeline that performs ingestion, schema enforcement, and graph traversal.

Practical implementation: Building GraphRAG in Python

We will build a pipeline that can ingest unstructured text data, build a knowledge graph, and query the graph via dynamic Cipher traversal.

Step 1: Install dependencies

You will need an active Neo4j database (local or cloud) and the following packages:

Bash
pip install neo4j neo4j-graphrag openai python-dotenv

Step 2: Initialize connections and production guardrails 

Production-ready pipelines require proper credential management and resilience.

GraphRAG ingestion is extremely expensive in terms of LLM API requests. Without configuring a rate-limit handler, your pipeline will fail with a RateLimitError on the first real document.

Python
import os
from neo4j import GraphDatabase
from neo4j_graphrag.llm import OpenAILLM
from neo4j_graphrag.embeddings import OpenAIEmbeddings 
from neo4j_graphrag.utils.rate_limit import RetryRateLimitHandler

# Initialize Neo4j Driver
neo4j_uri = os.environ.get("NEO4J_URI")
neo4j_user = os.environ.get("NEO4J_USERNAME")
neo4j_password = os.environ.get("NEO4J_PASSWORD")

if not all([neo4j_uri, neo4j_user, neo4j_password]):
raise ValueError("Missing required Neo4j environment variables (NEO4J_URI, NEO4J_USERNAME, NEO4J_PASSWORD).")

driver = GraphDatabase.driver(neo4j_uri, auth=(neo4j_user, neo4j_password))
driver.verify_connectivity() 

# Initialize LLM with strict rate limiting for heavy extraction tasks
llm = OpenAILLM(
    model_name="gpt-4o",
    model_params={"temperature": 0},
    rate_limit_handler=RetryRateLimitHandler(max_attempts=5, min_wait=2.0)
)
embedder = OpenAIEmbeddings(model="text-embedding-3-small")

Step 3: Enforcing schema in the knowledge graph builder 

An unstructured graph is nothing more than an unorganized database. The moment you do not provide a proper schema, your LLM will hallucinate entity labels, duplicating nodes of “Acme Corp”, “Acme Corporation,” and “Ame.” We will build our architecture beforehand and feed it into the pipeline.

“An unstructured graph is nothing more than an unorganized database.”

Python
from neo4j_graphrag.experimental.pipeline.kg_builder import SimpleKGPipeline

# Define explicit schema to prevent LLM hallucinations during extraction
schema = {
    "node_types": ["Organization", "Person", "Technology"],
    "relationship_types": ["ACQUIRED", "LEADS", "DEVELOPS"],
    "patterns": [
        ("Organization", "ACQUIRED", "Organization"),
        ("Person", "LEADS", "Organization")
    ]
}

# Define the pipeline enforcing the schema
kg_builder = SimpleKGPipeline(
    llm=llm,
    driver=driver,
    embedder=embedder,
    schema=schema
)

sample_text = """
Acme Corp acquired BetaTech in 2022. 
In 2023, Sarah Connor was appointed CEO of BetaTech to drive AI initiatives.
"""

import asyncio

# Execute extraction and graph population
asyncio.run(kg_builder.run_async(text=sample_text)) 

Engineering note: Under the hood, this pipeline still splits your text using a FixedSizeSplitter. The difference is that it uses the LLM to extract the semantic architecture from those chunks before saving them.

Step 4: The vector Cipher retriever 

This is where multi-hop reasoning happens. An ordinary vector retriever will provide us with nodes. To travel across the graph, we use a VectorCypherRetriever. We will embed the query, identify the semantic entry point, and finally perform a Cipher query on all the 1-hop connected neighbors.

Python
from neo4j_graphrag.retrievers import VectorCypherRetriever

# Define the traversal logic: Find the vector match, then expand 1-hop to get relationships
traversal_query = """
MATCH (node)[r](neighbor)
RETURN "Entity: " + coalesce(node.id, '') + " (Info: " + coalesce(node.text, '') + ") | " +
"Relationship: " + type(r) + " | " +
"Neighbor: " + coalesce(neighbor.id, '') + " (Info: " + coalesce(neighbor.text, '') + ")" AS text,
score
""" 

# Initialize a VectorCypherRetriever to enable actual GraphRAG
retriever = VectorCypherRetriever(
    driver=driver,
    index_name="entity_vector_index",
    embedder=embedder,
    retrieval_query=traversal_query
)

query = "Who leads the company that Acme Corp acquired?"

# Retrieve connected context based on the query
retrieved_context = retriever.search(query_text=query, top_k=3)

print(retrieved_context)

The retrieved context includes BetaTech, its acquisition by Acme Corp, and the fact that Sarah Connor leads it, since we provided a specific schema for retrieving relations. The LLM now has all the facts it needs to provide an accurate response.

Lessons learned: What engineers must know before adopting GraphRAG

If you want to upgrade your RAG system, consider the following trade-offs:

  1. Ingestion cost is high, whereas RAG costs are low because embedding models are inexpensive. To ingest each document chunk in GraphRAG, you need an LLM like GPT-4o to extract entities from the document. Filter your enterprise data carefully.
  2. Schema design is not optional: As shown in step 3, you shouldn’t ignore data modeling. The schema is the difference between an efficient retrieval engine and a mess of nodes.
  3. Observability is superior: To debug your naive RAG, you need to examine big floating-point arrays. To debug GraphRAG, you only need to look at your Neo4j dashboard to view the nodes. It is easy to see if the LLM has hallucinated.

Standard vector search is a very powerful solution, but treating it as a one-size-fits-all fix for enterprise AI is an easy way out. Real-world use cases need explicit data modeling and deterministic search capabilities.

“Standard vector search is a very powerful solution, but treating it as a one-size-fits-all fix for enterprise AI is an easy way out.”

GraphRAG demands more intentionality in its design, strict guidelines to follow, and greater upfront computational effort. However, the rewards are immense: GraphRAG gives you the power to do something truly valuable—multi-hop reasoning over data, which is exactly what enterprises desire. 

The post Why basic RAG fails at multi-hop reasoning (and how GraphRAG fixes it) appeared first on The New Stack.

  •