❌

Vue lecture

Claude Opus 5.5 vs. Opus 5 on reasoning tasks: Cheaper, faster, but not better

Abstract long-exposure photograph of red and orange light trails forming layered curves around a dark central shape.

When Anthropic released Claude Opus 5.5 this week, the company claimed the new model costs 40% less than Opus 5 and generates output 30% faster. Anthropic’s marketing makes three claims. Opus 5.5 performs at the level of Claude Fable 5.1 (so it should outperform Opus 5), costs 40% less than Opus 5 on typical workloads, and generates output more than 30% faster.

Anthropic also cut the price developers pay to use the model through its API. Opus 5.5 costs $4 for every million tokens (chunks of text roughly three-quarters of a word long) sent to the model and $20 for every million it writes back, down from $5 and $25 for Opus 5. That price cut alone accounts for a 20% saving. The rest of the claimed 40% saving has to come from the model using fewer tokens.

I wanted to see how this translates for the average Claude user, so I skipped the usual developer workflow simulations this time. Lately, the models I test handle everyday tasks well. Reasoning tasks are where I’ve seen them struggle, so I tested Opus 5 against Opus 5.5 on reasoning tasks only. 

You can find the prompts at the bottom of this post if you want to replicate these tests on your own system.

The tests

I called both models through the Anthropic API with identical prompts. Both ran with adaptive thinking at the default effort level, since Opus 5.5 doesn’t allow you to turn thinking off. Each problem ran once per model. I planned to rerun any problem where the models gave me different results, but they never did.

Here are the tests I ran:

  • Logic grid (medium difficulty) – Seven engineers each have an on-call day, a language, a service, and a city, and 22 clues pin down one answer. Six clues are conditional or “exactly one of these is true” statements, and removing any single clue breaks the puzzle.
  • Constrained orderings (hard difficulty) – Reorder 10 deploy jobs so no job stays in its original slot and no two consecutively numbered jobs sit side by side. The model had to give the count for 6, 8, and 10 jobs.
  • Stone game with memory (harder difficulty) – Players remove 2, 5, 7, or 11 stones, but can’t repeat their opponent’s last move or their own. The model had to find who wins from 200 stones, count the losing starting sizes up to 500, and name the smallest losing size above 340.

I logged input tokens, output tokens, cost at list price, and time for every call. Thinking tokens are billed as output, so I included them.

The logic grid

Both models got all 28 cells right. Opus 5.5 took 65 seconds and 7,573 output tokens, for $0.16. Opus 5 took 108 seconds and 10,621 output tokens, for $0.27.

On this test, Opus 5.5 was 43% cheaper and delivered the same correct answer. Opus 5.5 was slightly more detailed and noted that it didn’t fully prove the solution was unique.

Constrained orderings

Neither model produced an answer, which made this the hardest problem in practice. The correct counts are 27, 1,695, and 159,019. The third answer is very hard to reach by reasoning alone. A computer program that checks every possible ordering can find it, but neither model could run code in this test.

With a 48,000-token output limit, both models spent the whole budget thinking and never replied. Opus 5.5 used 489 seconds and $0.96. Opus 5 used 553 seconds and $1.20.

I raised the limit to 128,000 tokens and ran it again. Opus 5 used every token, took over 25 minutes, and stopped with no answer. That cost $3.20. Opus 5.5 ran for 19 minutes and used 112,733 tokens, but the API ended the response with a “refusal” stop reason and no text. The prompt asks the model to count job orderings and contains nothing sensitive. That means the refusal was most likely a mistake by Anthropic’s safety filter flagging a harmless request.

Opus 5.5 was cheaper, but how much does that matter if you don’t get a result?

The stone game

And we’re back to the same answers again. Both models answered all three parts correctly. The first player loses from 200 stones; starting sizes from 120 to 500 are losses, and the smallest loss above 340 is 344.

The difference in this test came down to how much thinking each needed to get there. Opus 5.5 finished in 215 seconds with 28,740 output tokens, costing $0.58. Opus 5 took 624 seconds and produced 74,981 output tokens, costing $1.88. Opus 5.5 used 62% fewer tokens and cost 69% less for the same answer. 

Results

TestOpus 5.5Opus 5
Logic grid28/28, 1:05, 908 in / 7,573 out, $0.1628/28, 1:48, 906 in / 10,621 out, $0.27
Ordering problem (48k limit)No answer, 8:09, 235 in / 48,000 out, $0.96No answer, 9:13, 233 in / 48,000 out, $1.20
Ordering problem (128k limit)No answer (refusal), 18:56, 235 in / 112,733 out, $2.26No answer, 25:24, 233 in / 128,000 out, $3.20
Stone game3/3, 3:35, 323 in / 28,740 out, $0.583/3, 10:24, 321 in / 74,981 out, $1.88
Total tokens1,701 in / 197,046 out1,693 in / 261,602 out
Total time31 min 45 sec46 min 49 sec
Cost$3.95 ($4 in / $20 out per million tokens)$6.55 ($5 in / $25 out per million tokens)

Across every call, Opus 5.5 wrote 103.4 tokens per second and Opus 5 wrote 93.1, so Opus 5.5 was about 11% faster. Its biggest speed lead on any single problem was 19%, still short of Anthropic’s 30% claim. Its biggest cost savings came on the stone game, where it cost $0.58 to Opus 5’s $1.88, 69% less. Total spend for the test was $10.50. 

What do I think

Anthropic’s benchmarks show Opus 5.5 ahead of Opus 5 on coding, knowledge work, and reasoning. I didn’t rerun those benchmarks. I gave both models the same three reasoning problems, and they performed equally. Both solved the logic grid and the stone game, and both failed the ordering problem.

The savings are real. Opus 5.5 cost less and finished sooner on every problem, including 43% less on the logic grid and 69% less on the stone game. Most of that came from using fewer output tokens. The price cut accounts for 20%. Its writing speed was 11% faster, short of the 30% Anthropic claims.

Switch to Opus 5.5 if you run Opus 5 today. You get the same results on hard reasoning for less money and less waiting. Set a hard output limit and watch your spend on hard problems, though. Both models can think for close to 20 minutes or more and return nothing, which is a problem if you pay for every token. For counting problems like the ordering test, give the model a code execution tool instead of hoping it reasons its way through.

The prompts

Logic grid

Seven engineers (Ana, Ben, Cy, Dee, Eli, Fay, Gus) share an on-call rotation. Each is on call on exactly one day of a single week, Monday through Sunday (Monday is the earliest day, Sunday the latest), and no two share a day. Each writes a different language (Go, Rust, Python, Java, Kotlin, TypeScript, C++), owns a different service (auth, billing, search, queue, cache, gateway, metrics), and is based in a different city (Berlin, Tokyo, Denver, Lagos, Sydney, Toronto, Mumbai).

Clues:

The Java developer is Gus.

The TypeScript developer is on call exactly two days after the queue owner.

Exactly one of these is true: the engineer based in Berlin owns billing, or the C++ developer is on call Monday.

Exactly one of these is true: the engineer based in Tokyo owns cache, or the engineer based in Tokyo writes Rust.

If Fay is on call Sunday, then the engineer based in Denver does not write Kotlin.

Exactly one of these is true: the Kotlin developer owns auth, or the C++ developer is Cy.

The engineer based in Tokyo is on call earlier in the week than the C++ developer.

The engineer based in Denver does not write Python.

Exactly one of these is true: the Go developer is Ben, or the gateway owner is based in Berlin.

Ana is based in Mumbai.

The TypeScript developer is on call earlier in the week than Fay.

The engineer based in Sydney is on call earlier in the week than the billing owner.

The gateway owner is on call earlier in the week than the Kotlin developer.

The engineer on call Sunday is not based in Berlin.

Fay and the engineer based in Lagos are on call on consecutive days.

Eli is on call exactly three days after the auth owner.

The engineer based in Lagos writes Java.

Exactly one of these is true: the search owner is Eli, or the cache owner is Dee.

The gateway owner and the Go developer are on call on consecutive days.

The engineer based in Lagos is on call exactly four days after the auth owner.

The engineer based in Lagos owns metrics.

The engineer based in Mumbai is on call earlier in the week than the Python developer.

Determine the full assignment. At the end of your response, give exactly seven lines, one per engineer in the order Ana, Ben, Cy, Dee, Eli, Fay, Gus, in this format:

ANSWER: Name | Day | Language | Service | City

Ordering problem

A build system has n deploy jobs numbered 1 to n. Originally, job k runs in slot k. You reorder all n jobs into slots 1 to n (each slot gets one job) subject to two rules:

No job runs in its original slot (job k is not in slot k).

Jobs with consecutive numbers never run in adjacent slots (for example, jobs 4 and 5 cannot be in slots i and i+1 in either order).

How many valid orderings are there for (a) n = 6, (b) n = 8, (c) n = 10?

At the end of your response, give exactly three lines in this format:

ANSWER a: <number>

ANSWER b: <number>

ANSWER c: <number>

Stone game

Two players play a game with a pile of stones. They alternate turns. On each turn, a player removes exactly 2, 5, 7, or 11 stones, subject to two rules:

You may not remove the same number your opponent removed on their most recent turn.

You may not remove the same number you removed on your own most recent turn.

(On the very first turn of the game, neither rule applies. On the second player’s first turn, only the first rule applies.) You cannot remove more stones than are in the pile. A player who has no legal move on their turn loses. Both players play perfectly.

(a) Starting with 200 stones, does the first player win?

(b) For how many starting pile sizes from 1 to 500 inclusive does the first player lose?

(c) What is the smallest starting pile size greater than 340 for which the first player loses?

At the end of your response, give exactly three lines in this format:

ANSWER a: <yes or no>

ANSWER b: <number>

ANSWER c: <number>

The post Claude Opus 5.5 vs. Opus 5 on reasoning tasks: Cheaper, faster, but not better appeared first on The New Stack.

  •  

AI spending can run negative. Qodo’s CEO built an ROI equation to fix it.

Five stacks of mixed copper, silver, and brass coins arranged left to right in ascending height, like a bar chart, against a neutral beige background.

Flush with the proceeds of a $70 million Series B raised earlier this year, you might expect Qodo to spend freely on internal AI. After all, the startup uses artificial intelligence to ensure AI-generated code meets customer quality and governance requirements. An upstart technology company using AI to improve AI outputs is AI-pilled by definition.

Instead, the company has limits on AI consumption. Qodo CEO Itamar Friedman tells The New Stack that his engineers can access $10,000 worth of tokens per month, a cap that he described as “generous.” Most Qodo developers never reach it. The ceiling wasn’t enacted to “restrict usage,” Friedman says, but instead to drive “visibility and efficiency” at the startup so that it can “scale without runaway costs.” Put another way, the cap exists to make somebody answer this question: “Which path of automation or usage will be the best [use] of our money?”

Qodo’s AI footprint is larger than its developer token budget. The startup’s AI infrastructure spend — the cost of running the product for customers rather than the cost of its own engineers using AI — is growing at “roughly 5x year over year,” the company tells TNS in an email, reflecting both “increased user adoption” and its agents taking on more, and longer tasks as they mature. Qodo says it is pushing the other direction at the same time, driving down the cost of reviewed pull requests through routing and inference efficiency.

What Qodo runs on

The company also dogfoods heavily, running its pull requests through Qodo. Friedman said the product powers its entire software development life cycle (SDLC). Around that sits a stack most engineering organizations would recognize: Slack and Notion and their constituent “bots,” a centralized knowledge base built to be agent-readable, AI inside Google Workspace, and models from several providers including Google.

Qodo’s own product sits alongside Claude Code and other leading coding assistants rather than replacing them. Claude Code still holds the crown internally, but OpenAI’s Codex has been taking share, with staff “shifting quickly towards Codex.” Friedman tracks this two ways. He polls his 130-person staff, spread across offices in several countries, on the tools they prefer, and compares those answers against what the usage data shows.

Friedman reports that Qodo sees “roughly double” the number of PRs “every couple of months” alongside “a decreasing amount of bugs and incidents.” Hold on to those two numbers. They’re important here.

The bottleneck moved

More PRs and fewer bugs indicate that Qodo is onto something with its focus on software testing and governance. Its technology helps developers deal with an increasingly common issue: What do with all the code that AI agents generate? Companies that adopt AI coding tools often find that they create more code with machines than their humans can assess. As a result, the SDLC bottleneck simply shifts one step down the process.

“We solved the speed of writing code,” Friedman argues, “we didn’t solve the velocity of creating software.” The difference between accelerating one part of a task and its entire arc is the difference between AI hype and AI ROI. 

The software development example shows that when we consider AI costs and benefits, we need to think broadly. If we focus too much on a single metric, we might spend our entire budget on Claude Code credits while shipping no more software than before. Alongside a massive bill.

The AI ROI Equation

Friedman recommends an equation-based approach. The Qodo perspective on AI ROI is similar to a popular equation for happiness: Personal joy is the distance between your expectations and reality. The greater the expectations, the harder it is to be happy. The lower the expectations, the greater the chance of being content. 

This can be expressed as either simple subtraction or as a ratio:

  • Reality/expectations = Happiness, where larger results indicate greater joy

Take the same mathematical approach to AI ROI, per Friedman: Compare the positives against the negatives, add up all the good, and set it over all the bad.

  • AI benefits/AI costs = AI ROI, where larger results indicate greater return

Friedman found the shape of the equation in The Phoenix Project, the 2013 DevOps novel that contrasts types of software development work and sorts them into good and bad buckets. Plug those terms in:

  • (Features + Infrastructure)/(Incidents + Bugs) = Software development velocity

Now, those two numbers from earlier. Qodo has seen more PRs and fewer bugs thanks to AI. In DevOps terms, it’s shipping more and fixing less, so the equation returns a larger, better result. Feed the same terms into the AI ROI version, and it produces more benefits over fewer costs, and a larger final calculation. 

The fraction is not a thought experiment at Qodo. It’s the shape of what the company says is already happening to it.

Terms that have nothing to do with software development work too. Qodo runs AI inside Google Workspace, Slack, and Notion, and those benefits and costs go into the same calculation. 

The Qodo approach to measuring total AI ROI is less specific than The Phoenix Project’s DevOps equation, but the difference is acceptable. Friedman argues that you have to start somewhere: “I know [the equation is] a simplification,” the CEO tells TNS. “But what you can’t measure, you can’t improve.”

His argument is that imperfect beats absent. “Don’t think about it too much,” he says. “Try to put any number [in the AI ROI equation] and start tracking.” Being told not to overthink an equation is a great soundbite, but the benefit is real: A rough calculation on paper beats holding the same information in your head without form. In this case, the journey is a large part of the destination.

Friedman reckons that startups should pick no more than six or eight terms for their own calculations. That’s an afternoon’s work. A start on what will prove to be an ongoing exercise. 

Negative ROI

The fraction runs backward, too. 

Recall Friedman’s point about a company writing more code faster but not accelerating its software development speed. Stuff those terms in:

  • (Faster code generation + other AI benefits)/(Slower code review and approval + agentic coding costs + other AI costs) = Smaller AI ROI

That’s how a company spends a king’s ransom on AI credits and winds up nowhere or nonexistent. 

Which is not hypothetical at Qodo either. AI doesn’t excel everywhere, and Friedman named email automation as an example. The company went all in on automating it, then pulled back, “mov[ing] from AI automation to AI enhancement” after discovering that AI struggled to match writing tone and intelligently extract tasks from messages. The retreat is the interesting part: Qodo’s stated approach to any task is to “go all in on complete automation,” and then “take a step back to human judgment.” 

The CEO says that automation falls short today in two areas: When human judgment is required and when context is missing. The second cuts across everything from software development to personal productivity to answering customer questions. Without timely context, what can AI do other than filibuster? Qodo’s service helps answer the context issue for software development, but collecting a company’s data and making it accessible, timely, and well-governed for general agentic usage is a massive undertaking, and one that a host of startups want to help solve. If they can, everyone’s AI ROI math should improve.

No mandate, high expectations

Qodo doesn’t require its staff to use AI. As Friedman puts it, you won’t get fired simply because you’re “not AI all the way,” or “eating AI for breakfast.” The company expects staff to complete their work as efficiently as possible and leaves the method to them.

Employees make their own decisions and execute their own work. If they start to fall behind on assigned tasks, they’re expected to reach for automation. It’s a balanced approach with high expectations: An employee who isn’t as efficient as they could be with AI could find themselves at risk.

Friedman has been on the unpopular side of an AI argument before. When he was building Qodo in 2023 and talking up agents, “agents” was a “bad word,” dismissed as little more than “fluff.” Three years and a $70 million Series B, the bet has paid off.

His advice to founders starting now looks like his past. Predict “what’s going to happen two years from now,” he says, then solve for it immediately, because whatever looks like two years tends to arrive inside of twelve months. The future “is coming faster” than you think, he says.

It’s a lot to ask of anyone working from an incomplete picture. Predicting the future is hard, he admits, “but you have to.”

The post AI spending can run negative. Qodo’s CEO built an ROI equation to fix it. appeared first on The New Stack.

  •  

OpenAI cut GPT-6 token prices in half. The bigger lever may be the cache.

Sam Altman, OpenAI CEO

OpenAI released GPT-6 Sol and Luna on Tuesday, essentially more affordable versions of GPT-6 Astra that come closer to Astra on alignment than GPT-5.6 Sol did, but still fall short of the flagship model. 

Most notably, the AI company slashed token prices, making the new GPT-6 models significantly cheaper to use. Beyond token prices, though, OpenAI says better caching can also help developers push costs down even more.

Per OpenAI: “Improvements in caching and inference let us serve these models at lower cost,” with API prices for Sol and Luna down 50% compared to their GPT-5.6 counterparts (58% lower for Luna output tokens).

What improvements? Namely, higher cache-hit rates by default, the ability to preserve earlier context even when reasoning effort and tool availability change, and new tools to monitor and diagnose caching performance. 

Reuse context without starting over

Prompt caching isn’t, of course, novel to the new GPT-6 models themselves. But the upgraded Sol and Luna come with improvements designed to keep more previously processed context reusable as the agent moves forward on a task. 

“We’ve improved prompt caching for GPT‑6 to deliver higher cache hit rates by default, helping agents reuse more context, respond faster, and benefit from discounts of 90% on cached input-token reads.”

That adds another opportunity to lower the already low API price tag, though the 90% cached-input discount matches GPT-5.6 pricing; what’s new is how often the cache gets hit. By using cached context to reuse work it’s already done, the model doesn’t have to process the same context again from scratch for every single call, thereby reducing latency — and token costs.

Beyond this higher default cache-hit rate, OpenAI says the new GPT-6 models offer more flexibility to optimize caching performance. 

The new models let developers adjust reasoning effort and tool availability without having to break the cache. This way, an agent can scale reasoning effort up and down based on how difficult a step is, then make different tools available depending on what the task requires without disturbing earlier cached context — again, a win for both speed and cost. 

See what gets cached and what doesn’t 

GPT-6 Sol and Luna also arrive with a Prompt Caching Dashboard, where OpenAI says developers can view caching performance to understand how much context is reused. 

Specifically, they can see how much input is cached and how that amount changes over time. The diagnostics tool then flags missed caching opportunities to help developers understand what could use more efficient caching. 

Rather than keeping cache performance largely hidden behind the scenes, the idea is to make it more visible so developers can actively measure and optimize cache reuse. 

Altogether, OpenAI says these caching improvements are already making a difference. Per the AI company, GitHub reports, “these improvements have reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests to OpenAI models.”

These results span the past “several months.”

Token prices aren’t the only way to make agents cheaper

OpenAI’s pricing cuts for GPT-6 Sol and Luna made the biggest splash, with the AI company significantly dropping API prices from GPT-5.6 levels.

Compared to the current prices for GPT-5.6 Sol and Luna, which stand at $4 and $0.20 per million input tokens and $20 and $1.20 per million output tokens, respectively (GPT-5.6 Sol’s rates are promotional pricing), the new GPT-6 models come in at just $2 and $0.10 per million input tokens and $10 and $0.50 per million output tokens, respectively.

With GPT-6 Sol and Luna’s caching improvements and lower token pricing, OpenAI is making the case for tackling agent costs from both sides: charging less for fresh processing and reducing how often the same context needs to be reprocessed.

But as more AI model providers compete aggressively on pricing, it’s becoming clearer that cheaper models alone won’t save your AI budget — and lower token prices aren’t the only way to make agents cheaper. 

With GPT-6 Sol and Luna’s caching improvements and lower token pricing, OpenAI is making the case for tackling agent costs from both sides: charging less for fresh processing and reducing how often the same context needs to be reprocessed.

As agents continue to work on longer and more complex tasks, there will likely be more pressure to do both. 

The post OpenAI cut GPT-6 token prices in half. The bigger lever may be the cache. appeared first on The New Stack.

  •  

Why an old caching trick is your secret to lower LLM costs

Server racks in a dark data center, their mesh doors revealing dense bundles of orange and teal cables looping between hardware lit by rows of small green and yellow status LEDs.

An LLM can answer the same question a thousand times and charge you each time. Before paying for another answer, check whether anything that could change it has changed: the request, its context, the model settings, or the underlying data. I fingerprint those inputs and dependencies to create an exact-match cache key. If that key points to an answer that’s still valid and safe to reuse, I return it without calling the model. The savings start with a simple decision: knowing when the work is already done.

I didn’t learn this lesson from an LLM job. In production data pipelines, I’ve encountered a recurring pattern: a nightly job recalculates aggregations that haven’t changed since the previous run. It passes all its checks and moves the results into production successfully, all while burning compute that could have been used elsewhere.

The waste hides in plain sight because nothing appears broken. It often surfaces during a cost review, when someone notices that a significant portion of upstream compute is re-answering a question whose inputs never changed. The fix is change detection: hash the upstream inputs that could change between runs, fingerprint the job’s dependencies, and skip recomputation when the fingerprints match. Done well, this significantly reduces the compute that job consumes.

The lesson is common, and it’s the same one we keep trying to drive home in LLM workloads. There, repeated requests can also produce repeated charges, since billing is by token.

The problem is simple enough to state, but the more you look into it, the more you need a framework to engineer a good answer. For most of the LLM calls in our codebase and infrastructure, we’re billed by tokens, and many APIs treat duplicate requests as new ones anyway. Duplicate sources are almost as inevitable as rain.

Upstream users converge on similar questions to answer with their LLM tools. Batch jobs dutifully repeat boring boilerplate every time they run. Prompt-engineering experiments in development and CI runs invoke the same prompt repeatedly. And tool-calling agents may hit the same knowledge-base tool many times in a single work day.

Native prompt caching is a different thing from the response caching I’m describing. In prompt caching, providers reuse cached prompt computation and charge eligible cache reads at reduced rates; output generation remains billable. In response caching, we try to skip the call entirely when an answer already exists in our own infrastructure.

Tier 1: exact match

The simplest approach is to normalize the model request body, run it through a cryptographic hash like SHA-256, then look up the hash in an in-memory store like Redis. If we find a match, we return the answer without waiting for model inference. An exact-match cache works best when we can expect our model requests to be bounded and predictable. That doesn’t sound exciting, but for most of our batch pipelines, CI runs, and boilerplate summarization tasks, it’s exactly what we need.

Tier 2: semantic match

For many workloads, exact match isn’t enough. We’d like to look up a response for a query that’s close but not identical. So we take the user’s query, run it through an embedding model, and store the resulting vector in a vector database. When a new query arrives, we run it through the same model and search for close matches by cosine similarity.

Close enough by what measure? A common starting point is a cosine-similarity threshold in the [0.90, 0.95] range, but treat that as a number to tune, not a default — the right value depends on your embedding model and your data, and you should test it against real queries. Note that vector stores differ in what they return: cosine similarity rises toward 1 for closer matches.

At the same time, some engines report a distance that falls toward 0, so confirm which your threshold is comparing against. Either way, a looser threshold raises the risk of wrong matches, where the system answers one query while the user was asking about another. (“What’s the weather in my town?” can’t be safely conflated with the same question about a different town just because the cosine similarity is high.)

Tier 3: hybrid

A common approach runs both tiers in sequence: check the exact-match store first, and run semantic search only on a miss. When semantic search returns a close-enough match, the result is promoted back into the exact-match store under the hash of the new query that triggered it, so the paraphrase and its answer are an exact hit next time.

This favors cheap exact matches on repeat traffic. The pseudocode below shows the full flow: normalization and SHA-256 for exact match; a Redis get followed by a set on a miss; embedding the query and searching the vector DB with top_k=1; checking cosine similarity against the per-category threshold; and setting the TTL before writing the response back into the exact store.

Both tiers key on more than the query text alone: the context and documents in the prompt, the model and its settings, the version of any retrieved source, and the caller’s access scope. Two identical questions asked against different documents, or by users with different permissions, must not share a cache entry.

def cached_completion(query, ctx):

    # ctx bundles everything that changes what the correct answer is:

    # the context/documents in the prompt, the model and its settings,

    # the source-version of any retrieved content, and the caller's access scope.

    key = sha256(normalize(query, ctx))

    # Tier 1: exact-key lookup on Redis (O(1)).

    # Correctness still depends on cache contents, request scope, and freshness.

    if (hit := redis.get(key)):

        return hit

    # Tier 2: semantic search, restricted to the same scope as the request.

    emb = embed(query)

    match = vector_db.search(emb, top_k=1, filter=scope_of(ctx))

    if match and same_scope(match, ctx) \

            and match.score >= threshold_for(category(query)):

        # Promote, but preserve the original freshness deadline.

        remaining = match.expires_at - now()

        if remaining > 0:

            redis.set(key, match.response, ttl=remaining)

            return match.response

    # Miss on both tiers: call the model, validate before writing back.

    resp = llm(query, ctx)

    if is_valid(resp):  # no errors, no empty payloads, no malformed JSON

        ttl = ttl_for(category(query))

        redis.set(key, resp, ttl=ttl)

        vector_db.insert(emb, resp, ttl=ttl, scope=scope_of(ctx))

    return resp


One threshold does not fit all. Code-like queries often need stricter thresholds, around 0.95 or higher, because small wording changes can produce entirely different results. Conversational queries can tolerate looser thresholds, in the 0.85 to 0.90 range. These numbers are starting points, not settled values — validate them for your own workload and embedding model before relying on them. Cache freshness works the same way, and the right TTL follows from how much staleness the use case can tolerate, not from the data type alone.

A cached market-data answer might be acceptable for only a minute or two, because a stale price can be actively misleading. An internal HR policy answer can often be reused for weeks, because the underlying document rarely changes and a slightly old answer is usually still correct. The interval is a judgment about acceptable staleness, not a fixed property of the content.

The math

For illustration, suppose a workload of 1,000,000 calls per month at $0.006 per call, roughly $6,000 with no caching. Say a hybrid cache gives about a 60% hit rate, whichever tier hits first, avoiding 600,000 calls to the model, and that embedding and vector-store costs come to about $150. That brings monthly spend closer to $2,550, a 57.5% reduction, plus the latency win of answering many questions without waiting on the model.

One caveat worth shouting: measure your hit rate before you project any savings.

The decisions

There’s more to this than the tiered framework. Tune your TTLs to the freshness each data type actually needs, and invalidate entries when you update the content behind them. A fine-grained approach assigns a per-category TTL based on how quickly each answer goes stale: a news summary might hold up for an hour, while a live sports score is worthless within seconds and shouldn’t be cached at all during a game.

Live scores require a freshness policy matched to the application. Verified final scores can support much longer caching, with invalidation for corrections. The distinction here is whether the underlying value is still moving. A blunter approach skips per-category tuning entirely and purges the whole cache whenever the source content changes. Either way, run the cache in shadow mode first, logging what you would have returned without changing behavior. Evaluate cached answers against verified reference answers or expert review. A fresh model response can help identify differences, but it is not ground truth.

Warm the cache from a historical set of common queries before you rely on it, and validate answers before writing them back, so you don’t poison the cache with errors, empty responses, or malformed content.

When not to cache? Skip it for requests with personal or account-specific data, to avoid leaking one user’s cached output into another’s request. Skip it for creative tasks, where you want a different answer each run. And skip it for genuinely real-time data like stock prices and live inventory, where an answer even a minute old may be too stale for the application.

The takeaway

The principle predates the web: Donald Michie described memo functions in 1968. When you can, fingerprint the question and store the hashed exact form alongside the semantic-variant form, so you avoid repeated model calls while a valid cached answer remains available.

The post Why an old caching trick is your secret to lower LLM costs appeared first on The New Stack.

  •  

Aider, Claude Code, and OpenClaw ran an identical model. Token use varied 70-fold.

Collage of a woman climbing progressively taller stacks of coins, with arrows tracing her upward path.

When teams price an AI coding agent, they tend to scrutinize the model. But three recent benchmarking efforts suggest the harness — the software that steers it through tasks — may matter just as much.

The reason will be familiar to any web developer: Each inference request needs context. The relevant history must either be supplied again or reconstructed by the serving system. As a result, the provider processes large, overlapping blocks of text on every turn, including the harness’s system prompt and tool descriptions.

In June, an independent benchmark compared 12 configurations across two models on the same 12 Python tasks. In August, Composio compared eight harnesses using DeepSeek V4 Flash on 30 enterprise workflows. Artificial Analysis, meanwhile, continuously tracks harness-model pairings in its coding-agent index.

What the three benchmarks measured

Composio reported thirty workflows spanning Airtable, Gmail, Google Calendar, Google Sheets, GitHub, Slack, and PostHog. Each task ran under a 900-second ceiling. A programmatic verifier, rather than an LLM judge, graded the outcome, using isolated fixtures seeded with decoys and near-identical keys.

Composio reported 240 runs, of which 129 workflows completed successfully. Cost per successful task ranged from $0.028 for Pi Agent up to $0.195 for Claude Code. DeepAgents matched Claude Code’s pass rate exactly while costing a quarter as much per success.

The control was not perfect, and Composio disclosed that transparently. Pi ran a different reasoning setting across two model providers, and Prime Agent produced only 24 gradable runs out of 30. Those caveats rule out calling this a clean single-variable experiment.

The June benchmark measured tokens rather than dollars, and it stretched further. Its author reported runs of Aider, Claude Code, Codex, Goose, Hermes, Kilo, Kimi Code, Nanobot, OpenClaw, Opencode and Qwen Code, counting Aider’s architect mode separately. All twelve configurations ran through OpenRouter on the same tasks, so each harness used the same API and model. The suite ran on DeepSeek V4 Flash and then Nvidia’s Nemotron 3 Ultra. The second model offers a free tier on OpenRouter so that you can rerun it at no cost.

The post reported tokens per solved task ranging from roughly 3,500 for Aider in architect mode to 292,000 for OpenClaw. What makes that range usable is its stability, since the ordering barely moved between two unrelated models. That points to the harness software rather than model behavior.

Artificial Analysis approaches the same question with more statistical weight. Artificial Analysis reported an index combining DeepSWE, Terminal-Bench v2.1 from the Laude Institute, and Scale AI’s SWE-Atlas-QnA. That comes to 326 tasks, with pass rates averaged across three attempts each. It reports cost per task, token use, and wall time for each pairing. It also publishes a controlled comparison that holds Claude Opus 4.7 fixed while swapping between Claude Code, Cursor CLI, and Opencode.

The startup tax

At its core, the spread stems from a single measurement: the June benchmark, the startup tax. Before a prompt does any work, the harness ships its own baggage. That baggage is the system prompt, the tool descriptions, and the environment setup. The benchmark reported around 700 tokens for Aider in architect mode, against around 26,000 for OpenClaw.

A 40x overhead would be tolerable if it were paid once, but the resend pattern makes it otherwise. The author noted that a harness carrying a 26,000-token floor through fifteen turns spends roughly 390,000 input tokens on scaffolding alone.

The math behind that holds up under testing. The startup tax multiplied by the turn count predicts tokens per solved task, with an R-squared of 0.99 across both models. Developers looking to cut agent spend should look first at the prompt floor and the turn count before touching anything more sophisticated.

One caveat belongs on that regression figure. The June benchmark ran a single pass per harness, task, and model combination, so it carries no variance estimates, and agent runs are stochastic. Artificial Analysis carries more weight there, since it averages three attempts per task across 326 tasks. Consider the June overhead finding as evidence for the mechanism, and the larger index as the better instrument for comparing current production-scale pairings.

The expensive harnesses are not hoarding context; they are simply carrying a heavier floor.

What is surprising is what does not explain the spread. Every harness grew its context at a similar rate of a few hundred tokens per turn, so growth is not what separates them. The expensive harnesses are not hoarding context; they are simply carrying a heavier floor.

Cached tokens reorder the leaderboard

The two harness experiments converged on a second mechanism, and Artificial Analysis treats it as substantive in its methodology. It is the one platform teams are most likely to get wrong.

Composio disclosed that Claude Code drew only 1.5% of its input tokens from cache, against roughly 70% for Codex and 57% for OMP. Fresh input costs about five times as much as cached input. Claude Code’s token consumption, therefore, was comparable to its rivals’, while its invoice was not.

The June benchmark found a similar asymmetry in its own setup. On the DeepSeek run, Codex billed over a million tokens for the suite. Cache reads made up 77% of that, charged at roughly a tenth of the normal rate. Priced the way an invoice actually arrives, Codex came out cheaper per solved task than Claude Code, a harness that used half as many raw tokens.

The author attributed Claude Code’s near-zero cache share to the serving path rather than to its prompts. In that setup, Claude Code was the only harness communicating with OpenRouter via the Anthropic-style messages endpoint. The gateway’s translation of that dialect appeared to reduce the cache hits that identical traffic would have earned through the OpenAI-style endpoint. Gateway behavior changes, so treat that as one observed path.

Artificial Analysis builds that assumption into its cost model rather than discovering it. It warns that prompt cache hit rates vary widely with provider routing. Its cost model prices cached input and cache writes separately, rather than billing every prompt token at the uncached rate. A benchmark that has to price cache writes separately is telling developers how much the serving path matters.

Where the heavy harnesses earn their cost

Explaining the cost gap is not the same as crowning the cheapest harness. The agentic loop of look, act, check, and repeat is a cost multiplier. It helps most with unfamiliar code, failing tests, and changes across several files. On a small, well-described edit, the loop mostly re-confirms what a single call could have assumed.

But the hard-task test did not go the way the scaffolding argument predicts. Four harnesses spanning the cost range ran ten SWE-bench Lite tasks with generous limits, and all four resolved exactly one task, the same one. The author reported that Aider arrived at 0.8 million tokens, while Codex spent 15 million tokens. Ten tasks on one model is a thin basis for generalization, so consider it suggestive.

Quality is where the harness-sets-the-price framing needs qualifying, because the benchmarks argue against it. Composio reported pass rates across the eight harnesses, ranging from 46.7% for OpenCode to 66.7% for Pi Agent, a 20-point spread on one model. Artificial Analysis publishes the same kind of gap on a larger task set. Harness choice moves cost by multiples and moves task success by percentage points, which are different magnitudes rather than different directions.

What should platform teams measure?

Enterprises are already finding that owning the harness does not settle the cost question. Teams that built their own coding agents still pay for the underlying inference. Cost control has moved into the platform layer rather than the model contract.

QuestionWhat to measureWhy the obvious metric misleads
Which harness is cheaper?Cost per verified outcomeToken counts ignore pass rate and cache tier
Are we paying list price?Cached share on live trafficThe discount depends on the gateway and endpoint, not the harness alone
Will it survive a large prompt?Bytes transmitted versus bytes sentSilent truncation reads as success in the output

Cost per successful task, not cost per task

A pass rate and a token count, read independently, will incorrectly rank harnesses. Claude Code and DeepAgents finished the same number of Composio workflows, yet a successful Claude Code run cost more than four times as much. Procurement teams should ask for cost per verified outcome and refuse per-token comparisons.

Cached share on the actual serving path

Cache discounts are not a property of the harness alone. They depend on the endpoint dialect, the gateway, and the provider.

Cache discounts are not a property of the harness alone. They depend on the endpoint dialect, the gateway, and the provider.

Platform teams can verify the cached fraction on their own traffic in an afternoon, and that exercise is worth more than a model migration.

Prompt fidelity under load

The June benchmark injected 100,000 tokens of irrelevant log noise before each task and reported five distinct behaviors. Seven of the twelve configurations transmitted the prompt faithfully. Kilo and Opencode dropped 83%-89% of it while still reporting success.

Silent truncation is the dangerous one, because it looks exactly like success in the harness output.

OpenClaw refused to run, Kimi Code crashed, and Claude Code sent everything and then performed poorly. Silent truncation is the dangerous one, because it looks exactly like success in the harness output.

What comes next

Anthropic, OpenAI, Google, and Microsoft have already split on how to charge for the harness layer. DeepSeek then made every component of its own runtime swappable under an MIT license. Developers can now compare harness-model pairings in a publicly available, continuously updated index. That is the condition under which model pricing became contested in the first place.

For enterprises standardizing an agent platform this quarter, the harness deserves the same scrutiny as the model. In these experiments, harness choice produced cost spreads large enough to rival big differences in model pricing on workloads where the model held roughly steady. Developers, platform teams, and finance owners now have something reproducible to argue from. That is more than the harness conversation offered three months ago.

The post Aider, Claude Code, and OpenClaw ran an identical model. Token use varied 70-fold. appeared first on The New Stack.

  •  

How telemetry pipelines keep AI agent costs under control

Abstract neon pink and purple angular pathways interlock against a dark geometric background.

As enterprises move from experimenting with AI to running autonomous agents in production, an infrastructure problem is emerging: rising telemetry costs. Non-deterministic, iterative, and capable of generating data at machine speed, agents are far harder to monitor — and their costs far harder to predict — than conventional applications.

Many companies are struggling to attribute and defend their telemetry bills. In fact, 59% of organizations have already terminated or delayed an agentic AI deployment due to monitoring costs, according to a survey of more than 300 enterprise IT decision-makers in North America and Western Europe, commissioned by Apica and conducted by Omdia/Informa TechTarget. 

The agents most affected are often in some of the most high-stakes deployments: think cybersecurity, compliance, and fraud detection. As monitoring bills explode, deployments aren’t necessarily getting killed by engineering teams. More often than not, it’s finance pulling the plug.

Andi Mann, chief product and technology officer at Apica, recently saw this play out at a large bank. The organization couldn’t pin down exactly what it was spending on its AI programs.

“They knew they couldn’t afford to keep going on the same trajectory, so they had no choice but to cancel certain AI programs,” Mann tells The New Stack. “It’s a pattern I have seen before, because AI projects are cannibalizing typical budgets.” 

“It’s a pattern I have seen before, because AI projects are cannibalizing typical budgets.”

The implications are huge. As people, funding, and monitoring resources are diverted toward new AI workloads, other parts of the business start to suffer. Mann says he’s seeing outages, downtime, penetration attacks, and DDoS protection compete for the same resources.

The problem is already showing up in research. Most (54%) enterprises have seen telemetry volume triple in the past year alone, with 43% of that growth coming from AI/ML workloads — by far the largest driver. Businesses are under pressure to address this crisis as their observability bills climb.

They report spending an average of $3.17 million on observability, with that figure growing 28% year over year and showing no obvious ceiling. No wonder then that 83% rank AI observability as a top priority for the year ahead.

The coming wave could be catastrophic

With a new cloud service, database, or application, you’d expect the monitoring burden and telemetry data to rise by a relatively predictable increment. But enterprises now foresee an average 9.5X increase in telemetry data within two years.

“Imagine if your credit card or grocery bill went up more than nine times — that is not a marginal amount,” Mann says. “This isn’t a gentle ramp; it’s a skyscraper, and it’s prompting panic.”

“This isn’t a gentle ramp; it’s a skyscraper, and it’s prompting panic.”

Some 44% of organizations expect their telemetry data to increase by 6X to 100X.

The reason is that an agent task isn’t the same as a single application request. A customer-support task might generate a top-level trace, several model calls, retrieval operations, tool calls, retries, and loops. But if the agent delegates work to another agent, that adds another branch to the trace. 

Each model can produce token, latency, cost, and provider data, while each tool call generates its own records for arguments, results, status, and downstream activity. Identifiers such as tool_name, agent_id, and trace_id also create high cardinality, making data harder to aggregate and more expensive to index, with costs compounding at every stage.

That creates an uncomfortable gap between AI ambition and infrastructure readiness. Although 35% of enterprises claim widespread agentic AI deployment, operating and managing those systems is very different. Nearly two-thirds are only somewhat prepared for the shift. Unlike conventional applications, agents can call multiple models and tools, repeat tasks, or expand a workflow in unpredictable ways, making both capacity and monitoring costs difficult to forecast.

From extreme telemetry costs to an upstream control layer

The answer, according to Mann, is to intervene earlier. “You can’t keep sending essentially useless data to an expensive central analytics or storage platform because there’s no point analyzing data which says everything is fine,” he says. “As early as possible, get the pipeline to use data collectors and manage those at source.”

Legacy observability platforms were built around collecting data, ingesting it, storing and indexing it, and then analyzing it. That model worked for human-driven, dashboard-queried workloads, but agentic AI requires decisions to be made before telemetry reaches the most expensive parts of the stack.

A pipeline-first architecture can sample repetitive successful events while retaining failures, retries, policy violations, and unusually slow traces. It can enrich records with agent, session, model, tool, token, and estimated-cost data, redact sensitive prompts and identifiers, and aggregate metrics and long-term records to destinations with different cost and retention profiles.

A compact metric or sampled span might represent a routine successful tool call, while a failed call retains its parent trace, error details, retry history, and security context. The aim is to preserve the information needed to explain an agent’s behavior while reducing redundant data and limiting what gets indexed.

Agents also need millisecond-level context for autonomous decisions. Processing telemetry close to its source allows organizations to quickly detect retry storms, excessive tool loops, or unusual token consumption, rather than waiting for data to be ingested and indexed centrally. 

Without upstream control, businesses risk feeding fragmented, unnecessary telemetry into platforms that charge for every additional gigabyte, index, and retained record.

Architecture that separates winners from cancellations

The payoff for rethinking the telemetry pipeline is hard to ignore. Enterprises with a telemetry pipeline are 50% more likely to be prepared for the growth of agentic AI data. Among mature agentic AI organizations, pipeline adoption is what sets them apart: these organizations are 80% more likely to have avoided the operational cost challenges hampering their peers.

Rather than treating the observability platform as a catch-all destination, enterprises can make decisions about telemetry before it gets there.

The answer is to move the intelligence upstream. Rather than treating the observability platform as a catch-all destination, enterprises can make decisions about telemetry before it gets there — filtering out noise, enriching what matters, and routing data according to its value and purpose. That means less data hitting expensive storage and analytics systems, while the information that does make it through is more useful and available in real time.

Crucially, this isn’t about ripping out the observability platforms enterprises already rely on. It’s about putting a smarter control layer in front of them: deciding what data deserves to be ingested, where it should go, and how much it should cost.

Existing observability platforms still have an important role to play. “The pipeline can’t do everything, but it can act as a first responder,” says Mann. “You still work with the big analytics platforms, but you’re saving money, reducing risk, and improving your compliance performance.”

The Apica study finds that its pipeline control, metrics foundation, and data readiness services, for example, can reduce the total cost of ownership by 40% compared to legacy observability platforms. Of course, though, the actual savings will depend on telemetry volumes, retention policies, sampling rules, routing decisions, existing contracts, and the proportion of data that can be processed before ingestion.

The window to rethink the architecture is now. Some 68% of enterprises plan to evaluate changes to their observability stack within the next six months, while almost a quarter say existing vendor relationships won’t be a significant factor in those decisions. The next phase of observability will be won by the ability to handle what agentic AI throws at the infrastructure.

That means the pipeline can no longer be treated as plumbing that moves telemetry from A to B. It’s becoming the control layer for an increasingly autonomous, data-hungry environment. 

Organizations that establish agentic-ready infrastructure will be better positioned to reduce observability costs and improve risk performance. Mann doesn’t think platform engineering and SRE teams have much choice. “This has already become a board-level decision,” he says. “Ultimately, it’s a choice about how smart you can afford to make your business.”

Download the Omdia research report: “The Agentic AI Telemetry Crisis: Are You Ready for What’s Coming?”

The post How telemetry pipelines keep AI agent costs under control appeared first on The New Stack.

  •