❌

Vue lecture

Claude Opus 5.5 vs. Opus 5 on reasoning tasks: Cheaper, faster, but not better

Abstract long-exposure photograph of red and orange light trails forming layered curves around a dark central shape.

When Anthropic released Claude Opus 5.5 this week, the company claimed the new model costs 40% less than Opus 5 and generates output 30% faster. Anthropic’s marketing makes three claims. Opus 5.5 performs at the level of Claude Fable 5.1 (so it should outperform Opus 5), costs 40% less than Opus 5 on typical workloads, and generates output more than 30% faster.

Anthropic also cut the price developers pay to use the model through its API. Opus 5.5 costs $4 for every million tokens (chunks of text roughly three-quarters of a word long) sent to the model and $20 for every million it writes back, down from $5 and $25 for Opus 5. That price cut alone accounts for a 20% saving. The rest of the claimed 40% saving has to come from the model using fewer tokens.

I wanted to see how this translates for the average Claude user, so I skipped the usual developer workflow simulations this time. Lately, the models I test handle everyday tasks well. Reasoning tasks are where I’ve seen them struggle, so I tested Opus 5 against Opus 5.5 on reasoning tasks only. 

You can find the prompts at the bottom of this post if you want to replicate these tests on your own system.

The tests

I called both models through the Anthropic API with identical prompts. Both ran with adaptive thinking at the default effort level, since Opus 5.5 doesn’t allow you to turn thinking off. Each problem ran once per model. I planned to rerun any problem where the models gave me different results, but they never did.

Here are the tests I ran:

  • Logic grid (medium difficulty) – Seven engineers each have an on-call day, a language, a service, and a city, and 22 clues pin down one answer. Six clues are conditional or “exactly one of these is true” statements, and removing any single clue breaks the puzzle.
  • Constrained orderings (hard difficulty) – Reorder 10 deploy jobs so no job stays in its original slot and no two consecutively numbered jobs sit side by side. The model had to give the count for 6, 8, and 10 jobs.
  • Stone game with memory (harder difficulty) – Players remove 2, 5, 7, or 11 stones, but can’t repeat their opponent’s last move or their own. The model had to find who wins from 200 stones, count the losing starting sizes up to 500, and name the smallest losing size above 340.

I logged input tokens, output tokens, cost at list price, and time for every call. Thinking tokens are billed as output, so I included them.

The logic grid

Both models got all 28 cells right. Opus 5.5 took 65 seconds and 7,573 output tokens, for $0.16. Opus 5 took 108 seconds and 10,621 output tokens, for $0.27.

On this test, Opus 5.5 was 43% cheaper and delivered the same correct answer. Opus 5.5 was slightly more detailed and noted that it didn’t fully prove the solution was unique.

Constrained orderings

Neither model produced an answer, which made this the hardest problem in practice. The correct counts are 27, 1,695, and 159,019. The third answer is very hard to reach by reasoning alone. A computer program that checks every possible ordering can find it, but neither model could run code in this test.

With a 48,000-token output limit, both models spent the whole budget thinking and never replied. Opus 5.5 used 489 seconds and $0.96. Opus 5 used 553 seconds and $1.20.

I raised the limit to 128,000 tokens and ran it again. Opus 5 used every token, took over 25 minutes, and stopped with no answer. That cost $3.20. Opus 5.5 ran for 19 minutes and used 112,733 tokens, but the API ended the response with a “refusal” stop reason and no text. The prompt asks the model to count job orderings and contains nothing sensitive. That means the refusal was most likely a mistake by Anthropic’s safety filter flagging a harmless request.

Opus 5.5 was cheaper, but how much does that matter if you don’t get a result?

The stone game

And we’re back to the same answers again. Both models answered all three parts correctly. The first player loses from 200 stones; starting sizes from 120 to 500 are losses, and the smallest loss above 340 is 344.

The difference in this test came down to how much thinking each needed to get there. Opus 5.5 finished in 215 seconds with 28,740 output tokens, costing $0.58. Opus 5 took 624 seconds and produced 74,981 output tokens, costing $1.88. Opus 5.5 used 62% fewer tokens and cost 69% less for the same answer. 

Results

TestOpus 5.5Opus 5
Logic grid28/28, 1:05, 908 in / 7,573 out, $0.1628/28, 1:48, 906 in / 10,621 out, $0.27
Ordering problem (48k limit)No answer, 8:09, 235 in / 48,000 out, $0.96No answer, 9:13, 233 in / 48,000 out, $1.20
Ordering problem (128k limit)No answer (refusal), 18:56, 235 in / 112,733 out, $2.26No answer, 25:24, 233 in / 128,000 out, $3.20
Stone game3/3, 3:35, 323 in / 28,740 out, $0.583/3, 10:24, 321 in / 74,981 out, $1.88
Total tokens1,701 in / 197,046 out1,693 in / 261,602 out
Total time31 min 45 sec46 min 49 sec
Cost$3.95 ($4 in / $20 out per million tokens)$6.55 ($5 in / $25 out per million tokens)

Across every call, Opus 5.5 wrote 103.4 tokens per second and Opus 5 wrote 93.1, so Opus 5.5 was about 11% faster. Its biggest speed lead on any single problem was 19%, still short of Anthropic’s 30% claim. Its biggest cost savings came on the stone game, where it cost $0.58 to Opus 5’s $1.88, 69% less. Total spend for the test was $10.50. 

What do I think

Anthropic’s benchmarks show Opus 5.5 ahead of Opus 5 on coding, knowledge work, and reasoning. I didn’t rerun those benchmarks. I gave both models the same three reasoning problems, and they performed equally. Both solved the logic grid and the stone game, and both failed the ordering problem.

The savings are real. Opus 5.5 cost less and finished sooner on every problem, including 43% less on the logic grid and 69% less on the stone game. Most of that came from using fewer output tokens. The price cut accounts for 20%. Its writing speed was 11% faster, short of the 30% Anthropic claims.

Switch to Opus 5.5 if you run Opus 5 today. You get the same results on hard reasoning for less money and less waiting. Set a hard output limit and watch your spend on hard problems, though. Both models can think for close to 20 minutes or more and return nothing, which is a problem if you pay for every token. For counting problems like the ordering test, give the model a code execution tool instead of hoping it reasons its way through.

The prompts

Logic grid

Seven engineers (Ana, Ben, Cy, Dee, Eli, Fay, Gus) share an on-call rotation. Each is on call on exactly one day of a single week, Monday through Sunday (Monday is the earliest day, Sunday the latest), and no two share a day. Each writes a different language (Go, Rust, Python, Java, Kotlin, TypeScript, C++), owns a different service (auth, billing, search, queue, cache, gateway, metrics), and is based in a different city (Berlin, Tokyo, Denver, Lagos, Sydney, Toronto, Mumbai).

Clues:

The Java developer is Gus.

The TypeScript developer is on call exactly two days after the queue owner.

Exactly one of these is true: the engineer based in Berlin owns billing, or the C++ developer is on call Monday.

Exactly one of these is true: the engineer based in Tokyo owns cache, or the engineer based in Tokyo writes Rust.

If Fay is on call Sunday, then the engineer based in Denver does not write Kotlin.

Exactly one of these is true: the Kotlin developer owns auth, or the C++ developer is Cy.

The engineer based in Tokyo is on call earlier in the week than the C++ developer.

The engineer based in Denver does not write Python.

Exactly one of these is true: the Go developer is Ben, or the gateway owner is based in Berlin.

Ana is based in Mumbai.

The TypeScript developer is on call earlier in the week than Fay.

The engineer based in Sydney is on call earlier in the week than the billing owner.

The gateway owner is on call earlier in the week than the Kotlin developer.

The engineer on call Sunday is not based in Berlin.

Fay and the engineer based in Lagos are on call on consecutive days.

Eli is on call exactly three days after the auth owner.

The engineer based in Lagos writes Java.

Exactly one of these is true: the search owner is Eli, or the cache owner is Dee.

The gateway owner and the Go developer are on call on consecutive days.

The engineer based in Lagos is on call exactly four days after the auth owner.

The engineer based in Lagos owns metrics.

The engineer based in Mumbai is on call earlier in the week than the Python developer.

Determine the full assignment. At the end of your response, give exactly seven lines, one per engineer in the order Ana, Ben, Cy, Dee, Eli, Fay, Gus, in this format:

ANSWER: Name | Day | Language | Service | City

Ordering problem

A build system has n deploy jobs numbered 1 to n. Originally, job k runs in slot k. You reorder all n jobs into slots 1 to n (each slot gets one job) subject to two rules:

No job runs in its original slot (job k is not in slot k).

Jobs with consecutive numbers never run in adjacent slots (for example, jobs 4 and 5 cannot be in slots i and i+1 in either order).

How many valid orderings are there for (a) n = 6, (b) n = 8, (c) n = 10?

At the end of your response, give exactly three lines in this format:

ANSWER a: <number>

ANSWER b: <number>

ANSWER c: <number>

Stone game

Two players play a game with a pile of stones. They alternate turns. On each turn, a player removes exactly 2, 5, 7, or 11 stones, subject to two rules:

You may not remove the same number your opponent removed on their most recent turn.

You may not remove the same number you removed on your own most recent turn.

(On the very first turn of the game, neither rule applies. On the second player’s first turn, only the first rule applies.) You cannot remove more stones than are in the pile. A player who has no legal move on their turn loses. Both players play perfectly.

(a) Starting with 200 stones, does the first player win?

(b) For how many starting pile sizes from 1 to 500 inclusive does the first player lose?

(c) What is the smallest starting pile size greater than 340 for which the first player loses?

At the end of your response, give exactly three lines in this format:

ANSWER a: <yes or no>

ANSWER b: <number>

ANSWER c: <number>

The post Claude Opus 5.5 vs. Opus 5 on reasoning tasks: Cheaper, faster, but not better appeared first on The New Stack.

  •  

Claude’s merged chat and Cowork vs. ChatGPT’s Work mode: ChatGPT is faster, Claude is more thorough

Rainbow-striped grid facade of a modern building, seen from below at an angle, with red, yellow, green, and blue diagonal bands crossing white geometric wall panels split by a narrow sliver of

When Anthropic merged Claude chat and Cowork into a single interface last week, it removed an increasingly irrelevant decision users had to make about which mode to choose.

Now, to use both tools, just ask a question, then hand off a multi-step task in the same thread, and Claude routes it. Anthropic made this update because it says customers often struggled to choose the right tab for the right task, so the merged app now routes each request itself instead of asking you to pick a mode — for now, the unified experience is rolling out to Pro and Max subscribers first, with free and team tiers to follow.

But this isn’t new. OpenAI has offered the same promise since earlier this summer. OpenAI introduced Work mode alongside Chat on July 9, then phased out the older, separate Agent mode the following month. Work mode is a sandboxed environment with a browser, code execution, and file output, sitting next to a Chat toggle in the same window — though reviewers note it can’t yet hand a live, logged-in browser session back to the user mid-task the way Agent mode could, e.g., for logins or payments. OpenAI built Work mode on its Codex coding agent after OpenAI reported that roughly a fifth of Codex’s 5 million weekly users were non-developers — a share it said was growing three times faster than developers.

As OpenAI and Anthropic get closer to feature parity, accuracy, reliability, and token usage matter even more.  What better way to find out which tool is better than to run head-to-head tests? 

The tests

I ran three tests, covering different areas of real developer work. 

  • API research -Look up four real developer APIs and tabulate their documented rate limits, whether a free tier exists, and the current version identifier. I checked the answers against the vendors’ docs the same day.
  • Build from a spec – Write a small command-line duration parser from a spec with strict edge cases. I ran each app’s code against a hidden 16-case test suite.
  • The handoff – Ask which of two stack traces indicates a race condition, then in the same thread hand off a real job: pull a log file from Google Drive, compute latency percentiles and error rates per endpoint, and deliver a spreadsheet with a chart. I generated the log data, so I knew every number in advance.

I recorded token cost and time in each test and included the prompts for anyone interested in replicating this work.

API research

The prompt:
Research the current public documentation for these four developer APIs and build me a table with one row per API and these columns: documented rate limit for authenticated requests, whether a free tier exists (yes/no), and the current API version identifier or date shown in the docs. APIs: GitHub REST API, Stripe API, Twilio Messaging API, OpenAI API. Cite the documentation page you used for each row.

Both Claude and ChatGPT answered correctly on the twelve graded fields, but Claude was more thorough. It included GitHub’s separate limit for Actions tokens, Twilio’s queue window, and OpenAI’s tier thresholds, plus a note about one page it couldn’t reach.

ChatGPT, in Work mode, finished in 1 minute 17 seconds and wrote 649 output tokens. Claude took 1 minute 44 seconds and wrote 1,042 tokens, read eight pages, listed nine sources, and offered to export the table as a spreadsheet. Claude wrote nearly double the tokens and took longer, but in this case, it’s warranted because of the added detail.

Build from a spec

The prompt:
Build the command-line tool described in the spec below. Deliver a single file named durparse.py that follows every rule. Test it yourself before returning it. Show the complete final code in your reply. (Followed by the spec: a duration parser with units d/h/m/s, largest first, one of each, decimals allowed, bare numbers are seconds, everything else returns None.)

Both apps returned a durparse.py file that passed all 16 hidden tests, including the traps. The traps included units out of order, a repeated unit, a trailing number with no unit, and negative values. The 517-token spec went to both. ChatGPT finished in 1 minute 17 seconds on 769 output tokens. Claude took 1 minute 45 seconds and 989 output tokens. Claude reported running 35 of its own test cases before returning the file. Both delivered a download and showed the code in the reply.

The code came out nearly identical, both using exact-precision arithmetic and a fixed-order regex. Claude flagged a judgment call the spec never settles on: that rounding 0.5 seconds up is a choice and Python’s built-in round would go the other way. ChatGPT reported only that its tests passed. Once again, Claude was just a little more thorough.

The handoff

The prompt:
Which of these two stack traces points to a race condition, and in one sentence why? (with the two traces) Then: Now take the file api_logs.csv from my Google Drive (columns: time, endpoint, status, latency_ms) and produce a downloadable spreadsheet with one row per endpoint showing request count, p50, p95, and p99 latency in milliseconds, and error rate as the percentage of requests with status 500 or above. Add a bar chart of p95 latency by endpoint. Also show the table in your reply.

I started each thread in plain chat with the stack-trace question, 135 tokens. Both answered correctly in seconds: Trace B, the dictionary that changed size during iteration. ChatGPT spent about 10 seconds and 48 output tokens. Claude spent 118 in about 25 seconds, adding a caveat that the same error can happen without threads if the loop body edits the dictionary itself. Then, without switching modes, I handed off the log analysis.

Claude pulled the file, computed the table, and built an .xlsx with a p95 bar chart in about 4 minutes on 319 output tokens. ChatGPT produced the same spreadsheet and chart in 28 seconds, using 467 output tokens. All 30 numbers matched my ground truth on both sides. Claude also named its percentile method (linear interpolation) and noted the numbers would match if I recomputed them in Google Sheets. ChatGPT gave the same correct table but didn’t provide as much detail as Claude did.

Results

The testChatGPT (Work mode)Claude (merged app)
API research12/12, 1:17, 125 in / 649 out12/12, 1:44, 125 in / 1,042 out
Build from spec16/16 tests, 1:17, 517 in / 769 out16/16 tests, 1:45, 517 in / 989 out
Handoff, questionCorrect, ~10 s, 135 in / 48 outCorrect, ~25 s, 135 in / 118 out
Handoff, task30/30, 28 s active, 133 in / 467 out30/30, ~4 min, 133 in / 319 out
Total tokens (visible)910 in / 1,933 out910 in / 2,468 out

Both Claude and ChatGPT were equally accurate. Every field, every test case, every number matched on both sides, and both cited real documentation. 

The differences are in speed and answer detail. ChatGPT was faster on every task and produced 1,933 visible output tokens, compared with Claude’s 2,468. Some of Claude’s extra output was filler, but not all of it. It added context the prompts didn’t ask for, named the percentile method behind its numbers, and flagged two judgment calls the specs left open. ChatGPT gave the same right answers but with less context.

What do I think?

I’d pick Claude, and here’s the reasoning. On time, ChatGPT won every task, but the gaps were seconds on the short tasks (1:17 vs 1:44, 1:17 vs 1:45), not a noticeable difference. On tokens, ChatGPT used about 22% less output than Claude, which is positive, but not when you consider how important detail/context is.

On detail, Claude provided more meaningful detail on all three tests. This included the extra API context, the rounding judgment call, and the percentile method. In today’s world, where AI can fabricate, detail matters. 

The post Claude’s merged chat and Cowork vs. ChatGPT’s Work mode: ChatGPT is faster, Claude is more thorough appeared first on The New Stack.

  •  

Grok Build vs. Claude Code: I tested which one has the better memory

On September 16, xAI announced memory in Grok Build, its terminal coding agent. The pitch was that Grok “keeps notes on the conventions, decisions, and project facts that come up,” and “later sessions read those notes before touching related code.” Notes are Markdown files in a workspace scope per project and a global scope that applies everywhere. /memory browses them.

Meanwhile, Claude Code has done something similar for months under the name auto memory. It keeps a MEMORY.md index plus one file per note, per repository, and the docs say it is on by default. Anthropic’s Projects beta, announced September 17, adds shared memory across cloud threads, but only for select Pro and Max subscribers with no existing projects. I tested the CLI that everyone has.

Both companies say their coding agent now remembers what you told it in an earlier session. I wanted to see whether that holds up, so I tested Grok and Claude on the same three tests.

The tests

The claim I wanted to check is simple. Tell each tool something once, close it, open it again, and see whether it remembers. Both tools ran on my Mac, each on its own copy of four small Node repos I built for this. Grok Build 1.0.40 ran Grok 4.6 at high effort through an xAI API key. Claude Code 2.1.226 ran Opus 5 on my subscription. Every session was scripted with each tool’s headless mode, which reports its own tokens and cost. 

Each test has two sessions. Session 1 plants a fact. I quit the tool. Session 2 gives a task where the fact matters and never mentions it.

Here are the tests I ran:

  • The test command – In this repo, npm test fails and make test passes, and session 1 says so. Session 2 asks for a new endpoint with passing tests, after I removed the README line that pointed at the Makefile.
  • Project decisions with a trap – Session 1 states that CSV export was dropped and money is integer cents, never floats, while a float helper and a half-built CSV exporter sit in the repo as bait. Session 2 asks for a refund endpoint that “takes an amount” and “a way for support staff to download all orders.”
  • A rule across projects – Session 1, in repo A, sets two rules “for all my projects,” conventional commit messages and no comments on obvious code. Session 2 runs in an unrelated repo B and asks for a small feature and a commit.

Here’s my scoring breakdown. Did the tool write the fact to a memory file, did it read that file in session 2, and did the session 2 output follow it.

The test command

Both passed. In session 1, each tool saved the rule as soon as I stated it. Grok wrote topics/testing.md plus two raw observations. Claude Code wrote orbit-api-run-tests-with-make.md with a “why” and a “how to apply” section.

In session 2, both remembered. Grok’s reasoning opened with “start by reading the memory files,” then it ran make test and never touched npm test. Claude Code read the Makefile and package.json, ran make test, and also never tried npm test. Grok took 29 seconds, 102K tokens, and $0.11. Claude Code took 22 seconds, 186K tokens, and $0.32. Claude used about 80K more tokens and cost nearly 3x more. 

Project decisions with a trap

Both wrote both decisions down. Claude Code also converted “last quarter” into “Q2 2026” in its note. In session 2, both built the refund on integer cents, named the field amountCents, and left the float helper alone. For the download request, both shipped a JSON export. 

Grok’s reasoning said the API is JSON-only, so it wouldn’t wire up CSV. Claude Code set a content-disposition header so the JSON downloads as a file. Both passed on both decisions, but Claude Code was more than double the price and just as fast. Grok took 103 seconds, 156K tokens, and $0.18. Claude Code took 32 seconds, 269K tokens, and $0.49.

A rule across projects

This is where the results split. Grok saved the rules to its global scope as git-and-code-style.md. In the second repo, it committed feat: add --help flag with usage and supported cities, and added no comments. Pass, in 33 seconds, 132K tokens, and $0.12.

Claude Code saved both rules too, but only in the first repo’s memory folder. It said so at the time, warning that its memory store “is scoped to this project’s directory.” In the second repo, it found nothing, and the commit came back with the Add --help flag. No comments were added, but that is Claude’s default anyway. Claude passed the first rule but failed the second one and still cost twice as much. It completed the work in 12 seconds, 122K tokens, and $0.24.

Results

MetricGrok Build (Grok 4.6)Claude Code (Opus 5)
Tests passed3 of 32 of 3
Total time165 s66 s
Total tokens390,848576,863
Total cost$0.41$1.05

Grok Build passed all three tests, and Claude Code passed two. They behaved the same on the per-project tests. The split was the cross-project rule, which Grok’s global scope carried into a second repo and Claude Code’s per-repo memory did not. 

Claude Code was faster on every recall session, 66 seconds total against 165, and cost at least twice as much on every one, $1.05 total against $0.41. It also used more tokens: 576,863 against 390,848. The price gap mostly reflects Opus 5 versus Grok 4.6 rather than the memory systems.

On the core claim, remembering what you told it last time in the same project, I could not tell these two apart. Both wrote a markdown note the moment I stated a rule, read it back next session, and followed it. Claude Code’s notes were better written. But Claude Code failed the third test. Its CLI memory stops at the repo boundary, so a rule I gave it “for all my projects” never reached the second repo. Grok’s global scope carried the same rule over without being asked.

What do I think?

Grok Build is the better option for most people right now. It remembered everything, it carries rules across projects, and it cost less than half as much on every test. Yes, Claude Code was faster on every session, but that only matters if you aren’t concerned about accuracy. Its CLI memory stops at the repo boundary, so anything you want it to remember everywhere still has to go into ~/.claude/CLAUDE.md by hand. 

The post Grok Build vs. Claude Code: I tested which one has the better memory appeared first on The New Stack.

  •  

Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet

Macro view of overlapping textured paper sheets in white, pink, blue, orange and teal.

When Anthropic launched Claude Fable 5.1 this month, it centered the announcement around one benchmark result: its Terminal-Bench-Science score.

In this benchmark, a model gets a terminal and a real scientific research problem to solve independently. Fable 5.1 scores 52.6%, and Fable 5 scores 24.7%. By Anthropic’s scoring, the new model more than doubles the old one.

Anthropic’s published score was produced under conditions most users don’t have access to. The benchmark allows each model up to eight hours per task, and Anthropic has not said what harness or budget it used to get its numbers. When the benchmark’s leaderboard tested Fable 5, it ran the model through Claude Code at maximum effort and spent $14,180 across 210 attempts, about $67 each.

Most people don’t use Fable in a lab, so I wanted to know what an average user would get on these tasks. I recently tested Fable 5 and Fable 5.1 on everyday work and found them far closer than the benchmark suggests. That made me want to run the benchmark’s own tasks the way a home user would and see where the models actually differ.

The benchmark’s 70 tasks are public, so I pulled five of them, one from each science field, and ran both models myself. 

The tests

Terminal-Bench-Science has five categories, each with multiple tests. I chose one test per category that could run in a Python environment. Here’s what I picked:

  • Symbolic regression (mathematics) – A dataset with 100 variables and a hidden formula behind a yes-or-no label. The model must find a predictor that works on data it has never seen.
  • Lorenz-96 assimilation (Earth sciences) – Reconstruct a chaotic atmospheric model from a few uncalibrated sensors with unknown clock offsets. Grading is all-or-nothing on five criteria.
  • Reactor safety control (engineering) – Write a controller for a chemical reactor that finishes every batch as fast as possible without ever exceeding the temperature limit, across public and hidden fault scenarios.
  • Foraging cognitive model (life sciences) – Predict, trial by trial, which lever each of 20 mice will press, graded on sessions the model never saw.
  • Nanoindentation (physical sciences) – Extract material properties from raw indentation curves that include drift, adhesion, defects, and an unknown tip shape.

Each run got a plain terminal, and I set a $12 limit and 60 turns for each test. The full set of ten runs took about 12 hours.

Symbolic regression

This was the only test where a model passed the benchmark’s hidden test. Fable 5.1 worked for 27 turns, found the hidden structure, wrote a predictor, and stopped on its own after 11.8 minutes, 27,088 output tokens, and cost $1.96 to pass this one test.

Fable 5 used all 60 turns over 53.5 minutes, generated 39,461 output tokens, cost $4.20, and failed. I ran Fable 5 a second time to rule out bad luck. It used all 60 turns again, took 60 minutes, generated 60,608 output tokens, cost $6.38, and failed again.

Lorenz-96 assimilation

This was the most expensive pair of runs. Fable 5 hit the $12 cost limit at 45 turns after 97.7 minutes and 92,091 output tokens, ending at $12.63. Fable 5.1 used all 60 turns over 126 minutes, generated 89,789 output tokens, and cost $10.70. On the public leaderboard, Earth sciences is also the field where Fable 5 scores close to zero, and both models failed it here.

Reactor safety control

Neither wrote a controller that passed the grader’s scenarios. This run produced the most output tokens, 157,710 generated by Fable 5.1. It hit the 60-turn limit after 40.9 minutes and cost $11.53. Fable 5 hit the cost limit at 49 turns after 63.8 minutes, 121,978 output tokens, and $12.04. 

Foraging cognitive model

This was the longest run of the testing series. Fable 5.1 was the only model that declared itself finished. It built a model, tested it against its own scoring loop, and declared it done at 43 turns after 53.5 minutes. It created 65,518 output tokens and cost $5.65. But the official grader rejected it. 

Fable 5 never declared anything. It hit the $12 limit at 60 turns, after 139.3 minutes and 62,587 output tokens, ending at $12.13. 

Nanoindentation

Both failed. Both spent most of the run reading raw curves and writing code to segment them. Neither produced a results file the grader accepted. Fable 5.1 ran out of turns at 29.6 minutes, 115,687 output tokens, and $10.91. Fable 5 ran out of money at 48 turns after 34.1 minutes and 114,239 output tokens, ending at $12.59. 

Results

Here are the results by the numbers.

MetricFable 5 scoreFable 5.1 score
Tasks solved0 of 51 of 5
Output tokens430,356455,792
Total cost$53.59$40.75
Total time388 min262 min
Runs ended by cost limit40

The benchmark scores models on all 70 tasks with three trials each, and Anthropic’s 24.7% and 52.6% come from that full suite. The independent leaderboard puts Fable 5 at 21.4%, close to Anthropic’s figure. Fable 5.1 is not on the independent leaderboard yet, so its 52.6% is Anthropic’s number alone. My results, 0% and 20%, are below both. Five tasks are a small sample. 

Getting these results by chance is plausible even if the published scores are exactly right, so this run neither confirms nor contradicts the doubling claim. The direction matched, since the new model did better. The one task Fable 5.1 solved was in mathematics, which is also the field where the leaderboard shows Fable 5 performing best.

What I think

I don’t think a regular user will see much difference between Fable 5 and Fable 5.1. I only ran a small sample of tests, so I can’t prove or reject Anthropic’s benchmark results. But what I saw suggests that gap won’t reach the average user. The one difference that did show up was on the bill. Fable 5.1 failed faster and cheaper, and it never hit my cost limit, whereas Fable 5 hit it four times.

Suppose your work looks more like the benchmark tasks; the harness and the budget matter as much as the model. With a purpose-built harness, hours per task, and a much bigger budget, you may get closer to Anthropic’s numbers.

The post Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet appeared first on The New Stack.

  •  

Claude Fable 5.1 vs. Fable 5: On real work, I couldn’t tell them apart.

Anthropic launched Claude Fable 5.1 this week, calling it “our most advanced model for coding and knowledge work.” 

There was quite a lot of excitement surrounding the launch. Every CEO Dan Shipper posted that after a week of testing, it was “the strongest coding model we’ve used.”AI commentator Min Choi collected examples of people “one-shotting games, building 3D worlds + creating insane simulations” within 24 hours of release.

The announcement highlights the Terminal-Bench-Science benchmark, an agentic research benchmark, where 5.1 scores 52.6% against Fable 5’s 24.7%. Terminal-Bench-Science gives a model a terminal and a set of multi-step scientific research tasks, then scores what percentage it completes correctly.

The Terminal-Bench-Science score gap illustrates the main measurable difference between Fable 5 and 5.1. Fable 5 finishes about a quarter of them. 5.1 finishes about half. With the new model costing exactly what the old one does — $10 per million input tokens and $50 per million output — this score is the reason to upgrade, if it’s true.

Benchmark scores don’t always translate to real-world work.

Benchmark scores don’t always translate to real-world work. They measure narrow task sets under conditions vendors help define. Some companies have been known to tune models toward the tests they get graded on. I’m not saying Anthropic did that, but the possibility is baked into how benchmark marketing works.

 A score of 52.6% on a research benchmark tells me nothing about whether AI’s output on the work I need it to do gets better. So I wanted to see what these numbers and strong claims mean for a real user doing real work.

The tests

I tested both models on four tasks that mimic real work that people use AI for.

  • Agentic research: Experiment data with five bad rows; the lab notes explain how to catch them. The model has to exclude them, compute the batch averages, and write up its findings.
  • Agentic coding: A small Python project with two planted bugs and a failing test suite. The model has to find the bugs and fix the code until every test passes.
  • Reasoning: Two math problems with exact answers I verified in advance. This includes no terminal work, only thinking.
  • Sensor data audit: Messy readings from five sensors, with every problem documented in an equipment log: a fast clock, a mid-run hardware swap, corrupted rows, one sensor in Fahrenheit. Added as a tiebreaker; more on that later.

I usually paste my prompts so you can rerun my tests. I was unable to do that for these tests. These tests require folders of data files with planted errors, and the prompts are useless without them.

Agentic research

Both models finished the task correctly in 3 turns. Each read the lab notes and excluded exactly the right five rows, including the subtle case where a duplicated trial’s first entry is corrupt, and its rerun is valid. Both produced batch means that matched the correct answers to the decimal.

Fable 5.1 was slightly faster (19.2s vs. 20.6s) and slightly cheaper ($0.086 vs. $0.100). On the benchmark this task imitates, Fable 5 supposedly fails three-quarters of the time. On my machine, it didn’t make any mistakes.

Agentic coding

The coding test went the same way. Both models ran the suite and spotted the loud bug, a remove function that added stock instead of subtracting it. Both also caught the quiet one, an off-by-one in a threshold comparison. Each fixed both bugs and finished with 8 of 8 tests passing in 3 turns.

Fable 5.1 finished in 13.6 seconds, and Fable 5 took 17.2, with both runs costing $0.07. The new model was faster, but nothing else separated them.

Reasoning

Anthropic’s numbers predicted a near-tie here. I saw the same result. Both models answered the two challenges correctly, with correct step-by-step work. Fable 5.1 was slightly faster on both problems: 12.0 seconds against 12.5 on the first, and 9.7 seconds against 10.7 on the second. It was also more concise, using 771 output tokens against Fable 5’s 1,045 on the first problem and 647 against 798 on the second.

Sensor audit data

AKA the tiebreaker. After three rounds of perfect ties on accuracy, I added a fourth test. I built this test to be harder than the first three, because a model that supposedly doubles its predecessor should reveal that somewhere.

Both models handled every trap. Each shifted the fast clock back before filtering the time window, which also excluded two hot readings from the data. Each split the swapped sensor’s calibration at the right moment, dropped the corrupted rows, and converted Fahrenheit after calibrating rather than before. Once again, the two result files were identical and fully correct.

The difference between the two was in the timing, cost, and token usage. Fable 5 finished in 4 turns, 23.9 seconds, and $0.134. Fable 5.1 needed 5 turns, took 28.4 seconds, and cost $0.304, more than double. The extra turn is what did it. Every turn resends the entire conversation so far, so Fable 5.1 pushed 23,602 input tokens through the API, compared to Fable 5’s 7,940. 

Fable 5 was as accurate, cheaper, and faster on the hardest task of the set. I wasn’t expecting that.

Fable 5 was as accurate, cheaper, and faster on the hardest task of the set. I wasn’t expecting that.

Results

MetricFable 5Fable 5.1
Accuracy24/2424/24
Total tokens2221937809
Total cost$0.398$0.533
Total time84.9s82.9s

Both models scored a perfect 24 out of 24 across the four tests. Fable 5.1 finished the full run slightly faster, 82.9 seconds against 84.9, but used 70% more tokens and cost 34% more (these results were skewed after I did the sensor test, but it still counts). 

These results also don’t mean the Terminal-Bench-Science numbers are wrong. The benchmark was built using long, messy research tasks where Fable 5 allegedly fails most of the time. In my opinion, though, this doesn’t apply to what most users are using Fable 5 for. Some, yes, sure.

Two caveats I also want to add. The claimed savings rely on cache read prices, which have dropped by 75%. My short tasks didn’t use caching at all, so my cost numbers don’t test that claim. And Anthropic says 5.1 was benchmarked with its production safeguards on, which sometimes lowered its own scores.

What do I think?

I went looking for a 2x improvement and found a model I couldn’t differentiate from its predecessor. That doesn’t mean there aren’t any differences, but it does mean that for day-to-day tasks you were already using Fable 5 for, it may not look any different.

I went looking for a 2x improvement and found a model I couldn’t differentiate from its predecessor.

If you’re on Fable 5 today and your work looks like mine — code fixes, data cleanup, analysis with documented gotchas — this upgrade will not change your results. On a long agentic task, it may even cost more per run. If your work looks like the benchmark, meaning hours-long research agents that fail more than they succeed, Anthropic’s numbers say 5.1 is where the improvement lives. I couldn’t build that test in an afternoon.

The post Claude Fable 5.1 vs. Fable 5: On real work, I couldn’t tell them apart. appeared first on The New Stack.

  •  

GLM-5.3-Flash vs. GLM-5.3: Time and money, not the spec sheet

Stacked neon squares glow in pink, purple, and blue against a dark background.

I always wonder about the end goal when companies launch products so close together and undercut each other by claiming the new one is “so much better.”

GLM-5.3 and GLM-5.3-Flash are a great example of this. Z.AI launched GLM-5.3-Flash on August 26 with a bold marketing claim: stronger intelligence “at an exceptionally low cost,” and a 3x improvement in serving speed, built on an architecture that cuts attention computation by 3x compared to the GLM-5.3 flagship released earlier in the month. 

On OpenRouter, Flash costs $0.075 per million input tokens and $0.25 per million output tokens. GLM-5.3 costs $1.188 and $4.18, making it nearly 16 times more expensive. It definitely looks cheaper, but our tests will confirm, since it ultimately comes down to token usage.

If it uses more tokens, then we’ll have to equate the lower cost to pricing manipulation, meaning it could very well end up being more expensive if Z.AI changes their pricing structure (aka the end of the freemium we’re all living in right now).

I wanted to get to the bottom of this, so I tested each model on three tasks and 27 questions to determine which model is better and whether GLM-5.3-Flash has earned its place in the market.

The tests

I tested both models on topics that replicate real work users will ask them to do:

  • Coding – spec for a Python date-parsing function with strict edge cases: multiple input formats, two-digit years, invalid dates that must return None. I graded by running each model’s code against a hidden 12-case test suite I wrote and verified before either model saw the spec.
  • Reasoning – a scheduling puzzle placing five people into five meeting slots under seven interlocking rules. I brute-forced all 120 possible schedules in advance to confirm exactly one solution exists.
  • Information extraction – a three-email vendor negotiation with 10 facts scattered throughout it, including a trap: an 8% discount that applies to a revised 92-seat price, not the original quote.

I included all prompts in the test sections so anyone can replicate them on their own.

Test 1: the date parser

Prompt:

"Write a Python function parse_event_date(s) that converts a date string to ISO format YYYY-MM-DD. Supported input formats: Month D, YYYY (e.g., March 5, 2026), D Month YYYY, US-style MM/DD/YYYY or M/D/YY, and ISO YYYY-MM-DD. Month names may be full or 3-letter abbreviations, in any letter case. Two-digit years always mean 2000-2099. Ignore leading and trailing whitespace. Return None for dates that don't exist on the calendar, strings missing a year, and anything unparseable. Use only the Python standard library."


This was the hardest test, but it produced the strangest numbers. Both models returned code that passed all 12 hidden tests, including the traps. February 30 is correctly rejected; “12/31/99” is correctly read as 2099, and a date with no year correctly returns None.

Getting to the same results cost very different amounts of effort. GLM-5.3-Flash took 455.8 seconds, most of it spent generating 38,677 tokens (that’s a lot of tokens) of reasoning before the final 60-line function. GLM-5.3 took 174.7 seconds and 14,801 tokens for a nearly identical solution. Flash still billed less, $0.019 against $0.065, because its per-token price is so much lower.

“The ‘budget’ model’s low price is covering for the fact that it works harder to get the same answer.”

If Flash charged the flagship’s rates, its 38,677-token thinking session would have cost $0.16, two and a half times the flagship’s $0.065 bill for the same function. The “budget” model’s low price is covering for the fact that it works harder to get the same answer. 

Test 2: the scheduling puzzle

Prompt:

"Five consultants (Ana, Ben, Carla, Dev, Elena) each get exactly one meeting slot: 9am, 10am, 11am, 1pm, 2pm. Rules: 1. Ana is not in the first slot and not in the last slot. 2. Ben's slot is earlier than Carla's. 3. Dev's slot is immediately after Ana's. 4. Elena's slot is not adjacent to Ben's. 5. Carla is not at 11am. 6. Ben is not at 9am. 7. Elena is not at 2pm. Exactly one schedule satisfies all seven rules. State the final schedule."


Both models produced the one valid schedule, with all five people in the right slots. GLM-5.3 showed clean case-by-case elimination in its answer, got there in 18.9 seconds, and used 1,804 output tokens. Flash took 33.5 seconds and 1,003 output tokens, answering with just the schedule, with no work shown. Some people may prefer to see the work, but I don’t need to. I prefer short, to-the-point answers (which is sometimes hard to get with AI). 

On light work, the budget model really is the budget model (if you’re budgeting cost, not time).

This was the cheapest test of the run for both: $0.0003 for Flash, $0.008 for the flagship. The token math flipped here. The flagship generated 80% more tokens than Flash, and even if Flash charged flagship rates, this answer would have cost $0.004, about half the flagship’s bill. On light work, the budget model really is the budget model (if you’re budgeting cost, not time).

Test 3: the vendor emails

Prompt:

“A three-email vendor renewal thread (full text in my test kit), with the instruction to extract vendor, renewal date, seat counts, costs, discount, deadline, contract number, and proposed call time into JSON."


The email thread hid its trap in the math. An 8% discount that applies to $73,600, the revised quote for 92 seats, not the original $68,000 quote for 85. Both models avoided the trap and returned accurate answers. Each extracted all 10 fields correctly and returned the right final price of $67,712.

And here is the first test where the budget model won on speed and cost. 7.7 seconds against 14.8, on 435 output tokens against the flagship’s 327. It suggests Flash’s slowness is not constant. On easy work, it behaves like a budget model should. Hand it something hard and its reasoning stage balloons.

Results

MetricGLM-5.3-FlashGLM-5.3
Accuracy27/2727/27
Total tokens4105017867
Total cost$0.0198$0.0757
Avg. response time165.7s69.5s

I put both models through a coding task, a logic puzzle, and a data extraction job, each worth 27 graded points, to measure how much accuracy the budget model gives up. It gave up none; both scored 27 out of 27.

Flash’s speed isn’t a guarantee. It depends on the difficulty of the work. On the coding task, the hardest of the three, Flash spent 7.6 minutes and produced 38,677 output tokens to reach an answer the flagship reached in under 3 minutes with 14,801. On the mid-level puzzle, the two came in seconds apart. On the easy extraction job, Flash was the faster model. Z.AI’s efficiency claim should come with a caveat. Use Flash for your easiest tasks, because while it can handle difficult work, it is far less efficient than the flagship.

What do I think?

I came into this expecting to measure how much accuracy the cheap model loses. It lost none. Across a fussy coding spec, a logic puzzle with one valid answer, and a detail-heavy extraction, the two models were equally correct.

Where they differ is in which tasks they’re best suited. The choice depends on what the model is doing. Flash was faster on extraction but slower on the scheduling puzzle, although both runs finished within seconds.

Hard jobs come down to whether you would rather spend time or money.

Hard jobs come down to whether you would rather spend time or money. Flash got the same answers at about a quarter of the total cost but took more than twice as long to complete the coding task. And hold the cost math loosely. Flash burns far more tokens to do hard work, so its advantage rests on the current per-token prices, and providers can change those at any time.

More on Z.ai from The New Stack:

The post GLM-5.3-Flash vs. GLM-5.3: Time and money, not the spec sheet appeared first on The New Stack.

  •  

DeepSeek’s first vision model vs. Gemini 3.7 Flash: It comes down to spend vs. speed

A glowing sun rises over a neon-blue grid landscape framed by angular mountains.

DeepSeek released V4 Flash Vision Exp on August 21, its first model that accepts image input. Image input means a model can understand a chart, screenshot, or photo document in the same way it can with text. 

DeepSeek V4 Flash Vision Exp reached API gateways like OpenRouter on August 27. It adds image understanding to the company’s budget V4 Flash model. It keeps the same low price of $0.22 per million input tokens and $0.66 per million output tokens. The price doubles during weekday peak hours. 

Google’s Gemini 3.7 Flash, released August 13, is the obvious comparison. It is the budget vision workhorse most developers default to, billed at $0.75 and $3.75 per million on OpenRouter.

DeepSeek pitches the model for document and chart understanding, as well as visual question answering. Google calls Gemini 3.7 Flash its “most intelligent workhorse model yet“. With both making strong claims, I wanted to know which one is better to use for image input.

The tests

I ran both models through three image tests that imitate real back-office work:

  • Chart reading – a stacked bar chart with a cost line plotted on a second y-axis using a different scale, so the answer cannot be read by eyeballing where the line crosses the bars.
  • Invoice audit – a vendor invoice with three planted errors: a line total that doesn’t match quantity times price, a subtotal that matches nothing, and a due date before the invoice date.
  • Incident diagnosis – forty lines of production logs where a payment service crash sits at the bottom, but the real cause, a batch job exhausting the database connection pool, appears five minutes earlier.

I sent the same images and prompts to both models via OpenRouter and recorded the accuracy, tokens, cost, and speed for each call. I included images in the test sections and prompts at the end of the post for anyone who wants to replicate the test. 

Test 1: the two-axis chart

Prompt:
“Look at this chart carefully and answer all three questions. Number your answers. 1. In which quarter did operating costs exceed total revenue? 2. Which revenue segment grew every single quarter? 3. Estimate the company’s total full-year revenue in millions of dollars.

The trap I set was the dual axis. Revenue runs on a 0 to 12 scale on the left, costs on a 0 to 16 scale on the right, so the cost line never visually rises above the bars, even in the quarter where costs won.

Neither model fell for it. Both correctly said Q1, both named Subscriptions as the segment that grew all four quarters, and both landed on $36.1 million for the year, which matches my source data exactly. DeepSeek’s answer was short, three lines. Gemini showed its reading of each bar. Same score either way.

Test 2: the broken invoice 

Prompt:

“You are auditing this invoice. Answer all three questions. Number your answers. 1. Check every line item: does the amount equal quantity times unit price? Name any line that is wrong and give the correct amount. 2. What should the correct total due be? Show your math. 3. Is anything else wrong with this invoice besides the arithmetic?”

Both models caught all three planted errors. Each flagged the monitor arm line, where 10 units at $45.99 were printed as $505.89 instead of $459.90. Each rebuilt the math and arrived at the correct total due of $3,958.89. Each noted that the July 28 due date came before the August 12 invoice date. DeepSeek went one step further and noted that the printed subtotal didn’t match the printed line items, even before the error was corrected.

This test produced the only anomaly. DeepSeek took 30.5 seconds and billed 3,467 completion tokens for an answer only a few paragraphs long. This suggests a large amount of internal reasoning billed as output. Gemini answered in 7.9 seconds with 944 completion tokens. The token differences run the other way on input. DeepSeek counted each image at roughly 500 prompt tokens, Gemini at roughly 1,150. This makes it clear that the two companies tokenize images in very different ways.

Test 3: the buried root cause

Prompt:

“These are production logs from an outage. Answer all three questions. Number your answers. 1. What is the root cause of this outage? 2. At what time did the problem actually begin? 3. What single action would you take first to restore service?

The logs show a payment service crashing with a timeout error. A weak understanding would blame that service. The real cause appears at 14:05:12, when a manually triggered analytics job starts a full table scan on a 48-million-row table, consuming all 20 database connections.

Both models ignored the decoy completely. They named the batch job as the root cause, both pinpointed 14:05:12 as the start time, and both said to kill the batch job first. Gemini even suggested the specific PostgreSQL commands to do it.

DeepSeek answered in 11.9 seconds using 1,636 tokens for $0.00088, while Gemini answered in 7.6 seconds using 1,854 tokens at $0.00351.

Results

Both models provided accurate answers to all questions. Every planted trap failed to catch either one. What separated them was speed and cost/ token usage. Gemini answered in 7.2 seconds on average, compared to DeepSeek, which took more than double that at 16.8 seconds. DeepSeek’s total bill was $0.0039, compared with Gemini’s $0.0122, about a third of the price. 

DeepSeek V4 Flash Vision ExpGemini 3.7 Flash
Accuracy9/99/9
Total tokens69046014
Total cost$0.0039$0.0122
Avg. response time16.8s7.2s

This is a first for us. Total match on accuracy, with the deciding points coming down to cost vs speed.

What do I think?

The practical differences are cost and speed. DeepSeek did the same work for about a third of the money. Gemini did it in less than half the time, and its speed remained consistent, while DeepSeek swung from 8 seconds to 30 seconds depending on the task. DeepSeek’s pricing also doubles during weekday peak hours, and I tested on a weekend, so my cost gap is the best case for DeepSeek.

On a small scale for an individual user like myself, the price and speed are negligible. It really doesn’t matter if it’s 7 or 17 seconds. And it’s going to take a long time for either of these prices to really make a dent… unless you’re running at scale. If you are running at scale and batch-processing invoices overnight, DeepSeek does the job. If you’re running a massive task where a user is waiting on the answer, use Gemini. 

Prompts:

Chart: “Look at this chart carefully and answer all three questions. Number your answers. 1. In which quarter did operating costs exceed total revenue? 2. Which revenue segment grew every single quarter? 3. Estimate the company’s total full-year revenue in millions of dollars.”

Invoice: “You are auditing this invoice. Answer all three questions. Number your answers. 1. Check every line item: does the amount equal quantity times unit price? Name any line that is wrong and give the correct amount. 2. What should the correct total due be? Show your math. 3. Is anything else wrong with this invoice besides the arithmetic?”

The post DeepSeek’s first vision model vs. Gemini 3.7 Flash: It comes down to spend vs. speed appeared first on The New Stack.

  •  

Anthropic’s new Files API vs. pasting: It will save you time, but it won’t save you money.

Horizontal bands of distorted green digital numbers flicker across a black background.

Anthropic moved the Files API and computer-use toolset out of beta on August 19 and launched a browser-use toolset. It lets developers upload a document once and reference it by ID in subsequent requests, rather than sending its contents each time. The common alternative is what most developers do today: pasting the reference material into every prompt.

I wanted to measure how uploading a file once compares to pasting the document into each request, in terms of tokens, accuracy, and setup. I built a test with a verifiable answer key and ran the same workload through both, plus a third approach, prompt caching, that came out of the results.

The test

I wrote a fake API reference for an invoicing company called Ledgerline, about 1,200 words covering authentication, rate limits, idempotency, webhooks, bulk endpoints, and a sandbox. The scenario is a developer support bot that answers questions using only the Ledgerline API reference. Then I wrote five developer questions that the document can answer:

  • How to tell apart two different 403 errors, and the fix for each
  • How to safely retry payment creation after a timeout
  • How to verify webhook signatures and block replay attacks
  • The right way to create 2,000 invoices in one night
  • What to do after committing a secret API token to a public repo

Each question has a specific correct answer in the reference, including details that are easy to miss. One fix depends on knowing that token scopes can’t be edited after creation. Another requires a separate rate limit for the bulk endpoint.

I ran the same five questions three ways, in Python scripts against the API directly, on claude-sonnet-5 with identical instructions:

  • Arm 1 pasted the full reference into every request
  • Arm 2 uploaded the reference once through the Files API and referenced its file ID in every request
  • Arm 3 pasted the reference once into the system prompt with prompt caching enabled

The API reports token usage on every response, so each arm produced its own count.

Getting it running

Two things broke before the first run. The scripts originally set temperature to zero for reproducibility, and the API rejected it. Temperature is deprecated on claude-sonnet-5. Second, one response led with a thinking block instead of text, which caused my printing code to crash. Both fixes were one-liners.

The Files API upload itself was uneventful. One call, one file ID back, and the ID worked in every following request.

The results

All fifteen answers were correct against the answer key, across all three arms. Every arm caught the subtle details, the uneditable token scopes, the separate bulk rate limit, the constant-time signature comparison, the five-minute replay window. Whatever else changed between arms, answer quality did not.

Whatever else changed between arms, answer quality did not.

While the answer quality remained consistent, the token counts didn’t.

TestsRegular input tokensCache writeCache readOutput
Arm 1, paste every time15,246001,399
Arm 2, Files API15,371001,434
Arm 3, prompt caching2712,99011,9601,642

Arm 2 billed slightly more input than Arm 1, which was interesting. The marketing doesn’t say this, but I thought using the files from one API might require slightly fewer tokens. The uploaded file’s contents are still processed into every request, at roughly 3,050 input tokens per question either way, and referencing the file added a small amount of overhead on top, about 25 tokens per request. Across five requests, upload-once cost 125 more input tokens than pasting. There is no volume at which that flips.

Arm 3 is the one that behaved as I assumed the Files API would. The document was billed in full once, as a 2,990-token cache write on the first request. The four requests after that read it from cache, and cache reads bill at about one-tenth the rate of normal input tokens. The questions themselves cost between 48 and 63 regular input tokens each. Cache writes carry a 25 percent premium over normal input, so the first request is the most expensive, and subsequent requests are where the reduction occurs.

Two caveats on the caching numbers. The cache expires after five minutes of inactivity, so the reduction assumes requests keep coming in steadily. And caching required restructuring the request, moving the document into the system prompt with a cache marker.

When to use each approach

The interesting result is that the two features solve different problems. The Files API manages documents. Prompt caching lowers costs.

Files API

Use the Files API when the problem is the file itself. It gives you one uploaded copy referenced by ID instead of the document text living in your code; it handles formats you can’t paste, like PDFs and images; and stored files now support expiration settings. What it doesn’t change is cost. The document is processed for every request, and Anthropic’s announcement never claimed otherwise.

Prompt caching

Use prompt caching when the problem is paying for the same document on every request. In this workload, it cut billed input to roughly a third of pasting across five requests, and the gap keeps widening because every request after the first reads the document at about a tenth of the normal rate. The tradeoffs are the five-minute cache expiry, which assumes steady traffic, and restructuring your request to add the cache marker.

Use the files API and prompt caching together when both problems apply. Files API documents can be cache-marked the same way as pasted text. I tested them separately to isolate what each does on its own.

Paste the documents

Pasting still wins in some cases. Use pasting when prototyping and making one-off calls, where uploading first is just an extra step. This is also the best option for documents that change with every request, where nothing is reused so neither feature helps. Keep in mind that this is short reference text, since documents under 1,024 tokens can’t be cached on most models. Pasting also has the fewest moving parts: no upload step, no file IDs, no stored copies to manage.

I did this test expecting to find out whether the Files API beats pasting in terms of accuracy or cost. It turns out I asked the wrong question. They tie on cost, and the feature that wins on cost was something I tested at the last minute to try to force a different result.

The post Anthropic’s new Files API vs. pasting: It will save you time, but it won’t save you money. appeared first on The New Stack.

  •