❌

Vue normale

Reçu avant avant-hierInfra

Q.ANT gives away the software for its light-powered AI chips in a CUDA-style bet on developers

24 septembre 2026 à 00:23

Q.ANT, a startup out of Stuttgart, Germany, builds processors that use light instead of electricity to do some of the math behind AI. The company pitches them as a way to run AI on a fraction of the power today’s chips need.

Now developers can start writing software for those chips without owning one. Q.ANT pushed a free, open-source software kit to GitHub this week that lets developers build and test programs on a normal computer, then run them on the real chips once they get access.

This is a move out of Nvidia’s playbook. Nvidia owes its lead in AI as much to CUDA, the software developers use to program its GPUs, as it does to the chips themselves. 

But with Q.ANT, the catch is the hardware. Q.ANT’s chips are running at a few research computing centers, and everyone else has to wait “the coming months” for cloud access through German provider IONOS or an on-site server from Q.ANT.

The kit, called the Q.ANT Native Computing Toolkit, is free on GitHub under a license that allows commercial use. Developers can work in Python or C. The key piece is a simulator that mimics the chip on a regular computer, with no Q.ANT drivers required.

What can it do today? The AI tools in this first version focus on running models that have already been trained. The examples read handwritten numbers, identify objects in photos and outline shapes in images. Training still happens on regular CPUs and GPUs.

The pitch for photonic computing is power. AI chips burn a lot of energy moving data back and forth between memory and the processor. Q.ANT’s chips do part of the math with light, specifically wave-shaped functions similar to a cosine, which regular chips calculate digitally. Q.ANT says AI models built around those functions get better results with fewer parameters, the settings a model learns during training. Fewer parameters means a smaller model, less data to move and less power. The kit includes examples comparing a standard model with one built Q.ANT’s way. Those comparisons are the company’s own.

“An ecosystem isn’t created by hardware alone. It emerges when the software layer is open and others can build on it,” said Michael Förtsch, Q.ANT’s founder and CEO. He calls the release the “Linux moment” of photonic computing.

Q.ANT is betting light can do the math itself. Lightmatter, one of the best-known companies in the field, now puts its focus on Passage, which uses light to move data between chips. The idea of light-based AI isn’t new, either. TNS covered MIT’s photonic processor for building optical neural networks back in 2017.

Q.ANT raised €62 million in July 2025 in a round led by Cherry Ventures, UVC Partners and imec.xpand. In March, it said its second-generation chips were running at the Leibniz Supercomputing Centre near Munich. The results it published from there compare the new chip with its old one: more than 50 times faster at the kind of math that does most of the work in AI models, and six times less energy on typical jobs, by the company’s numbers. Its bigger claims, like up to 30 times better energy efficiency, don’t say what they’re measured against.

Good software alone won’t carry a new chip. Nvidia has been building CUDA for nearly 20 years and is still adding to it, including deeper native Python support last year. Graphcore, the British AI chip startup, had its own software kit and still ended up being sold to SoftBank in 2024.

Q.ANT calls this the first openly available software kit for programming a photonic processor. That depends on how you count. Xanadu has offered free, open software for its light-based quantum computers since 2018. For now, developers can play with the simulator. What they can’t do yet is test Q.ANT’s power-saving claims on their own models. That has to wait until the chips open up.

The post Q.ANT gives away the software for its light-powered AI chips in a CUDA-style bet on developers appeared first on The New Stack.

This week’s news from Zed, Anthropic, and OpenRouter shows why better harnesses matter more than better models

19 septembre 2026 à 12:00
Loose computer cables with black connectors spread across a white background.

Our five most-read stories this week covered code collaboration, a model router, a UI change, a benchmark, and a caching tutorial. Five different stories about the same problem: Turning a model into something users can actually use.

That’s the harness: The software around the model that supplies context, connects tools, routes work, and checks results.

Zed was the big news this week and is rebuilding how people review agent-written code. OpenRouter is giving companies more control over where requests are processed. Anthropic is rebuilding its stack to be less confusing. The benchmark shows where coding agents still struggle, and the tutorial explains when you can skip a model call entirely.

Late in the week, Vercel’s AI Gateway reported that the average price per token fell 23.2% in August, the third straight monthly decline. Inference is getting cheaper. Companies are buying the harness around inference now, and Zed, OpenRouter, and Anthropic all spent this week selling it.

Zed and Anthropic are removing blockers

On Wednesday, two companies took aim at a familiar problem: getting people to organize work before they can do it.

Zed launched Delta in public beta, building code collaboration around shared threads instead of pull requests. Anthropic began rolling Claude Chat and Cowork into one interface, removing the upfront decisions about which mode a task belongs in. 

Paul Sawers reported the Zed story for The New Stack. It was our runaway piece of the week, drawing more than double the traffic of anything else. Zed is clearly onto something.

Zed CEO Nathan Sobo said it best: “It seems like everyone is in a race to replace GitHub right now.”

His argument is that a diff shows you where an agent landed, but leaves much of the conversation that got it there somewhere else. Delta keeps that conversation attached to the code, while DeltaDB records changes at edit-level granularity. 

Zed says 33 of its team members landed 570 changes to Delta’s main branch without opening a single pull request. That’s Delta’s own repository; the public Zed editor repository still accepts conventional pull requests. 

Meanwhile, GitHub reported 2.9 billion monthly commits in August, up from 1.4 billion in April. This was despite GitHub crashing for nearly eight hours in August. Sobo thinks the thread will become software development’s fundamental unit. And Zed is not on this quest to replace GitHub alone. Cursor’s Origin and GitLab’s Project Switch are other takes on rebuilding code collaboration for agents.

Anthropic is tackling a different user handoff. This week, it announced a unified interface bringing Cowork’s capabilities into Claude Chat. The goal is to reduce confusion and merge capabilities. The rollout starts with Pro and Max users, with other plans following.

“People used both, and told us the frustrating part was deciding where a task belonged,” the company said, as Amanda Caswell reported.

I’ve been using it since the switch, and removing the choice was the right call. Now, users should more easily grasp the capabilities and workflow possible with Cowork.

The economics make the harness matter even more

OpenRouter’s US in-region routing is now generally available to business and enterprise customers. Requests sent through its US endpoint are decrypted, processed, and served inside the country, or rejected if that guarantee can’t be met.

Sawers wrote that story too. One number explains the demand: Open-weight models accounted for roughly 60% of OpenRouter’s US-originating token consumption in August, with Chinese models making up most of the volume.

DeepSeek V4 Pro, Kimi K3, GLM 5.2. These are models developed in China and available through providers running them in US data centers

The models’ country of origin and the location where your data gets processed are different questions. OpenRouter is selling control over the second one. That matters, since Deloitte’s global survey cited in the story found that 77% of companies factor an AI solution’s country of origin into vendor selection. Stripe’s announced acquisition of OpenRouter, reportedly worth about $8 billion, adds another measure of the interest in this layer.

Vercel’s September AI Gateway Production Index, published September 17, shows the economics from another angle:

Open weights took the volume. Closed models kept the money.

Share of tokens against share of spending on Vercel’s AI Gateway, August 2026.

Models Share of tokens Share of estimated spend
Open-weight models 56% 14%
Closed-weight models 44% 86%
Anthropic, all models n/a 64%
GPT-6 Astra, first 12 days after Sept. 3 launch n/a 7.7%

Source: Vercel AI Gateway Production Index, Sept. 17, 2026. Gateway traffic only; spending is estimated from labs’ published list prices, so actual bills may differ.

Open weight crossed into the majority of Vercel’s gateway token volume for the first time, up from a reported 7% in December 2025. Vercel says its current open-weight classification is broader than the one used in earlier reports. Among teams running more than ten million tokens in both comparison months, the median team paid 7.6% less per token, following July’s 2.9% decline.

So what justifies the premium? 

Boris Renski, CEO of AI agent integration company Apelogic, argues that much of it buys enterprise plumbing: identity integration, connectors, and observability. In Adrian Bridgwater’s July reporting, CNCF executive director Jonathan Bryce described paying ten times more for a four-month capability lead as “a very expensive form of lock-in.”

Closed models still command most of the estimated spending, and customers may be paying for capabilities that cheaper alternatives don’t reliably deliver.

But cheaper inference raises the pressure on everything around it, and Anthropic spent its week on exactly that. Instead of making headlines with a new model, it shipped a merged interface, documents and slides — features arguably more important to most users.

The harness has to earn its keep

Amanda Caswell covered Real-SWE, a benchmark from Y Combinator-backed Specific Labs built on private codebases from real companies. 

The best setup tested, Claude Fable 5.1 running through Claude Code, succeeded 38.8% of the time. GPT-6 Astra through Codex CLI reached 33.8%; Gemini 3.8 Flash through Gemini CLI reached 31.2%. None broke 40%. 

The benchmark is small: 10 tasks, with eight attempts per model on each. And it tests models with their coding tools, so the harness is already in the score. Still, no model solved the analytics stream reducer across 64 attempts. Zero.

Solutions touched a median of 11 files, compared with six in the public benchmarks cited in the reporting. Real work is spread out. These agents struggled to follow it. 

Fable’s leading failure categories were missed requirements and integration errors. That doesn’t mean the information was missing. An agent can have it and still overlook it, misunderstand it, or skip the check.

The engineering problem is getting context, tools, and verification to work together. So is knowing when the model doesn’t need to run at all. 

Abhilash Rao Mesala, a senior data engineer at Meta, wrote a practical guide to LLM response caching for us: reuse an answer when the request, context, permissions, and underlying information make it valid. His example starts with a million calls a month at $0.006 each, or $6,000. A 60% cache hit rate, plus $150 in embedding and vector-store costs, brings that to $2,550. A 57.5% reduction from a principle older than the transformer.

The hard part is knowing when a stored answer is still the right answer.

Pair Mesala’s piece with Ida Silfverskiöld’s on Towards Data Science guide to saving on tokens, which covers prompt caching, model routing, and keeping unnecessary materials out of an agent’s context. Both are worth your time. 

It’s hard to ignore the doom and gloom hovering over AI at the moment. But there’s good news out there, too. Each week, the price of inference is dropping, and companies are improving harnesses – both are critical improvements to push AI into the mainstream.

The post This week’s news from Zed, Anthropic, and OpenRouter shows why better harnesses matter more than better models appeared first on The New Stack.

Jacob Coxon warns AI could kill us all. Anthropic’s own report exposes safety gaps. 

12 septembre 2026 à 13:00
Abstract teal composition with overlapping translucent geometric shapes and scattered flecks of red, yellow, white, and black paint.

I’m Matt Burns, Chief Content Officer at Insight Media Group. Each week, I round up the most important AI developments, explaining what they mean for people and organizations putting this technology to work. The thesis is simple: workers who learn to use AI will define the next era of their industries, and this newsletter is here to help you be one of them.


Former OpenAI and Anthropic researcher Jacob Coxon made a lot of noise this week with a viral thread on X. He has since talked to Wired and Axios. His warning deserves attention, but another revelation this week gives developers something concrete to investigate.

In retrospective testing of the Mythos 5 incident, Anthropic’s offline monitor flagged around 1% of actions. When researchers removed the model’s chain of thought and showed the monitor only user turns, tool calls, and tool results, that rose to about 50%. 

The model’s written reasoning is what talked the monitor out of flagging that harmful behavior. The offline monitor kept concluding it was operating in a simulation.

The lesson is to test whether your monitoring accepts an agent’s justification a little too easily. An AI’s explanation can help a safety monitor reveal a problem, or persuade it to essentially convey, “move along, nothing to see here.”

An AI’s explanation can help a safety monitor reveal a problem, or persuade it to essentially convey, “move along, nothing to see here.”

Coxon is warning about the pursuit of self-improving superintelligence. Developers are already pointing agents at real credentials during everyday work. Those problems differ substantially, but both deserve attention.

Forecasts aren’t checkable. Incident reports are.

Jacob Coxon’s thread is short and hard. He says neither of his former employers is acting responsibly, and that both are “racing straight to self-improving superintelligence and gambling with our lives.”

Maxwell Zeff pressed him on the end-of-humans mechanism twice in an interview with Wired. Coxon described shutdown avoidance, then biological and cyber threats. His answers explain the concern, but leave substantial questions about how AI will kill us all.

Annie Jacobsen, whose books on this stuff I wholeheartedly recommend, spelled out the scary biological scenario: A system that can hack anything, turned loose on the world’s 3,600-plus BSL-3 and BSL-4 labs, is how you get from a software problem to a biological one. For me, Jacobsen’s commentary is the most concrete, real-world implication anyone has given of what Coxon is actually worried about. 

Two incident reports published in the last few weeks are more useful anyway. OpenAI’s report on the Hugging Face incident, published August 26, traces the whole thing from May 12 to July 20. The agents found a hole in an internal package server and used it to reach the open internet that the same server was supposed to keep them off of. OpenAI rebuilt the server; they got back out. By July 11, they were far enough into Hugging Face to run their own commands on its machines.

The model doing most of that was never released. It’s a research version OpenAI runs in-house, with fewer safety checks than the ChatGPT everybody else uses. When the company ran the same work again with those checks switched on, it found that this kind of behavior could drop by more than 100 times.

That context matters. 

Amanda Caswell covered this in August. The incident demonstrates a serious failure under particular conditions. It also illustrates why the surrounding safeguards belong in any assessment of an agent’s behavior.

The monitor was talked out of it

Anthropic’s assessment, published on September 9, covers four incidents and analyzes the first three in a scan of roughly 141,000 transcripts, identified a fourth in August while assembling material for METR, then widened the search to roughly 481 million transcripts and reported no additional incidents of similar or greater severity. In each case, the model reached the open internet through a misconfiguration rather than by breaking out of a properly isolated sandbox. That’s an important detail. It makes the evaluation infrastructure part of the story. 

If you’re a developer, you should look at the monitoring results. Mythos 5 kept reasoning its way to the conclusion that it was operating in a simulation, and the offline monitor read and accepted that reasoning. The limits matter here too: These were cybersecurity evaluations running without released cyber safeguards, Anthropic reports no coordination between models and no concealment, and its independent METR review was planned rather than finished. This is not the same failure as the OpenAI incident. Don’t equate them as similar incidents

But yet the shape of it is familiar, and it’s what Adrian Bridgwater reported on last week in his piece on agent observability, which is a useful read if you’re producing agents (who isn’t). Jacob Krell of Suzu Labs put it about as well as it can be put: “The model can recite your constraints and reason past them in the same breath.” Liran Hason of Coralogix added the result: “System guardrails help, but a guardrail only stops what the developer already thought of.”

What developers can take from Steven Adler’s recommendations

Largely overshadowed this week by Coxon, Steven Adler’s op-ed in The New York Times asks AI companies for four things: disclose incidents including near misses, the way aviation does; keep tamper-evident records of model behavior; never let a model cut power to its own alarm systems; and formally swear off training techniques that undermine researchers’ ability to detect deception. 

Adler worked for OpenAI from 2020 to 2024 and later founded a nonprofit called Guidelight AI Standards. It has since graded six control practices at the frontier labs using public information and gave out a top mark of C-plus.

Guidelight also publishes a Control standard for frontier labs, recently revised in August. It wants monitors to access the chain of thought and scan both reasoning and actions. 

But those asks of frontier labs don’t apply to most developers. Here are five you can check on your own setup right now.

Five checks you can run on your own agent setup

Steven Adler’s asks of the frontier labs, applied one layer down.

Check What passes
Someone reads the near misses An agent crosses a line and nothing breaks. There’s still a record, a person who looks at it, and a copy the agent can’t reach.
The agent can’t edit its own history Logs are append-only and stored outside the agent’s reach. If its credentials can delete the log, the log is decorative.
The agent can’t touch its own oversight Separate tokens help. Then check whether it can shut off monitoring, rewrite an approval rule or get around the kill switch.
Rule-breaking gets blocked and logged In a test environment, hand it a job it can only finish by breaking a rule. Judge the monitor on what it does with the action, not with the agent’s account of it.
The explanation doesn’t change the answer Run the same rule-breaking action past your monitor twice. Change nothing but the explanation, one flat and one reassuring. If the reassuring version gets flagged less, you’ve found the problem.

Sources: Steven Adler, The New York Times; Anthropic alignment assessment, Sept. 9, 2026.

These checks do not establish that frontier alignment is solved. They give a team evidence about the system it operates on and a clearer account of where its protections need work.

Anthropic’s report is why I’d include that last test in a deployment review. Hold the action and its permissions constant, change the explanation, and see whether your monitor changes its judgment.

That won’t settle Coxon’s warning. Will AI kill us all? I don’t know. But I know developers need to test their agents’ permissions, protect their logs, and make sure a convincing explanation can’t talk a safety monitor into overlooking an unauthorized action.

The post Jacob Coxon warns AI could kill us all. Anthropic’s own report exposes safety gaps.  appeared first on The New Stack.

OpenAI will sell you Astra, but not the system that scored 98.6% on ARC-AGI-3

4 septembre 2026 à 17:25

Investor Matt Turck, whose fantastic podcast has hosted the people who built ARC-AGI, summed up Astra’s blockbuster benchmarks with three words: “This is wild.” Then he added four more in parentheses: “w/ its native harness.”

ARC Prize ran GPT-6 Astra through its own standard harness, and the model scored 62.7%. When run through OpenAI’s Provider Adapter, the same model in the same reasoning setting scored 98.6%. The model didn’t change, but the software around it did, adding 36 points to the score. And the better-performing system cost less: $17,332 with OpenAI’s adapter versus $26,098 with ARC Prize’s.

That’s why harness engineering is becoming as important as model selection. The software around the model can change both what it accomplishes and what it costs.

The benchmark measures the system, not the model

ARC-AGI was built to resist the brute-force scaling that eats so many other benchmarks. ARC-AGI-3 raised the bar again this year by dropping a model into interactive environments with no instructions, no stated goal, and no stated rules, then scoring how efficiently it learns to operate. When ARC Prize launched it this year, humans scored 100%. Frontier AI scored 0.51%.

On Thursday, Amanda Caswell covered OpenAI’s improving score on the ARC-AGI-3 benchmark. As she noted, Astra ran under different settings from competing models. ARC Prize is specific about what those settings do. Its standard harness lets a model carry forward notes it chooses to keep. OpenAI’s adapter preserves the opaque reasoning state between requests and compresses longer conversations, so the model can resume its own thinking instead of reconstructing it. 

ARC Prize published every reasoning level.

Same model, two harnesses

GPT-6 Astra on ARC-AGI-3, by reasoning effort. At every setting, the run inside OpenAI’s Provider Adapter scored higher and cost less than the same model inside ARC Prize’s standard harness.

Reasoning effort ARC Prize standard harness OpenAI Provider Adapter
Max 62.7% for $26,098 98.6% for $17,332
XHigh 59.3% for $37,317 98.4% for $18,147
High 54.8% for $40,705 99.9% for $18,817
Medium 38.6% for $48,090 98.4% for $19,285
Low 17.5% for $38,166 98.0% for $21,298
None 35.2% for $49,791 96.7% for $23,457

Source: ARC Prize.

Astra inside OpenAI’s harness with no reasoning effort at all scored 96.7% for $23,457. The same model at max reasoning within ARC Prize’s standard harness scored 62.7% and cost $26,098. The harness beat the reasoning dial outright. I’ve been arguing the harness matters for months. I didn’t expect the result to be this lopsided.

The score wasn’t the only gap. Across the 167 game-reasoning pairs both harnesses solved, ARC Prize clocked the Provider Adapter runs at 49% fewer tokens and roughly 3.66x faster.

OpenAI isn’t hiding the details: The adapter runs on documented Responses API capabilities anyone can call. What you can’t buy is the assembled system that scored 98.6%.

That’s also why OpenAI President Greg Brockman’s claim during a press briefing this week — “I think it’s not unreasonable to feel that we are now in the AGI era” — lands harder than it should. Brockman is describing a benchmark result produced by a particular system, not establishing that the underlying model is AGI. Frederic Lardinois’ launch coverage for The New Stack gets the distinction right: The framing goes well beyond what the evidence establishes. Maybe we’re in the AGI era. But this benchmark doesn’t prove it.

The harness is becoming the product

On coding, the frontier models now cluster inside a few points. Artificial Analysis scores its Coding Agent Index by running each model inside a harness rather than on its own: Astra in Codex at 67, Opus 5 and Fable 5 in Claude Code at roughly the same, Muse Spark 1.3 in Muse Code alongside them, and Fable 5.1 in Claude Code leading at 70. The unit being measured is already the pair. 

The labs already figured this out. Back in April, Janakiram MSV documented the four-way split: Anthropic, OpenAI, Google, and Microsoft all treat the harness as a product to sell, and disagree only on how to charge for it. Anthropic meters Managed Agents at $0.08 per session hour, in addition to token rates. OpenAI gave its Agents SDK away with no runtime fee at all. Google and Microsoft bill sessions, memory, code execution, and observability as separate line items.

Nobody is treating the harness as a free accessory to the model.

Other companies are moving the same way. Stripe paid a reported $8 billion for OpenRouter in August, as Paul Sawers reported for TNS, acquiring a gateway that routes 10 trillion tokens a day across more than 400 models for 10 million developers. Patrick Collison’s framing was that tokens are the central currency for companies building with AI. Stripe bought the routing layer that sits in front of the models. Nvidia built a harness of its own.

We covered the proof two weeks ago. Adrian Bridgwater reported that Claude Opus 5 scores 30.2% on ARC-AGI-3’s public set on its own. Wrapped in Nvidia’s AVO, which gives it persistent memory and programmatic supervision that steps in when progress stalls, it cleared all 183 levels across 25 environments. That one isn’t the clean A/B that ARC Prize ran on Astra. Nvidia changed the memory, supervision, and context management at once. More than one variable moved.

Nvidia said it beautifully here: “Model capability matters enormously, but the surrounding system determines how effectively that capability can be converted into sustained autonomous progress.”

Harness engineering is the job

Janakiram MSV found token usage varying 70-fold across Aider, Claude Code, and OpenClaw running an identical model. Cache hit rates swung from about 70% down to 1.5% depending on the serving path. No model choice explains a spread like that.

The work itself is pretty ordinary. Deciding what an agent remembers and what it forgets. What it’s allowed to touch, and when it has to stop and check with a person. Jeremy Daly’s piece on our site is the version with the engineering, and it’s the one to read if you’re the one building.

I argued in June that model triage was the skill worth hiring for. I’d revise that. Picking the model is the easy half, and it gets easier every quarter as the frontier converges. A year from now, I think the people running agents will spend less of the week picking models and more of it building what goes around them. 

The post OpenAI will sell you Astra, but not the system that scored 98.6% on ARC-AGI-3 appeared first on The New Stack.

Nvidia is paying $12.9 billion to keep open models on its chips 

28 août 2026 à 17:43

I’m Matt Burns, Chief Content Officer at Insight Media Group. Each week, I round up the most important AI developments, explaining what they mean for people and organizations putting this technology to work. The thesis is simple: workers who learn to use AI will define the next era of their industries, and this newsletter is here to help you be one of them.


Nvidia agreed to pay $12.9 billion for Hugging Face this week, The Information reported. The purchase feels similar to when Microsoft bought GitHub. 

Nvidia followed developers. It bought the place they already go for AI models. Meanwhile, Anthropic and OpenAI are selling a premium into a market that’s increasingly treating free as the default.

It’s getting easier to run those open models, too. Earlier this week, Ollama shipped a release that lets Claude Desktop run Qwen, DeepSeek, and Kimi. You open Anthropic’s app, click the model picker, and pick Kimi K3 instead of Opus 5. 

Open models moved into the tools developers already use

Our own Paul Sawers covered the Ollama release for The New Stack. Version 0.33.0 launched on August 21 and uses a local proxy to get around Claude Desktop’s block on non-Anthropic model IDs. Flip “Use Ollama models” in the Mac menu bar, and whatever you’re running, locally or on Ollama Cloud, shows up in Anthropic’s model picker. “Open models should be easy to run, easy to build with, and available wherever people need them,” Ollama CEO Jeffrey Morgan said. The company raised a $65 million Series B in July.

The hardware is catching up, too. Apple launched new Mac Mini and Mac Studio configurations this week that seem purpose-built to run larger models locally. Frederic Lardinois wrote up Alibaba’s Qwen3.8-27B two weeks ago. The 4-bit build is 16.1GB and runs on a Mac with 32GB of unified memory; Alibaba’s own benchmarks put it at or past Opus 4.6 Max for coding and computer use. Those are Alibaba’s benchmarks, and early users say the model overthinks. It still fits on a laptop.

The cost of switching dropped with it. TNS writer Janakirm MSV started covering this in May, when OpenCode passed Claude Code on GitHub stars, from 157,000 to about 122,000. The real decision in front of most developers, he writes, is “whether their environment can tolerate a single-vendor harness at all.” Jason Calacanis put the commercial version on X on Wednesday: “Open source is solving 90% of startup use cases right now, so frugal founders aren’t paying for Fable.” Five hundred dollars per employee per month feels fine, he writes, but “when you start hitting five figures, CFOs start clenching.” Founders are the early signal. CFOs are the broader one.

Nvidia is paying for the place developers get their models

So what does $12.9 billion buy Nvidia? A website with about $150 million in revenue. Nvidia already builds open models of its own and gives them away, so it isn’t short on weights.

Microsoft paid $7.5 billion for GitHub in 2018 because that’s where code lived, and owning it made GitHub the surface for everything Microsoft wanted developers to do next. Hugging Face is that for weights. Every open model release and every fine-tune lands there. It’s the default place developers go to figure out how to run one.

Product analyst Aakash Gupta did the math to explain the price. Nvidia just reported $96.2 billion in quarterly revenue, so $12.9 billion is about 12 days of sales. And Nvidia’s largest customers are all building escape routes: OpenAI designing chips with Broadcom, Anthropic training on Amazon’s Trainium, Google a decade into its own TPUs. 

Open models are the counterweight, because a downloaded model gets fine-tuned and served on Nvidia CUDA by default. As long as developers keep choosing open models, they’re also choosing Nvidia’s software stack. “Nvidia spent 12 days of revenue to make sure the open-source rival to its own customers never dies,” Gupta writes.

It works because open models run on Nvidia hardware. Download Qwen, fine-tune it, and every step runs on CUDA unless you go out of your way. Broadcom and Amazon can build a chip. Neither of them can make ten thousand repos target it.

So the thing you did to get out from under one vendor’s pricing put you further under another’s. 

Hugging Face is also the leaderboard, the datasets, and the transformers library a good chunk of the industry uses. Nvidia will soon own all of it. The stalwart venture capitalist Bill Gurley posted the reason on Wednesday: “Open-models are the inevitable outcome of high stakes software competition,” he writes. “The more at stake, the more likely open wins. This is water running downhill.” 

He made the longer case in The Washington Post in July, under a headline that names the two companies preparing to go public: Open-model AI is good competition for Anthropic and OpenAI.

The obvious objection is that Hugging Face downloads aren’t production traffic, and that most real work still runs against a closed API. It’s fair. But Chinese open models have taken more than 30% of U.S. token usage on OpenRouter every week since February, peaking at 46%. Right now, that’s the corner of the market where switching is cheapest. 

The lesson for AI-native developers isn’t to replace Claude with Qwen or stop paying Anthropic. Closed models will still be the right answer for plenty of work. The lesson is to stop assuming today’s default will still be tomorrow’s.

Every AI-native application has defaults: the model, the registry, the harness, the API. Those become dependencies, and dependencies become leverage. Nvidia reportedly agreed to spend $12.9 billion on the place developers already go to find and fine-tune models.

That changes the job for AI-native developers. Five years ago, the best developers learned Kubernetes. Today, they spend their time comparing benchmark scores and arguing about which model is smartest. And that’s becoming the less important skill. The hard one is building systems that survive when the underlying defaults change. 

Because it will. 

The post Nvidia is paying $12.9 billion to keep open models on its chips  appeared first on The New Stack.

❌