❌

Vue normale

Reçu avant avant-hierThe New Stack

Q.ANT gives away the software for its light-powered AI chips in a CUDA-style bet on developers

24 septembre 2026 à 00:23

Q.ANT, a startup out of Stuttgart, Germany, builds processors that use light instead of electricity to do some of the math behind AI. The company pitches them as a way to run AI on a fraction of the power today’s chips need.

Now developers can start writing software for those chips without owning one. Q.ANT pushed a free, open-source software kit to GitHub this week that lets developers build and test programs on a normal computer, then run them on the real chips once they get access.

This is a move out of Nvidia’s playbook. Nvidia owes its lead in AI as much to CUDA, the software developers use to program its GPUs, as it does to the chips themselves. 

But with Q.ANT, the catch is the hardware. Q.ANT’s chips are running at a few research computing centers, and everyone else has to wait “the coming months” for cloud access through German provider IONOS or an on-site server from Q.ANT.

The kit, called the Q.ANT Native Computing Toolkit, is free on GitHub under a license that allows commercial use. Developers can work in Python or C. The key piece is a simulator that mimics the chip on a regular computer, with no Q.ANT drivers required.

What can it do today? The AI tools in this first version focus on running models that have already been trained. The examples read handwritten numbers, identify objects in photos and outline shapes in images. Training still happens on regular CPUs and GPUs.

The pitch for photonic computing is power. AI chips burn a lot of energy moving data back and forth between memory and the processor. Q.ANT’s chips do part of the math with light, specifically wave-shaped functions similar to a cosine, which regular chips calculate digitally. Q.ANT says AI models built around those functions get better results with fewer parameters, the settings a model learns during training. Fewer parameters means a smaller model, less data to move and less power. The kit includes examples comparing a standard model with one built Q.ANT’s way. Those comparisons are the company’s own.

“An ecosystem isn’t created by hardware alone. It emerges when the software layer is open and others can build on it,” said Michael Förtsch, Q.ANT’s founder and CEO. He calls the release the “Linux moment” of photonic computing.

Q.ANT is betting light can do the math itself. Lightmatter, one of the best-known companies in the field, now puts its focus on Passage, which uses light to move data between chips. The idea of light-based AI isn’t new, either. TNS covered MIT’s photonic processor for building optical neural networks back in 2017.

Q.ANT raised €62 million in July 2025 in a round led by Cherry Ventures, UVC Partners and imec.xpand. In March, it said its second-generation chips were running at the Leibniz Supercomputing Centre near Munich. The results it published from there compare the new chip with its old one: more than 50 times faster at the kind of math that does most of the work in AI models, and six times less energy on typical jobs, by the company’s numbers. Its bigger claims, like up to 30 times better energy efficiency, don’t say what they’re measured against.

Good software alone won’t carry a new chip. Nvidia has been building CUDA for nearly 20 years and is still adding to it, including deeper native Python support last year. Graphcore, the British AI chip startup, had its own software kit and still ended up being sold to SoftBank in 2024.

Q.ANT calls this the first openly available software kit for programming a photonic processor. That depends on how you count. Xanadu has offered free, open software for its light-based quantum computers since 2018. For now, developers can play with the simulator. What they can’t do yet is test Q.ANT’s power-saving claims on their own models. That has to wait until the chips open up.

The post Q.ANT gives away the software for its light-powered AI chips in a CUDA-style bet on developers appeared first on The New Stack.

Intel squeezed a 1.58-bit LLM down to 1.485 bits without changing a single weight

17 septembre 2026 à 22:51
Digital void

The 1.58 in a 1.58-bit language model sounds like a hard limit, but Intel researchers pushed a ternary model below it by changing how its weights are stored rather than changing the model itself.

Their new BITCOS format compressed one checkpoint to 1.485 bits per weight and improved decoding throughput by as much as 18% on CPUs and 27% on GPUs.

The key is that the familiar 1.58-bit figure assumes a model uses its three possible weight values equally, while real ternary models contain far more zeros than that calculation accounts for. BITCOS stores the location and sign of each nonzero weight separately, allowing zeros to take up less space without retraining the model or altering its output — the equivalent of packing the same contents into a smaller box.

Where the 1.58-bit figure comes from

Ternary models use only three weight values — -1, 0, and +1 — and 1.58 bits is the theoretical minimum needed to represent three equally likely options. That number is cleaner than the reality of storing the weights, where the standard approach fits five ternary values into an eight-bit byte for an average of 1.6 bits each. Models commonly store weights in blocks of 128; however, this leaves the final byte partly unused and pushes the actual rate to 1.625 bits per weight.

When Intel’s researchers measured the distribution of weights across 29 checkpoints from seven ternary model families, they found that zeros accounted for between 29.7% and 51.5% of the weights. In 26 of those checkpoints, there were enough zeros for BITCOS to beat five-trit packing.

The sparsest was a ternary version of Qwen3-1.7B produced with CAT-Q post-training quantization, where 51.48% of the weights were zero, and BITCOS brought the storage cost down to 1.485 bits per weight.

Models commonly store weights in blocks of 128; however, this leaves the final byte partly unused and pushes the actual rate to 1.625 bits per weight.

How zeros save space

BITCOS stands for “BITmap and COmpacted Signs” and divides a model’s weights into two streams. The first assigns one bit to every weight to record whether it is zero or nonzero, while the second assigns a sign bit only to nonzero weights.

A positive or negative weight therefore consumes two bits, but a zero needs only the presence bit because it has no sign to record.

If z is the proportion of zero weights, BITCOS uses 2 − z bits per weight, dropping from 1.6 bits at 40% zeros to 1.485 bits at 51.5%. Because it changes only the storage format, unpacking restores the original -1, 0 and +1 values without affecting accuracy.

BITCOS becomes smaller than five-trit packing once more than 37.5% of a model’s weights are zero, a threshold reached by 26 of the 29 checkpoints Intel examined.

BITCOS becomes smaller than five-trit packing once more than 37.5% of a model’s weights are zero, a threshold reached by 26 of the 29 checkpoints Intel examined.

Making smaller weights run faster

Built for token-by-token decoding with small batch sizes, the format reduces the weight data moving through memory. Intel developed separate unpacking kernels for AVX-512 and AVX2 CPUs as well as Xe2 GPUs, joining other efforts to fit compressed models into faster inference pipelines for AI agents.

On AVX-512 hardware, the kernel uses the presence bitmap as a mask and pdep to scatter the compacted sign bits across the nonzero weight positions. Because Xe2 GPUs lack an equivalent instruction, Intel implemented the same operation with a 2KB lookup table.

Benchmarks across five systems

Compared with the 2-bit kernels, BITCOS ran 10% to 18% faster on the 64-core Xeon server and 2% to 15% faster on the 24-core Core Ultra 9. Performance improved by 9% to 22% on the integrated Arc 140V and by 2% to 27% on the discrete Arc Pro B70. These results measure decoding after the model has loaded, separate from efforts to cut GPU inference cold starts from minutes to seconds.

The smaller format did not win everywhere

On the eight-core Lunar Lake CPU, Intel’s fixed 2-bit kernel beat BITCOS on every model because the system had enough bandwidth to make unpacking the bottleneck. BITCOS remained faster on the GPUs, although decoding overhead limited the gains. Computer scientist and AI infrastructure author Chip Huyen has made the same point about inference more generally, arguing that the right optimization depends on whether compute, memory or bandwidth is holding back the workload.

On the eight-core Lunar Lake CPU, Intel’s fixed 2-bit kernel beat BITCOS on every model because the system had enough bandwidth to make unpacking the bottleneck.

Format limits and open questions

The paper has not been peer-reviewed; all five test systems used Intel hardware, and the end-to-end benchmarks covered seven models at batch size one. Intel has yet to test the format on Nvidia, AMD, or Arm hardware.

The post Intel squeezed a 1.58-bit LLM down to 1.485 bits without changing a single weight appeared first on The New Stack.

Perplexity’s new agent runs entirely on your GPU — with one expensive catch

14 septembre 2026 à 20:21
Abstract server

Running an LLM on your PC is easy enough, but putting an agent to work there is a different story. Portable Computer, the local version of Perplexity’s Computer agent, is now available inside the Perplexity app for Windows on compatible Nvidia GeForce RTX and RTX PRO GPUs.

That’s the good news; the catch is, you’ll need an Nvidia GPU with at least 24GB of VRAM.

The Windows launch gives Perplexity three platforms in less than three weeks. Portable Computer debuted on Linux and Nvidia DGX Spark on August 25, followed a week later by hybrid compute for Apple silicon, which splits tasks between local and cloud models on Macs. Now Windows joins the mix, but bringing Portable Computer over took more than simply porting the app. Perplexity had to adapt the model runtime, orchestration, security, and hardware integration for each platform while keeping the user experience the same.

you’ll need an Nvidia GPU with at least 24GB of VRAM to use it.

Orchestration beyond the model

Portable Computer bundles those pieces together. On Windows, it supports PPLX 27B — Perplexity’s post-trained model — and Qwen 3.8 27B, both optimized for RTX GPUs, alongside a built-in browser, tool calling,  and Perplexity’s proprietary SPACE sandbox.

It’s a different lane from LM Studio or Ollama, which make running models locally as painless as possible but stop well short of giving a model autonomy over multistep work. DeepSeek’s recent hiring spree of roughly 150 new roles, nearly all of them focused on agent infrastructure rather than the model, hints at how much engineering sits between a capable model and a capable agent.

DeepSeek’s recent hiring spree of roughly 150 new roles, nearly all of them focused on agent infrastructure rather than the model, hints at how much engineering sits between a capable model and a capable agent.

Connectors blur local boundaries

Perplexity ships connectors for Microsoft Outlook, OneDrive, and Word, plus Google Drive, Gmail, Slack, and GitHub — which tells you something about what “local” actually means here.

The agent can reach external services because it’s not air-gapped. Locally completed tasks can process files without sending documents to a cloud model. Once an agent has access to both local files and remote APIs on the same machine, figuring out which resources it actually needs and where to find them gets harder.

Hybrid cloud as fallback

Perplexity isn’t pretending that a 27-billion-parameter model running on a desktop GPU can handle everything, which explains the hybrid architecture. When the agent determines that a task needs more reasoning power than the local model can deliver, it can escalate to Perplexity’s cloud models.

According to Nvidia, the agent identifies when cloud support would help and asks the user for permission before sending any data off the machine.

For organizations handling sensitive or regulated data, that split can make all the difference. A local agent can grind through source code or financial records without uploading them to a hosted model for basic processing. There’s a cost angle too, since tasks completed locally don’t burn Perplexity Computer credits.

High VRAM floor limits reach

Portable Computer is available with Perplexity Pro ($20/month) and Max ($200/month), across individual and enterprise plans, with Nvidia DGX Station support coming later. The real challenge is taking local agents from developer passion projects to enterprise-ready tools. By baking this into Windows, it immediately gets in front of the scale of users needed to make that happen.

The real challenge is taking local agents from developer passion projects to enterprise-ready tools.

The post Perplexity’s new agent runs entirely on your GPU — with one expensive catch appeared first on The New Stack.

Your Mac is now part of Perplexity’s AI infrastructure

1 septembre 2026 à 21:21
abstract

Perplexity wants its AI agents to use more of the computing power already sitting inside a Mac. The company launched Hybrid Compute on Tuesday, a new feature that lets Perplexity Computer move parts of the same task between powerful cloud models and smaller models running directly on Apple silicon.

The timing is significant. Hybrid Compute arrives on the first day of John Ternus’ tenure as Apple’s CEO. Ternus, who previously led Apple’s hardware engineering organization and played a key role in the company’s transition from Intel processors to its own Apple silicon, succeeds Tim Cook after 15 years as CEO. Perplexity is now betting that same hardware can become part of the infrastructure behind autonomous AI agents.

The task starts in the cloud. When a step involves sensitive information, Computer can move that part of the job to a model running on the Mac without having to start over.

When a step involves sensitive information, Computer can move that part of the job to a model running on the Mac without starting over.

Privacy Gate checks first

A Perplexity-trained Privacy Gate runs locally on the Mac and looks for sensitive information such as names, addresses, account numbers, and secrets. When it flags something, the user can decide whether that part of the task should stay on the device.

Users can review what the system wants to keep local before work begins and ensure it hasn’t missed anything they don’t want sent to the cloud. They can also choose which model handles the local work.

At launch, users can choose between Gemma E4B, Qwen3.6 35B-A3B, and a version of Qwen3.6 35B that Perplexity post-trained itself, with more models planned for later. Perplexity also handles the installation through the desktop app, so users don’t have to open a terminal or set up the model themselves.

Local tokens cost nothing

Once a task is running, the app displays local CPU, GPU, and memory utilization, as well as the number of tokens consumed. Users aren’t charged for tokens generated by models running locally on their Mac. That gives Perplexity another reason to move work onto the device beyond privacy, since running models locally can also cut inference costs. (The question of who bears inference costs and when is becoming a competitive issue across the industry.)

There is a compromise. Perplexity uses its most capable models in the cloud, while smaller models handle local work, so keeping more of a task on the Mac can improve privacy and cut costs at the expense of some capability. Hybrid Compute leaves that choice to the user.

Perplexity uses its most capable models in the cloud, while smaller models handle local work, so keeping more of a task on the Mac can improve privacy and cut costs at the expense of some capability.

Context crosses the boundary

The trickier bit is what happens to context when part of a task moves onto the Mac. The cloud model still needs to know enough about what happened locally to continue the job, without getting access to the private information that was supposed to stay there.

Perplexity says Computer can move a step from the cloud to a local model without restarting the task or losing context, and ultimately combine the cloud and local work into a single result. It does not detail in its announcement exactly what context is passed between those environments or how information produced by the local subagent is filtered before it returns to the broader workflow.

The local subagent can work with private files and data and take actions on the Mac. Users can also start a task on an iPhone and hand off local work to their Mac without having to start over.

For enterprise customers, Perplexity adds company-wide rules for what stays local and a record of what leaves each device, bringing the same governance questions facing AI agents down to the device level.

DGX Spark starts local

Hybrid Compute reverses the approach Perplexity introduced for Nvidia’s DGX Spark last week. DGX Spark starts locally and reaches out to frontier cloud models only with permission, while the Mac version starts in the cloud and moves work onto the device when needed.

In both cases, the agent harness decides where each part of the job runs. That orchestration — deciding which tools and context an agent actually needs — is becoming an increasingly difficult engineering problem as agents gain access to more systems and data.

Apple silicon becomes part of the agent stack

Running this much of an agent locally still requires a fairly powerful Mac. Perplexity recommends at least 32GB of unified memory, Apple silicon, and macOS 15. The feature is available to Pro and Max subscribers as well as enterprise customers.

Those requirements show the limits of local AI at the moment. Smaller models can run on plenty of Macs, but giving an agent enough compute to handle meaningful work still requires relatively high-end hardware.

That will likely change as Macs get better at running larger models. For Ternus, the bigger question is how much of the AI work now happening in the cloud will eventually move onto the machines Apple sells.

For Ternus, the bigger question is how much of the AI work now happening in the cloud will eventually move onto the machines Apple sells.

The post Your Mac is now part of Perplexity’s AI infrastructure appeared first on The New Stack.

SpaceX designed an orbital Vera Rubin. Radiation comes next.

31 août 2026 à 20:55
NVIDIA Vera CPU

SpaceX and Nvidia say they are adapting the Vera Rubin NVL72 rack-scale AI platform for orbital use, with SpaceX targeting a first launch in the fourth quarter of 2027. 

The dream of an AI data center in space lives on in SpaceX and Nvidia’s August 24 announcements that the platform for Low Earth Orbit (LEO) Starmind AI satellites will be based on the Vera Rubin NVL72 chip family and architecture.

This proposed system would form the computing core of SpaceXAI’s first-generation Starmind AI satellite and extend Nvidia’s architecture from terrestrial AI data centers into space. 

SpaceX CEO Elon Musk posted on X the same day, “SpaceX, in partnership with Nvidia, has designed a space-optimized Vera Rubin NVL72 system for launch to orbit in Q4 next year, with significant scale in 2028.”

SpaceX, in partnership with Nvidia, has designed a space-optimized Vera Rubin NVL72 system for launch to orbit in Q4 next year, with significant scale in 2028 https://t.co/qdDq8YBkzl

— Elon Musk (@elonmusk) August 24, 2026

Musk’s post came after he said during SpaceX’s Q2 earnings call, “Going forward, we’ve decided to build exclusively on Nvidia because we think the Vera Rubin architecture is the best architecture.” Musk continued, “This is not some sort of far-future, distant thing; we expect to start launching these next year. We think the design of the NVL72 VR computer is a much better design than, say, having a standard rack -style design. So we expect to deploy this on the ground as well as in orbit, because we think it’s going to be a radical simplification of the standard NVL72 rack. It will cost less. It will be more effective. If we’re going to put it in space, why not want to put it on the ground? I think that’s going to be pretty cool.”

On Earth, the Vera Rubin NVL72 is Nvidia’s rack-scale AI design that combines 72 Rubin GPUs and 36 Vera CPUs, alongside high-speed networking components such as ConnectX-9 SuperNICs. Nvidia says SpaceXAI’s planned Starmind satellite will be based on an optimized version of that system.

A conventional NVL72 rack assumes gravity, technicians, stable grid power, a building-scale liquid loop and frequent replacement of failed parts. Orbit removes each of these assumptions.

The idea is more ambitious than putting a conventional edge-AI accelerator aboard a spacecraft. Nvidia and SpaceXAI are proposing to bring a modified architecture used in AI data centers into orbit, while altering it for orbital operational requirements.

Getting that working in orbit, though, is easier said than done. 

As Curtis Pyke, founder of Kingy AI, writes, “A conventional NVL72 rack assumes gravity, technicians, stable grid power, a building-scale liquid loop and frequent replacement of failed parts. Orbit removes each of these assumptions.“

“Cooling is unforgiving. Space is cold, but vacuum does not carry heat away through convection.”

In particular, Pyke continues, “Cooling is unforgiving. Space is cold, but vacuum does not carry heat away through convection. Heat must travel from the chips to the radiator surfaces and then leave as infrared radiation. SpaceX says AI1 can avoid chillers, cooling towers and fans and reduce cooling overhead by an order of magnitude.”

SpaceX explains that AI1 would instead use closed-loop liquid cooling inside the spacecraft and large deployable radiators to send heat directly to space as infrared radiation. While the claimed reduction is physically plausible in principle, there’s no proof yet that these AI satellites’ cooling systems can deliver. 

Another major problem that remains unaddressed is how to make the orbital rack radiation-tolerant. Making Vera Rubin NVL72 radiation-tolerant means far more than putting an ordinary NVL72 rack in a shielded satellite enclosure. It would require a system-level redesign of its GPUs, CPUs, memory, networking, power, cooling, firmware, and operations around a specified orbit and mission life.

LEO orbit is not benign. NASA cites typical trapped-particle dose rates of 100 to 1,000 rad(Si) per year for low-inclination LEO spacecraft below 500 km. That level of radiation is not an immediate death sentence for electronics, but over a multiyear mission it will cause cumulative degradation. Radiation-qualified space hardware can deal with that. Commercial Off-The-Shelf (COTS) electronics are another matter. A true radiation-hardened Rubin GPU would also require design changes at the transistor and circuit levels. 

Even were Nvidia to make such a chip, for a high-density AI system such as the SpaceX design, the concern isn’t simply whether one processor survives a 5- or 10-year dose. The satellite contains numerous radiation-sensitive elements, such as GPU logic, SRAM caches, register files, system memory, and memory controllers. With thousands of cores and billions of memory storage cells, the aggregate fault rate — not the behavior of an individual component — drives the design.

The most realistic near-term answer would be a radiation-tolerant, fault-managed Rubin-derived orbital system, not a fully radiation-hardened NVL72 in the traditional military-space sense. It could use selected commercial Nvidia parts, substantial shielding, ECC and data integrity mechanisms, redundant controllers and power paths, aggressive fault detection, software recovery, and reduced-performance operating modes.

The post SpaceX designed an orbital Vera Rubin. Radiation comes next. appeared first on The New Stack.

This duck will teach you reinforcement learning — and pick up your socks

27 août 2026 à 18:20

You could have a mechanical duck waddling through your home before Christmas. 

Hugging Face‘s Pollen Robotics on Thursday opened pre-orders for the Microduck, a $399 (introductory) bipedal duck-adjacent robot that can walk, waddle, and use roller-skates(!), with a beak to pick up objects. And when it falls, it can get back up, too.

The company expects to make the first deliveries of the Microduck before Christmas. It’s available in North America and Europe.

Microduck is the follow-up to Reachy Mini, the desktop robot Hugging Face and Pollen launched last year, which was stationary and focused on interacting with humans. Pollen also still sells Reachy 2, a far larger and pricier humanoid aimed at research labs. Microduck, the team writes in its announcement, is meant to focus on action.

Credit: Pollen Robotics.

“How do you teach a robot to move? How do you train a behavior in simulation, transfer it to real hardware, see what went wrong, and try again? What changes when the robot can leave the desk, carry something, fall over, and recover? It is an ideal platform for developers who want to train physical behaviors, experiment with reinforcement learning, and test how AI moves from simulation into the real world,” the team writes.

Not just a toy

And indeed, Microduck is not just a 25cm-tall toy. It’s an open-source platform with an SDK, virtual training environment, and reinforcement learning scripts to help developers train the robot to perform new tasks. The code is Apache 2.0, though the hardware design files are licensed non-commercially, so nobody is building and selling a clone.

There’s also a full simulator for those of us who just want to play with a duck robot, but if you do buy one, you’ll also get a game controller to control the robot on the fly, too.

Credit: Pollen Robotics.

But if you want to go deep, you can use the physics simulator to teach the robot new movements. On the project’s GitHub page, the team notes how to train the duck’s walking policy across 4,096 virtual ducks in parallel, for example, which results in a usable gait in one to two hours. The repo registers 13 task families in all, including a forward roll and six built around a set of passive wheels that go under the feet.

To train the robot, you’ll need an Nvidia GPU, or you can train it on Hugging Face’s own infrastructure. That’s not incidental. The simulator the ducks run in is MuJoCo Warp, built on Nvidia’s Warp framework, and mjlab, the training framework underneath, reimplements the API of Nvidia’s own Isaac Lab.

Out of the box, the robot comes with seven trained moves, including walking, sitting and standing, kicking, grabbing objects with its beak, roller skating, and getting back up off the ground.

Inside the hardware

As for the hardware, the robot will weigh in at about 800 grams and will be powered by a Rockchip RK3566 with AI accelerator. That’s basically a quad-core Arm Cortex-A55 with a Mali GPU. But now that Nvidia is reportedly in talks to acquire Hugging Face, in a deal first reported by The Information, I would expect a future version to use a slightly more powerful Nvidia-made chip.

Credit: Pollen Robotics.

It features 1GB of on-board memory and 32 GB of storage.

What’s more important, though, is its set of sensors. There’s a single front camera, a small LiDAR sensor with an 8×8 time-of-flight matrix, and 2 inertial measurement units. The camera and LiDAR let it see and place objects around it, while the IMUs keep track of its orientation and balance.

The team is still working out what the final camera resolution and LiDAR range will be.

Pollen Robotics will sell a few accessories as well, including, for example, a Charger Pack for $39 with two batteries and a charger, and a Dev Pack for $119 with three spare motors, five motor cables, two batteries, a dual charger, ten NFC tags, Hugging Face credit, screws, and a screwdriver.

Credit: Pollen Robotics.

The robot also comes with microphones and a speaker. One interesting note here: when you first turn the robot on, it generates its own signature sound that is different from any other Microduck. The robot doesn’t speak, though, as the team notes, the Microduck “communicates through weird little sounds, closer to a creature than an assistant.”

Everything is better with more ducks

The team says having several of them together is what really makes the robots come alive. “Races, football, or simply robots reacting to one another immediately make the experience feel more alive. For developers, it also creates a practical way to explore multi-robot behaviors without a room full of expensive hardware,” they write.

At the end of the day, this is also just a fun project, and in this depressing world, we all deserve some ducking fun every now and then.

The post This duck will teach you reinforcement learning — and pick up your socks appeared first on The New Stack.

Nvidia’s $12.9B Hugging Face deal has an open-source problem

27 août 2026 à 17:35
Abstract yellow path

Nvidia has reportedly agreed to buy Hugging Face for $12.9 billion, putting one of the biggest names in AI hardware in charge of a platform developers rely on to find and run open models.

The Information first reported the deal Wednesday, citing a person familiar with the agreement. Nvidia and Hugging Face had not publicly confirmed it as of publication.

Hugging Face doesn’t push developers toward one chipmaker, which is what makes the acquisition interesting. Its Optimum libraries work with Nvidia’s TensorRT-LLM and also support hardware from AMD, Intel, and AWS. Projects such as Optimum AMD and Optimum Intel let developers run Transformers and Diffusers models on non-Nvidia hardware.

The company also plays a role in what happens after a developer chooses a model, including how easily they can get it running on the hardware they want to use. The problem for Nvidia might be this: It is buying a platform whose value depends on openness and hardware neutrality, but if the purchase means Nvidia hardware is favored, that value might diminish.

The company also plays a role in what happens after a developer chooses a model, including how easily they can get it running on the hardware they want to use.

Hugging Face already sits between the model and the chip

Hugging Face has expanded well beyond file hosting. With Inference Endpoints, developers can deploy a model from the Hub while Hugging Face handles the underlying infrastructure.

Those hosted deployments can run on AWS, Microsoft Azure or Google Cloud, but most of the GPU options Hugging Face lists are Nvidia chips, including the T4, L4 and A100. That gives developers a wider choice of hardware through Hugging Face’s open-source libraries than through its hosted services.

If the deal goes through, Nvidia would own both sides of that experience.

Deployment defaults favor Nvidia

NIM (Nvidia Inference Microservices) already works with models hosted on Hugging Face. Developers can point NIM to an hf:// repository path and pull the model directly from the Hub.

Owning Hugging Face would give Nvidia more room to bring NIM and CUDA-optimized containers directly into the deployment experience. Nvidia hasn’t announced plans to make NIM the default, and support for AMD and Intel could remain exactly where it is.

Owning Hugging Face would give Nvidia more room to bring NIM and CUDA-optimized containers directly into the deployment experience.

The bigger question is what happens over time. Nvidia could provide earlier support for new models on its own hardware or make deployment easier. At the same time, AMD, Intel, and AWS may have to reconsider how much engineering work they want to contribute to integrations maintained within a competitor-owned platform. Some of that work could eventually move elsewhere.

For developers, the difference may come down to which path requires less work. A competing chip doesn’t have to disappear from Hugging Face to become less appealing if an Nvidia model deployment takes fewer steps. We’ve seen a similar fight over the layers between AI models and the developers using them, with Cloudflare building more of that infrastructure itself.

Open models counter custom chips

That tension also helps explain why Hugging Face could be worth considerably more to Nvidia than its revenue alone would suggest.

Nvidia has been expanding its own Nemotron family of open models while investing heavily across the AI ecosystem. At the same time, some of its biggest customers are working to reduce their dependence on Nvidia hardware.

Google has its TPUs, AWS has Trainium, and Microsoft has Maia. OpenAI and Anthropic are also developing their own AI server chips. The push extends beyond the hyperscalers. Earlier this month, five European companies committed to purchasing AI compute built around non-Nvidia hardware that hasn’t been manufactured yet — a sign that the appetite for alternative accelerators is strong enough to attract forward contracts.

OpenAI this week published results from its new Jalapeño accelerator, which showed 1.5 to 1.9 times more work per watt while cutting end-to-end latency by up to 3.6 times on large open-weight models — although the chip has not yet been deployed at anything approaching Nvidia’s scale. A strong open-model ecosystem gives Nvidia a counterweight to that trend.

Open models are often expected to run in very different environments, and Hugging Face helps developers make that possible. A model found on the Hub might end up running on Nvidia hardware, an AMD GPU, or a cloud accelerator.

That flexibility is part of what Nvidia would be buying. Pushing Hugging Face too heavily toward its own hardware could make the platform less useful to developers who rely on it to work across different systems.

Interest in those models is also growing. Models from companies including DeepSeek, Moonshot AI and Z.ai have narrowed the gap with proprietary systems. At the same time, Hugging Face CEO Clément Delangue told The Information in June that the company had doubled its number of paying subscribers during the first six months of 2026. Delangue later said the company was “close to profitability.”

The Information puts Hugging Face’s annualized revenue at about $150 million. Against a $12.9 billion price tag, that’s a multiple of roughly 86.

So Nvidia would be paying for much more than Hugging Face’s current business. It would be buying a place developers already turn to when they want to work with open models, including models that don’t have to run on Nvidia hardware.

It would be buying a place developers already turn to when they want to work with open models, including models that don’t have to run on Nvidia hardware.

Community resists easy forking

Buying Hugging Face wouldn’t give Nvidia control over everything developers find there. Libraries such as transformers and diffusers are open source, and models on the Hub remain subject to their own licenses. Openly licensed models can still be hosted elsewhere, while the underlying libraries can be forked.

Much harder to recreate is the community Hugging Face has built around them. Developers already know where to look for models and have built workflows around the Hub and its integrations.

The post Nvidia’s $12.9B Hugging Face deal has an open-source problem appeared first on The New Stack.

OpenAI’s Jalapeño chip tackles a problem AI agents make worse

25 août 2026 à 16:00
silicon wafer up close

When OpenAI unveiled Jalapeño, its first custom inference chip, in June, the company made some big promises. The chip, developed with Broadcom, was built from scratch for large language model inference, with OpenAI saying early testing showed substantially better performance per watt than existing accelerators. At the time, though, OpenAI didn’t release the detailed performance results to back that up.

On Tuesday, OpenAI published its first results from working Jalapeño silicon across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5. The results show what OpenAI was aiming for with Jalapeño: higher throughput without the longer response times that can come with it.

“Agents need to complete many steps in sequence, so delays can compound across an entire task.”

Agents compound inference delays

An agent may call a model over and over as it works through a task, using tools and deciding what to do next based on the results, which means a delay that barely registers during a single inference can become much more noticeable when it happens repeatedly over the course of a longer task.

“Agents need to complete many steps in sequence, so delays can compound across an entire task,” OpenAI said.

Jalapeño was designed with those delays in mind. Different parts of running a large language model place different demands on the hardware, with the initial prompt requiring more compute and the response generation putting more pressure on memory bandwidth. Every time data has to move between cores and chips, that can add even more waiting.

Jalapeño takes a different approach, cutting down on that waiting without optimizing one part of the process at the expense of another.

“Agents need to complete many steps in sequence so that delays can compound across an entire task,” OpenAI said.

That helps explain some of the choices OpenAI made with Jalapeño. Running a large language model puts different demands on the hardware at different points: processing the initial prompt requires a lot of compute, while generating the response token by token relies more heavily on memory bandwidth. There’s also time lost whenever data has to move between cores and chips, leaving parts of the system waiting for what they need.

The idea is to reduce that waiting without optimizing one part of the process at the expense of another. Model state, including the KV cache used while generating a response, can be kept local, while Jalapeño’s networking allows more of the workload to stay within the same connected system. That means less time spent moving data around as the workload shifts between compute and memory.

Jalapeño’s first public benchmarks

OpenAI put Jalapeño through InferenceX, SemiAnalysis’ public benchmark for AI inference, using GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. Across the three models, Jalapeño handled 1.5 to 1.9 times more work per watt while cutting end-to-end latency by 1.7 to 3.6 times. On highly interactive workloads, OpenAI says it was 2.1 to 4.1 times faster than the systems it compared against. 

The differences become particularly large when Jalapeño is compared at the previous best time-between-tokens operating point. OpenAI reported between 8.6 and 104.3 times more work per watt, depending on the model.

OpenAI based the power-efficiency comparisons on each accelerator’s published power rating. Jalapeño is rated at 700 watts, although the company says it never drew more than 550 watts during these tests. The bigger point is that OpenAI isn’t trying to improve throughput at the expense of response time, which is often the tradeoff with inference. 

Batching more work can make infrastructure more efficient, but it can also mean making an individual user wait longer. OpenAI’s argument with Jalapeño is that an inference system increasingly needs to be good at both — particularly as the company continues cutting the cost of API access while also needing to keep interactive workloads responsive.

The bigger point is that OpenAI isn’t trying to improve throughput at the expense of response time, which is often the tradeoff with inference. 

AI-generated code runs faster

OpenAI used its own models throughout Jalapeño’s development, helping the hardware team move from initial design to tapeout in nine months by exploring implementations and shortening design, measurement, and verification cycles. The work didn’t stop once the chip was built.

The company says AI-generated implementations of selected GPT-OSS attention and mixture-of-experts blocks ran 1.5 to 1.8 times faster than versions written by its own experts. That doesn’t mean the entire model ran that much faster, but it does show what OpenAI is trying to do with Jalapeño: make the chip straightforward enough for AI, not just humans, to program and optimize.

Engineers describe work using local tensors, explicit communication and predictable synchronization, giving AI a way to help determine how that work should be mapped, placed and scheduled across the system. That could make it faster to adapt the chip as new models come along, although OpenAI says each new model family still requires its own kernels and optimizations.

Custom silicon meets model roadmap

Using Codex with GPT-Astra and earlier OpenAI models, the hardware team brought three open-weight models that weren’t part of Jalapeño’s original production plan to high performance within two months. That fits with OpenAI’s broader plans for Codex, which the company has said is still early in its development, and shows how it could eventually play a role well beyond writing code.

OpenAI plans to start using Jalapeño in its own infrastructure by the end of the year, and it’s already working on the next two generations. The company will continue to use accelerators from Nvidia and other partners, but building its own chips gives OpenAI more control over how the hardware evolves alongside its models.

OpenAI plans to start using Jalapeño in its own infrastructure by the end of the year, and it’s already working on the next two generations.

The post OpenAI’s Jalapeño chip tackles a problem AI agents make worse appeared first on The New Stack.

❌