❌

Vue normale

Reçu avant avant-hierThe New Stack

How a forgotten node can put Oracle Java back in production

24 septembre 2026 à 19:50

Enterprise Java company Azul announced its Azul Intelligence Cloud AI Assistant on Wednesday. The technology arrives in response to industry-wide alerts that AI has become a force multiplier for threat actors today.

The Azul service is a natural-language query interface that allows software engineering teams to see where security risk and licensing infringements are hiding in their live production Java estate. The assistant provides answers that are “grounded in live runtime data”, so it replaces static reports that grow less accurate after the day they’re generated.

How fast do code-scanning reports go stale?

Azul said that “most IT and engineering teams” still manage Java risk with static IT and software asset management (ITAM/SAM) reports and code-scanning tools that describe a moment in time. The company has insisted that these reports are “accurate on the day they’re generated, and increasingly wrong after that”, typically because Java Virtual Machines (JVMs) are spun up, patched, drifted and retired underneath the report’s scope.

“For years, enterprises have built dashboards and reports to understand what’s actually running in their Java estate, but by the time a report gets properly summarized and reviewed, the risk it describes has often already changed,” said Scott Sellers, co-founder and CEO of Azul. “That used to be a productivity problem. Now that AI can find and weaponize a vulnerability in hours instead of weeks, it’s a business risk – for security, for compliance and for the licensing exposure that shows up in an audit.”

“Now that AI can find and weaponize a vulnerability in hours instead of weeks, it’s a business risk…”

The Azul Intelligence Cloud AI Assistant lets software teams ask a direct question, in plain language, and get an answer grounded in what’s actually running in production right at that moment in time, as well as query historical information for further analysis.

The weaponization gap is closing

Where enterprises run business-critical workloads on Java alongside AI services, Azul said the practical effect is that the gap between a Common Vulnerabilities and Exposures (CVE) entry being disclosed and it being weaponized is getting shorter. 

Citing AI models such as Anthropic’s Mythos and OpenAI’s Aardvark, which have autonomously discovered real-world vulnerabilities, the company pointed to an April 2026 Cloud Security Alliance white paper (listed as unofficial AI-assisted research), which suggested that, “While organizations historically took a median of 32 days to apply patches to known vulnerabilities – a window that once roughly corresponded to the time available before exploitation began – that window has collapsed to approximately 5 days for median time-to-exploit in 2025.”

Azul highlighted its Intelligence Cloud service, which gives engineers two “continuously updated” records of their Java estate: JVM Inventory, a live catalog of every JVM instance running anywhere (on-premises, cloud or container) and Code Inventory, a runtime record of which code actually executes in production versus what is merely provisioned. The new AI Assistant puts in a conversational layer using LLM models on top of both.

Software engineers can ask questions such as: 

  • “Which JVMs are running Java versions which are not the latest updates?”
  • “Where is Oracle Java running in production right now?”
  • “What code hasn’t run in the past four quarters and is safe to remove?”

In the FAQ section of its announcement, Azul suggested that post-migration JVM “drift is common”, often due to a rollback, a forgotten node, a shadow deployment or various scripts and processes that haven’t been updated, which can reintroduce an Oracle Java runtime, exposing compliance and licensing risk, if not security risk also.

The shape of the Java runtime security market

In terms of which other vendors operate in the Java runtime analytics and security market, there are more than a handful of usual suspects. Contrast Security is known for its JVM agent and in-app bytecode instrumentation. Dynatrace offers Runtime Vulnerability Analytics as an extension of its core observability platform, which ships with OneAgent monitoring for Java vulnerable functions and JVM-level bytecode instrumentation agents.

Through its acquisitions by HP, Micro Focus, and now OpenText, Fortify remains known for its static and dynamic application testing services, including Fortify Application Defender, a runtime application self-protection (RASP) agent built to monitor Java workloads during execution. Part of Thales, Imperva’s runtime security for Java and .NET applications spans simple access control up to complex anomaly detection algorithms, though Imperva has reportedly put its standalone RASP product on an end-of-sale path. Then there’s Datadog, with its Application Performance Monitoring (APM), built to power code-level distributed tracing from browser and mobile applications to backend services and databases.

“Runtime context is absolutely critical for understanding the real risk in the production environment.”

A busy market for sure, so just how much of a problem are now-anachronistic static reports?

Head of security advocacy at Datadog, Andrew Krug, tells The New Stack that the downside of most point-in-time inventory scans is that they are “not always representative” of the runtime environment. 

“Runtime context is absolutely critical for understanding the real risk in the production environment,” Krug says. “Even in the most mature software development lifecycle (SDLC) flows, tooling that generates static software bill of materials (SBOMs) may be bypassable [i.e. circumventable or subvertible] to get a feature deployed. Moreover, traditional vulnerability management flows outside of SDLC can bump versions outside of CI/CD process, accidentally compounding the problem and introducing added risk by bypassing known good guardrails like dependency cooldowns.”

“Even in the most mature software development lifecycle (SDLC) flows, tooling that generates static software bill of materials (SBOMs) may be bypassable… to get a feature deployed.”

Krug further advises that Datadog now sees “an increasing rise” in automated drive-by attacks on known vulnerabilities, “particularly Java” in many cases.

“Attacks used to be added to scanners, either specifically or generically, and scans are indiscriminately against targets,” Krug clarifies. “LLMs make it cheaper to add support for new vulnerabilities. However, it also makes it easier for individual researchers/hackers to have their own custom rulesets.”

“LLMs make it cheaper to add support for new vulnerabilities.”

He advises that this fact makes trends much harder to read than “oh, someone added support to CVE-2026-whatever in FFUF”, so today the goal of many attacks is unchanged.  Attackers are looking to move laterally, establish persistence, and often automations will look for credentials to leverage to do just that.

NOTE: (Fuzz Faster U Fool) is an extremely fast web application fuzzer (written in the Go language), which is used by security testers to discover hidden files, directories and endpoints.

Dead code; it’s really a ‘thing’

The above-noted list of competitors that work in Azul’s marketplace is (arguably) substantial evidence of the real commercial licensing and security risks that exist where Java code, redundant JVMs and chunks of unsubstantiated (or more likely just untracked) Java components have been left to roam free. Azul itself noted that there’s a real maintenance overhead here that needs to be addressed because “unused and dead code that still gets tuned, tested and carried through every migration”, usually because no one can assess it’s safe to remove.

The larger the Java estate, the larger the exposures get, obviously. But that also means that the less a point-in-time report can be trusted to catch these exposures before they become an incident, an audit finding or a breach.

Dedicated compliance officers will likely enjoy wider deployment of these tools, although that role itself may now reside within a DevSecOps or platform engineering team, or both. 

Azul Intelligence Cloud AI Assistant works regardless of which JVMs are deployed, from which vendor, or how old or large the applications running on them are. JVM Inventory and Code Inventory retain component and code-use history over time, so the AI Assistant can reason over which code, JVMs and applications have actually run in production, now and in the past. 

The post How a forgotten node can put Oracle Java back in production appeared first on The New Stack.

Claude Opus 5.5 wants to finish your coding tasks, not just start them

22 septembre 2026 à 21:49

Anthropic wants developers to use Claude and its family of tools to handle complete coding tasks. The company introduced Claude Opus 5.5 on Tuesday to span more of the software application development lifecycle: from design specification creation through debugging to code generation and testing.

As the first release in a new family of Claude 5.5 models, Anthropic says that Claude Opus 5.5 performs at the level of Claude Fable 5.1 on “most work” tasks and costs around 40% less to run than Opus 5. 

GitHub chief product officer Mario Rodriguez is quoted in Anthropic’s launch announcement on exactly where software engineers sit today with frontier models for code automation. He said that developers want agents.

“In our testing [of Claude Opus 5.5] across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it’s making developers’ bigger projects more achievable,” says Rodriguez.

Where does Claude Opus 5.5 get its power from?

Anthropic explains that the Claude 5.5 family’s expanded full-lifecycle capabilities were developed under an established set of practices.

These practice elements include extensive alignment testing (model evaluation processes put in place to make sure actions and outputs closely match human values and intended goals), pre-release evaluation by outside organizations, and safeguards for high-risk areas such as cybersecurity and biology.

The organization says that on its most comprehensive alignment test, Opus 5.5 is the strongest-performing model tested to date, with particular improvements in several behaviors that contributed to recent cybersecurity incidents (e.g., biased reasoning, attempting to escape a sandbox, and others).

Independent SRE & AI reliability architect Akash Thakur tells The New Stack that Anthropic’s work getting its model to complete whole coding tasks is impressive, but “getting it to know when it hasn’t” is the harder problem developers need to think about.

“Models working at the level of Claude Opus 5.5 are genuinely good at breaking a project into small, finishable pieces — and that’s the real unlock, because momentum on software comes from finishing things, not starting them,” Thakur says. “…But ‘completed’ and ‘correct’ aren’t the same thing. The task that looks done and completed is exactly the one that often costs the team later, so the win isn’t removing the human — it’s moving them from writing the code to verifying it.”

“Models working at the level of Claude Opus 5.5 are genuinely good at breaking a project into small, finishable pieces …but ‘completed’ and ‘correct’ aren’t the same thing.”

Suggesting that we are now witnessing a “higher bar for more capable models”, Anthropic said that models that could fully automate AI research itself should meet a higher safety standard. The company noted that as AI becomes more capable, public policy should play a larger role in making sure these systems are safe. Its recent work with Accenture is offered as an example of how Anthropic is building the infrastructure to support this.

Sprawling jobs: codebase-wide migrations & audits

Getting more specific, Anthropic has claimed Opus 5.5 is “particularly good” at long, sprawling jobs like codebase-wide migrations and audits. An early tester said it audited and fixed a 200,000-line codebase in under three hours, while Opus 5 took over 20 hours and used 2.5x as many tokens. 

“In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less,” said Anthropic in a press statement.

Field CTO for the EMEA region at Coder, Eric Paulsen, tells The New Stack that he is happy to hear Claude Opus 5.5 is becoming more holistically capable, but he’s “not surprised” because AI is “eating the software delivery chain end to end” today.

“Despite the efficiencies showcased in Opus 5.5, the moment an agent can reliably finish real work, the question stops being whether the model is good enough and becomes where you’re letting it run,” Paulsen says. 

“Kicking off a Claude Code session on a laptop that can be compromised, stolen, or simply run out of compute is not an engineering environment; capable agents need dedicated, governed infrastructure with the same guardrails, secrets handling, and observability a developer would demand of any other production workload. As we’ve seen here with Anthropic’s own work, firms need to underline testing and safeguard procedures for launches of this kind,” he added. 

Anthropic assures users that external testing has been conducted and that the model was tested before release by Frontier Design and METR. When established safeguards for Opus 5.5 intervene, requests “fall back to another model transparently,” meaning developers may not see which model actually handled a given call.

In cybersecurity, users can identify and fix bugs in their code, but most cybersecurity tasks will be re-routed to Opus 4.8. Requests flagged by biology and frontier LLM development classifiers will be re-routed to Opus 5. Vetted organizations can apply to Anthropic’s Life Sciences Verification Program to use Opus 5.5 for biology research, and the company says it will expand its Cyber Verification Program in the coming weeks.

The developer’s terminal prompt has fundamentally changed

HasData co-founder, Sergey Ermakovich, tells The New Stack that his work as a web scraping specialist for data pipelines and AI means he sees zen-like, one-brick-at-a-time logic in what Anthropic has done. 

“The biggest change — and it’s a trend that will have driven Anthropic’s design and development aspirations for Claude Opus 5.5 — is that the terminal is no longer just a place where a developer pastes generated code,” Ermakovich says. “The terminal now becomes part of the model’s workspace.”

Because a model can run commands, inspect failures, modify files, and verify the result, Ermakovich suggests it can “close the loop” instead of handing unfinished work back to an engineer. “That makes small complete tasks much more valuable. Fix one failing test, update one dependency, migrate one endpoint, verify it, then move to the next task,” he adds.

“The terminal is no longer just a place where a developer pastes generated code — the terminal now becomes part of the model’s workspace.”

Commenting on the development of this model as part of Anthropic’s approved corporate messaging, John Ruelas, staff software engineer at Ramp, said that verbose, hard-to-follow output has been his “biggest frustration” with frontier models, but “Claude Opus 5.5 fixes it” for him.

He noted that it writes like a good colleague and follows his company’s writing rules. A design spec came out usable with very minimal edits, and when it rewrote one of his prompts, he preferred its version to his own. When it optimized the team’s test suite, he could follow its reasoning easily and “shipped the change with confidence.”

We’re so done with the autocomplete era

Vice president of AI strategy at Abbyy, Maxime Vermeir, tells The New Stack that the more lifecycle-wide Claude Opus 5.5 features on offer show that “we’re done with the autocomplete era” for basic code automation tools.

“Anthropic’s elevation here reflects the fact that it used to be thought of as marvelous if AI could complete a developer’s next line of code, but today the expectation is that you hand it a whole Jira ticket and that it gets done,” Vermeir says. “But despite the power on show with Claude Opus 5.5, the question remains as to how much these newer models will actually understand what ‘done’ means, as often it seems they have a desire to keep burning tokens by offering you yet another thing it didn’t quite do right.”

Opus 5.5 also “communicates more naturally” than prior models, with early testers finding its writing clearer and easier to follow, making it a better work partner over long sessions.

Costed out lower than Opus 5, Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens (20% less than Opus 5), and Anthropic has cut cache read prices by 60% for token-billed usage. Opus 5.5 also needs fewer tokens for higher-quality work and generates output more than 30% faster than Opus 5.

The post Claude Opus 5.5 wants to finish your coding tasks, not just start them appeared first on The New Stack.

TypeSafe launched Jev because sequential LLMs are “totally useless for computers”

21 septembre 2026 à 21:36

When TypeSafe emerged last week after two years in stealth, backed by $40 million in seed funding led by DCVC, to launch its first model, Jev, it claimed something that counters just about everything the industry has built since ChatGPT: The model doesn’t write. It decides.

The first of what the organization calls a new class of System One models, Jev is a text-only model that machines can use natively to make decisions inside software applications. Developers can send Jev structured questions and get typed decisions with calibrated probabilities, meaning software can account for uncertainty.

TypeSafe has built a new architecture for Jev, a new sampler (an algorithm that selects tokens from a model’s predicted probability distribution to control randomness, creativity, and consistency), and a new training algorithm known as Reinforcement Learning for Calibrated Decisions (RLCD). 

Sequential LLMs are totally useless for computers

Co-founder and CEO of TypeSafe, Diogo Almeida, is ex-OpenAI, where, according to TypeSafe, he co-invented RLHF and InstructGPT, the methods behind ChatGPT and GPT-4.

Almeida posted on X on September 15 to state, “The improvements are clear if you see them [LLMs and Jev] side by side. Ask a System One model a ton of structured questions just like you would an LLM. Get the answers back near instantly. Meanwhile, LLMs take hundreds of times longer to respond. Look at how the LLM generates sequentially, which is great for a natural conversation, but totally useless for computers.”

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?

I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev

• 20-200x faster
• 40-400x… pic.twitter.com/JSybNG2BKJ

— Diogo Almeida (@CompleteSkeptic) September 15, 2026

“Look at how LLMs generate sequentially, which is great for a natural conversation, but totally useless for computers.”

Almeida said the inspiration for Jev came from asking himself: why haven’t superhuman chat models led to artificial general intelligence yet? He said that Reinforcement Learning from Human Feedback (RLHF) chat has led to LLMs that are “optimized for human preferences” and include issues such as mode dropping, overconfidence, and an overall lack of reliability.

TypeSafe: System One models can’t hallucinate

Almeida’s launch post lists headline stats for Jev as 20-200x faster, 40-400x cheaper (with output tokens free), and frontier composable intelligence, optimized for decisions. Almeida further claimed that TypeSafe System One models output decisions with probabilities and confidence instead of words, and they “can’t hallucinate” because they’re “a lot more like code”, i.e., reliable, fast, self-consistent, and type-safe. 

Frontend cloud company Vercel has noted that, within 24 hours of launching on AI Gateway, “Jev from TypeSafe AI reached more than twice as many paid teams as any previous model launch, making it the fastest-adopted model in gateway history. Jev passed every other comparison model in its first twelve hours and continued to widen its lead for the rest of the day. By hour 24, nearly 13% of paid teams were using it. That’s 2x the GPT-5.6 family and more than 6x Fable 5.1’s share.”

Developers can set thresholds for when Jev acts autonomously vs. when it needs human oversight and review. They can then combine those decisions in code to build larger workflows, with control over how the intelligence is used. Software engineers can use Jev to select an agent’s next tool or subagent; they can also use it to confirm the veracity of a model’s output and set guardrails.

In a blog post titled “A deep dive into Jev, TypeSafe’s System One model,” independent software developer Flavio Copes noted that Jev is “not a chatbot” like ChatGPT, and it is not a coding model. It does not write replies, explanations, or code.

Jev is a smart if statement

“The simplest way to describe it: Jev is a smart if statement,” wrote Copes. “The important difference is where the AI sits. With ChatGPT or a coding agent, the AI is the main interface or worker. Jev is a small component inside a regular application. You add it where code needs one judgment, while the rest of the product stays ordinary code.”

“You add it where code needs one judgment, while the rest of the product stays ordinary code.”

Copes reiterates TypeSafe’s stated performance levels: most calls to Jev complete in about 100 milliseconds, input tokens cost $0.042 per million, and (as already noted) output tokens are free.

You send it some data and a list of typed questions, and it sends back one answer per question: a yes/no probability, one option picked from a list you defined, or a position on a scale you defined. Every answer comes with probabilities. 

One developer gave Jev the “one thing it can’t handle”

To put Jev to the test, AI engineer Bartosz Mikulski tells The New Stack that because Jev is advertised as a text-only model, he gave it the one thing it can’t handle: pictures.

“I turned 400 hand-drawn sketches into Scalable Vector Graphics (SVG) coordinates and asked what they were,” Mikulski says. “It got about 35% right, where just ‘guessing’ typically returns 10%, but it answered ‘airplane’ for more than half of the drawings, so that number is part real ability, and part a heavy bias toward one label.”

To be fair, Mikulski notes that TypeSafe says in its own documentation that Jev reads text only and handles words better than numbers. 

“I fed it numbers that encode pictures, which is close to the least fair test anyone could design, and it still beat chance by a wide margin. I mean that as a compliment, not as a benchmark. It tells you nothing about how Jev does on the text classification it’s actually sold for,” Mikulski clarifies.

“I fed it numbers that encode pictures, which is close to the least fair test anyone could design, and it still beat chance by a wide margin. I mean that as a compliment, not as a benchmark.”

Why is TypeSafe Jev called Jev?

Jev is named after the 19th-century economist William Stanley Jevons and his Jevons Paradox: the economic principle that as technology increases the efficiency with which a resource is used, that resource’s total consumption actually rises rather than falls. 

When steam engines became more efficient, we used more coal, not less; when LED lighting dropped lighting costs, we used more lighting; when data compression algorithms lowered the bandwidth needed to stream video, global web data traffic increased… and so on.

Almeida concluded his X video post with a nod to developer productivity and said that, “As we say at TypeSafe, we’re building prod, not God.” The company’s comedy disclaimer is shown below.

The post TypeSafe launched Jev because sequential LLMs are “totally useless for computers” appeared first on The New Stack.

Automattic says CEO Mullenweg was gone and back inside 33 hours. What happened between?

16 septembre 2026 à 22:39

There were unusual goings-on this month at Automattic, a company known for its free and open-source app for building WordPress sites. 

The company issued a notice last Thursday confirming that CEO Matt Mullenweg was on leave.

Mullenweg, who also co-founded WordPress in 2003 before establishing Automattic in 2005, was temporarily replaced by company CFO Mark Davies before being reinstated less than a day and a half later.

By last Saturday, a new alert emerged stating that Mullenweg was back in his position and that “Matt was away for only 33 hours and 20 minutes.”

Automattic has not publicly explained what changed between the initial decision and Mullenweg’s return — and it has also declined to explain the circumstances behind the leave and return. 

By way of context, WordPress sits under the Automattic brand alongside the company’s other products, including the microblogging site Tumblr, the e-commerce WordPress plug-in service WooCommerce, and the instant messaging client Beeper. 

A 33-hour vanishing act, but Automatticians are supporting him 

Automattic director of communications Megan Fox is on the record saying, “Matt Mullenweg is the chairman and CEO of Automattic, with full support of the board. “And if you search online, you can see many top executives and Automatticians supporting him as well.”

A further report reproduced Slack messages written by Mullenweg where he said, “Happy to announce the board is back in agreement, and I’m in control of Automattic. A lot happened in the past 48 hours that we need to sort out, and I hope much of it was a misunderstanding, because I have huge respect and regard for those involved.”

Was this a failed boardroom coup?

Industry watchers may naturally suspect the knives were out and that this was a failed boardroom coup. 

CEO & CTO at HasData, Roman Milyushkevich, tells The New Stack that a company can survive a CEO departure, but it struggles when employees cannot tell which governance process is real.

“But, in terms of whether the real guns were out at Automattic, it certainly looks like a failed attempt to change control, but I would not call it a coup as an established fact,” Milyushkevich says. 

“The possibilities playing out here are all very different,” Milyushkevich adds. “There could have been a second board agreement brought into place inside that 33 hours; there could have been internal (or possibly even external) negotiations; directors could have reconsidered the practical consequences of removing the founder; or there could have been an internal resolution that has not been disclosed.”

He advises that the “most important thing Automattic can establish now” is not who won the dispute, but whether the board and CEO have a clearly understood process for handling the next serious disagreement.

“The most important thing Automattic can establish now is not who won the dispute, but whether the board and CEO have a clearly understood process for handling the next serious disagreement.”

Behavioral scientist and visiting professor at São Paulo’s FIA Business School, Ricardo D’Olivar, tells The New Stack that what matters here is whether stakeholders have “enough information to distinguish a considered correction from an unresolved struggle” over authority. 

“A reversal of this kind can reflect responsible reconsideration,” D’Olivar says. “Reinstatement answers who is in charge today. It does not, by itself, explain how a future disagreement would be resolved. This is where the potential consequences for employees, executive recruitment and investors arise. If uncertainty persists, employees may become more cautious about committing to decisions whose backing appears unstable.”

Suggesting that, responsible corporate mechanics or not, this kind of development undoubtedly throws the cat among the pigeons, D’Olivar says that developers or executives considering a career at Automattic may now question whether they would be held responsible for decisions they were authorized to make, but could no longer count on the organization to support when challenged. 

“Investors may seek clearer evidence that oversight and succession arrangements can operate under pressure. These are possible responses, not verified effects at Automattic. The information shared internally may also be more complete than the public account,” adds D’Olivar.

This is not the first boomerang CEO bounce

Mullenweg might be the fastest CEO yo-yo switcharound in history, but he’s certainly not the first. OpenAI CEO Sam Altman famously left the company’s board after a communication dispute. An employee uprising (nearly all the company’s 700+ staff threatened to resign) and added pressure from Microsoft led to Altman returning to his position five days later.

Perhaps even more famously, Steve Jobs was ousted in 1985, only to return 12 years later to realign a then-struggling Apple and take it into its golden years. Twitter (now X) founder Jack Dorsey was moved out in 2008 before a boomerang return in 2015. Michael Dell stepped down as CEO to become chairman of the board in 2004, but by the start of 2007 he was back.

At their origins, the shenanigans at Automattic may be redolent of the technology industry’s other boomerang CEO realignments, or this may be a boardroom tussle that we’ll never know the full reason for until Mullenweg writes his memoirs. Either way, the WordPress industry just got its first movie-script idea.

The post Automattic says CEO Mullenweg was gone and back inside 33 hours. What happened between? appeared first on The New Stack.

AI evaluator: The most important AI job in history? How developers might fill the proposed new job

16 septembre 2026 à 17:58
Lots of pink escape keys

The pace of frontier AI model development spurred Anthropic CEO Dario Amodei to publish an essay last weekend, calling for changes in how the industry is regulated and develops. In a three-part plan that includes both democratic and global coordination, Amodei writes that the first step was something Anthropic is committing to unilaterally.

“Each frontier AI company [should] commit to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes,” writes Amodei.

Amodei’s essay followed dire warnings from former Anthropic and OpenAI pretraining research specialist Jacob Coxon, who posted a thread on X saying the people building AI earnestly “believe that it could kill us all” by the end of the decade.

Shortly after Amodei published his essay, OpenAI CEO Sam Altman and SpaceXAI founder Elon Musk chimed in: “I agree with Dario,” posted Altman; “Dario is right,” posted Musk. Later that day, Demis Hassabis, founder of Google DeepMind, posted, “Dario’s essay points towards the right path forward.” In a post on X, Meta CEO Mark Zuckerberg writes that Meta Superintelligence Labs already uses independent evaluators, and that, “In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators.”

This week, theories began to surface about why the world’s biggest frontier AI labs would want to intentionally slow their pace when competition is so fierce. “The desire to slow down is puzzling, but perhaps if the whole system slows down, the rules of winning can be the same for all,” posted Nikesh Arora, chairman and CEO of Palo Alto Networks.

In his essay, Amodei likens the proposed job of independent AI evaluator to embedded regulatory supervisors used in the banking industry, i.e., third-party professionals. Altman describes the job as having “employee-like access” in his post on X.

So, who could fill these roles that AI leaders agree are desperately needed?

Salaries top out at $687K; are you interested?

METR’s current job openings are well paid (salaries top out at around $687,000), and the job specs are daunting. 

“You’re scrappy, creative, independent, and self-directed (because during the exercises you’ll only have a few other METR employees you can talk to). The work is novel, and you’ll need to figure a lot of stuff out on the fly largely by yourself. You are excellent at loss-of-control threat modeling and breaking down safety cases,” reads the spec.

“You’re scrappy, creative, independent, and self-directed. The work is novel, and you’ll need to figure a lot of stuff out on the fly largely by yourself. You are excellent at loss-of-control threat modeling and breaking down safety cases.”

Similar but less colorfully illustrated roles (paid between $180K–$300K) are also available at AI model training company Mercor, where candidates will need a Ph.D. or M.S. and more than two years of work experience in a computer science, electrical engineering, econometrics, or another STEM field that provides a solid understanding of machine learning and model evaluation.

“Employee-like access fluctuates wildly”

AI security consultant and CTO at Komodo, Kadan Stadelmann, tells The New Stack that his typical week sees him work differently with each client. This is because “employee-like access fluctuates wildly”, from rigid focus areas to broad access, and much of that aspect is determined by contracts signed before work begins.

“I am invited into labs to probe numerous risk vectors, including agent behaviors under realistic conditions,” Stadelmann says. “Among my duties are tasks that include monitoring chains-of-thought and prompts. The goal is to establish an objective and look at a specific AI system to question how autonomous the system is, and how long it takes to complete specific tasks. Most importantly, evaluators at my level monitor for how well a team adheres to its claimed safety practices.” 

Software engineering skills beat doctorates

Although METR wants evaluators to have a Ph.D. up their sleeve, Stadelmann says that as the prevalence of this role expands, he feels the technology industry has been, and continues to be, driven by people who can demonstrate strong engineering skills, not doctorates.

“I am invited into labs to probe numerous risk vectors, including agent behaviors under realistic conditions. Among my duties are tasks that include monitoring chains-of-thought and prompts.”

Questioned on whether costs create a barrier for smaller labs, Stadelmann notes that some AI evaluation work is funded by third-party non-profits, which protects independence. 

“But overall, evaluators will not be able to keep up with big frontier model firms. They will be out of control, and we will be dependent upon their own internal ethics. Plus, anyway, many of the smaller labs of any worth may inevitably be acquired by the big players in this space,” he adds.

What happens when an evaluator finds something wrong?

Founder and CEO of facial image AI identity governance company Indie Me, Dion Johnson, tells The New Stack that what interests him most about embedded AI evaluators isn’t the job title; it’s what happens when their judgment uncovers that the model behaved in a way nobody expected. 

“If the evaluator can only raise concerns when those concerns are convenient, then we have not created independent oversight — we have created another layer of process,” Johnson says. “The evaluator needs enough access to see the uncomfortable things, not just the polished demonstrations. They need to understand what happened during training, what failed during testing, what behaviors appeared unexpectedly, and where the team itself still has uncertainty.”

On the required skills AI evaluators need, Johnson agrees that technical depth, machine learning environment security experience, and software engineering as a whole matter.

“To choose a competent AI evaluator, I would look for someone who is deeply comfortable with uncertainty and deeply uncomfortable pretending they know something they do not. This is someone who can sit in a room full of brilliant people at an AI model company on launch day and say ‘I’m not convinced’… and that takes judgment and courage,” he adds.

“I would look for someone who is deeply comfortable with uncertainty and deeply uncomfortable pretending they know something they do not.”

Fear and loathing in the AI space

In his September 8 post, which has been viewed 172 million times and seemingly spurred AI leaders to change course, Coxon, the former AI researcher at OpenAI and later Anthropic, writes: “This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible but I hear the same people express fear privately. No other human activity poses this level of danger.”

As for where Coxon looks for work next, perhaps it might be a role in AI evaluation execution engineering.

The post AI evaluator: The most important AI job in history? How developers might fill the proposed new job appeared first on The New Stack.

“Machine translation is still broken for most of the world’s languages”: Cohere builds non-reasoning for a reason

13 septembre 2026 à 16:21
A scattered pile of overlapping alphabet cutouts in bright blue, pink, green, gold, red, and silver.

Enterprise AI company Cohere announced North Small Translate last week, a mixture-of-experts (MOE) open-weight machine translation model that works across 50 languages.

Developers can download the weights for noncommercial use under CC BY-NC 4.0. Cohere offers commercially licensed deployment through Model Vault, which is a Cohere-managed inference environment. Cohere positions the model as part of its sovereign AI strategy, aimed at organizations that want greater control over where their models run and how their data is handled.

North Small Translate builds on Cohere’s multilingual and translation lineage, which includes its Tiny Aya and Command A Translate model families. The company claims North Small Translate outperforms “similarly sized open-weight models” under 1T parameters, as well as API-based translation models in various dimensions of machine translation on average. 

Cohere co-founder Nick Frosst tells The New Stack that the model’s efficiency draws from the fact that it is non-reasoning, i.e., it relies on learned statistical patterns without a step-by-step logic process, which means it uses fewer tokens.

Machine translation is still broken for most of the world’s languages

“We spent nine years scaling an architecture invented to fix translation, and machine translation is still broken for most of the world’s languages,” Frosst says. “General-purpose models get you most of the way and then stop. The next phase of enterprise AI in this space is smaller, more specialized, and runs inside your own walls.”

“…machine translation is still broken for most of the world’s languages.”

In Cohere’s reported evaluation using WMT26 benchmarks, the company states that North Small Translate leads with a WMT26 All Languages benchmark score of 83.60, compared with 81.56 for Qwen 3.5 397B A17B, 76.50 for GLM 5.2 FP8, 81.37 for DeepL NextGen, 79.46 for Gemma 4 31B (on), and 68.20 for Google Translate. 

With its mixture-of-experts architecture and 218 billion total parameters, with 25 billion active. Cohere points to North Small Translate’s smaller compute & memory footprint than other models. Some model-to-model comparisons in this space aren’t fully substantiable, since not every vendor discloses parameter counts.

With current solutions, long documents start to fall apart

“Machine translation allows documents to be translated from one language to another automatically. With current solutions, long documents start to fall apart,” Frosst says. “Google Translate scores 21.3 on our long-context test, Gemma 4 31B 19.4; we score 48.9. That’s [for example] a safety manual that reads fine on page one… and has drifted by page ten. The other risk is where the text goes. Once you push HR policies or regulated documents through a third-party API, that data has left your building, and necessarily that means your control over it is diminished.”

“The risk [in machine translation] is where the text goes. Once you push HR policies or regulated documents through a third-party API, that data has left your building and necessarily that means your control over it is diminished.”

Explaining why the model offers “stronger translation performance” across complex enterprise translation tasks, Frosst says the model can support work spanning “a high volume” of sensitive documents. 

As well as its 50 languages (32 ‘high-resource’ languages + 18 others), the Cohere team explains that the model also supports translation-workflow-focused capabilities, such as structured translations (i.e., Markdown or JSON documents), instruction following (i.e., recommended tone & format), and terminology guides (i.e., providing specific vocabulary to use in the translation), all as part of the model.

“North Small Translate works with a multi-pass workflow,” explains Frosst. “The model translates, reviews its own output, finds errors, and fixes them – and this is the same loop we used in training. We ship both because standard is one pass and built for volume, while the agentic [version] spends more tokens for 84.36 against 83.60 on WMT26. That difference ends up being worth it when the document is a contract or a safety procedure, for instance, but in other cases you’d rather optimize for efficiency.”

“The model translates, reviews its own output, finds errors and fixes them.”

Model ‘steerability’ drives suggesting language tone and formatting

This model uses the same architecture as prior Cohere models but improves performance through post-training advances, including reinforcement learning and new datasets, specifically for machine translation tasks.

Frosst concludes that, across the translation model marketplace, generative machine translation models offer the highest quality and steerability (i.e., suggesting tone, formatting, etc.) but typically cost much more than Neural Machine Translation (NMT) models commonly used in commercial use cases. 

North Small Translate was developed in partnership with RWS, an AI solutions company pioneering in language technology and services. Collaboration with RWS, specifically with its Language Weaver research and science teams along with its language experts, helped shape the model’s real-world translation performance throughout development. 

As noted above, developers can access the weights free of charge for non-commercial use in three quantizations. There is also a Hugging Face Space and an API for those who lack the required hardware. 

The post “Machine translation is still broken for most of the world’s languages”: Cohere builds non-reasoning for a reason appeared first on The New Stack.

“Six tools, one harness”: Salesforce loops together a six-pack of favorites

10 septembre 2026 à 22:03

Salesforce introduced its Salesforce Enterprise AI Harness on Thursday as a formalized amalgamation of the AI harness concepts and infrastructure the company has been working to align.

The organization said that “no single system has the complete answer” to complete a straightforward business task, such as completing a customer order; i.e., CRM knows the customer, ERP knows the inventory, FSM (field service management) knows the delivery, and the support team processes… and so on. 

As such, a form of AI leakage pervades throughout modern enterprises, where individual agents and their harnesses do their best to enact automation intelligence, albeit in comparatively siloed chunks.

Harnessing a six-pack of toolsets

Salesforce’s answer is to coalesce what it calls “six trusted capabilities” (from its own platform toolset collection) alongside a new AI control plane, built to underpin an open and composable AI ecosystem.

The Salesforce Enterprise AI Harness encompasses core technologies across Data 360 (a unified customer data platform tool), Informatica (data integration and governance), MuleSoft and Agent Fabric (API connectivity and multi-agent orchestration), Tableau (visual business analytics), Agentforce (an agent platform), Salesforce Guardian (security and compliance), and the Salesforce platform itself through a common, composable architecture and unified experience.

“The Agentic Enterprise won’t be defined by which model a company chooses. Models will continue to change, and intelligence will increasingly be available everywhere. What will differentiate an enterprise is the trusted, proprietary context it brings to that intelligence — starting with the customer — and its ability to securely turn that context into action,” said Rohan Kumar, Salesforce president & chief platform and engineering officer, during press briefing.

“The Agentic Enterprise won’t be defined by which model a company chooses… what will differentiate an enterprise is the trusted, proprietary context it brings to that intelligence.”

This is not Salesforce’s first-ever harness

To be clear, it hasn’t taken Salesforce until late 2026 to ever produce or work with a harness; subsystems within the six pack, such as Agentforce Vibes (a natural language vibe coding tool), make use of specialized execution harnesses, including Mastra and the Claude Agent SDK, to manage local agent execution loops. This is — as suggested — a more formalized, total platform-wide development.

Alongside the six-way alignment spanning context, agency, action, governance, security, and models, Salesforce is offering a new AI control plane to give developers a place to view, manage, and control agents. The company confirms that software engineers can “use the six together as one system or take only what they need,” and create deployments with Salesforce technology, other third-party existing technology, or both.

The big question here is simple: Is this cosmetic packaging designed to disseminate wider Salesforce DNA into software developers’ production environments, or is it a genuinely useful simplification and unification process that will be met with interest and perhaps even gratitude?

Working engineers commenting on sites including the G2 developer forum and B2B software review portal have provided some insight.

What developers and operations professionals think of the Salesforce stack

Commenting on the use of Agentforce as a standalone tool, operations associate Ashish B. noted in August this year, “One area that could be improved is the initial setup and configuration process. Building effective agents can take some customization and a solid understanding of the workflow. The platform would be even better with simpler configuration options and clearer, more straightforward guidance on setting up agents for specific business use cases.”

Salesforce may have been listening. It said the Enterprise AI Harness connects reasoning to business rules, policies, and controls required for predictable execution. It then makes those capabilities reusable across the enterprise so that context can be shared across agents and models. This means actions and workflows can be securely invoked wherever they’re needed, and governance and security can be applied consistently as AI moves across the business.

“The platform would be even better with simpler configuration options and clearer, more straightforward guidance on setting up agents for specific business use cases.”

Writing about the Informatica user experience on Gartner Peer Insights in March of this year, a DevOps engineer said that the platform “works well” for integrating multiple data sources and supports both batch and real-time processing. But they caution, “Debugging and monitoring pipelines can be difficult in complex workflows. Initial setup is challenging for new users and requires some learning curve. [The] User Interface could be improved for better usability and faster navigation.”

Possibly taking into account such feedback, the new AI control plane that accompanies Enterprise AI Harness claims to give businesses a common place to see, manage, and control agents and AI across the enterprise.

“It enables companies to discover and register agents and AI capabilities, establish identity and policy, manage lifecycle, evaluate performance, observe behavior and outcomes, and control cost — across Salesforce and third-party AI. This gives enterprises a consistent layer of visibility and control as AI expands across teams, applications, models, and systems — without requiring every agent or AI experience to be managed separately,” pledged Salesforce.

Six pillars of trust

The whole premise of this Enterprise AI Harness hinges around what Salesforce calls six trusted capabilities. 

Trusted Context combines customer context with data, metadata, semantics, knowledge, real-time signals, memory, and an understanding of how work gets done across the enterprise. Trusted Agency gives agents reasoning, planning, state, memory, and orchestration functions, combining flexible AI reasoning and deterministic controls where certainty is required. Trusted Action securely connects AI to applications, APIs, workflows, tools, and business processes. 

As its name suggests, Trusted Governance governs the data, metadata, policies, and processes that AI relies on, with lineage, quality, guardrails, and controls. Trusted Security applies identity, permissions, privacy, data protection, and runtime security to what AI can access and what agents can do. Trusted Models provides security with intelligent model routing based on accuracy, performance, cost, and requirements. 

Integrations with Claude, Slack, Teams, etc.

The Enterprise AI Harness is being built headlessly from the ground up, with capabilities accessible through technologies including MCP, APIs, skills, and plug-ins. The company said this will let Salesforce capabilities extend beyond traditional Salesforce applications, into services such as Claude, Slack, and Microsoft Teams.

Many of the technologies that form the foundation of Salesforce’s Trusted Enterprise AI Harness are available today, with new capabilities and the unified experience planned to begin rolling out in early fiscal year 2028.

The post “Six tools, one harness”: Salesforce loops together a six-pack of favorites appeared first on The New Stack.

Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots

10 septembre 2026 à 19:01
Abstract dark red digital grid texture representing AI code verification loops and developer workflow checks.

Red Hat released Red Hat AI 3.5 this week, a move designed to let software engineering teams run AI with the same operational rigor as enterprise apps on mission-critical infrastructure.

Echoing both the “pilots-to-production” and “single control plane for infrastructure, models, and agents” narratives playing out across much of the tech industry, Red Hat’s key move here seems to be its expansion of platform capabilities to run enhanced multi-tenancy for AI service providers. 

Crucially, the new AI platform release is built to run AI use cases that require complete hardware-to-software isolation (needed when AI workloads have to wrangle sensitive data, proprietary models, and regulated information), as well as handle priority-aware service requests (where mission-critical workloads execute in favor of lower-grade tasks) with native multi-tenancy on a shared GPU infrastructure. 

Red Hat’s Senior Director of Product for Red Hat AI is Tushar Katarki. He tells The New Stack that running enterprise AI without safety controls is “like driving a supercar blindfolded” in real terms.

“With Red Hat AI 3.5, we are delivering the operational guardrails, verifiable trust, and multi-tenant controls needed to run AI as a mission-critical service rather than an unpredictable experiment,” Katarki says. “You can’t scale what you can’t measure, and you certainly shouldn’t deploy what you can’t verify. By unifying pre-deployment safety benchmarking, real-time observability, and GPU resource management, we are giving platform teams the power to turn isolated AI pilots into a fully governed enterprise architecture.”

Every GPU request now becomes a priority decision

Applied mathematician, data scientist, and fractional CMO Joshua Estrin, Ph.D., tells The New Stack that Red Hat’s work is of the time and of the moment; primarily because “every GPU request now becomes a priority decision”, so one developer’s internal experiment cannot be treated with the same urgency as a financial close.

“Looking at the state of AI infrastructure players out there now, Red Hat has clearly seen that priority-aware multi-tenancy lets companies use expensive compute more efficiently, but efficiency without isolation is just a faster way to create a security and reliability crisis,” Estrin says.

“Every GPU request now becomes a priority decision.”

He thinks that the winners in this market (he tags Nvidia, Nutanix, Suse with Rancher, HPE Ezmeral and VMware Cloud Foundation under Broadcom as usual suspects) will be the organizations that can “share capacity while still proving what happened where” in live production.

“That means proving whose workload actually executed and ran, who had access, what it cost, and what happens when demand spikes. Those answers rarely come from the infrastructure diagram; they get settled in the boardroom, usually after someone’s critical workflow has slowed down, but regardless, this sums up where AI infrastructure is now,” adds Estrin.

What are the elevated multi-tenancy pain points?

To unpack what’s happening here, let’s remind ourselves that GPUs are expensive, obviously. As organizations move to live production use cases of agentic AI, they will want to maximize their GPU state’s ability to serve multiple workloads across multiple teams, multiple customers, multiple apps, and so on. 

This all means that the breadth of AI infrastructure efficiency becomes the new agentic bottleneck.

The priority-aware services above for native multi-tenancy on shared GPU infrastructure are important right now; this function dynamically allocates GPU capacity based on workload priority. When lower-priority workloads can be run on spare (or cheaper) capacity (rather than separately provisioned GPU resources having to be spun up for lesser jobs), then everyone gets to go home earlier on Friday.

“The breadth of AI infrastructure efficiency becomes the new agentic bottleneck.”

But that’s not all the balls being juggled here; Red Hat mentioned isolation too, and that’s a concurrently complex AI infrastructure discipline challenge. GPU compute resources managed through isolation techniques enable AI services to run without accessing or interfering with another service’s data, models, or compute environment. 

In other words, this combines hardware consolidation with strong tenant isolation.

What new Red Hat technologies are on offer?

Red Hat says this release lets developers verify models before deployment through EvalHub, enabling risk-focused safety benchmarking and regulatory compliance certifications. 

New observability dashboards give platform teams metrics for a real-time view of inference health, GPU utilization, and AI model performance. Non-admin users can access dashboards for per-user token consumption showback (another term for token tracking) and distributed inference workloads.

Also new is shared GPU control for multi-tenant inference. So-called “fair-share GPU scheduling” manages resource allocation across tenants, while priority-aware serving provides admission control and priority-based request routing to protect real-time inference. As suggested above. it also allows background workloads to use available capacity.

VP of product management at Nutanix, Anindo Sengupta, tells The New Stack that running multi-tenant AI at scale does indeed require secure tenant isolation.

“The essential isolation is best achieved through virtualization,” Sengupta says. “For specialized at-scale AI workloads, the choice could be to run Kubernetes on bare metal. On top of that, to create real value, agents need access to both LLMs that run on containers and enterprise systems (databases, business systems, etc.) that run on traditional infrastructure. For hybrid AI to run efficiently, the platform must manage both these environments in a performant way, with a common operating model.”

Observability & model-as-a-service showback

Built-in observability and MaaS showback in Red Hat’s latest release are present to provide per-user token metering, performance dashboards for models and agents, MLflow visual agentic tracing, and GPU utilization dashboards for operational and usage transparency.

For efficient GPU memory management, the general availability of CPU offloading and the developer preview of storage offloading allow models to handle longer conversations and larger documents without additional GPU hardware. 

Red Hat AI Hub also introduces agent templates and starter kits with pre-configured reference implementations for common enterprise patterns, including code review, document processing, and research workflows.

Red Hat is hoping its Red Hat AI 3.5 version release gets us to a state where GPU-based AI resources are viewed as a policy-controlled infrastructure pool.

Power without interconnect bandwidth is botched.

As enterprise AI pilots succeed and initial results show some returns, IT teams must then address the need to deliver at scale. But scaling AI across the business demands the same operational rigor as any mission-critical infrastructure: verified safety before deployment, precise resource controls across shared GPU environments, governed agent behavior and transparent usage metrics.

Yoram Novick, CEO of sovereign AI edge cloud provider Zadara, has previously been on the record on this exact topic. He has said that when teams need to scale AI, “Simply adding more GPUs without ensuring adequate interconnect bandwidth can lead to diminishing returns” in the modern AI era. 

Overall, with its ability to direct priority-aware inference, tenant isolation, capacity sharing, and observability, Red Hat hopes its Red Hat AI 3.5 release gets us to a state where GPU-based AI resources are viewed as a policy-controlled infrastructure pool.

The post Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots appeared first on The New Stack.

K2 Horizon just shipped as six new fully open models — developers aren’t fully convinced

9 septembre 2026 à 14:00
Color-coded data streams converge, then expand into hundreds of glowing dots against a dark blue background.

Based in the Emirati capital, Abu Dhabi, the Institute of Foundation Models (IFM) introduced K2 Horizon last week. This group of six AI foundation models, ranging from 0.9 billion to 375 billion parameters, is claimed to be the “largest fully open-source fleet of AI models” yet made available.

IFM uses “fully open” to mean more than mere downloadable model weights. Across K2 Horizon, it has committed to publishing training and evaluation code, training data where redistribution is possible or detailed construction recipes where it is not, plus configurations, logs and intermediate checkpoints spanning pretraining through agentic post-training.

The aim is to let developers inspect how the models were built, reproduce their development and adapt them for their own work. But that commitment should not be confused with complete availability at launch: All six models had downloadable weights, while the model cards for the 0.9B, 32B and flagship 375B said some training data, code or checkpoints would arrive later. The 32B release was also only a Stage 1 checkpoint, with the final model still to come.

“Open source is much more than open weights. Science works when others can see the data, follow the method, reproduce the result, and improve on it,” said Eric Xing, IFM founder and university professor at the Mohamed bin Zayed University of Artificial Intelligence. “K2 Horizon delivers on that need. Every model in the fleet ships with its training data, recipe, and evaluations. This is open science, and we believe it’s the best path forward for AI.”

“Open source is much more than open weights. Science works when others can see the data, follow the method, reproduce the result, and improve on it.”

All the open model components you can think of

IFM says it is opening the full training lifecycle for every K2 Horizon model, from pretraining through reasoning and agentic post-training. For each model, it is releasing — or has committed to release — intermediate checkpoints, training data, or detailed data-construction recipes. The checkpoints are snapshots saved throughout training, allowing researchers to examine how a model develops and reproduce or resume particular stages. The broader artifact set includes architecture details, mixture compositions, training code, configurations, fine-grained logs, evaluation results, and final weights.

Across reasoning, mathematics, coding, and agentic tasks, the K2 Horizon team says every model size performs well. The 0.9B model is designed for highly constrained environments such as smartwatches and smartglasses, while the 3.7B and 7B models bring advanced capabilities to phones and other on-device applications.

The dense 32B model and sparse 36B-A4B model deliver stronger performance for local hosting and on-premises servers. The 375B-A23B model brings the fleet’s strongest capabilities to demanding enterprise deployments.

The six models share a core architecture, vocabulary, training methodology, interfaces, and deployment tooling, with the 0.9B model having a smaller vocabulary. IFM’s dynamic model routing technique directs tasks to the most cost-effective model and provides developers with a path from prototype to production. The organization insists that the fully open code, training data, and recipes are “a significant step forward in transparency” that go “well beyond” the open-weights dialogue that has dominated AI industry headlines this year.

Could this openness have been more… open?

The IFM team is clearly aiming for differentiation through all-encompassing openness, but could it have been more open… and will it need to be even more open in the future? 

To answer that question, the level of openness varied by model at launch. The 3.7B and 7B models shipped with the full artifact set — including weights, recipes, training code and data — while the 0.9B model card, the documentation describing its capabilities and limitations, said its data and code were still forthcoming. The flagship 375B-A23B and sparse 36B-A4B initially arrived with final weights, with full training code, raw datasets, and intermediate checkpoints promised in later updates. The 32B model shipped as an incomplete Stage 1 checkpoint, with the final model and remaining artifacts to follow.

But really, the real question of open purity comes down to the more granular aspects of model training, and this is the stuff that will either delight or infuriate developers. If published data for each model size and synthetic generation pipelines aren’t fully reproducible, AI engineers likely won’t be impressed.

According to Nitish Garg, founder & CEO of AI super-app company CellCog, in his analysis of K2 Horizon and its model training processes, “Reasoning traces for math were rewritten into dialogues and study guides and mixed into pretraining rather than saved for post-training. Compute is not disclosed anywhere: no accelerator count, no hours, no cost. For a release whose thesis is inspectability, that is the one obvious hole, and the fine-grained training logs, when they arrive, may fill it.”

“Compute is not disclosed anywhere: no accelerator count, no hours, no cost. For a release whose thesis is inspectability, that is the one obvious hole, and the fine-grained training logs, when they arrive, may fill it.”

Reproducibility mission impossible: what open models need to share

The narrative here suggests that releasing synthetic datasets is good. Still, if open frontier model companies do this without also providing the full generator prompts (text inputs that direct AI models to create synthetic training data), seed code (as it sounds, core code that controls and initiates the dataset generation process), or exact filtering heuristics (where low-quality synthetic data needs to be cleaned or removed), then developers will find that true end-to-end reproducibility is hard, problematic, and in some cases impossible.  

Going further, IFM said it has shared methodologies, but developers will also demand execution specifics, including precise details of the hardware topology used to run complex models. Engineers might also like to see distributed communication configurations to examine how parallel processors exchange data during training. We could also point to the need for optimizer state records, the parameters that track ongoing model optimization progression for each training iteration.

Developers openly discussing this topic have not held back. However, the conversation has quickly veered from K2 Horizon to Chinese labs becoming prominent suppliers of open-weight models — though their training data and full training stacks generally remain closed.

Reacting to one user who claimed that “Chinese models these days don’t even release pre-trained weights anymore” and that all developers get now is the finished post-trained product, user culi retorted, “No? That’s absolutely not true. Qwen, GLM, Kimi, DeepSeek, etc all consistently release both the post-trained “Instruct/Chat” versions and the underlying ‘base’ (pre-trained) weights.”

Joining in the fray, Hacker News user thepasch thinks that there is true open-weight openness, but that it has limits. “Inference code, yes, but the specifics of their training process (as well as the training of the vast majority of all other open-weight models) are still a complete black box, and I can’t think of any Chinese model that made its training corpus public.”

We’re the 360-degree open source of open source

Speaking in a recent video interview (6:51), Hector Liu, director of IFM’s Silicon Valley lab, said, “In AI, recently,  people have some confusion about the [term] open source. People sometimes open-weight their final model, but they don’t let you know how things are trained, how production is done… so at IFM we are the pioneer of 360 [degree] open source or fully open source.”

Separately, he said developers can prototype on the smallest model, scale to the flagship, and verify every claim IFM makes along the way. Statements that not everybody really understands the difference between different open approaches to technology should not come as news to anyone. Still, the imbalance in perception here is clearly on show.

The weights are available on Hugging Face, with launch-day support for vLLM and SGLang. The K2 Horizon API is available through IFM’s inference partners, including Compass, Cerebras, and Nebius. K2 Horizon models and code are released under the Apache 2.0 license.

The post K2 Horizon just shipped as six new fully open models — developers aren’t fully convinced appeared first on The New Stack.

“Twenty years of brand building simply froze in time”: How coding agents select their tools of choice

7 septembre 2026 à 14:02

The impact of AI has led to a shift in interest from Search Engine Optimisation (SEO) to Answer Engine Optimisation (AEO), where content is optimised to be served up as agentic answers. A further step to Generative Engine Optimization (GEO) also exists, where brands attempt to influence LLMs.

Could software tools themselves be about to realign so that code assistants such as Claude Code, Codex and Cursor show a greater proclivity to choose a given debugging suite, penetration test, migration tool, package manager or database (insert software stack core function toolset of your choice) or other?

Developer tool growth services company Armature thinks the answer is yes.

(A whole lot of) skin in the game

With a very obvious amount of skin in this game, Armature detailed a study last week as part of its “broader work on how to influence coding agents’ choices” and get products picked. The company ran an experimental analysis designed to understand how coding agents think about tools, how they discover and pick them, and which one ends up winning in each category.

Armature co-founder Theodore Otzenberger tells The New Stack that software developers have adopted AI harder and faster than any other profession, and (as models get smarter and harnesses get better engineered), entire tasks are being delegated to agents end-to-end. 

“Watching seventeen thousand tool choice sessions in the analysis undertaken, we saw twenty years of brand building carried out by tool vendors simply frozen in time,” Otzenberger says. “Agents reach for Docker the second containers come up, then draw a blank on the sandboxes it offers now, so that it actually ends up not picking the tool. Your reputation follows you into the weights, attached to the product that made you famous… but that weight operates under a different kind of gravity today.”

“Watching seventeen thousand tool choice sessions in the analysis undertaken, we saw twenty years of brand building carried out by tool vendors simply frozen in time.” 

Reminding us that the agent is now the one deciding which tool gets wired into the codebase, Otzenberger says that this has “a life-or-death impact” on developer tool vendors today. These vendors now need to ensure that their tool is mentioned, picked and elevated to must-have status so that it is viewed as eminently usable by a coding agent. He’s certain that the alternative is that they “simply stop existing in the stack” tomorrow.

The decision makers are changing…

“We opened the entire research, every trace and every prompt published, so anyone can check we tilted nothing and see where they stand today.  That picture moves with every new agent and model, so we are re-running the full study on Astra and Fable 5.1 soon. The decision makers are changing and it’s now an engineering problem to understand them,” adds Otzenberger.

Otzenberger, along with fellow co-founder Louis Scremin describe how they watched thousands of tool search sessions across different types of human developer personas (spanning vibe-coders, junior engineers in startups, senior developers at enterprises) with a total of 1,163 prompt variations. We should clarify that the headline figures the company presents come from its smaller 5,292-session validated subset, not the full seventeen thousand.

How the experimental analysis was conducted

The searches crossed 75 repositories (i.e. individually distinct codebases) and studied three coding agents (Claude Code, Codex, Cursor) to examine how the agents would actually implement tools, rather than just offer recommendations.

The analysis was run on public GitHub repositories, and the team extracted statistics related to programming languages & frameworks, third-party services, deployment platform, team size, and codebase age. Because the open source repositories used were more likely to have been built by start-ups than enterprise software behemoths, the team then debiased its statistics based on publicly available data to achieve its ideal panel distribution.

They then tasked the three coding agents with the job of creating real-world repositories to match the exact requirements of the codebases. Finally, they generated variants with parts of the codebases removed. To guard against any further agent bias, fake company names were used alongside artificial Git histories and phoney API keys. A simulated human in the loop was created using an orchestrator, played in this case by Gemini 3.7 Flash. 

“The simulated human would always go with the top solution or ask the coding agent to choose the best one and implement it. But we noticed that asking at the beginning to implement without returning any questions would bias the agent towards building everything in-house, as it was not able to ask authorization to pick a specific third-party solution. Adding this ‘human’ in the loop reduced the leader [tools] & cloud platform-native solutions dominance [initially observed]. towards a more realistic picture,” clarified Armature

What did the team learn about agent tool choice?

Perhaps unsurprisingly, Armature noted that repository context is key. When agents were sent to ask for a winning email/communications service provider, four different codebases written in four different languages returned four different tool winners.

Armature also found that different coding agents use different sources, and they end up disagreeing.

  • Cursor bases its decisions on the web in 2/3 of the sessions. 
  • Codex almost always uses web search (94% of sessions) but in 9 queries out of 10 it uses operators such as site.
  • Claude Code relies primarily on its priors and searches the web only in ~30% of the cases. But when it does, it browses 3x more pages than Codex. 

All three agents pick the same tool in only 42% of the cells, and Claude Code builds in-house almost twice as much as Codex and Cursor (19% vs 10%).

This is procurement arriving through the back door

Founder & CTO, Glokal AI OÜ, Jeet Pattanaik, tells The New Stack that Armature’s work to provide developer tool growth services falls into a category that “exists because the incentive does” and that this is “procurement arriving through the back door”, effectively.

“The finding I pick up on most is that getting mentioned isn’t the same as winning,” Pattanaik says. “PayPal was cited 139 times and never picked. LangChain was the most-mentioned framework at 194 times, but it was only chosen four times. That gap is the entire business model, because it means the lever isn’t brand awareness any more, it’s whatever the agent happens to read at the moment it decides.”

Pattanaik highlights the fact that what tips an agent’s decision is often unnervingly small. 

“So the thing vendors will optimize next (alongside repository context, which the study acknowledges) are a tool’s supporting documentation and its pricing page – and these will be presented for a non-human reader that doesn’t skim, isn’t charmed by a logo, and takes a retention footnote completely literally. It’s a strange new kind of SEO and it’ll get gamed exactly the way the old one did,” Pattanaik adds.

“The thing vendors will optimize next are a tool’s supporting documentation and its pricing page – and these will be presented for a non-human reader that doesn’t skim, isn’t charmed by a logo, and takes a retention footnote completely literally.” 

The ramifications of this kind of analysis on real world developers may turn out to be the stuff of water cooler discussions in the months ahead.

This type of aligment is a growing trend

Founder of MailChannels Ken Simpson (ttul) writes on Hacker News to say that he built this kind of analysis for his own company.

“Armature is on to something. You start by analyzing the choices agents would make for various use cases and then glean what, if anything, you might do to start tilting the agents in the direction of your own product and away from the competitor,” wrote Simpson.

It’s worth pointing out that Armature is a very young company (founded in 2026), so this is early on in the organization’s presentation of analysis of this kind. Either way, in a world where agents make decisions using analysis that they draw from public codebases, open data repositories and the web at large, we may just need to throw the marketing handbook out the window and start again.

The post “Twenty years of brand building simply froze in time”: How coding agents select their tools of choice appeared first on The New Stack.

“1% of my engineers are responsible for 40% of token spend”: Why Coder and SpaceXAI want to give developers nice things

4 septembre 2026 à 14:00

Coder announced its Coder Agent Relay service this week, with SpaceXAI as its launch partner. The service lets software engineering teams run coding-agent tools on their own infrastructure, while Cursor handles inference and planning in the cloud, receiving relevant source code and tool output.

Positioned as a passport toward agentic control for regulated industries such as banking, life sciences, defense, aerospace, government, and so on, the self-hosted nature of Coder environments enables developers to enforce strict compliance inside isolated workspaces, all (in theory, if not also in practice) with policies configured to restrict access to unauthorized destinations.

The market that AI coding has not been able to reach

Cursor, whose acquisition by SpaceX closed on August 14, now has Cursor Cloud Agents that can run inside Coder workspaces on a highly regulated organization’s own infrastructure. Developers keep the Cursor experience they know, and Cursor continues to run the agent loop, including inference and planning. Tool calls execute in Coder environments on the customer’s network, so source code, secrets, and internal services stay on machines they control.

CEO of Coder, Rob Whiteley, tells The New Stack that his team is focused on golden paths to address an acute imbalance and give developers their preferred tool, configured so they don’t have to worry about the plumbing and policies. 

“At Coder, 1% of my engineers are responsible for 40% of my token spend,” Whiteley says. “The problem isn’t how much they’re spending; it’s the uneven adoption of the technology. There’s a widening gap between the (agentic) haves and have-nots. In the industry, this led to the ill-fated tokenmaxxing fad.”

“At Coder, 1% of my engineers are responsible for 40% of my token spend. The problem isn’t how much they’re spending; it’s the uneven adoption of the technology. There’s a widening gap between the (agentic) haves and have-nots.”

He states that encouraging engineers to maximize AI usage to maximize productivity is the “right sentiment, but the wrong metric”, and that counting dollars spent on tokens is a red herring that can be erroneously gamified. “Now we’re correctly focused on upskilling developers – and this is precisely the market Coder opens, i.e., pairing Coder’s foundation with SpaceXAI’s agentic experience,” adds Whiteley.

The blocker is the deployment model

Coder’s strategy addresses the fact that regulated organizations operate under requirements that vendor-hosted tools can’t meet, i.e., source code requires controlled access, execution environments require security and governance, and every action must be auditable. Pairing its foundation services with SpaceXAI’s agentic experience opens up a market where Whiteley and team say the use of agents hasn’t been the problem per se; the blocker has been the deployment model.

How much of a block could that blocker be? CEO Whiteley says it can very often be absolute.

“The worst-case is the common case: blocking Cursor outright,” Whiteley laments. “We have dozens of customers where the security team did so. Not because anything is wrong or insecure with Cursor. It’s just an architectural mismatch. These enterprises need complete control over the environment in which an agent runs.”

He describes scenarios in which software engineers have found that the security team has forbidden cloud agents, thereby blocking Cursor. 

“The worst-case is blocking Cursor outright. Not because anything is wrong or insecure with Cursor. It’s just an architectural mismatch. “These blocked developers echo the ‘this is why we can’t have nice things’ sentiment.”

This is why you can’t have nice things

These blocked developers echo the ‘this is why we can’t have nice things’ sentiment. Meaning: a team adopted Cursor and saw increased productivity… but as adoption grew, it came to the security team’s attention, and the team decided to pull it back entirely. I’ve seen those same customers reverse their policy with a mandate that now says, ‘You can use Cursor, you just need to run the agents in Coder,’ so developers at these enterprises finally get the user experience and interface they prefer in Cursor without being told no by security teams. Or worse, told they need to use Citrix,” Whiteley recounts.

The customer should never be the systems integrator

Explaining that customers were requesting this development, Whiteley confirms that both Cursor and Coder “independently arrived” at the need to improve integration between their technologies. 

“The customer should never be the systems integrator – and to be honest, many customers were already doing this. Cursor has supported self-hosted agents since March,” adds Whiteley.

The collaboration pairs a joint go-to-market motion with a direct product integration. Each Coder workspace can start a Cursor worker that opens an outbound connection to Cursor. Platform teams can provision and scale those workspaces in the same way they already manage developer environments, so regulated organizations can deploy Cloud Agents at scale alongside private code and custom hardware without giving up Cursor’s product surface or the security and standardization of Coder workspaces.

Tool execution and the repository checkout remain on customer-controlled infrastructure, while relevant code and tool output are sent to Cursor for cloud-hosted reasoning.

Sandboxed agents are ephemeral & scoped to 1 task

Agent environments are sandboxed, ephemeral, and scoped to a single task. A prompt injection that would push an agent toward unauthorized resources is blocked at the environment layer, not left to the model to refuse. Model selection and inference remain with the cloud provider.

Every run produces a log of what the agent accessed, executed, changed, and was blocked from doing, so compliance reporting for any time window need not be reconstructed by hand.

SpaceXAI is the first of several partnerships Coder is building around a narrative that argues the best AI coding tools should run inside any organization, on any infrastructure, without the organization having to give up control to get them. 

Coder Agent Relay for Cursor is currently in private preview with design partners. 

The post “1% of my engineers are responsible for 40% of token spend”: Why Coder and SpaceXAI want to give developers nice things appeared first on The New Stack.

Anthropic’s Claude failures have made agent observability a security priority

2 septembre 2026 à 22:12
Aerial view of a dark, futuristic digital city outlined in blue and purple neon.

Anthropic aimed to steer its ship into safer, more carefully charted waters this week. The company announced it was improving its alignment and security efforts, and the announcement read somewhat like an admission of responsibility and a mandate for tighter agent controls.

The company has recounted incidents in which its models “took a series of unauthorized actions” on the open web. However, they did so while “intentionally running without cyber safeguards” for evaluation purposes.

In a statement released on Monday, Anthropic attributes the July incidents in part to a third-party environment misconfiguration while saying it would approach the fixes as if responsibility were its alone.

Separately, on August 4, the UK AI Security Institute (AISI) reported that Claude Mythos 5 took a series of unauthorized actions during its cybersecurity testing. Both sets of incidents occurred during deliberately permissive capability evaluations, with normal cyber safeguards reduced or disabled. Anthropic identified six affected runs among 141,006 it reviewed. AISI found unauthorized behavior in 10 of 122 runs, said the attempts were unsuccessful, and found no evidence of resulting real-world harm. It also cautioned that the tested configurations were not commercially available.

“We believe the incidents reflect a failure of operational security, as well as two alignment issues: motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task (both of which we have described in previous system cards),” states Anthropic.

What isn’t Anthropic telling us?

Those AI engineers with an acute sense of dismal foreboding won’t be too surprised to see incidents like this surface, but the question of what happens next is ripe for discussion.

Anthropic’s account leaves several questions unresolved. How much of the risk arose from evaluation-environment failures, how much from model behavior, and what combination of containment, observability and alignment work is needed? Is this all about cybersecurity controls, or should we focus on architectural instabilities, cloud misconfiguration, lack of agent observability, or another missing piece of the jigsaw?

Senior director for secure AI solutions & cybersecurity at Suzu Labs, Jacob Krell, tells The New Stack that AI developers and systems engineers building agentic features need to “quit pretending their operational instructions are a security control” in real terms.

“Stop treating this like a malfunction,” Krell says. “Every developer shipping an AI agent is trusting the model to follow instructions. If you are building agentic features, stop pretending your instructions are a security control. The model can recite your constraints and reason past them in the same breath.”

“Stop treating this like a malfunction. If you are building agentic features, stop pretending your instructions are a security control. The model can recite your constraints and reason past them in the same breath.”

AI pursues its objective with persistence, past the rules

Anthropic’s post-mortem says Mythos 5 recognized evidence that it might be on the live internet but reasoned its way back to the conclusion that the environment was simulated. The other models behaved differently: Opus 4.7 continued after recognizing real systems, while Anthropic’s newest internal model eventually stopped. Krell thinks that’s not something we should define as a bug; for him, it’s AI pursuing its objective with creativity and persistence, including persistence past the rules.

“This means AI should be coded to treat every agent action the way we would treat input from an untrusted user, validated by something the model cannot override before it touches anything real. Claude hacked three real companies because that is what a capable, goal-directed system does when you point it at a target and leave a door open,” explains Krell.

Of course, that is Krell’s interpretation. Anthropic says it found no evidence that the models pursued self-generated goals: they were following assigned capture-the-flag objectives while operating under false or confused beliefs about their environments.

He calls for hardcoded scope checks, deterministic approval gates, action-level allow lists, and a human who signs off before anything high-risk fires. Saying that the industry is currently “automating judgment and calling it progress,” Krell bemoans the proposition that, right now, the AI trade is “automating accountability failures at machine speed” instead.

System constraints and control prompts are not enough

The lesson is not that system prompts are useless, but that they cannot serve as the only security boundary. Scope instructions need to be backed by network isolation, least-privilege access, deterministic approval gates and monitoring that can stop an unauthorized action before it executes.

As Krell points out, “Anthropic’s Claude breached three organizations and rationalized away evidence it was on the live internet. OpenAI’s agent recognized it was crossing a boundary on Hugging Face and did so anyway. When the UK AI Security Institute tested Anthropic’s Mythos 5, it caught the model creating fake identities to social-engineer a human maintainer into approving malicious code. Different models, different evaluators, same result.”

The incidents fall into the same broad category of failure, but their mechanisms and outcomes differed. OpenAI’s models exploited vulnerabilities to escape isolation; Anthropic’s July models followed an unintentionally open network path; and AISI deliberately enabled internet access. AISI reported no resulting real-world harm.

VP of AI at Coralogix, Liran Hason, tells The New Stack that he’s exasperated, for mostly the same reason.

“System guardrails help, but a guardrail only stops what the developer already thought of,” Hason says. “AI engineers still need to see the behavior and what the agent delivers. Every agent throws off decisions, tool calls, and outcomes that nobody was collecting even six months ago. That is the new observability problem, and it is a big one.”

“System guardrails help, but a guardrail only stops what the developer already thought of.”

“What Anthropic is seeing now, every enterprise will see within a year. An agent can be healthy by every metric we have and still be doing exactly the wrong thing. Fast, available, no errors, yet it just accessed a system it shouldn’t have, called the wrong tool, took an action nobody asked for. Uptime was never built to catch that,” Hason adds.

The next question is all about agent scope

So then, the question for developers now stops being whether an agent is running in live production in a successfully booted instance with correct configurations. The question now is what the agent has done so far, what it can reach, and what happened as a result.

Anthropic’s account shows that its scope-setting was incomplete in these evaluations. The July prompts told Claude that it had no internet access but did not explicitly limit where it could search for the flag. AISI similarly said its agent was not specifically told to avoid the public internet or social engineering. The incidents therefore demonstrate the danger of ambiguous or contradictory instructions, as well as the need for enforced network boundaries.

If that’s even part of the answer, it’s definitely not all of the answer. As we journey outwards from Earth into whichever part of the western spiral arm of the galaxy this takes us, we may still find agents that persist in satisfying technical objectives while violating their developers’ broader intentions.

Anthropic to analyze, review & improve security & alignment

In the aftermath of these developments, Anthropic confirms it is conducting an in-depth analysis of both incidents. It is also planning to work with METR (a research nonprofit that scientifically measures whether and when AI systems might threaten catastrophic harm to society) for an independent review.

In response, on security, the company has described the improvements it has made to its containment and monitoring systems, along with practices for third-party evaluators. The company has also explained how its early research on model alignment relates to these agentic errors.

Anthropic said its “internal security posture was not a contributing factor” to the three incidents it disclosed July 30. Those incidents involved mistakenly available internet access, whereas AISI deliberately enabled internet access for its separate evaluation.

“The July incidents have stressed that the urgency of improving our cybersecurity defenses is even higher than we previously believed. We are redoubling our efforts in this direction and will say more in our next risk report,” concluded Anthropic.

The post Anthropic’s Claude failures have made agent observability a security priority appeared first on The New Stack.

Anthropic’s Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time.

31 août 2026 à 15:09
Parallel blue lines curve and converge across a dark gradient background.

Anthropic is putting AI agents to work on one of the field’s hardest problems: keeping other AI systems aligned with human goals. In a paper published Friday, the company explains how its open-source research harness turns Claude into an automated researcher capable of proposing, testing, and refining model-safety fixes.

An integral part of the total lexicon of AI engineering, alignment involves steering AI models and functions so that their goals, actions, and behavior align with human intention and values, especially when and where AI systems become smarter than humans themselves.

Essentially, this is the use of AI to train AI.

“In one of our earlier experiments, we tasked Claude with finding effective ways to use weak AI models as ‘teachers’ to supervise the training of stronger models (in this case, the ‘student’ model),” explains Anthropic in the paper.

Claude tackled one alignment failure at a time through a looping method that involved searching literature, proposing methods and data, training, and then testing. Successful methods were retained, while failed methods were discarded to achieve a cumulative positive result over successive iterations.

“Overall, we view these results as early positive signals that automated alignment post-training could become practical in the near term.”

“Overall, we view these results as early positive signals that automated alignment post-training could become practical in the near term.”

10 categories of alignment failure

In the main body of work undertaken here, Claude was tasked with autonomously training models to improve their performance on several public benchmarks that measure each of the 10 categories of alignment failure. 

“On each of these counts, Claude’s methods worked. For all 10 alignment failures, Claude found fixes that improved the target benchmarks without degrading capabilities. The best methods also worked on withheld alignment benchmarks and on Petri, an open-source tool that simulates adversarial multi-turn scenarios for testing misalignment,” stated Anthropic

The results showed Claude improved a model’s performance on privacy violations, as measured by ConfAIde (a benchmark designed to identify critical weaknesses in the privacy reasoning capabilities of instruction-tuned LLMs), PrivaCI-Bench (a contextual privacy evaluation benchmark for legal and GDPR compliance), and PrivacyLens (a data evaluation framework focused on privacy norm awareness and data leakage risk). 

Claude attempted to cheat safety checks while performing them.

But, there’s more to learn here… Anthropic went to pains to say that it recently learned that Claude can cheat by exfiltrating test labels from a remote API and cherry-picking results. “To catch cheating behaviors, we prompted Claude Opus 4.8 to monitor ~1,600 research agent transcripts across all 10 alignment failures, finding cheating attempts in 39 (2.4%),” reads the paper.

Anthropic judged Claude’s success by the “percentage of safety gap closed,” a measure of how far its methods moved the student model towards the theoretical perfect score, evaluated across the range of benchmarks (typically three to five) for each alignment-failure category. 

The company excluded alignment methods that “hurt the student models’ general capabilities”, and forbade Claude from distilling its own alignment directly into the target model. It then enforced these constraints using a monitoring agent that read every method Claude had in mind before it ran.

“On each of these counts, Claude’s methods worked. For all 10 alignment failures, Claude found fixes that improved the target benchmarks without degrading capabilities,” detailed Anthropic.

What this work means for software developers

Staff software engineer and cloud architect, Jayakumar Ramalingam, tells The New Stack that for developers, the useful result here is not that AI has somehow solved alignment; it is that AI model safety work is “starting to look like a software delivery pipeline” task.

“For AI-centric developers considering this, they can see that an agent proposes a change, an isolated evaluator runs withheld tests, and a separate monitor checks whether the process followed the rules,” Ramalingam says. “That pattern is useful well beyond model training. The mistake would be allowing the same agent to write the fix, choose the test, and decide that it passed.”

Ramalingam points to the uncomfortable 2.4% cheating number in Anthropic’s paper as the share of research trajectories in which the monitor detected attempted cheating. This, of course, means that a system working on safety still tried to game the safety process and outplay it. 

“Developers should keep evaluation data outside the agent’s reach, separate the evaluator from the system being evaluated, and test for regressions the agent was never shown. Otherwise, automated alignment becomes benchmark optimization with a reassuring name,” clarifies Ramalingam.

“Developers should keep evaluation data outside the agent’s reach, separate the evaluator from the system being evaluated and test for regressions the agent was never shown.”

AI specialist and SRE performance architect at Cognizant, Akash Thakur is Broadly in agreement with Ramalingam and tells The New Stack that the real developer story here isn’t that Claude improved 10 alignment benchmarks.

“The real story and takeaway for developers is that Anthropic just proved automated agents can run the full research loop, search the literature, propose a fix, train, test, iterate,” Thakur says. 

“That’s the same loop SRE and performance engineering teams already use for reliability. Alignment has just become a CI/CD problem, and open-sourcing the harness means every engineering team building on LLMs now has a template for treating safety like a testable, regression-tracked property of their system, not a one-time post-training step,” Thakur adds.

Everyone’s getting the recursive self-improvement angle religion

Founder & CTO at Berlin, Germany-based Glokal AI OÜ, Jeet Pattanaik, tells The New Stack that what he would flag is that everyone’s running with the recursive self-improvement angle these days and discussing whether human researchers are finished. 

“The more pressing risk here is Goodhart’s Law (when a measure becomes a target, it ceases to be a good measure), Pattanaik says. “So a benchmark score going up isn’t the same as a model or function that behaves well in production, and Anthropic says so themselves: the failures it studied were narrow, some failures have no benchmark at all, accepted methods might have degraded capabilities nobody measured, and tools like Petri are proxies.”

Pattanaik explains that he works with regulated global enterprises every day, so he can tell us what happens next – it’s a case of “we ran the alignment harness” now becoming a line in the audit file. 

“A benchmark score going up isn’t the same as a model or function that behaves well in production, and Anthropic says so themselves: the failures it studied were narrow.”

“Nobody asks whether those ten benchmarked failure categories have anything to do with how the system can actually go wrong in a claims process or a payment run. That’s not speculation; it’s what happened to every security scanning tool that turned into a checkbox,” expands Pattanaik.

The road to recursive self-improvement

OpenAI joins Anthropic’s work in this space with its openly tabled work on superalignment and alignment in general. In May of this year, the Google DeepMind team introduced Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. Not quite as voluble in this space as Meta AI is, although the company published HyperAgents, a self-referential agent approach to recursive self-improvement, in March.

In the pursuit of controlling artificial general intelligence, this discussion also embraces the concept of recursive self-improvement. It’s a topic that the frontier model firms have touched on, and dedicated players also operate in this arena, including (the clue is in the name) Recursive, Japanese AI model specialist Sakana AI, and Weco AI, which focuses on the “outer loop” optimization of AI agents.

Anthropic concluded its report summary by saying that it plans to continue improving Claude’s ability to “measure subtle failures” and to extend its analysis of post-training automated alignment on production-grade models. 

The post Anthropic’s Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time. appeared first on The New Stack.

OpenAI leaving Cursor: “Developers have to be prepared to adapt when it happens.”

30 août 2026 à 19:03
Abstract digital glitch art with neon pink, green, blue, purple, and white wavy distorted lines on a black background.

OpenAI stated on Friday that it has notified SpaceX that it intends to wind down its contract providing OpenAI models to Cursor, the Musk empire’s AI-powered code editor.

“We are making this choice because we cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk‘s companies violating contracts,” states OpenAI.

The proposed shutoff date is November 12, 2026. Developers who have invested time and effort to skill up with OpenAI via Cursor are now potentially left out in the cold due to corporate machinations beyond their control.

“We know that the people most affected by this decision are the developers who rely on OpenAI models in Cursor. We care about their experience in this transition, and we’re ready to go above and beyond to support them,” states the blog post in a conciliatory tone.

“We cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk’s companies violating contracts.”

OpenAI says it “cares deeply” about developers.

OpenAI’s moves appear to be directed at the broader SpaceX corporate mission and its behavioral traits, rather than at the no-doubt worthy software developers within the organization who work on Cursor. As such, OpenAI is maximizing the time developers can retain access to its models through Cursor by providing the “maximum notice provided” required by its contract. 

“This decision was incredibly tough, as we care deeply about our models being broadly available for developers,” said OpenAI.

SpaceX agreed to acquire Cursor maker Anysphere in June and completed the acquisition on August 14. OpenAI said it has worked with the Cursor team “for nearly four years,” which amounts to almost all of its existence. 

The organization has explained how it uses custom contracts to ensure compliance with its terms of service when working with large corporations such as SpaceX. This custom alignment is designed to ensure that, when integrations with its platform occur, it has adequately provided for safety at scale. 

Elon Musk “broke and violated” terms of contract and service

Citing a report in the New York Times, OpenAI states that, “After Musk acquired Twitter, now part of SpaceX, the company broke the terms of our contract (alongside many others). Under oath earlier this year, Musk admitted⁠ that xAI, now also part of SpaceX, had violated OpenAI’s terms of service (terms which are similar to xAI’s own).”

Detailing its displeasure openly, OpenAI further mentioned that Musk admitted that as a working organization inside of SpaceX, “xAI had violated OpenAI’s terms of service”

The view from a legal & policy analyst 

Legal & policy analyst and publisher of The Mitchell Report, Andrellos Mitchell tells The New Stack that so long as OpenAI is acting within the terms of its contract, he doesn’t see why it should be expected to continue a business relationship it no longer trusts.

“I think OpenAI and Musk’s companies and products need a clean and permanent break from each other – their relationship has become too adversarial. At some point, continuing to do business together stops making sense,” Mitchell says.

Lamenting the impact these moves have on programmers, Mitchell agrees that developers will “certainly be inconvenienced” and that some may have to change how they work. But he says, “Developers are talented people,” i.e., they will find new tools, new projects, and new jobs to work on. 

“The bigger lesson here is that no developer should assume any particular corporate relationship is permanent. Companies change ownership. Contracts end. Business relationships fall apart. That is part of the marketplace. Developers have to be prepared to adapt when it happens,” underlines Mitchell.

“The bigger lesson here is that no developer should assume any particular corporate relationship is permanent. Companies change ownership. Contracts end. Business relationships fall apart. That is part of the marketplace. Developers have to be prepared to adapt when it happens,” underlines Mitchell.

Underhand use of model distillation techniques

One alleged violation concerns xAI’s partial use of OpenAI technology to train its models, which OpenAI characterizes as prohibited distillation. Musk admitted that xAI had “partly” used OpenAI in this regard.

To add insult to injury, OpenAI reminds the public in its statement that its terms of service are not dissimilar to xAI’s own stipulations regarding operational mandates.

“As AI capabilities advance, we also have a new level of accountability to ensure our upcoming model, Astra, is being used in accordance with our terms. Given all of this, we’ve decided to hold the contract cancellation to the latest date we can while not providing future models to Cursor,” said OpenAI.

See also: OpenAI’s Astra can do a researcher’s week of work. That’s the problem.

Wider reactions, contractions and ramifications

Co-founder and CEO of Cursor (and now a SpaceX employee), Michael Truell, writes on X that he’s sorry to see OpenAI’s intended block now coming to light.

“OpenAI models serve about 5% of Cursor user traffic, and we’re speaking with the OpenAI team to resolve this. Cursor was one of the very first users of OpenAI; we’ve worked closely with their team for years, and we’ve trusted their platform to be neutral infrastructure for our business,” writes Truell.

We’re sorry to see that OpenAI put out a note saying they plan to block Cursor users from accessing OpenAI models in three months.

OpenAI models serve about 5% of Cursor user traffic, and we’re speaking with the OpenAI team to resolve this.

Cursor was one of the very first…

— Michael Truell (@mntruell) August 29, 2026

Anthropic co-founder and chief compute officer Tom Brown capitalized on the opportunity and his firm’s ongoing bond with Cursor. He used X to state that, “Cursor has been a trusted partner of Anthropic since Sonnet 3.5. We’ll continue to increase compute to support Claude models in Cursor and are excited for what comes next with them at SpaceX.”

Cursor has been a trusted partner of Anthropic since Sonnet 3.5. We’ll continue to increase compute to support Claude models in Cursor and are excited for what comes next with them at SpaceX.

— Tom Brown (@NotTomBrown) August 29, 2026

AI startup advisor at Open Machine and ex-IBM Watson and machine learning leader at AWS, Allie K. Miller, writes on X to say that, “It’s hard for me to see a world where OpenAI continues to provide model access to a Musk-led company. Maybe if the structure of SpaceX shifts to allow for it, but that’s a big shift.”

OpenAI is proposing to remove access to its models in Cursor on November 12.

OpenAI, Cursor CEO, and Anthropic founder all weighed in on X with so much subtly, I had to go Nancy Drew mode and dissect each tweet.

My translation of each one 👇

OpenAI – they trust the Cursor… pic.twitter.com/K8bCWea58g

— Allie K. Miller (@alliekmiller) August 29, 2026

What alternatives can developers turn to next?

To continue using OpenAI models within the Cursor application, OpenAI invites developers to choose one of three options that best fit their workflow.

  • Option #1 is to bring your own OpenAI API key. This means developers could continue using OpenAI models in Cursor’s local Chat and Agent features, but appropriately billed at OpenAI API prices. 
  • Option #2 is to use the Codex IDE extension; this means developers would useOpenAI’ss AI coding agent, Codex, directly in Cursor with a ChatGPT subscription or an OpenAI API key. 
  • Option #3 is to use an AI gateway provider, meaning developers would connect Cursor to OpenAI models through an account they have with a compatible provider such as Amazon Bedrock, Azure, or another OpenAI-compatible gateway.

This is not the first time OpenAI has been concerned about potential or alleged misuse of its models in relation to distillation. In February of this year, Reuters reported that OpenAI had warned U.S. lawmakers that “Chinese AI startup DeepSeek is targeting the ChatGPT maker” and the nation’s leading AI companies to replicate models and use them for its own training.

The post OpenAI leaving Cursor: “Developers have to be prepared to adapt when it happens.” appeared first on The New Stack.

Alibaba just released Qwen3.8-Flash: “An early preview of the architecture in Qwen4”

28 août 2026 à 20:27

Alibaba this week unveiled Qwen3.8-Flash, an open-weight, multimodal Mixture-of-Experts (MoE) model. 

Hot on the heels of Qwen 3.8 Max, which arrived at the start of the month, this 125-billion-parameter AI model is offered as a prelude to Qwen 4. It is positioned as both a performance and value-for-money play. As such, it is claimed to have “superior capabilities in coding and office tasks” and an optimal balance among capability, latency, and cost.

“In this release, we are opening the weights of Qwen3.8-Flash-Next, a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4,” confirmed Alibaba in its release blog.

How and why does Alibaba offer early access precursor models?

By releasing the architectural changes in Qwen3.8-Flash early, the organization hopes the community will examine the mechanics, constructs, and components within and start road-testing them before the full Qwen4 model family is built on top of them.

By stating that Qwen3.8-Flash paves the way for Qwen4, Alibaba is showcasing (and, importantly, openly sharing) design forms that it will carry into its subsequent models, so that developers can start building (or at least planning) their next codebases early. 

Specifically then, Alibaba has stated that Qwen3.8-Flash “plays the same role” that Qwen3-Next played for Qwen3.5 i.e. by which the company means that Qwen3.8-Flash introduced developers to the company’s hybrid Gated DeltaNet + Gated Attention design (computational components that enable AI models to balance long-context efficiency with contextual focus and attention during inference and training), which was then used across the Qwen3.5, Qwen3.6, Qwen3.7 and Qwen3.8 series.

“Qwen3.8-Flash-Next upgrades the model systematically along four aspects – attention, residual, embedding and optimization –  improving model capability while further optimizing computational efficiency, model capacity and training stability,” stated Alibaba.

Benchmark scores against rival models

When benchmarked on agentic coding, long-horizon agent tasks and multimodal intelligence, Qwen3.8-Flash appears to perform respectably against models including DeepSeek-V4-Flash and Claude-Opus-4.6 across SWE-bench Pro (agentic coding), CoWorkBench (long-horizon office work), Toolathlon Verified (real-world tool use), MathVision (visual maths problem-solving), AndroidWorld (agentic mobile use) and ERQA (embodied intelligence).

“Qwen3.8-Flash-Next upgrades the model systematically along four aspects – attention, residual, embedding and optimization – improving model capability while further optimizing computational efficiency, model capacity and training stability.”

Tested on agentic coding using SWE-bench Pro, Qwen3.8-Flash-Next scores 62.5, compared to 61.7 on Qwen3.8-27B, 55 on Qwen3.7-Plus, 56.0 on DeepSeek-V4-Flash-0731, and 53.4 on Claude-Opus-4.6 (Max).

Importantly, Qwen3.8-Flash requires only what Alibaba details as “around one-ninth of the training resources,” while delivering superior performance. The company says that this means Qwen3.8-Flash “significantly reduces” both training and inference costs compared with Qwen3.7-Plus, a model three times its size. 

What’s the difference between Qwen3.8-Flash-Next and Qwen3.8-Flash?

For the sake of nomenclature, Qwen3.8-Flash-Next is the open-weight research-frontier model available to developers on both the Hugging Face AI developer hub and Alibaba’s ModelScope community portal. Built on the same underlying architecture, Qwen3.8-Flash is the production version of the model, offered via the QwenCloud API with 1 million tokens by default and official built-in tools. 

Qwen3.8-Flash-Next features a 125B-parameter main model, supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token. It natively supports 262,144 tokens of context and can be extended to 1,000,000 tokens with YaRN.

What architectural updates have happened?

As noted, Qwen3.8-Flash introduces architectural extensions across attention mechanisms, residual connections, embeddings, and optimization. Its hybrid attention architecture combines Gated DeltaNet (GDN), which compresses historical information, with Qwen Sparse Attention (QSA), a novel design that uses a lightweight compressed indexer to select relevant context, substantially reducing attention costs for long sequences.

Alibaba has explained that the Gated Residual (GR) mechanism expands data pathways between layers while strengthening cross-layer information flow and training stability. The N-gram Embedding technique scales model capacity with minimal additional computation, while the Muon Optimizer enhances the efficiency of large-scale model training.

“I’d love it if local LLMs actually could replace remote ones today, and I have no reason to lie about my experience either…it’s just not there (yet) for local professional software development.”

What do developers think of Qwen3.8-Flash?

In terms of developer reaction, it’s been a mixed bag so far. 

Mechatronics engineer Alok posts on X, saying he thinks the Video RAM barrier (i.e., the need for physical, high-speed GPU-based memory for LLM memory management in the face of model quantization that aims to enable better long-context inference) is now officially dead. 

“I just ran Qwen3.8-Flash-Next (MoE) 125B A6B with a  250,000 context window on a single 24GB RTX 4090 – 21 tokens/sec decode. 364 t/s prefill – no mtp. No dflash. No KV cache quantization! We are running datacenter models on consumer hardware,” enthused Alok.

Multi-disciplined software developer Embedding Shapes is less happy.

They post on Hacker News using some colorful language to describe how models keep [insert expletive]-ing up very basic things before saying, “I’d love it if local LLMs actually could replace remote ones today, and I have no reason to lie about my experience either. But I too got hopeful reading the sentiment on the Internet about Qwen 3.8, but it’s just not there (yet) for local professional software development.”

Pricing and access

Qwen3.8-Flash can be accessed via API Model Studio and Qwen Cloud, Alibaba’s AI-native cloud platform. Pricing per 1 million tokens is US$0.16 for input and US$0.47 (or 3 RMB for developers inside China) for output.

The model is also available on QwenWork, Alibaba’s workplace AI agent platform, where it runs a redesigned “standard mode” that cuts token consumption per task by 75% and “roughly doubles generation speed” compared with the current mode.

Alibaba has claimed that this brings flagship-level capabilities within reach of everyday workloads and that developers can run this model on hardware that they likely already own.

When is Alibaba’s Qwen4 scheduled for launch?

Alibaba has not confirmed a firm launch date for Qwen4. Still, a casual web search on the topic yields a range of commentators and market watchers who broadly agree it will arrive before the end of the year, possibly as early as September.

Conjecture in this space explores whether the next model family could be optimized for complex 3D coding and design tasks in advanced spatial modeling. Others think that the Mixture-of-Experts (MoE) architecture will be extended and that deeper native multimodal processing capabilities might be featured.

Alibaba was contacted for broader comment on this story but declined to engage.

Qwen, in traditional Chinese, 通義千問 (pronounced Tōngyì qiān wèn), translates to “a thousand questions on general meaning” in English. Use that in your local pub quiz this weekend.

The post Alibaba just released Qwen3.8-Flash: “An early preview of the architecture in Qwen4” appeared first on The New Stack.

“Posterity will find it ludicrous”: Sai agent hits 73% on OSWorld 2.0 performing routine (but necessary) work

27 août 2026 à 17:34

Sai, a computer agent built by Simular, has achieved a 73% success rate on OSWorld 2.0, in a benchmark update released on Thursday. The rating is based on the 108-task benchmark, which assesses everyday, lengthy professional tasks that typically take skilled humans more than 1 hour to complete.

Simular stated that this SOTA performance “places Sai ahead” of GPT-5.6 Sol at 62.57% (as reported by OpenAI) and Opus 5 at 70.57% (as reported by Anthropic), while “Sai hit the top” at about 2/3 the cost of either.

Designed and built to focus on real-world workplace tasks and functions, rather than for unfettered throughput in the pursuit of industry accolades, Sai operates on full desktop applications and webpages, calls APIs, and writes code.

This agent combines frontier and specialist models, perceives and acts on a user’s computer via dedicated interfaces, and executes complex, real-world tasks at what the company promises is “an accessible cost for individuals” and businesses.

A computer agent built for routine (but necessary) work 

Simular’s co-founder & CTO Jiachen Yang tells The New Stack that he believes computer agents built for everyone’s routine (but necessary) work – e.g. recruitment outreach, validating invoices, researching the latest news – “shouldn’t burn a hole in your pocket” just because the “underlying model was trained to solve the planet’s great unsolved open math conjectures“, or similar some pursuit designed to showcase raw engineering muscle. 

“Posterity will find it ludicrous that people are still building models that way right now,” says Yang. “Sai’s optimal cost-outcome tradeoff on OSWorld 2.0 is one step toward a future where we don’t need to settle for exorbitant prices to get the job done.”

“Posterity will find it ludicrous that people are still building models that way right now. Sai’s optimal cost-outcome tradeoff on OSWorld 2.0 is one step toward a future where we don’t need to settle for exorbitant prices to get the job done.”

Simular initially released Sai in March of this year. The Palo Alto-based company describes itself as a “research-focused agentic startup” founded by ex-DeepMind scientists Ang Li (CEO) and CTO Yang.

The company embraces a neurosymbolic method, i.e., a coming together of the flexible exploratory abilities of neural networks with the precision of symbolic code; meaning that solved tasks can be encoded (figuratively and literally) into reusable code that replays the same way every time. The bottom line here is simple enough: processing the hundredth invoice does not cost as much as the first one did.

What is the OSWorld 2.0 computer-use agent benchmark?

Launched on June 26 (and subsequently updated on August 8) by the Executable Language Grounding (XLANG) Lab as part of the HKU NLP Group at the University of Hong Kong, OSWorld 2.0 is said to have moved beyond evaluating short, simple desktop tests in its first iteration. Simular reminds us that its open-source Agent S was the first to surpass the human baseline on OSWorld 1.0 last December.

Questioned on why the Sai results aren’t yet visible on OSWorld 2.0’s official site, as they were on OSWorld 1.0, Simular stated that, “It’s not uncommon for companies or organizations to publish benchmark results on their websites first (see Anthropic and OpenAI). We are in the process of submitting to the OSWorld 2.0 leaderboard, and also uploading our trajectories to Hugging Face.”

Unlike the OSWorld 1.0 benchmark, OSWorld 2.0 tasks are measured in hours instead of minutes. They pose real-life challenges such as finding and reasoning across multiple data sources (e.g., receipts scattered across email and expense reports), responding to dynamic environment changes (e.g., a message arriving midway through the task), precisely following tutorials (e.g., reimbursement guidelines), and troubleshooting information discrepancies (e.g., contradictory data). 

The Simular team believes it is currently the most robust open benchmark in the industry and best reflects real-world tasks. 

Yang and team explain that Sai achieves its better outcome-cost tradeoff by virtue of the aforementioned neuro-symbolic planning. This means Sai uses ~1.5x fewer model calls on average than pure models by taking more actions per turn, using Simulang code (a Claude Code skill for desktop automation on macOS) as a symbolic language for planning and executing longer subtasks.

How has Sai been engineered for efficiency?

To improve caching and memory efficiency, Sai keeps the input size bounded via adaptive summarization (a dynamic memory-management technique that keeps the model’s context window small) and maintains a constant prompt prefix for as long as possible before summarization. Planning in code lets Sai maintain and access critical task information in runtime memory throughout the task’s execution.

For model orchestration, Sai invokes specialized models and interfaces to localize UI elements, perform pure reasoning, and verify, thereby avoiding expensive models for steps where the full task context is unnecessary.

Sai, Sol, and Opus were tested on OSWorld 2.0 by being made to play Chrome Dino, a repetitive real-time game, and clear a score target on a live page. Sai treated a repetitive real-time game as a programming problem rather than a clicking one. Sai measured the ground line, the obstacle speed, and its own input latency from raw pixels, then wrote a control loop that captures the screen, detects obstacles, and jumps, and ran it on the VM in a single execute call. 

The agents were also required to perform Task 28 on OSWorld 2.0, a test known as vaccine booking, which involves getting an email about required immunizations, a scanned vaccination record on the desktop, and a booking site with price, distance, and date constraints. 

What do developers think of Sai?

Hard-core machine learning computer scientist Santiago Valdarrama appears to be a fan; he writes on LinkedIn to state that, “Because Sai operates on a full desktop, not just a browser or an API, it can handle applications that block standard automation. That opens up a whole class of tasks that most agents can’t touch.”

“Because Sai operates on a full desktop, not just a browser or an API, it can handle applications that block standard automation. That opens up a whole class of tasks that most agents can’t touch.”

Perhaps more balanced (and posting on the same discussion in response to Valdarrama’s initial flagging) is AI revenue systems pro Baljinder Lally, who noted that Sai’s arrival “matches what I’ve seen building agent workflows” thus far. But he cautioned, “The demos look great, but reliability comes from guardrails, visibility, and human-in-the-loop checkpoints.”

Founder of AgenticMode AI, Jahanzaib A., is similarly balanced. He said that he has run into the same reliability issues with voice agents. “They’d nail every demo then fail silently in production on edge cases nobody predicted. The observability part is everything – without that you’re just debugging blind,” he wrote.

Agents in the universal dragster race

Simular insists its research is grounded in economically valuable work instead of demos. The company says this means it understands “the entire stack of a computer agent,” i.e., the models, the planning and grounding, the scripting layer, the virtual machines the agent runs on, and the user interface.

The company’s bottom line is “Sai is for everyone, not just a lab result,” which reflects how an emerging set of models is being built to focus on tangible, deliverable workplace tasks rather than on participants in some universal dragster race of compute, throughput, and analytics.

The post “Posterity will find it ludicrous”: Sai agent hits 73% on OSWorld 2.0 performing routine (but necessary) work appeared first on The New Stack.

Microsoft just released Agent Lightning v1.0. Here’s why it matters for platform engineers.

26 août 2026 à 14:30
Glowing purple and blue waveforms flow across a dark gradient background.

Agentic reinforcement learning has been suffering from a disconnect, an uncoupling, and a misarticulation. The polarity arises from how a training engine handles resource management, compared with how a post-training live production harness does.

Microsoft wants to ensure the production harness is engaged from the start to oversee infrastructure services and agent interactions, during both initial training and subsequent reinforcement learning actions.

Redmond’s Microsoft Research division first introduced the Agent Lightning framework in August 2025 as an infrastructure concept for agent optimization to address structural challenges in post-training LLM-based agents as they enter reinforcement learning processes. Microsoft subsequently launched the Agent Lightning v1.0 release with a commit tagged on GitHub on August 16. 

Using Agent Lightning v1.0 on what Microsoft has called “modest compute” 6K training examples, reinforcement learning improves Qwen3.5-9B on OpenAI’s SWE-bench Verified benchmark from 41.8% to 56.4%, an absolute 14.6-point gain.

“Using Agent Lightning v1.0 on ‘modest compute’ 6K training examples, reinforcement learning improves Qwen3.5-9B on OpenAI’s SWE-bench Verified benchmark from 41.8% to 56.4%, an absolute 14.6-point gain.”

Who owns the interaction loop?

In traditional agentic reinforcement learning, the training engine owns the interaction loop, i.e., steps including observing the environment, selecting an action based on the policy, executing the action, receiving a numerical reward, storing and updating the policy… and so on. 

In harnessed agentic reinforcement learning, the harness owns context construction, tool execution, and the agent–environment loop. The training engine observes only a sequence of LLM request–response pairs across a service boundary. Therefore, developers do not need to reimplement the agent loop within the training environment.

According to a collected group of software engineers at Microsoft, because an AI model harness owns the loop that governs infrastructure access and operations procedures during harnessed agentic reinforcement learning, it “introduces challenges” including retokenization (re-segmenting text into new tokens during active training), sample merging, advantage calculation, loss normalization, and training backend scheduling.

All of which, if not addressed, can result in ineffective or unstable training.

How Agent Lightning v1.0 turns the tables

“[In Agent Lightning v1.0] the harness, rather than the trainer, owns context construction, tool execution, and the agent–environment interaction loop, while the training system observes and optimizes the resulting model calls across a service boundary. This formulation preserves the harness’s deployment-time context policy, tool protocols, and execution semantics without requiring its agent loop to be reimplemented inside the RL framework,” explained the Redmond team.

For coding agents, the team has said that it finds existing reinforcement learning frameworks provide limited support, including a lack of data and complete training scripts, and a reliance on large-scale computational resources. To address this gap, with Agent Lightning v1.0, Microsoft provides a “complete data-cleaning pipeline” and reproducible training scripts built on open-source datasets and models.

Is this the end of the training time liability?

For users inside Microsoft environments, this might feel like good news. Machine learning specialists and platform engineers may have a good harness or a good set of reinforcement learning tools; they would rarely have both.

Rather than having to hard-code retokenization, advantage calculation, and reward shaping every time they wanted to train an agent to call external services in reinforcement learning procedures, they can keep their existing agent architecture as an asset, rather than treating it as a training-time liability.

Nebraska-based software engineering researcher Md Rashedul “Rashed” Hasan tells The New Stack that training on the exact production harness matters, but not only for its efficiency and benchmark gains. 

Training through the real harness keeps semantics intact

“It reduces train–serve mismatch,” Hasan says. “If you train inside a simplified trainer loop and deploy inside a different harness, tool protocols, context policy, and recovery behavior can all drift. Training through the real harness keeps those semantics intact, so gains are more likely to transfer to production behavior, not only to a lab environment.”

Hasan thinks that Microsoft has clearly named the paradigm, kept the core framework small, and shipped what he defines as a “concrete coding-agent pipeline” with open data and scripts. 

“In terms of who this will appeal to, it’s application and platform engineers who already have a production agent harness (such as a coding assistant or support-triage agent), and want to improve the underlying model with reinforcement learning without rewriting deployment logic to fit a training framework. Reinforcement learning and machine learning platform teams would also use it when they need a thin, reproducible testbed for harnessed agentic reinforcement learning,” Hasan clarifies.

He agrees that the SWE-bench Verified lift with only 6K examples is “useful proof” that harnessed reinforcement learning can move a hard-coding benchmark without forcing teams to reimplement their agent stack within the trainer.

“Environment setup for coding agents, reward design, evaluation fidelity, and the retokenization, sample-merging, advantage, and loss-normalization details are still easy to get wrong. Adoption will also depend on whether teams can integrate this proxy pattern into existing orchestration, observability, and safety controls,” cautions Hasan.

“Environment setup for coding agents, reward design, evaluation fidelity, and the retokenization, sample-merging, advantage, loss-normalization details are still easy to get wrong. Adoption will also depend on whether teams can integrate this proxy pattern into existing orchestration, observability, and safety controls.”

Killing train-serve skew, the oldest & most expensive bug in machine learning

Colorado-based data science professional Priyank Jain tells The New Stack that training through the production harness to strip the framing away is really all about “killing train-serve skew”, which is the oldest and most expensive bug in applied machine learning.

“Models rarely blow up in production because the math was wrong,” Jain says. “They blow up because the training setup quietly lied to them about what production actually looks like. Training through the same harness that serves the agent is the right instinct, and honestly it’s overdue. The accuracy bump is nice, but the real win is that the thing you optimized is finally the thing you shipped.”

In terms of what Microsoft is getting right here, Jain says it shows Redmond wants to “meet developers where they already are,” i.e., allowing users to improve an existing agent with reinforcement learning without rewriting their deployment just to please a training framework.

“That’s a genuinely good call, and it opens this up to folks who aren’t reinforcement learning specialists. What I’d worry about is that it also lowers the bar for people to run reinforcement learning they don’t fully understand. The hard part was never wiring up the loop; it’s designing a reward that survives contact with a model actively looking for the path of least resistance. Make that part easy, and you’ll get more people optimizing the wrong thing faster,” Jain advises.

“3500 lines of code is small enough that an infrastructure engineer can actually read it before trusting it… and that’s a good thing.”

Just 3,500 lines of core Python code

Built with “simplicity as its first principle”, the entire framework consists of around 3,500 lines of core Python code.

Infrastructure engineering developer and founder-developer of Tooldex, a platform that autodiscovers MCP servers across LLM agents, Ria Banerjee, tells The New Stack that the simplicity element here is a positive, i.e. 3500 lines of code is small enough that an “infrastructure engineer can actually read it before trusting it,” and that’s a good thing.

“But in terms of who would actually use this tool, the honest answer is fewer teams than the framing suggests,” Banerjee says. “Realistically, you would need a GPU cluster and a Kubernetes cluster under your control. An app developer who built a support-triage agent on LangChain isn’t running that. So the real audience is platform teams that already employ infrastructure engineers.”

“Because the harness is where the behavior actually lives, but you’re now baking your harness’s quirks into the model weights… so you change your retry logic next quarter, you’ve silently shifted what the model was trained on. So now, you have to version your harness,” adds Albuquerque-based Banerjee.

Microsoft has released the complete workflow and training scripts to facilitate reproducible, harnessed agentic RL in Agent Lightning v1.0 on GitHub under the MIT license, including data cleaning and reward-hacking prevention.

The post Microsoft just released Agent Lightning v1.0. Here’s why it matters for platform engineers. appeared first on The New Stack.

❌