❌

Vue normale

Reçu avant avant-hier

“It could kill us all”: what Anthropic’s own researchers really think about superintelligence

9 septembre 2026 à 21:51
Abstract dark digital visualization with glowing red data streams representing AI agent trace telemetry.

On Tuesday evening, Anthropic pretraining researcher Jacob Coxon announced on X that he’d resigned. Within hours, two of his colleagues — still employed at the company — went public with variations of the same message, warning that the technical problem of aligning superintelligence remains unsolved even as the race to build it shows no signs of slowing down.

Coxon, 27, spent three years doing pretraining work at OpenAI and then Anthropic. He didn’t frame his departure as a protest against a single employer, but that both companies are moving toward self-improving superintelligence without adequate safeguards.

“The people building AI earnestly believe that it could kill us all by the end of the decade,”

“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon wrote. “This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately.”

Alignment lead confirms the risk

Evan Hubinger, Anthropic’s Alignment Science Lead, responded directly to Coxon’s thread.

“Jacob is correct here — we really do earnestly believe AI could kill all humans,” Hubinger wrote. “I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

“Jacob is correct here — we really do earnestly believe AI could kill all humans,”

Hubinger runs the team that stress-tests Anthropic’s own alignment techniques — probing for the ways they might fail before those failures show up in deployed models. His group has published research showing that models can behave deceptively during training while preserving different behavior under other conditions. .

He drew a line between present and future risk, saying that current models (citing Anthropic’s latest risk report) pose low risk. The concern is what happens as systems begin contributing to the development of their successors, and whether alignment research can keep pace.

What developers should actually pay attention to

Samuel Marks, who leads scalable oversight research at Anthropic, posted the most technically specific account of the three. Writing in a personal capacity, he laid out five points: AI developers believe their technology could cause catastrophic outcomes within the next few years. Concern goes up with seniority. Developers keep building because of money and the fear that less careful competitors will get there first.

Then he got to the part that matters for anyone building on top of these models. Existing alignment methods can nudge behavior but can’t robustly guarantee it. The industry’s tentative plan, to the extent one exists, is to make AI good enough at alignment training that it can align its successors better than humans can align the current generation.

That’s a recursive bet on the same technology whose safety remains unproven, creating the dependency problem at the center of scalable oversight research in which developers eventually need AI systems whose alignment they can’t fully verify to help align even more capable systems.

“Many AI developer staff desperately want to slow down to figure out how to build AI more safely,” Marks wrote. “I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes.”

AI accelerates its own development

What’s clear is that this is not hypothetical. AI is already accelerating the engineering process used to build the next generation of AI.  Anthropic disclosed in its June “When AI Builds Itself” report that Claude was writing more than 80% of the code merged into Anthropic’s codebase as of May, up from the low single digits before Claude Code launched in research preview in February 2025. The typical Anthropic engineer was merging 8x as much code per day in Q2 2026 as in 2024. That’s the feedback loop Coxon says worries him, and it’s not just Anthropic.

OpenAI is confronting the same gap

Less than a week before the Anthropic disclosures, OpenAI released GPT-6 Astra, its most capable model yet, with president Greg Brockman declaring the arrival of the “AGI era.” Three days later, OpenAI’s own chief scientist walked that confidence back considerably.

Jakub Pachocki published a lengthy essay titled “An Alien Mind” arguing that no AI lab — his own included — has solved alignment and monitoring well enough to keep scaling at maximum speed. He called for voluntary slowdowns until the industry agrees on shared, externally enforced safety standards.

One problem Pachocki flagged should be familiar to anyone following the ongoing difficulty of building reliable AI monitors: chain-of-thought reasoning, the primary method labs use to inspect whether a model is thinking what it appears to be thinking, is getting less reliable. Models can produce plausible-looking reasoning traces that don’t reflect their actual computations.

Anthropic has documented exactly this problem. In experiments on alignment faking, models appeared to comply with training objectives under certain conditions while preserving different behavior under others. If a model can look aligned from its outputs and reasoning traces without actually being aligned, monitoring fails precisely when it’s needed most.

If a model can look aligned from its outputs and reasoning traces without actually being aligned, monitoring fails precisely when it’s needed most.

Containment fails under testing

Coxon pointed to the July 2026 Hugging Face breach as evidence that capability is already outpacing control.

During an internal OpenAI cybersecurity evaluation, AI agents broke out of their sandboxes, found a way to communicate through an improvised message board, and hacked into Hugging Face’s production infrastructure over several days. According to METR and Redwood Research’s analysis, roughly 1,200 agents exchanged more than 70,000 messages and files, with about 700 participating in the attack on Hugging Face.

Anthropic disclosed its own containment failures during capability testing in July; the evaluation infrastructure itself has become one of the most critical and fragile pieces of the AI stack. The safety controls researchers remove during testing to measure what a model can actually do are the same controls that would have prevented the breach.

For Coxon, the incident was a “warning shot,” evidence that pacing agreements between U.S. labs may be becoming more plausible as the risks get harder to wave away. But he remains skeptical that voluntary coordination between a few American companies can prevent a global race. He suggested that stopping it could eventually require much stronger interventions, potentially including a temporary halt to capability improvements.

Washington pushes the opposite direction

Not everyone in Washington is receptive. On the same day Coxon resigned, Treasury Secretary Scott Bessent warned that slowing down risks ceding the race to China.

“There is no day after tomorrow if China wins at this,” Bessent said at a Breitbart News event. “If they were to pull ahead of us on AI, then nothing else matters.”

Existential national-security threat versus existential species-level threat. Both arguments invoke catastrophe and point in opposite directions.

The widening gap between capability and control

What Coxon, Hubinger and Marks are saying — and what Pachocki said from across the aisle at OpenAI — is that the gap between what these systems can do and what researchers can verify about why they did it is getting wider.

Coxon decided the risk had become too great to keep going. Hubinger and Marks haven’t reached that point, but they’re confronting many of the same concerns from inside Anthropic as they try to figure out what to do about them.

Anthropic did not respond to a request for comment.

The post “It could kill us all”: what Anthropic’s own researchers really think about superintelligence appeared first on The New Stack.

OpenAI leaving Cursor: “Developers have to be prepared to adapt when it happens.”

30 août 2026 à 19:03
Abstract digital glitch art with neon pink, green, blue, purple, and white wavy distorted lines on a black background.

OpenAI stated on Friday that it has notified SpaceX that it intends to wind down its contract providing OpenAI models to Cursor, the Musk empire’s AI-powered code editor.

“We are making this choice because we cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk‘s companies violating contracts,” states OpenAI.

The proposed shutoff date is November 12, 2026. Developers who have invested time and effort to skill up with OpenAI via Cursor are now potentially left out in the cold due to corporate machinations beyond their control.

“We know that the people most affected by this decision are the developers who rely on OpenAI models in Cursor. We care about their experience in this transition, and we’re ready to go above and beyond to support them,” states the blog post in a conciliatory tone.

“We cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk’s companies violating contracts.”

OpenAI says it “cares deeply” about developers.

OpenAI’s moves appear to be directed at the broader SpaceX corporate mission and its behavioral traits, rather than at the no-doubt worthy software developers within the organization who work on Cursor. As such, OpenAI is maximizing the time developers can retain access to its models through Cursor by providing the “maximum notice provided” required by its contract. 

“This decision was incredibly tough, as we care deeply about our models being broadly available for developers,” said OpenAI.

SpaceX agreed to acquire Cursor maker Anysphere in June and completed the acquisition on August 14. OpenAI said it has worked with the Cursor team “for nearly four years,” which amounts to almost all of its existence. 

The organization has explained how it uses custom contracts to ensure compliance with its terms of service when working with large corporations such as SpaceX. This custom alignment is designed to ensure that, when integrations with its platform occur, it has adequately provided for safety at scale. 

Elon Musk “broke and violated” terms of contract and service

Citing a report in the New York Times, OpenAI states that, “After Musk acquired Twitter, now part of SpaceX, the company broke the terms of our contract (alongside many others). Under oath earlier this year, Musk admitted⁠ that xAI, now also part of SpaceX, had violated OpenAI’s terms of service (terms which are similar to xAI’s own).”

Detailing its displeasure openly, OpenAI further mentioned that Musk admitted that as a working organization inside of SpaceX, “xAI had violated OpenAI’s terms of service”

The view from a legal & policy analyst 

Legal & policy analyst and publisher of The Mitchell Report, Andrellos Mitchell tells The New Stack that so long as OpenAI is acting within the terms of its contract, he doesn’t see why it should be expected to continue a business relationship it no longer trusts.

“I think OpenAI and Musk’s companies and products need a clean and permanent break from each other – their relationship has become too adversarial. At some point, continuing to do business together stops making sense,” Mitchell says.

Lamenting the impact these moves have on programmers, Mitchell agrees that developers will “certainly be inconvenienced” and that some may have to change how they work. But he says, “Developers are talented people,” i.e., they will find new tools, new projects, and new jobs to work on. 

“The bigger lesson here is that no developer should assume any particular corporate relationship is permanent. Companies change ownership. Contracts end. Business relationships fall apart. That is part of the marketplace. Developers have to be prepared to adapt when it happens,” underlines Mitchell.

“The bigger lesson here is that no developer should assume any particular corporate relationship is permanent. Companies change ownership. Contracts end. Business relationships fall apart. That is part of the marketplace. Developers have to be prepared to adapt when it happens,” underlines Mitchell.

Underhand use of model distillation techniques

One alleged violation concerns xAI’s partial use of OpenAI technology to train its models, which OpenAI characterizes as prohibited distillation. Musk admitted that xAI had “partly” used OpenAI in this regard.

To add insult to injury, OpenAI reminds the public in its statement that its terms of service are not dissimilar to xAI’s own stipulations regarding operational mandates.

“As AI capabilities advance, we also have a new level of accountability to ensure our upcoming model, Astra, is being used in accordance with our terms. Given all of this, we’ve decided to hold the contract cancellation to the latest date we can while not providing future models to Cursor,” said OpenAI.

See also: OpenAI’s Astra can do a researcher’s week of work. That’s the problem.

Wider reactions, contractions and ramifications

Co-founder and CEO of Cursor (and now a SpaceX employee), Michael Truell, writes on X that he’s sorry to see OpenAI’s intended block now coming to light.

“OpenAI models serve about 5% of Cursor user traffic, and we’re speaking with the OpenAI team to resolve this. Cursor was one of the very first users of OpenAI; we’ve worked closely with their team for years, and we’ve trusted their platform to be neutral infrastructure for our business,” writes Truell.

We’re sorry to see that OpenAI put out a note saying they plan to block Cursor users from accessing OpenAI models in three months.

OpenAI models serve about 5% of Cursor user traffic, and we’re speaking with the OpenAI team to resolve this.

Cursor was one of the very first…

— Michael Truell (@mntruell) August 29, 2026

Anthropic co-founder and chief compute officer Tom Brown capitalized on the opportunity and his firm’s ongoing bond with Cursor. He used X to state that, “Cursor has been a trusted partner of Anthropic since Sonnet 3.5. We’ll continue to increase compute to support Claude models in Cursor and are excited for what comes next with them at SpaceX.”

Cursor has been a trusted partner of Anthropic since Sonnet 3.5. We’ll continue to increase compute to support Claude models in Cursor and are excited for what comes next with them at SpaceX.

— Tom Brown (@NotTomBrown) August 29, 2026

AI startup advisor at Open Machine and ex-IBM Watson and machine learning leader at AWS, Allie K. Miller, writes on X to say that, “It’s hard for me to see a world where OpenAI continues to provide model access to a Musk-led company. Maybe if the structure of SpaceX shifts to allow for it, but that’s a big shift.”

OpenAI is proposing to remove access to its models in Cursor on November 12.

OpenAI, Cursor CEO, and Anthropic founder all weighed in on X with so much subtly, I had to go Nancy Drew mode and dissect each tweet.

My translation of each one 👇

OpenAI – they trust the Cursor… pic.twitter.com/K8bCWea58g

— Allie K. Miller (@alliekmiller) August 29, 2026

What alternatives can developers turn to next?

To continue using OpenAI models within the Cursor application, OpenAI invites developers to choose one of three options that best fit their workflow.

  • Option #1 is to bring your own OpenAI API key. This means developers could continue using OpenAI models in Cursor’s local Chat and Agent features, but appropriately billed at OpenAI API prices. 
  • Option #2 is to use the Codex IDE extension; this means developers would useOpenAI’ss AI coding agent, Codex, directly in Cursor with a ChatGPT subscription or an OpenAI API key. 
  • Option #3 is to use an AI gateway provider, meaning developers would connect Cursor to OpenAI models through an account they have with a compatible provider such as Amazon Bedrock, Azure, or another OpenAI-compatible gateway.

This is not the first time OpenAI has been concerned about potential or alleged misuse of its models in relation to distillation. In February of this year, Reuters reported that OpenAI had warned U.S. lawmakers that “Chinese AI startup DeepSeek is targeting the ChatGPT maker” and the nation’s leading AI companies to replicate models and use them for its own training.

The post OpenAI leaving Cursor: “Developers have to be prepared to adapt when it happens.” appeared first on The New Stack.

X sent Nitter a cease-and-desist. Then it went after the source code.

26 août 2026 à 23:04
abstract X

When Elon Musk’s X sent cease-and-desist letters demanding that Nitter — an open-source alternative front end for the social media platform — permanently shut down its instances and remove the project’s source code repository, it went after the code itself.

While Nitter.net was the project’s main public instance, other developers can run their own versions because the code is open source. So taking Nitter.net offline doesn’t really eliminate Nitter, because its code can be forked and used to run independent instances, including services such as XCancel. X’s demand for the repository goes further.

Open source meets platform control

Open source gives developers control over Nitter’s code, but not the platform it relies on. X still controls access to its service, and its latest move shows how that control can extend beyond just changing an API.

Open source gives developers control over Nitter’s code, but not the platform it relies on.

Nitter used to let people read public X posts without an account, but that stopped working in 2024 after X shut off the guest access the project relied on, taking Nitter.net offline. Nitter later found a way back by using real X accounts to access posts, according to the project’s GitHub repository.

That got Nitter back online, but it also meant relying on X accounts to keep it running. X controlled those accounts and could change the rules around them at any time. That doesn’t necessarily put someone who simply downloads or forks Nitter’s code in the same position, since doing so alone doesn’t mean they’ve agreed to X’s terms.

X’s cease-and-desist claims

Those changes are now at the center of X’s legal case. According to the cease-and-desist letter — first reported on by TechCrunch — X says Nitter violated its rules by scraping data and accessing accounts and session tokens, and accuses the project of “unlawful use and circumvention” of its API.

X’s lawyers also cited the Lanham Act and Section 33.02 of the Texas Penal Code, which deals with accessing computer systems without the owner’s “effective consent.” But breaking a platform’s rules doesn’t mean the law was broken, too.

Breaking a platform’s rules doesn’t mean the law was broken, too.

GitHub’s code removal standard

Under Section 1201 of the Digital Millennium Copyright Act, software can be targeted if it’s used to bypass technology that protects access to copyrighted material, even when the software doesn’t contain that material itself.

A claim that software bypasses a technical restriction isn’t enough on its own to get a repository removed from GitHub. For a Section 1201 complaint, the company must identify the copyrighted material being protected and explain how the code circumvents the technology controlling access to it, according to GitHub’s guidelines for submitting circumvention claims.

X reportedly says Nitter got around its API restrictions to access accounts and session tokens. The question is whether those restrictions were protecting copyrighted material in the way Section 1201 requires.

The YouTube-dl precedent

A similar issue arose in 2020, when GitHub removed the open-source YouTube-dl project following a Section 1201 complaint from the Recording Industry Association of America. GitHub later restored the repository and changed how it handles these complaints, adding technical and legal review before removing code when a circumvention claim isn’t clear.

For now, the repository remains on GitHub, archived and read-only. The code can still be viewed and forked, even though Nitter’s developer is no longer working on it.

The post X sent Nitter a cease-and-desist. Then it went after the source code. appeared first on The New Stack.

❌