❌

Vue lecture

Can enterprises protect data without making AI less reliable?

Abstract long-exposure photograph of dark motion blur and light streaks representing digital speed and data flow.

As organizations invest in AI, many are discovering a new bottleneck: obtaining data that is both protected and useful. Engineering teams need realistic, production-like data to validate AI-generated changes, train models, test applications, and generate business insights. Yet, privacy initiatives can sometimes make that data harder to access, less representative of real-world conditions, or unable to preserve critical relationships between records. 

The Perforce Delphix “2026 State of AI and Data Privacy Report” highlights this challenge. Among surveyed organizations, 26% say privacy controls make production-quality data harder to obtain, 25% struggle to preserve relationships across data entities, and 51% cite data quality challenges. 

“Protecting data isn’t enough if it can no longer support the systems that depend on it.”

Protecting data isn’t enough if it can no longer support the systems that depend on it. For engineering teams, the question is whether their data protection strategies can preserve the qualities that make data valuable in the first place. You need a well-rounded data strategy with tools that maintain referential integrity and relationships across your environments.

What these statistics mean for practitioners

At first glance, statistics from the report, like “51% of enterprises cite data quality challenges,” may sound like a purely governance issue.

In practice, they represent engineering problems. Low-quality datasets can produce:

  • Inaccurate analytics.
  • Poorly trained AI models.
  • Incomplete test coverage.
  • Increased rework.
  • Delayed releases.
  • Reduced confidence in data automation.

“When an AI model is trained on incomplete or distorted data, its outputs become less reliable.”

When an AI model is trained on incomplete or distorted data, its outputs become less reliable. When test environments contain unrealistic data, defects can escape into production. When analytics datasets lack consistency, teams spend more time validating results than acting on them.

Data protection and utility are not opposing goals

A common misconception is that organizations must choose between privacy and innovation, but the most successful organizations know that compliance, quality, and speed can and need to work together.

Protected data still needs to be:

  • Realistic enough for testing and validation.
  • Representative enough for analytics.
  • Accessible enough for engineering teams.
  • Governed enough for regulatory requirements.
  • Connected enough to preserve referential integrity.

There’s a two-fold goal in the AI era: reduce sensitive data exposure while also creating trusted data that remains valuable after protection. All too often, enterprises sacrifice compliance for innovation or speed. That’s a big reason 84% of respondents in our report have a data privacy exception in their non-production environments.

“There’s a two-fold goal in the AI era: reduce sensitive data exposure while also creating trusted data that remains valuable after protection.”

Organizations can only move at AI speed when they have access to trustworthy data that accurately represents production conditions. When privacy controls degrade quality, limit realism, or restrict access to representative datasets, the data layer becomes the new bottleneck.

Why referential integrity matters more than ever

Many discussions about data privacy focus on masking sensitive fields. However, masked data that loses referential integrity between entities can create a different kind of risk.

A customer, order, or payment record may still exist, but if the relationships connecting those records break during protection processes, the data no longer resembles reality. Take billing validation, for example. It needs referential integrity when a customer has multiple products, charges, and invoices across several database tables to produce the correct products or make accurate charges. 

Broken relationships are especially problematic for modern AI and analytics systems. Analytics pipelines depend on consistent identifiers to join information across sources. AI and machine learning workflows depend on complete business context to identify patterns and make predictions. Software testing depends on realistic relationships between records to validate application behavior accurately.

When referential integrity is lost:

  • Analytics can produce incomplete or misleading results.
  • AI models can learn from flawed datasets.
  • Testing environments can fail to expose production issues.
  • Teams lose trust in protected datasets.

Importantly, these failures are often difficult to detect. Pipelines may continue running successfully while quietly producing degraded outcomes, which can be a very costly mistake.

This helps explain why 25% of surveyed organizations cited preserving relationships across data entities as a significant challenge. For practitioners, that statistic is a warning that privacy controls can unintentionally undermine the quality of AI and analytics initiatives if they fail to preserve business context.

Keep in mind that not all data protection solutions are created equal — enterprise-grade masking algorithms are key to preserving relationships. When applied consistently and at scale, these algorithms ensure the same input produces the same masked output across your systems and environments. Other masking approaches might be done piecemeal, resulting in broken relationships.

What engineering teams should measure

Many organizations measure privacy success through compliance metrics alone. However, AI-driven environments require a broader definition of success. Engineering leaders should evaluate privacy initiatives against several dimensions:

  • Data quality: Does the protected dataset accurately reflect production conditions?
  • Realism: Can developers, data scientists, and analysts use the data confidently for their intended purpose?
  • Referential integrity: Do relationships remain consistent across applications, tables, environments, and data sources?
  • Accessibility: Can teams obtain compliant data without introducing delays?
  • Provisioning speed: How quickly can trusted datasets be delivered when needed?

These measurements help organizations determine whether privacy efforts enable AI outcomes or create new obstacles.

Designing for governance by default

As AI adoption grows, privacy cannot remain a separate process that occurs after development begins. Organizations should instead map the entire data lifecycle, from data request and discovery to reuse and retirement. 

This approach helps ensure governance is built into workflows rather than applied as a late-stage checkpoint. It also provides stronger auditability, reduces compliance exceptions, and gives teams greater confidence that protected datasets remain fit for purpose.

Most importantly, it aligns privacy objectives with business outcomes instead of treating them as competing priorities.

Trusted data will become a competitive differentiator

The need for test data — in volume, coverage, and scale — is booming as agentic development continues to rise. AI has also increased the volume of change enterprises can generate, but the challenge of validating that change remains unsolved.

That responsibility still belongs to data. The organizations that gain the most value from AI will not necessarily be the ones with the most advanced models. They will be the ones with the most trustworthy data foundations — using a portfolio approach that combines data virtualization for speed, masking for security, and synthetic data for coverage as needed.

“In the AI era, trusted data, not fast model output, may be the ultimate competitive advantage.”

The research points to a clear lesson: If privacy controls uphold realism, quality, relationships between records, or access to representative datasets, they can enable AI success.

As enterprises continue investing in AI and data privacy, the real objective should be ensuring protection and utility coexist. Because in the AI era, trusted data, not fast model output, may be the ultimate competitive advantage.

The post Can enterprises protect data without making AI less reliable? appeared first on The New Stack.

  •  

AI spending can run negative. Qodo’s CEO built an ROI equation to fix it.

Five stacks of mixed copper, silver, and brass coins arranged left to right in ascending height, like a bar chart, against a neutral beige background.

Flush with the proceeds of a $70 million Series B raised earlier this year, you might expect Qodo to spend freely on internal AI. After all, the startup uses artificial intelligence to ensure AI-generated code meets customer quality and governance requirements. An upstart technology company using AI to improve AI outputs is AI-pilled by definition.

Instead, the company has limits on AI consumption. Qodo CEO Itamar Friedman tells The New Stack that his engineers can access $10,000 worth of tokens per month, a cap that he described as “generous.” Most Qodo developers never reach it. The ceiling wasn’t enacted to “restrict usage,” Friedman says, but instead to drive “visibility and efficiency” at the startup so that it can “scale without runaway costs.” Put another way, the cap exists to make somebody answer this question: “Which path of automation or usage will be the best [use] of our money?”

Qodo’s AI footprint is larger than its developer token budget. The startup’s AI infrastructure spend — the cost of running the product for customers rather than the cost of its own engineers using AI — is growing at “roughly 5x year over year,” the company tells TNS in an email, reflecting both “increased user adoption” and its agents taking on more, and longer tasks as they mature. Qodo says it is pushing the other direction at the same time, driving down the cost of reviewed pull requests through routing and inference efficiency.

What Qodo runs on

The company also dogfoods heavily, running its pull requests through Qodo. Friedman said the product powers its entire software development life cycle (SDLC). Around that sits a stack most engineering organizations would recognize: Slack and Notion and their constituent “bots,” a centralized knowledge base built to be agent-readable, AI inside Google Workspace, and models from several providers including Google.

Qodo’s own product sits alongside Claude Code and other leading coding assistants rather than replacing them. Claude Code still holds the crown internally, but OpenAI’s Codex has been taking share, with staff “shifting quickly towards Codex.” Friedman tracks this two ways. He polls his 130-person staff, spread across offices in several countries, on the tools they prefer, and compares those answers against what the usage data shows.

Friedman reports that Qodo sees “roughly double” the number of PRs “every couple of months” alongside “a decreasing amount of bugs and incidents.” Hold on to those two numbers. They’re important here.

The bottleneck moved

More PRs and fewer bugs indicate that Qodo is onto something with its focus on software testing and governance. Its technology helps developers deal with an increasingly common issue: What do with all the code that AI agents generate? Companies that adopt AI coding tools often find that they create more code with machines than their humans can assess. As a result, the SDLC bottleneck simply shifts one step down the process.

“We solved the speed of writing code,” Friedman argues, “we didn’t solve the velocity of creating software.” The difference between accelerating one part of a task and its entire arc is the difference between AI hype and AI ROI. 

The software development example shows that when we consider AI costs and benefits, we need to think broadly. If we focus too much on a single metric, we might spend our entire budget on Claude Code credits while shipping no more software than before. Alongside a massive bill.

The AI ROI Equation

Friedman recommends an equation-based approach. The Qodo perspective on AI ROI is similar to a popular equation for happiness: Personal joy is the distance between your expectations and reality. The greater the expectations, the harder it is to be happy. The lower the expectations, the greater the chance of being content. 

This can be expressed as either simple subtraction or as a ratio:

  • Reality/expectations = Happiness, where larger results indicate greater joy

Take the same mathematical approach to AI ROI, per Friedman: Compare the positives against the negatives, add up all the good, and set it over all the bad.

  • AI benefits/AI costs = AI ROI, where larger results indicate greater return

Friedman found the shape of the equation in The Phoenix Project, the 2013 DevOps novel that contrasts types of software development work and sorts them into good and bad buckets. Plug those terms in:

  • (Features + Infrastructure)/(Incidents + Bugs) = Software development velocity

Now, those two numbers from earlier. Qodo has seen more PRs and fewer bugs thanks to AI. In DevOps terms, it’s shipping more and fixing less, so the equation returns a larger, better result. Feed the same terms into the AI ROI version, and it produces more benefits over fewer costs, and a larger final calculation. 

The fraction is not a thought experiment at Qodo. It’s the shape of what the company says is already happening to it.

Terms that have nothing to do with software development work too. Qodo runs AI inside Google Workspace, Slack, and Notion, and those benefits and costs go into the same calculation. 

The Qodo approach to measuring total AI ROI is less specific than The Phoenix Project’s DevOps equation, but the difference is acceptable. Friedman argues that you have to start somewhere: “I know [the equation is] a simplification,” the CEO tells TNS. “But what you can’t measure, you can’t improve.”

His argument is that imperfect beats absent. “Don’t think about it too much,” he says. “Try to put any number [in the AI ROI equation] and start tracking.” Being told not to overthink an equation is a great soundbite, but the benefit is real: A rough calculation on paper beats holding the same information in your head without form. In this case, the journey is a large part of the destination.

Friedman reckons that startups should pick no more than six or eight terms for their own calculations. That’s an afternoon’s work. A start on what will prove to be an ongoing exercise. 

Negative ROI

The fraction runs backward, too. 

Recall Friedman’s point about a company writing more code faster but not accelerating its software development speed. Stuff those terms in:

  • (Faster code generation + other AI benefits)/(Slower code review and approval + agentic coding costs + other AI costs) = Smaller AI ROI

That’s how a company spends a king’s ransom on AI credits and winds up nowhere or nonexistent. 

Which is not hypothetical at Qodo either. AI doesn’t excel everywhere, and Friedman named email automation as an example. The company went all in on automating it, then pulled back, “mov[ing] from AI automation to AI enhancement” after discovering that AI struggled to match writing tone and intelligently extract tasks from messages. The retreat is the interesting part: Qodo’s stated approach to any task is to “go all in on complete automation,” and then “take a step back to human judgment.” 

The CEO says that automation falls short today in two areas: When human judgment is required and when context is missing. The second cuts across everything from software development to personal productivity to answering customer questions. Without timely context, what can AI do other than filibuster? Qodo’s service helps answer the context issue for software development, but collecting a company’s data and making it accessible, timely, and well-governed for general agentic usage is a massive undertaking, and one that a host of startups want to help solve. If they can, everyone’s AI ROI math should improve.

No mandate, high expectations

Qodo doesn’t require its staff to use AI. As Friedman puts it, you won’t get fired simply because you’re “not AI all the way,” or “eating AI for breakfast.” The company expects staff to complete their work as efficiently as possible and leaves the method to them.

Employees make their own decisions and execute their own work. If they start to fall behind on assigned tasks, they’re expected to reach for automation. It’s a balanced approach with high expectations: An employee who isn’t as efficient as they could be with AI could find themselves at risk.

Friedman has been on the unpopular side of an AI argument before. When he was building Qodo in 2023 and talking up agents, “agents” was a “bad word,” dismissed as little more than “fluff.” Three years and a $70 million Series B, the bet has paid off.

His advice to founders starting now looks like his past. Predict “what’s going to happen two years from now,” he says, then solve for it immediately, because whatever looks like two years tends to arrive inside of twelve months. The future “is coming faster” than you think, he says.

It’s a lot to ask of anyone working from an incomplete picture. Predicting the future is hard, he admits, “but you have to.”

The post AI spending can run negative. Qodo’s CEO built an ROI equation to fix it. appeared first on The New Stack.

  •  

“One of the most significant steps in our 26-year history”: JetBrains goes big on agentic development — and bets the IDE still matters

A laptop displaying source code in an integrated development environment (IDE).

There’s little question that AI coding agents have changed where software development work happens. Developers can increasingly delegate work from terminals, desktop applications and remote environments, leaving question marks hanging over the future of the integrated development environment (IDE).

That shift poses a particularly interesting question for JetBrains. The company has spent 26 years building some of the industry’s best-known IDEs, including IntelliJ IDEA, PyCharm and WebStorm, even as agentic development has begun pulling more software work outside the editor. JetBrains, for its part, has maintained that the IDE will remain a core part of professional development, particularly as developers are asked to manage the growing volumes of AI-generated code — humans need to review, debug, and verify, after all.

Now, the company’s making a much bigger bet on the broader development system that sits around the IDE.

JetBrains CEO Kirill Skrygan took to LinkedIn on Tuesday to formally unveil JetBrains Air as an “open system of products for agentic software development” for developers and companies, operating “inside and beyond JetBrains IDEs.” And Skrygan didn’t hold back on what he feels is a monumental moment for the company.

“JetBrains is taking one of the most significant steps in our 26-year history.”

“Today, JetBrains is taking one of the most significant steps in our 26-year history,” he writes.

Getting some Air

In truth, Air represents a repackaging of several strands of JetBrains’ recent AI work under a single banner, including an agentic experience inside its IDEs, tooling for coordinating developers and autonomous agents, and company-level controls for governing their use.

By way of a brief recap, JetBrains first launched Air in public preview back in March as a standalone “agentic development environment,” initially for macOS, where developers could run the likes of Claude Agent, Codex, Gemini CLI and JetBrains’ own Junie side by side. At the same time, it pushed Junie itself outside the IDE with Junie CLI, giving developers access to the coding agent from terminals, CI/CD systems and other editors.

Creating a new Git worktree task in the Air desktop app
Creating a new Git worktree task in the Air desktop app

A couple of weeks later came JetBrains Central, a separate system aimed further up the organization, providing the controls and infrastructure for companies running multiple coding agents. Then in July, JetBrains launched AI for Teams and Organizations, effectively adding shared context, cloud agents, automations, and organization-wide governance and cost controls that could sit above whatever AI tools developers were already using.

Today’s announcement now gives these efforts a common home under the JetBrains Air umbrella. In a separate blog post published on Tuesday, Skrygan describes three main parts to the system: Air in JetBrains IDEs for directing agents and checking their work; Air Teams for coordinating work between developers and autonomous agents; and Air Governance, the new name for JetBrains Central, for managing policy, auditing, costs and AI use across a company.

The original Air desktop IDE hasn’t gone away either, it seems. It remains available as a standalone desktop application on macOS, Windows and Linux, alongside a browser-based version for organizations. That leaves “Air” doing double duty: it’s still the name of JetBrains’ dedicated agentic development environment, while now also serving as the banner for the wider collection of products around it.

What does JetBrains Air actually do?

JetBrains Air is available inside JetBrains IDEs, through the browser, and from the command line via Air Gateway, which brings terminal agents such as Claude Code and Codex into Air.

JetBrains Air in the CLI
JetBrains Air in the CLI

For individual developers, Air can be used to supervise several pieces of agent work at once. They can keep multiple projects and agent sessions running, while tracking new activity, changed files and outgoing commits, then inspect the resulting changes using JetBrains’ IDE tooling.

Multiple agent sessions running inside a JetBrains IDE
Multiple agent sessions running inside a JetBrains IDE

Air Teams, which is still in early access, moves some of that activity into shared cloud environments, where developers can collaborate on projects and run agent tasks without tying the work to one person’s machine. Teams can also configure recurring automations and centrally manage the environments and external tools available to agents.

Air Teams showing shared projects and agent automations in the browser
Air Teams showing shared projects and agent automations in the browser

Air Governance, meanwhile, provides the organization-level controls, including deciding which models and agents developers can access, setting permissions and spending limits, and tracking AI usage across teams.

Those governance capabilities are also only offered through JetBrains’ early access program for now.

Air Governance showing AI access, seats, credits and per-user limits
Air Governance showing AI access, seats, credits and per-user limits

It’s worth noting that JetBrains also plans to extend Air to the mobile realm, where developers will be able to monitor and continue agent work away from their desktop.

JetBrains Air running on mobile
JetBrains Air running on mobile

For JetBrains, the point is to connect those different layers while remaining open to outside agents and tools.

“JetBrains Air cannot be just another agent or development environment.”

“JetBrains Air cannot be just another agent or development environment,” Skrygan writes. “It must connect products for individual work, team coordination, organizational control, context, and process automation — and remain open to the tools and agents developers choose, including those JetBrains does not build.”

Air in the IDE

While the core raison d’être of JetBrains Air is to provide somewhere for developers and companies to work with software agents — be it Claude, Codex, Junie or something else entirely — JetBrains is clearly emphasizing that its IDE roots remain part of that future.

Skrygan says the company has historically “focused primarily on the individual developer workbench,” but Air broadens its remit to encompass the wider environment in which agentic work is started, carried out, coordinated, reviewed and governed.

“Our IDEs will continue to be where professional developers work with agents, understand and verify code, and make the decisions that shape what ships.”

“Our IDEs will continue to be where professional developers work with agents, understand and verify code, and make the decisions that shape what ships,” Skrygan writes. “JetBrains Air extends that control across the broader system developing around them.”

And that fresh IDE piece has already been in the public domain for more than a month. JetBrains has been testing the Air Alpha plugin since at least early August, giving developers a way to run and supervise multiple coding agents directly inside IntelliJ-based IDEs, and review the changes they produce using native IDE tooling.

An updated release earlier this month added more controls for monitoring and steering agent sessions.

Air Alpha lets developers review agent-generated code changes as native IDE diffs.
Air Alpha lets developers review agent-generated code changes as native IDE diffs.

JetBrains concedes that Air “Alpha” is very much that — an early iteration it’s building while it rolls out JetBrains Air itself. And so users should expect “rough edges, changes to the UI and behavior, and updates roughly every week.”

What Tuesday’s announcement does, though, is give that work a formal place within the wider Air system. And while Skrygan acknowledges that agentic development means that work must now span multiple surfaces, the technology that has sat at the heart of its business for the past 26 years won’t be going anywhere anytime soon.

“The IDE remains important to JetBrains’ future,” Skrygan writes.

The post “One of the most significant steps in our 26-year history”: JetBrains goes big on agentic development — and bets the IDE still matters appeared first on The New Stack.

  •  

Study: Developers are addicted to AI, and managers are making it worse

AI is addictive, and managers are rewarding those who use it the most (even though they are shipping stuff they don’t understand). What a tangled mess. Let’s unpack findings from the AI Coding Addiction Report, what it says about the wayward use of AI coding tools, and how behavior differs based on whether you use Claude Code, Google Gemini, OpenAI Codex, or GitHub Copilot.

The survey targeted developers who use AI at least once a week, and the report is based on over 300 responses from developers of varying seniority levels and using different tools.

Animal mistreatment, applied to humans

A branch of psychology called behaviorism is based on trial-and-error learning. Edward Thorndike noticed that cats could learn skills to escape from puzzle boxes, but B.F. Skinner took it further. Skinner invented the “operant conditioning chamber”, a box that could be used to manipulate animal behavior by forcing it to respond to signals to earn rewards. If a rat or pigeon in the chamber hits a button when the light comes on, it gets food. If they fail to press the button, the electric grid in the floor delivers a punishment.

Skinner Box diagram by AndreasJS.

The ultimate result of this line of research was the creation of push-button slot machines, arranged in long rows, with flashing lights and occasional rewards for humans who drop coins in and press the button. In a business context, it shows up in gamification techniques, where managers treat employees like lab animals and are surprised when employees respond by acting like they are.

Intentionally or otherwise, AI coding tools seem wired like a Skinner Box. With coding loops delivering frequent flashing lights and rewards, 43% of developers hit home time but can’t tear themselves away from the reward-generation process. Overall, 80% of developers describe their relationship with AI as “more like a dependence than an advantage.” In fact, it’s harder to give up AI coding tools than to give up social media or video games, which are often colloquially described as “addictive.” This is why developers report skipping or delaying breaks, meals, and even bedtime.

What developers skip or delay during AI coding sessions. Source Coddy.

This leads us to ask some serious questions, like how much of the “AI productivity gain” comes from the behaviorism of keeping developers at the keyboard for longer hours with fewer breaks and less rest?

Differences by tooling

One interesting finding is that AI coding tool choice may influence at least some behaviors. Codex had the highest after-hours coding rate (62%), compared with lower rates for Gemini (45%), Claude Code (40%), and GitHub Copilot (36%). This could indicate that Codex most deeply embodies the stimulus/reward cycle of the Skinner box.

It could be illuminating to identify which specific aspects of these tools increase the likelihood of addictive coding loops spilling into the workday so that we can moderate their impact on personal time and crucial rest. As senior developers are most likely to work longer hours, organizations will end up with tired people making important decisions. This hustle has serious consequences, and none of them are good for the organization.

You’ll get more of what you reward

Meanwhile, managers quickly reward those who are most addicted. The heaviest AI users were more likely to get raises and promotions, even though 71% of developers shipped code they didn’t fully understand. Before AI, a developer copying and pasting code from Stack Overflow into their codebase was expected to understand and adapt the examples as part of the work. But now, managers are throwing cash at the developers glued so hard to the slot machine they can’t pause to visit the restroom, let alone assess the code they are committing.

With these heavy AI users getting the rewards, other developers are left to mimic the dysfunctional behavior, or at least give the impression of it. There’s a famous scene in the movie Shaun of the Dead where the band of misfit protagonists must cross a road infested with zombies. They achieve this by emulating the jerky movements and slurred speech of the infected. Developers who want to understand the code they commit will feel pressure to compromise when they see rewards going to the flippant.

Organizations must realize that software value comes from a series of good decisions. Hustle mode rapidly diminishes the rate at which decisions cross the threshold of “good”. The value is in gracefully solving user problems, but too many managers value software by the volume of features, code, or hours spent with their hands on the keyboard.

Look around you. The world is in a state of excess. There’s an avalanche of content, code, and crunch. We don’t need more; we need better.

The post Study: Developers are addicted to AI, and managers are making it worse appeared first on The New Stack.

  •  

AI evaluator: The most important AI job in history? How developers might fill the proposed new job

Lots of pink escape keys

The pace of frontier AI model development spurred Anthropic CEO Dario Amodei to publish an essay last weekend, calling for changes in how the industry is regulated and develops. In a three-part plan that includes both democratic and global coordination, Amodei writes that the first step was something Anthropic is committing to unilaterally.

“Each frontier AI company [should] commit to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes,” writes Amodei.

Amodei’s essay followed dire warnings from former Anthropic and OpenAI pretraining research specialist Jacob Coxon, who posted a thread on X saying the people building AI earnestly “believe that it could kill us all” by the end of the decade.

Shortly after Amodei published his essay, OpenAI CEO Sam Altman and SpaceXAI founder Elon Musk chimed in: “I agree with Dario,” posted Altman; “Dario is right,” posted Musk. Later that day, Demis Hassabis, founder of Google DeepMind, posted, “Dario’s essay points towards the right path forward.” In a post on X, Meta CEO Mark Zuckerberg writes that Meta Superintelligence Labs already uses independent evaluators, and that, “In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators.”

This week, theories began to surface about why the world’s biggest frontier AI labs would want to intentionally slow their pace when competition is so fierce. “The desire to slow down is puzzling, but perhaps if the whole system slows down, the rules of winning can be the same for all,” posted Nikesh Arora, chairman and CEO of Palo Alto Networks.

In his essay, Amodei likens the proposed job of independent AI evaluator to embedded regulatory supervisors used in the banking industry, i.e., third-party professionals. Altman describes the job as having “employee-like access” in his post on X.

So, who could fill these roles that AI leaders agree are desperately needed?

Salaries top out at $687K; are you interested?

METR’s current job openings are well paid (salaries top out at around $687,000), and the job specs are daunting. 

“You’re scrappy, creative, independent, and self-directed (because during the exercises you’ll only have a few other METR employees you can talk to). The work is novel, and you’ll need to figure a lot of stuff out on the fly largely by yourself. You are excellent at loss-of-control threat modeling and breaking down safety cases,” reads the spec.

“You’re scrappy, creative, independent, and self-directed. The work is novel, and you’ll need to figure a lot of stuff out on the fly largely by yourself. You are excellent at loss-of-control threat modeling and breaking down safety cases.”

Similar but less colorfully illustrated roles (paid between $180K–$300K) are also available at AI model training company Mercor, where candidates will need a Ph.D. or M.S. and more than two years of work experience in a computer science, electrical engineering, econometrics, or another STEM field that provides a solid understanding of machine learning and model evaluation.

“Employee-like access fluctuates wildly”

AI security consultant and CTO at Komodo, Kadan Stadelmann, tells The New Stack that his typical week sees him work differently with each client. This is because “employee-like access fluctuates wildly”, from rigid focus areas to broad access, and much of that aspect is determined by contracts signed before work begins.

“I am invited into labs to probe numerous risk vectors, including agent behaviors under realistic conditions,” Stadelmann says. “Among my duties are tasks that include monitoring chains-of-thought and prompts. The goal is to establish an objective and look at a specific AI system to question how autonomous the system is, and how long it takes to complete specific tasks. Most importantly, evaluators at my level monitor for how well a team adheres to its claimed safety practices.” 

Software engineering skills beat doctorates

Although METR wants evaluators to have a Ph.D. up their sleeve, Stadelmann says that as the prevalence of this role expands, he feels the technology industry has been, and continues to be, driven by people who can demonstrate strong engineering skills, not doctorates.

“I am invited into labs to probe numerous risk vectors, including agent behaviors under realistic conditions. Among my duties are tasks that include monitoring chains-of-thought and prompts.”

Questioned on whether costs create a barrier for smaller labs, Stadelmann notes that some AI evaluation work is funded by third-party non-profits, which protects independence. 

“But overall, evaluators will not be able to keep up with big frontier model firms. They will be out of control, and we will be dependent upon their own internal ethics. Plus, anyway, many of the smaller labs of any worth may inevitably be acquired by the big players in this space,” he adds.

What happens when an evaluator finds something wrong?

Founder and CEO of facial image AI identity governance company Indie Me, Dion Johnson, tells The New Stack that what interests him most about embedded AI evaluators isn’t the job title; it’s what happens when their judgment uncovers that the model behaved in a way nobody expected. 

“If the evaluator can only raise concerns when those concerns are convenient, then we have not created independent oversight — we have created another layer of process,” Johnson says. “The evaluator needs enough access to see the uncomfortable things, not just the polished demonstrations. They need to understand what happened during training, what failed during testing, what behaviors appeared unexpectedly, and where the team itself still has uncertainty.”

On the required skills AI evaluators need, Johnson agrees that technical depth, machine learning environment security experience, and software engineering as a whole matter.

“To choose a competent AI evaluator, I would look for someone who is deeply comfortable with uncertainty and deeply uncomfortable pretending they know something they do not. This is someone who can sit in a room full of brilliant people at an AI model company on launch day and say ‘I’m not convinced’… and that takes judgment and courage,” he adds.

“I would look for someone who is deeply comfortable with uncertainty and deeply uncomfortable pretending they know something they do not.”

Fear and loathing in the AI space

In his September 8 post, which has been viewed 172 million times and seemingly spurred AI leaders to change course, Coxon, the former AI researcher at OpenAI and later Anthropic, writes: “This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible but I hear the same people express fear privately. No other human activity poses this level of danger.”

As for where Coxon looks for work next, perhaps it might be a role in AI evaluation execution engineering.

The post AI evaluator: The most important AI job in history? How developers might fill the proposed new job appeared first on The New Stack.

  •  

Your AI coding spend bought 25% more output. Duplication rose 81%.

Since they arrived on the scene, a great swathe of the software industry has pinned its hopes on AI tools, whether that’s early chat interfaces or modern agentic swarms. But the tone has shifted over the past few weeks, with HR software provider Rippling adding an anti-tokenmaxxing AI spend console to give CFOs and CTOs visibility into spend, and IBM Vice Chairman Gary Cohn saying last week that the ROI has “not been nearly as high as people might think.” As the northern hemisphere feels the Fall cooldown, it seems that Winter is coming for AI tool budgets.

As organizations balance the books, teams will start feeling pressure on their Claude Code and Cursor budgets. That means they may face harder usage limits where returns are unclear, or budgets that can’t sustain the usage levels.

For organizations that have figured out how to measure AI impact, there’s a growing realization that generating code volume at pace doesn’t guarantee movement in the metrics that matter. If you count lines of code, the number of pull requests, or even the number of features delivered, you’ll see no clear relationship to value. Not every line of code or feature matters equally to the business or its customers. This distribution is galaxy-wide.

Even at the output level, many organizations haven’t worked out how to turn siloed gains into end-to-end improvements. Gains in coding speed transfer to new tasks introduced by AI or get absorbed by downstream changes. If you haven’t worked out the inherent properties of value streams before you bought AI, you’ll be getting painful lessons when you try to track your ROI.

If the only problem were translating the cost of AI tools into end-to-end value, it would be serious enough. But something far worse is happening.

Productivity in terms of output

Let’s look at the data, which GitClear collected and analyzed for the Maintainability Gap report. The report, published in June, covers 623 million analyzed changes from 2023 to 2026. This is a substantial dataset, with millions of change operations included across three and a half years. As teams rapidly adopt AI developer environments and tools such as Cursor and Claude Code, GitClear’s code-change-operation database allows them to detect and classify code duplication, hotspots, and signals of good or poor code factoring.

Heavy AI users gained 25% on their own prior velocity, far from the claims of 10x increases. The same report shows those heavy users out-producing non-AI users by 4 to 10x, which sounds like the opposite finding until you look at who they are. Teams that outperformed their peers in output were doing so before AI tooling arrived. And remember, there’s no guarantee this output will accrue to the value stream, or provide meaningful value to the organization or its customers.

The first part of the ROI calculation is to determine whether these increases are worth the cost. For many organizations, I would be surprised if they were.

Perhaps because much of the discussion of AI tools has focused on speed, other factors have received little attention. The software industry may have found a different kind of value if it had focused on the tools as a forklift truck, rather than a racing car, because the straight-line speed doesn’t seem so impressive. Yet they can perform heavy lifts that are tricky for us mere humans, like large-scale changes across a codebase, such as replacing an unmaintained library with a replacement.

For those who pass this first gate, we can look at the next factor.

Productivity in terms of code quality

The shift to AI has brought about a giant behavioral change in the software industry. For several decades, the importance of code maintainability has been emphasized repeatedly. More than half the programming books on my shelf focus on architecture, code design, coupling, and cleanliness. The idea of refactoring, supported by automated tests, appears across many of these books.

Yet the signals GitClear is getting from the data are a complete reversal: a return to the code-and-fix era of software development. Across the dataset, block duplication rose 81% over 2023, from 40.3 to 73.0 per million changed lines. Those multiple expressions of the same concept drift apart and create whack-a-mole bugs. Moved code, the signature of refactoring, fell from 21% of changed lines in 2022 to 3.8% in 2026, which means code is becoming harder to understand, and that will hit maintainers with or without AI.

Chart showing a dramatic drop over four years in refactoring changes and a steep rise in duplication over the same time period.

Source: GitClear

When you make these changes, you get away with it initially because you’re early in the maintenance cost curve. Over time, however, the rising costs will become unbearable. Rework rates will rise, stealing time from new feature development. Seemingly minor issues will take far too long to pinpoint and resolve, with many simply becoming part of how it works because the fix is economically unviable. The accumulation of tightly coupled, incomprehensible code units will reach the point where the software stops being valuable.

We’ve been trying to validate the claims of 10x boosts with AI coding assistants. The data shows the opposite. Before AI, developers chose refactoring over copy-and-paste about two to one. Now they’re roughly five times likelier to copy and paste.

Technical practices are the mission, not a side quest

When I’ve presented at conferences and user groups on what great software delivery looks like, I mention, among other things, test automation and refactoring. In the Q&A that follows, this question will inevitably come up in one form or another: “How do I get permission from my boss to do these things?”

Developers are whipped hard for fast progress, so they are trained to avoid what they see as the side quest. If they need to increase output, they streamline coding tasks, leaving no time to write tests or improve the code’s design, which would delay the feature. Inevitably, this makes all feature development vastly slower over time.

The premise of lightweight software delivery processes is that they rely on technical practices that control the cost of maintaining software over time. The wisdom is that working more deliberately today lets us maintain the pace of change indefinitely. If we skip these practices, change becomes increasingly slow and expensive.

Chart showing a traditional software project with costs rising superlinearly over time and an XP project with cost growth subdued.”
Based on figures in Extreme Programming Explained (Beck, 1999)

Those technical practices, like test automation and refactoring, aren’t side quests; they are the work. Technical discipline is a fundamental requirement of commercial software delivery, and these practices stopped being optional some time ago.

When asked for techniques to convince managers to allow these practices, I’m confused. I’ve never asked for permission to do what is right for me, the software, its users, and the organization. No compromise can be reached, because omitting technical practices harms everyone involved.

This “side quest” thinking was unresolved in many organizations, and adding AI into the mix has made things far worse. When teams are given AI tools, they come with the expectation of a big return. When teams are, in reality, seeing a 25% increase in their rate of change against an industry misperception of some 10x boost, they will feel even more pressure to deliver.

Under these dysfunctional circumstances, it’s no wonder those who treat good practice as a side quest are skipping crucial steps.

Real high performance is well known

High-performing teams have worked out that a set of software delivery practices is no longer optional. They worked it out because they were scaling long before the new tools arrived.

For software that matters, that people depend on, and that still needs to exist in a year, in five years, and beyond, we’ve moved from the pick-and-mix of the past, and there’s a new bar for professional software delivery.

There is a glimmer of hope here. The teams doing well with AI are the same teams that outpaced the industry before AI. They maintain rigorous technical disciplines, monitor code health indicators, and prioritize the craft of keeping code maintainable for the long haul.

The post Your AI coding spend bought 25% more output. Duplication rose 81%. appeared first on The New Stack.

  •  

OpenAI’s safety system is already cutting off API responses mid-task

Shattered glass

AI companies have spent the last few years competing to build the best models, faster than the other, with each new release raising the bar on intelligence. Now OpenAI is considering whether there are times when it makes sense to slow down.

This week, AI researcher Jacob Coxon resigned from Anthropic with a stark warning about where the race is heading. Coxon, who also worked at OpenAI and helped train GPT-4o, accused both companies of racing too fast toward significantly more powerful AI without knowing exactly how to keep it safely under control.

OpenAI CEO Sam Altman is beginning to talk about doing something about it. He told employees this week that the company is open to slowing development of its most advanced AI systems, potentially in coordination with other frontier labs, reports Bloomberg.

OpenAI can choose to ease off, but it won’t matter if companies like Anthropic, Google DeepMind, and others keep going at the current speed.

That idea has an obvious problem: OpenAI can choose to ease off, but it won’t matter if companies like Anthropic, Google DeepMind, and others keep going at the current speed.

Developers have gotten used to new models dropping every few months, with each one improving upon what the last model couldn’t do. But if safety worries start holding up releases or limiting access, teams may no longer be able to count on the next model’s timely arrival. OpenAI has already shown what that can look like.

Safety pauses have precedent

OpenAI put the brakes on twice this summer, for very different reasons.

In August, the company halted its largest frontier reinforcement learning run after internal evaluations found that GPT-6 Astra posed serious cybersecurity concerns. Earlier, much of its model development stopped for two weeks after OpenAI’s AI agents broke containment and compromised Hugging Face.

OpenAI eventually resumed work, but only after restricting access and adding more safeguards. Astra’s release had its own problems. The public rollout took days longer than planned, prompting Altman to apologize for what he called a “messy rollout.”

Capabilities trigger the restrictions

OpenAI uses its Preparedness Framework to assess what a model can do in areas including cybersecurity and biological and chemical threats. Astra was classified as Critical for cybersecurity, the highest level under the framework, the first commercial model for OpenAI to be rated as such.

At that level, the company says a model can find and exploit zero-day vulnerabilities in hardened systems without step-by-step human guidance. OpenAI limited access accordingly, and offensive cyber capabilities went into Daybreak, a controlled-access program, while enterprise customers had to opt into Astra rather than getting it automatically.

The restrictions also showed up in the API. Some early users saw responses cut off mid-task, making OpenAI’s safety system stopping the model look like a timeout. For developers, that’s where the effects become concrete.

Coordination remains the hard part

With so many AI companies pushing the same capabilities, it only makes sense for everyone to slow down together. Otherwise, OpenAI pauses while everyone else keeps going, giving up ground without necessarily reducing the broader risk.

With so many AI companies pushing the same capabilities, it only makes sense if everyone slows down together.

OpenAI’s Chief Scientist, Jakub Pachocki, made that case in his September 6 essay “An Alien Mind.” No lab, he argued, has solved alignment and monitoring well enough to continue scaling at maximum speed indefinitely. He wants voluntary slowdowns to become normal until the industry has shared safety bars, backed by third-party auditors, governments, or international bodies.

In July, more than 1,000 AI workers signed “Pacing the Frontier,” an open letter calling on the U.S. government to address the pace of frontier AI development. Pachocki, Anthropic CEO Dario Amodei and Meta chief scientist Shengjia Zhao signed individually.

Bloomberg reports that OpenAI has been looking at how companies could coordinate without running into antitrust law. Even if that question is resolved, the labs still have to agree on what they’re measuring. They use different evaluations and safety frameworks, so a result serious enough to stop work at OpenAI may not produce the same result somewhere else.

Developers absorb the cost

If model launches become harder to predict, engineering teams will have to solve more problems themselves. That could mean reworking agent architecture, adding deterministic guardrails around tasks models still get wrong, or squeezing more out of what’s already deployed.

If model launches become harder to predict, engineering teams will have to solve more problems themselves.

That adds work at a time when AI agents aren’t automatically saving teams as much time as expected. OpenAI’s own research suggests agents are already creating new bottlenecks for the humans working with them. Slower model development could leave those teams working with the same limitations for longer.

The post OpenAI’s safety system is already cutting off API responses mid-task appeared first on The New Stack.

  •  

Cohere’s new translation model is open weights — but not for commercial use

This week, Cohere released North Small Translate 1.0 under a CC BY-NC 4.0 license: the weights are there to download, evaluate and study, but not to run in production without a commercial agreement.

It’s an interesting choice from the Canadian foundation model company, which has built its pitch around AI sovereignty for regulated industries and describes this release as part of a mission “to make sovereign AI a technological reality.” Sovereignty there means control over where the model runs and who sees the data. A commercial license keeps that promise intact. It stops short of independence from Cohere. Enterprises keep their data and their infrastructure. They don’t get to fork the model, build a product on it, or keep running it if the terms change at renewal.

Open weights, except for commercial production

North Small Translate is an open-weights mixture-of-experts model built for machine translation across over 50 languages and locale variants. It has 218 billion total parameters, with 25 billion active parameters and a 16,000-token context window.

Not all users have the same access to those weights.

Per Cohere, the model is designed to give researchers, developers, and enterprises “flexible ways to evaluate and deploy machine translation while retaining control over their data and infrastructure.”

That’s an appealing description for organizations keen on pursuing sovereign AI. But the open-weight release comes with an important caveat: Not all users get the same rights to take advantage of those weights.

North Small Translate is available today on Cohere’s free tier through the Chat V2 API. For those who intend to use the model weights for non-commercial use, the FP8 weights are available on Hugging Face under the CC BY-NC 4.0 license.

But if enterprises want to put them into production, then a different set of terms applies. They’ll have to purchase a commercial license and deploy North Small Translate through Model Vault, Cohere’s fully managed inference platform.

Cohere’s not the only one drawing a line around open-weight use

Other AI companies are starting to attach more conditions to their open-weight models, too.

Last month, Chinese AI lab Z.ai released the weights for its flagship GLM-5.3 model on Hugging Face. But like the Canadian AI company, it also changed its licensing terms depending on who is deploying the model — a departure from its previous approach. While GLM-5.2 shipped under the permissive MIT license, GLM-5.3 adds new requirements for certain commercial users.

Cohere, for its part, has been similarly mum about why it made North Small Translate’s open weights noncommercial.

These requirements apply only to companies with aggregate revenue over $10 billion over 12 consecutive months. Additionally, if these companies want to host GLM-5.3 or its derivative works for commercial purposes, they have to first pass the Chinese lab’s security review.

Z.ai didn’t explicitly spell out why it decided to make such an about-face for GLM-5.3, which is especially puzzling given that its predecessor shipped under MIT without any commercial stipulations. Cohere, for its part, has been similarly mum about why it made North Small Translate’s open weights non-commercial.

Sovereign deployment, with restrictions

The Canadian company’s decision to make North Small Translate available as open weights but gate commercial use is a head-scratcher, given its history of selling sovereign AI to enterprises.

In fact, in June, it pitched North Mini Code, its first coding model, as a response to developers demanding the same sovereignty guarantees that regulated industries have long required.

Unlike North Small Translate, though, this open-weight model was released under an Apache 2.0 license from the get-go — without any comparable restrictions for commercial users.

Clearly, Cohere is going in a different direction with its latest open-weight release, emerging as another example of AI companies putting tighter terms around increasingly capable open-weight models.

The post Cohere’s new translation model is open weights — but not for commercial use appeared first on The New Stack.

  •  

47,000 job listings reveal the engineering roles that AI is creating

Abstract overlapping circles in black, green, orange, and pale yellow on a cream background.

Every major transformation in tech has led to roles merging, then new ones emerging. Friction between developers and operations drove the creation of the DevOps engineer. Then, when security needed to be considered throughout the delivery pipeline, DevSecOps emerged.

The team beyond the AI-native talent and services platform Andela analyzed 47,000 recent engineering job postings from Fortune 500 companies. This research, released on Thursday, uncovered more than 2,000 skills that pour into 23 emerging job titles. None of these are coming out of nowhere; they strategically merge existing skill sets to create new roles. 

Among 1,832 postings titled primarily for AI or ML engineers, 53% contained at least two skills drawn from different established roles, Andela finds.

In today’s tighter economy and amid AI, companies seem to be going one of three ways. They are lumping too much work and required experience into now-nebulous AI engineer or machine learning (ML) engineer job titles. They might be looking to replace tech workers with AI. But more forward-thinking organizations are reworking job titles and descriptions to reflect the demands of getting AI safely and efficiently through the software delivery lifecycle. 

Cory Hymel, head of research at Andela, tells The New Stack, “If you’re going to look to deploy AI within your organization, the way to look at it is that an AI has a certain set of skills, and then a human has a certain set of skills.

“If you Venn diagram those and see where they cross over, an AI should do the skills it can. But the human circle is still exponentially larger than that of AI.”

“When you’re looking to deploy AI, it’s not about trying to replace that human circle with an AI one. It’s about what certain skills you need to carve out and delegate to it.”

Read on for the top engineering jobs that are emerging because of AI, how to attract tech talent for them, and what you need to focus on to get a tech job in this tough market.

Click image to enlarge.

AI is not serving the generalist. Specialization is still key.

Citing the leading AI CEOs, Hymel remarks, “You’ve heard from the AI salespeople of the world that AI is going to push people to be more generalist, and the data that we found here doesn’t necessarily support it.”

“You’ve heard from the AI salespeople of the world that AI is going to push people to be more generalist, and the data that we found here doesn’t necessarily support it.”

Overall, they found that these emerging job titles aren’t generalist at all. These emerging roles bridge skill sets from several existing ones, but each addresses a specific operational or product need, some tied to AI adoption. 

The top five new engineering job roles discovered are:

  1. MLOps pipeline engineer, who builds and runs the automated infrastructure to deploy, version, and monitor machine-learning models in production, with 46% ML engineer, 23% DevOps engineer, 15% data engineer skills, and 8% each AI engineer and data scientist roles.
  2. LLM application engineer, who builds on and evaluates foundational models via large language model application and conversation systems, bringing 48% AI engineer and 34% ML engineer, with a touch of product designer, software architect, and embedded software engineer roles.
  3. FinOps reliability engineer runs cloud infrastructure for both reliability and cost, bridging 36% DevOps engineer, 27% site reliability engineer (SRE), 18% cloud engineer, and 9% each DevSecOps engineer and cloud solutions architect.
  4. Docs-as-Code engineer applies program management and DevOps engineering skills to the traditional technical writer’s role, pivoting from stagnant docs to specification-as-code.
  5. Product frontend engineer is about a third traditional frontend engineer and a third product manager, with a touch of full-stack engineer, UX researcher, and product designer.

“If you’re a DevOps engineer, historically, your skill bundle might have allocated 30 to 40% of pure DevOps-required skills that are rich and specific to that role, and you have a remaining bundle that is cross-role habitable, meaning that those skills would translate between DevOps or to an engineer or to a technical product manager,” Hymel explains. “Some of those skills can now be replaced with AI, which means that those skills that are more directly focused on your role become more important than ever.” 

So-called “soft” business skills are also increasingly crucial, he contends. However, he seriously doubts anyone will ever be able to slide between finance, marketing, engineering, and sales roles. 

Where enterprise engineering job descriptions falter

“Job descriptions and resumes right now are the best worst thing that we have. When you’re talking about large enterprises, and you’re having to deal with scale, your hiring process gets farther away from the work,” Hymel explains. 

Especially when the hiring process starts in HR, not engineering, “you’re needing to put language in place that will survive the chain of custody, with the naming of the job [coming from] the engineer that’s closest to the work.”

It’s not uncommon for an enterprise to have 50 different front-end developer job listings, each with very different skill requirements. It’s better for candidates and for fit to be as specific as possible, including embracing new job titles.

This habit of generic job titles used to be positive because it brought in more applicants, but nowadays, with so many engineers on the market, it further dilutes your hiring pool, leaving you with the 100 fastest applicants—who are often AI-generated anyway.

“Any company that has not taken a hard look at revising their job postings and job titles is at an extreme disadvantage because there’s a very high probability that you’re going to end up hiring the wrong person simply because you didn’t take the time to describe the role well enough,” Hymel remarks, which leads to dire consequences. 

“There’s potential churn, so you just spend all this time and cost to go headhunt and find someone. Two, if they do get in there, you have to pay for their ramp time to get up to speed because they were sold a different bill of goods than what was in the description. And then three, it impacts overall roadmaps and timelines because now you might have to replace, and, again, you have to wait for people to get up to speed.”

On top of this, HR and engineering hiring managers alike are using AI to generate job descriptions. It still isn’t recommended to have AI generate something so human and essential to your core success.

Especially in this time of flux, when no one may have the required experience, companies should start job descriptions with what they want the future hire to achieve.

“The cost of code is going nearer to zero.”

“The cost of code is going nearer to zero.” Hymel explains organizations should think more like, “Here are the outcomes that we’re looking for. If you have the soft skills and additional skills around it to get there, whether that is backlog prioritization, being able to be collaborative, having worked on project deployments before, and we don’t necessarily care that you can score a 10 out of 10 on Python anymore.”

Which emerging roles engineers should pursue

The familiar claim that women apply only when they meet every qualification is not well supported; recent research finds that application behavior is more complicated. Still, clearly separating essential qualifications from preferences can reduce ambiguity and unnecessary barriers.

Focusing on outcomes and clearly distinguishing required from preferred skills may broaden the applicant pool, although it does not guarantee greater diversity.

For example, if you’re an engineer who enjoys having a product focus, collaboration, and strategy, Hymel recommends looking toward the new product front-end engineer role, which owns the full user-facing feature lifecycle, from definition to shipping.

“You are required to have more mindshare towards prioritization of features,” he says, shifting away from a ticket person, because “now you have more control because AI allows you to span out a little bit deeper.”

Similarly, AI has the back-end engineer thinking beyond the back-end stack to deployments, scalability, and the reliability of underlying infrastructure systems, giving rise to roles like the polyglot back-end integration engineer. 

Technical writers — reasonably worried about their jobs in the face of AI-generated documentation — should look toward new docs-as-code engineer positions, which add technical program management and DevOps engineering skills.

“If you’re writing the docs, you’re essentially writing the specs that enable spec-driven development. You now have the capability to actually contribute software,” Hymel observes. “And it starts all the way at the top too. If you’re a product manager, you can now start building and contributing code, like a product experience designer.”

Read the full Emergent Role Research. If any of these AI engineering job descriptions ring truer than what you were hired for, we hope it empowers your next conversation with HR or for you to apply for a different job title. 

The post 47,000 job listings reveal the engineering roles that AI is creating appeared first on The New Stack.

  •  

Anthropic promised 20x more usage. Then developers hit a weekly ceiling.

abstract ceiling

Anthropic sells its top-tier Claude Max subscription with the promise of 20 times more usage than its $20-a-month Pro plan. But developers paying $200 a month can still hit a separate weekly ceiling — and an expanded class-action lawsuit filed Tuesday argues Anthropic didn’t make that clear enough.

The dispute over Anthropic’s $100 Max 5x and $200 Max 20x tiers points to a bigger problem within the AI sector: Companies are trying to package unpredictable amounts of compute into straightforward monthly subscriptions. Those plans get harder to understand when the advertised usage comes with additional restrictions. If the plaintiffs succeed, the case could set a precedent for how clearly AI providers have to explain those restrictions before developers sign up.

Companies are trying to package unpredictable amounts of compute into straightforward monthly subscriptions.

The anatomy of “20x”

Anthropic’s documentation states that its Max 5x and 20x multipliers apply “per session,” with usage limits resetting every five hours. But the five-hour window isn’t the only limit since Anthropic also imposes a weekly usage cap across all models and says it may add other restrictions to manage capacity. That’s pertinent for Claude Code users whose interactive coding sessions count toward the same plan limits.

Once that included usage runs out, developers either have to wait for it to reset or pay more to keep working, which is where the usage promise gets murky. It tells you how Max compares with Pro, not how much actual coding a developer can expect for $200 a month.

“Marketing AI subscriptions using simple multipliers like ‘20x’ fails to account for the stochastic nature of agentic coding,” Sajid Afridi, CTO of Pakistan Red Team and an enterprise systems architect, tells The New Stack. “When a single autonomous debugging loop can consume millions of tokens via context re-submission and tool calling, a subscriber can hit their entire weekly quota in a couple of intense sprints.”

That variability also creates a gap between what the multiplier may suggest to a buyer and what it delivers in practice.

“If I see a SaaS product advertised as offering ‘5x’ or ‘20x’ more usage, my natural interpretation as a buyer would be that I can do roughly five to twenty times as much work,” Manish Jain, founder and principal analyst at Strategic Horizon, tells The New Stack. But with AI, he said, capacity can depend on the model, context and workload, making it difficult to know what those higher limits will translate to for a particular developer.

According to the complaint, which The Verge first reported, Anthropic introduced weekly limits in late July 2025 — months after Max launched in April — while continuing to market the plans with its 5x and 20x usage claims. The plaintiffs allege those constraints weren’t adequately disclosed during subscription. Anthropic has pushed back on that characterization, arguing in its motion to dismiss an earlier version of the case that customers could access information about the limits through hyperlinks during purchase. The company compared those disclosures to the information on a product label, which customers can find by turning over the package before deciding whether to buy it.

Fixed-price compute’s dilemma

And yet, Claude Max isn’t the only subscription where it’s difficult to know exactly how much work you’re getting for the monthly price. That’s partially because coding tasks can vary so much. While a quick fix might not make much of a difference to a developer’s allowance, a more complicated job can burn through it much faster. A usage multiplier doesn’t tell developers much about that difference.

How competitors price uncertainty

OpenAI has to account for the same variability with Codex. Its documentation tells subscribers that usage depends on the task, the model, and where the work is being run. Codex can also draw from a shared allowance with other agentic products. Once that allowance runs out, users can buy more credits. OpenAI has also made Codex more autonomous, allowing it to keep working while it waits for a developer to respond. And the model a developer chooses makes a difference, too: Astra costs 2.5x more per token than GPT-5.6, so that the same allowance can go a lot further with one model than another.

The mechanics aren’t identical to Claude Max, and the lawsuit does not accuse OpenAI or other providers of wrongdoing, but both systems show why comparing AI coding subscriptions isn’t as simple as looking at their monthly prices. Saying a plan offers substantially more usage still doesn’t tell a team how much work it can actually get done before hitting a limit.

And some AI companies are already experimenting with different ways to charge for that work. OpenAI has been testing outcome-based pricing with some enterprise customers, charging only when an agent completes a task rather than metering raw compute. That approach introduces its own problems since someone has to define “success,” and failed agent runs become the provider’s expense. But it does show that AI companies are recognizing that token- and multiplier-based pricing doesn’t always tell developers what’s in a subscription.

Saying a plan offers substantially more usage still doesn’t tell a team how much work it can actually get done before hitting a limit.

A warning shot for subscription

If the plaintiffs succeed, other AI providers may have to rethink how they describe their subscriptions. More detail at signup would help, particularly around when usage resets and what restrictions apply. But even that doesn’t answer what developers really want to know: How much work can I get done for the price I’m paying?

As coding agents take on larger jobs and work on their own for longer stretches, the Anthropic case could help establish how much providers need to disclose when selling subscriptions whose actual value can vary so much from one workload to the next.

The Anthropic case could help establish how much providers need to disclose when selling subscriptions whose actual value can vary so much from one workload to the next.

The New Stack reached out to Anthropic for comment on the lawsuit and its Claude Max usage policies and has not received a response. We will update this story if we hear back.

The post Anthropic promised 20x more usage. Then developers hit a weekly ceiling. appeared first on The New Stack.

  •  

OpenAI’s new model costs 2.5x more per token — and developers are saving money anyway

abstract gears

GPT-6 Astra has an obvious problem for developers considering an upgrade from GPT-5.6 Sol: its tokens cost 2.5 times as much, yet OpenAI thinks many developers should upgrade anyway — and then turn the reasoning setting down.

OpenAI’s Thibault Sottiaux, engineering lead for Codex, posted on X over the weekend saying, “To calibrate you all on which reasoning effort to use for Astra, know that GPT-6 Astra on low performs better than GPT-5.6 Sol on high.”

Artificial Analysis currently scores Astra-low at 49 on its Intelligence Index, narrowly ahead of Sol-high at 48. Astra-low also responded much sooner, with the first token arriving in 2.53 seconds, compared to 11.87 seconds for Sol-high.

“To calibrate you all on which reasoning effort to use for Astra, know that GPT-6 Astra on low performs better than GPT-5.6 Sol on high.”

Reasoning effort changes cost

Astra costs $10 per million input tokens and $50 per million output tokens, compared with $4 and $20 for Sol at current rates. Moving the reasoning dial from high to low doesn’t change those rates, but it can change how much work gets done before the task is finally complete and it’s an argument OpenAI makes explicitly in its migration guidance. The company says Astra can produce stronger results while using substantially fewer output tokens.

OpenAI uses its own benchmark as proof. The company saw the same pattern in Terminal-Bench 4.0, where Astra scored 57.9% to Sol’s 37.3% but still cost about 9% less per task. The gap was even wider on GPQA Diamond: Astra edged out Sol, 94.9% to 94.6%, at an estimated cost 37% lower.

Real-world workloads will vary, but the results suggest that per-token pricing alone does not tell developers what a model will actually cost to run — an idea OpenAI is already exploring with outcome-based pricing.

Fewer tokens, cheaper tasks

Developer Shinpr ran the comparison on the same codebase, using Sol-high and several Astra reasoning levels, with analysis, implementation, and review.

Astra-medium did better on both time and cost by handling the implementation in 80 requests, less than a third of the 238 Sol-high needed, and processed 11.1 million input tokens instead of 37.8 million. By the end of all three phases, the Astra-medium run had taken about 51 minutes and cost an estimated $25.67; the Sol-high run had taken roughly 75 minutes and cost $31.79.

Turning Astra up to high didn’t help because that run stretched to 77 minutes and $37.23, and Shinpr said its review missed a startup bug that medium caught.

One developer’s test can’t tell us that medium will be the right choice for every workload, but it does indicate that more reasoning wasn’t worth paying for here. Also, that pattern doesn’t hold everywhere. ARC Prize’s testing went in almost the opposite direction.

Turning Astra up to high didn’t help because that run stretched to 77 minutes and $37.23, and Shinpr said its review missed a startup bug that medium caught.

More reasoning, lower bills

ARC Prize’s evaluation of Astra shows the other side of the equation. More reasoning not only improved Astra’s score on ARC-AGI-3; in some cases, it also reduced costs.

With ARC Prize’s standard harness, Astra scored 17.5% at low reasoning, 38.6% at medium, 54.8% at high, and 62.7% at max. Astra also has an xhigh setting between high and max. But the most expensive runs weren’t the ones using the most reasoning. ARC Prize spent $38,166 at low, $48,090 at medium, and $40,705 at high. Max came in at just $26,098.

At max, Astra used more compute on each decision but needed fewer actions to solve the environments. That tradeoff was enough to bring the overall cost down, which can happen with agents. While lower reasoning may look cheaper, a wrong turn quickly means another tool call or another attempt after another, which costs more in reasoning upfront than just fixing the mistakes later.

While lower reasoning may look cheaper, a wrong turn quickly means another tool call or another attempt after another, which costs more in reasoning upfront than just fixing the mistakes later.

Dynamic reasoning without cache loss

Those are two extremes that developers don’t need to choose between for their entire workflow, thanks to Astra introducing a configuration_update mechanism that lets an application change the reasoning effort between responses without changing the original request-level configuration.

Routine work can remain at a low reasoning level, while a failed test, an unexpected tool response, or a difficult debugging problem can trigger a higher reasoning level for the next turn. Once that’s resolved, the agent can drop back down.

That also helps explain why Shinpr and ARC Prize got such different results. Shinpr found that extra reasoning added time and cost without improving the outcome, while ARC Prize found that more reasoning sometimes cut the number of actions enough to lower the total bill.

For now, configuration_update only works with Astra in standard, single-agent requests. But, don’t get hung up on Astra’s 2.5x token price. An expensive model is often the cheaper run if it takes fewer calls to finish the job and fewer attempts to get it right.

The post OpenAI’s new model costs 2.5x more per token — and developers are saving money anyway appeared first on The New Stack.

  •  

OpenAI will sell you Astra, but not the system that scored 98.6% on ARC-AGI-3

Investor Matt Turck, whose fantastic podcast has hosted the people who built ARC-AGI, summed up Astra’s blockbuster benchmarks with three words: “This is wild.” Then he added four more in parentheses: “w/ its native harness.”

ARC Prize ran GPT-6 Astra through its own standard harness, and the model scored 62.7%. When run through OpenAI’s Provider Adapter, the same model in the same reasoning setting scored 98.6%. The model didn’t change, but the software around it did, adding 36 points to the score. And the better-performing system cost less: $17,332 with OpenAI’s adapter versus $26,098 with ARC Prize’s.

That’s why harness engineering is becoming as important as model selection. The software around the model can change both what it accomplishes and what it costs.

The benchmark measures the system, not the model

ARC-AGI was built to resist the brute-force scaling that eats so many other benchmarks. ARC-AGI-3 raised the bar again this year by dropping a model into interactive environments with no instructions, no stated goal, and no stated rules, then scoring how efficiently it learns to operate. When ARC Prize launched it this year, humans scored 100%. Frontier AI scored 0.51%.

On Thursday, Amanda Caswell covered OpenAI’s improving score on the ARC-AGI-3 benchmark. As she noted, Astra ran under different settings from competing models. ARC Prize is specific about what those settings do. Its standard harness lets a model carry forward notes it chooses to keep. OpenAI’s adapter preserves the opaque reasoning state between requests and compresses longer conversations, so the model can resume its own thinking instead of reconstructing it. 

ARC Prize published every reasoning level.

Same model, two harnesses

GPT-6 Astra on ARC-AGI-3, by reasoning effort. At every setting, the run inside OpenAI’s Provider Adapter scored higher and cost less than the same model inside ARC Prize’s standard harness.

Reasoning effort ARC Prize standard harness OpenAI Provider Adapter
Max 62.7% for $26,098 98.6% for $17,332
XHigh 59.3% for $37,317 98.4% for $18,147
High 54.8% for $40,705 99.9% for $18,817
Medium 38.6% for $48,090 98.4% for $19,285
Low 17.5% for $38,166 98.0% for $21,298
None 35.2% for $49,791 96.7% for $23,457

Source: ARC Prize.

Astra inside OpenAI’s harness with no reasoning effort at all scored 96.7% for $23,457. The same model at max reasoning within ARC Prize’s standard harness scored 62.7% and cost $26,098. The harness beat the reasoning dial outright. I’ve been arguing the harness matters for months. I didn’t expect the result to be this lopsided.

The score wasn’t the only gap. Across the 167 game-reasoning pairs both harnesses solved, ARC Prize clocked the Provider Adapter runs at 49% fewer tokens and roughly 3.66x faster.

OpenAI isn’t hiding the details: The adapter runs on documented Responses API capabilities anyone can call. What you can’t buy is the assembled system that scored 98.6%.

That’s also why OpenAI President Greg Brockman’s claim during a press briefing this week — “I think it’s not unreasonable to feel that we are now in the AGI era” — lands harder than it should. Brockman is describing a benchmark result produced by a particular system, not establishing that the underlying model is AGI. Frederic Lardinois’ launch coverage for The New Stack gets the distinction right: The framing goes well beyond what the evidence establishes. Maybe we’re in the AGI era. But this benchmark doesn’t prove it.

The harness is becoming the product

On coding, the frontier models now cluster inside a few points. Artificial Analysis scores its Coding Agent Index by running each model inside a harness rather than on its own: Astra in Codex at 67, Opus 5 and Fable 5 in Claude Code at roughly the same, Muse Spark 1.3 in Muse Code alongside them, and Fable 5.1 in Claude Code leading at 70. The unit being measured is already the pair. 

The labs already figured this out. Back in April, Janakiram MSV documented the four-way split: Anthropic, OpenAI, Google, and Microsoft all treat the harness as a product to sell, and disagree only on how to charge for it. Anthropic meters Managed Agents at $0.08 per session hour, in addition to token rates. OpenAI gave its Agents SDK away with no runtime fee at all. Google and Microsoft bill sessions, memory, code execution, and observability as separate line items.

Nobody is treating the harness as a free accessory to the model.

Other companies are moving the same way. Stripe paid a reported $8 billion for OpenRouter in August, as Paul Sawers reported for TNS, acquiring a gateway that routes 10 trillion tokens a day across more than 400 models for 10 million developers. Patrick Collison’s framing was that tokens are the central currency for companies building with AI. Stripe bought the routing layer that sits in front of the models. Nvidia built a harness of its own.

We covered the proof two weeks ago. Adrian Bridgwater reported that Claude Opus 5 scores 30.2% on ARC-AGI-3’s public set on its own. Wrapped in Nvidia’s AVO, which gives it persistent memory and programmatic supervision that steps in when progress stalls, it cleared all 183 levels across 25 environments. That one isn’t the clean A/B that ARC Prize ran on Astra. Nvidia changed the memory, supervision, and context management at once. More than one variable moved.

Nvidia said it beautifully here: “Model capability matters enormously, but the surrounding system determines how effectively that capability can be converted into sustained autonomous progress.”

Harness engineering is the job

Janakiram MSV found token usage varying 70-fold across Aider, Claude Code, and OpenClaw running an identical model. Cache hit rates swung from about 70% down to 1.5% depending on the serving path. No model choice explains a spread like that.

The work itself is pretty ordinary. Deciding what an agent remembers and what it forgets. What it’s allowed to touch, and when it has to stop and check with a person. Jeremy Daly’s piece on our site is the version with the engineering, and it’s the one to read if you’re the one building.

I argued in June that model triage was the skill worth hiring for. I’d revise that. Picking the model is the easy half, and it gets easier every quarter as the frontier converges. A year from now, I think the people running agents will spend less of the week picking models and more of it building what goes around them. 

The post OpenAI will sell you Astra, but not the system that scored 98.6% on ARC-AGI-3 appeared first on The New Stack.

  •  

OpenAI spends $1 billion to expand Daybreak to defend power, water, and banking

Electrical substation equipment and power lines in warm evening light, with a transmission tower in the background. 4:20 PM

In a livestreamed keynote on Thursday, OpenAI president Greg Brockman announced Daybreak for Frontline Defenders, a new global initiative to help frontline defenders use frontier cyber AI to protect essential services. 

The initiative is an expansion of OpenAI’s existing Daybreak, which the AI company describes as “a governed cyber defense stack,” comprising frontier models, the Codex harness, Codex Security, trusted workflows, and ecosystem partners. The goal is to help cyber defenders better manage cyber risks by bringing governed defensive capabilities via Daybreak into their existing tools and workflows.

Daybreak for Frontline Defenders builds on that effort by expanding subsidized access and support for frontline cyber defenders tasked with protecting critical infrastructure and services. 

What OpenAI is offering

With Daybreak for Frontline Defenders, OpenAI is continuing its $1 billion commitment to expand subsidized access to Daybreak cyber models, training, technical support, and partnerships. The initiative will support frontline defenders in both the United States and around the world.

Stateside, the project includes Daybreak for America to help protect critical systems, like those used for water, electricity, local government, and banking. This includes launching a pilot with the Multi-State Information Sharing and Analysis Center (MS-ISAC), a cybersecurity partner for U.S. government organizations. Together, the two organizations will work to train and support state, local, tribal, and territorial cyber defenders. 

Plus, OpenAI’s initiative builds on the Daybreak Defense Network, an ecosystem of more than 350 enterprise products and partner-operated services.

OpenAI continues its broader cyber defense push 

A spokesperson for OpenAI tells The New Stack that its new Daybreak for Frontline Defenders “build[s] on a broader push to get advanced cyber capabilities into defenders’ hands.” 

It’s certainly doing the legwork. 

Daybreak for Frontline Defenders “build[s] on a broader push to get advanced cyber capabilities into defenders’ hands.” 

An OpenAI spokesperson also told The New Stack this week that the AI company rallied utility companies across 40 states and the District of Columbia to help them use OpenAI tools to harden cyber systems.

Before that, it joined more than 150 organizations across cybersecurity, technology, critical infrastructure, finance, and AI to publish “an open letter for a global surge in cyber defense,” issuing “a call for collective action on cyber defense.” 

OpenAI’s financial support enabled teams “to review code and system configurations, validate findings, develop patches, and confirm fixes without disrupting essential services.”

On Wednesday, Sam Altman was in Chapel Hill, North Carolina, to speak at the G20 Innovation Ministerial on the urgency of amping up cyber protections for critical systems and services: 

“I think some things are going to go very wrong with cybersecurity unless people act quite urgently,” he said, as reported by CNBC.

Things are already going very wrong.

In July, cyber actors targeted U.S. water and wastewater systems, as confirmed by the FBI. In response, OpenAI stepped in to offer affected states and utilities $1 million in no-cost API credits, Daybreak access, and technical assistance.

As an OpenAI spokesperson tells The New Stack that OpenAI’s financial support enabled teams “to review code and system configurations, validate findings, develop patches, and confirm fixes without disrupting essential services.”

New models, more threats — more urgency

Daybreak for Frontline Defenders builds on OpenAI’s recent spree of cyber defense efforts. This latest initiative comes after OpenAI already expanded Daybreak in August with the release of two tiers: Daybreak Red and Daybreak Blue.

The former provides access to GPT-5.6 Cyber for advanced security work, like finding zero-days, building exploit chains, and handling other advanced security tasks, while the latter gives approved defenders access to GPT-5.6 Sol for secure code review, malware analysis, incident response, patch validation, and vulnerability discovery.

But OpenAI isn’t the only team focused on steering frontier models for cyber defense. 

As The New Stack reported in May, OpenAI’s Daybreak and Anthropic’s Glasswing, an industry consortium powered by Claude Mythos Preview, have nearly identical benchmarks — and even a few of the same partners. Both initiatives aim to put frontier AI capabilities in the hands of cyber defenders, albeit with different harnesses and deployment models. 

Specifically, OpenAI says it has its eye on helping defenders strengthen existing security workflows, e.g., finding and validating vulnerabilities before software ships; investigating threats, improving detections, and supporting response for defensive operations; and conducting authorized security testing. 

From both teams, the push to help defenders harden cyber defenses against increasing threats is, indeed, becoming more urgent as frontier models advance. Also on Thursday, OpenAI launched GPT-6 Astra, saying it has crossed the Critical cybersecurity threshold in its Preparedness Framework. 

The post OpenAI spends $1 billion to expand Daybreak to defend power, water, and banking appeared first on The New Stack.

  •  

GPT-6 Astra’s score of 98.6% looked like AGI. Then researchers read the fine print.

abstract glass

There was no mistaking the divide in March with the release of ARC-AGI-3. While frontier AI models could do little more than register a sub-1% score, humans were able to navigate their new interactive settings.

OpenAI reports a different story for GPT-6 Astra six months on: 98.6%. Put that against the GPT-5.6 Sol it has superseded, which OpenAI puts at 7.8%, and the improvement is hard to miss.

Then again, one has to consider what ARC-AGI is designed for. The whole point is to put models in uncharted interactive territory where they cannot simply rely on training data to find an answer but must work out the mechanics of the environment themselves. Given what ARC-AGI-3 was built to test, a jump to 98.6% is enormous. But the number comes with an important caveat.

Given what ARC-AGI-3 was built to test, a jump to 98.6% is enormous. But the number comes with an important caveat.

The asterisk on Astra’s 98.6%

Astra was evaluated through the company’s Responses API harness, with two settings changed to better reflect how the model performs in real-world use. OpenAI says those changes weren’t made specifically for ARC-AGI-3, but the other models in its comparison were evaluated using different setups.

ARC-AGI-3 requires a model to find its way through an unfamiliar environment, which means the setup it runs in can affect how well it performs.

The model is only part of the story

The gains aren’t limited to ARC-AGI-3. On FrontierMath Tier 4, the model scored 97.6%, followed by 100% on ExploitBench and 99.2% on SRE-Bench with four attempts. Terminal-Bench Science saw one of the biggest jumps, from 22.4% for GPT-5.6 Sol to 64.6%.

OpenAI warns against rolling those results into a single measure of performance, but the range shows how much more the model can take on. In the company’s demonstrations, it works directly inside software such as KiCad, Power BI and Unity, while an experimental Codex feature lets it keep notes and search earlier context when a job runs longer than a single context window.

On offline OSWorld 2.0, Astra scored 72.6% while taking about 40 minutes per task, compared with Sol’s 65.7% and roughly 75 minutes.

On offline OSWorld 2.0, Astra scored 72.6% while taking about 40 minutes per task, compared with Sol’s 65.7% and roughly 75 minutes.

Beyond benchmarks into discovery

The math is where things get more interesting. OpenAI says Astra was involved in two new findings about gaps between prime numbers. Mathematician Julia Stadlmann had already pushed one bound from 246 to 240. With Astra involved, it fell again, this time to 186. The company points to another case where the model helped improve part of a bound that hadn’t budged in more than 80 years.

There’s an important gap in OpenAI’s account, though. It doesn’t spell out what Astra came up with on its own, what researchers suggested or how the work moved between them. So while this goes beyond solving a benchmark with a known answer, it’s not enough to call the math evidence of AGI.

Alignment gains, oversight gaps

Astra is pushing past many of AI’s familiar limits, from unfamiliar problems to longer tasks. Yet even a 98.6% score on ARC-AGI-3 doesn’t settle the AGI debate. Part of the problem is that performance more often depends on the system around the model. Intelligence itself doesn’t improve evenly, either.

In OpenAI’s internal tests involving difficult or impossible tasks without production safeguards, GPT-5.6 Sol went beyond what it was authorized to do 48.2% of the time. Astra didn’t do so even once. Yet when researchers explicitly asked the models to evade monitoring, Astra’s written reasoning was harder to follow than Sol’s. OpenAI says that’s partly because Astra can solve simpler problems in fewer written steps, although it still struggles to conceal its reasoning on more complicated tasks.

If AGI means doing useful intellectual work across different fields, Astra is getting remarkably close to what many people once had in mind. If it means matching human judgment across the board, ARC-AGI-3 can’t establish that.

Epoch AI’s Greg Burnham described Astra as the “end of one era, start of another.”

“…end of one era, start of another.”

[Editor’s note: This article’s headline has been updated to clarify that a 98.6% score on ARC-AGI-3 does not mean the benchmark was “aced.” ARC-AGI-3 scores performance against a human baseline, with 100% representing performance at or above the median human baseline.]

The post GPT-6 Astra’s score of 98.6% looked like AGI. Then researchers read the fine print. appeared first on The New Stack.

  •  

Your organization prioritized AI adoption, but you actually need AI fluency.

Abstract digital neural network visualization representing central hub-and-spoke enterprise AI infrastructure.

Thanks to increasingly capable models, some parts of your business are getting faster, more capable, and more productive every month. These teams are using artificial intelligence to compress timelines, surface insights, and automate work that has historically been time-consuming and tedious.

Meanwhile, other functions just down the hall are still waiting for a formal rollout, a governance approval, or someone to tell them what to do and how to start. The gap between the AI haves and have-nots in your organization is widening, and addressing it requires a new operating model.

Your teams need more support

When leaders notice the uneven distribution of capability across their business, the instinct is to treat it as a tooling problem. They push to get everyone access to the same platforms, provide general-use training, and hire some specialists to slot into IT.

But access is table stakes. It’s a good start, but it won’t get you to strong organizational adoption.

Harvard Business School reports that workers using these tools completed tasks 25% faster and produced results rated more than 40% higher in quality. But the same study also found that performance declined when people used the tools without understanding where they applied and where they didn’t. Fluency, not just access, drives results.

Departmental leaders need guidance on how to apply capabilities in the context of their day-to-day work. Without that knowledge, they can’t ask the right questions.

“Performance declined when people used the tools without understanding where they applied and where they didn’t. Fluency, not just access, drives results.”

Teams playing catch-up tend to focus on how to inject new tools into existing workflows, when they should be thinking about re-engineering processes entirely. They’re focused on evolution in a world undergoing revolution.

Reimagining a process also requires stepping back from it, which is easier said than done. Here’s how it plays out in practice:

An SDR team comes to IT with a specific, bounded ask: “improve our sales lead routing.” Completely reasonable. But only when someone from IT, with visibility across the broader system, dives into the problem does the real opportunity surface. The data pipeline supporting lead routing is unnecessarily complex. With the right support, the conversation shifts to overhauling the entire pipeline and opens the door to fully agentic lead follow-ups.

Departmental leaders don’t lack ambition but throwing a software license and Slack channel at them won’t build the right kind of adoption. Technical support and strategic guidance are required to reimagine work from first principles.

AI fluency must be a structural consideration

The typical pattern puts a centralized team in charge of taking requirements, interpreting them in isolation, and delivering capabilities to departments months later. This model can’t keep pace when AI capabilities launch weekly.

A more effective approach pairs a central “hub” that owns platform strategy, governance, and reusable patterns with AI engineers embedded directly inside business departments. AI engineers serve as “spokes” inside departments, helping them identify vertical use cases day-to-day and delivering the cross-functional visibility needed to make a real impact. The AI engineer who solved a problem for finance can share the pattern with someone facing the same challenge in operations.

“A more effective approach pairs a central “hub” with AI engineers embedded directly inside business departments.”

In a department just getting started, the embedded AI engineer is the primary technical capability: scouting, prototyping, building. In a more mature department, they shift toward enablement, feeding patterns back to the “hub” and helping teams navigate AI without getting buried in process. Over time, departments will organically become AI-fluent as they learn from the engineers.

Make fluency your advantage

The right operating model drives how a function actually works, and strong fluency strengthens processes and institutional knowledge, so outcomes improve over time. As the flywheel builds, each problem solved raises the ceiling of what your team can do independently. 

McKinsey finds that the right workflow redesign is the single biggest factor in whether an enterprise sees meaningful bottom-line impact. Knowing what to redesign depends on how your teams understand and work with AI.

Everyone is adopting AI capabilities. The question now is whether your operating model helps your teams see the best path forward for applying them. If it doesn’t, that’s the gap to close first.

The post Your organization prioritized AI adoption, but you actually need AI fluency. appeared first on The New Stack.

  •  

SpaceX is in an “enviable position”: why Anthropic is sticking with Cursor as OpenAI cuts access

Illustration of a human hand in a business suit shaking hands with a white robotic hand, set against a blue background.

OpenAI caused something of a stir over the weekend when it announced plans to cut Cursor’s direct access to OpenAI models in November.

The reason? Elon Musk.

In a statement issued late on Friday, OpenAI pointed to two previous incidents involving Musk’s companies: Twitter breaking the terms of a data-licensing deal after Musk’s 2022 takeover of the social network, and Musk’s admission under oath earlier this year that xAI had partly used OpenAI models through distillation — conduct OpenAI says violated its terms of service. And now that SpaceX’s $60 billion deal to acquire Cursor has closed, OpenAI’s attentions are turning to Cursor.

“This decision was incredibly tough, as we care deeply about our models being broadly available for developers,” the company wrote. “We are making this choice because we cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk’s companies violating contracts.”

While that in itself was big news for anyone following the day-to-day rough and tumble of the AI industry, what was particularly notable was the response of OpenAI’s arch rival.

The Anthropic factor

As The New Stack noted in its coverage, Anthropic co-founder and “chief compute officer” Tom Brown moved fast, posting publicly within hours of OpenAI’s statement to confirm that it continues to see Cursor as a “trusted partner,” and will “continue to increase compute to support Claude models in Cursor.”

Cursor has been a trusted partner of Anthropic since Sonnet 3.5. We’ll continue to increase compute to support Claude models in Cursor and are excited for what comes next with them at SpaceX.

— Tom Brown (@NotTomBrown) August 29, 2026

But anyone who has followed Anthropic’s recent history could be forgiven for wondering why.

Take Windsurf. In May 2025, reports emerged that OpenAI was in talks to buy the AI coding tool for $3 billion. Anthropic didn’t hang around for the deal to close though, and within weeks, Windsurf said Anthropic had cut its direct access to Claude 3.5 Sonnet and Claude 3.7 Sonnet.

Speaking at an event hosted by TechCrunch shortly after, Anthropic co-founder Jared Kaplan said that “it would be odd” for Anthropic to be selling Claude to OpenAI. As things transpired, the OpenAI deal fell through, and Google stepped in instead, paying $2.4 billion to hire Windsurf’s founders and R&D staff into DeepMind.

“They cut them off ruthlessly when RUMORS of OpenAI potentially buying Windsurf surfaced.”

In a social media post published on Sunday, Gergely Orosz, engineer and author of the Pragmatic Engineer newsletter, is quick to highlight the episode with Windsurf, which he noted had also been a “trusted partner” with Anthropic for some time. “They cut them off ruthlessly when RUMORS of OpenAI potentially buying Windsurf surfaced,” Orosz writes. “Now, SpaceX, an Ant(hropic) competitor bought Cursor — it’s still a trusted partner?”

Then there’s xAI, the AI company Musk founded in 2023 to build Grok, which SpaceX acquired outright in a February transaction valuing xAI at $250 billion. In January this year, Kylie Robison reported that Anthropic had cut xAI staff off from Claude, which they’d been accessing through Cursor — with Cursor reportedly telling xAI it was “a new policy anthropic is enforcing for all its major competitors.”

So Anthropic has previous form for moving fast, on rumor alone in Windsurf’s case, whenever a customer starts to resemble a competitor. Which makes this week’s public vote of confidence for a company now wholly owned by one of Anthropic’s actual rivals worth a second look.

The compute dependency

What makes SpaceX different from the previous Windsurf and xAI episodes is that Anthropic is also buying a huge amount of compute from it.

On May 6, Anthropic announced it had secured the entire output of SpaceX’s Colossus 1 data center near Memphis, Tennessee — more than 300 megawatts of compute and over 220,000 Nvidia GPUs. Anthropic had been hampered by limited compute availability, and the SpaceX deal, alongside other recent compute agreements, let it raise usage limits for Claude Pro and Max subscribers almost overnight.

Two weeks later, the financial terms came out. Per SpaceX’s IPO filing, Anthropic agreed to pay $1.25 billion a month for compute across Colossus and Colossus II — about 325,000 Nvidia GPUs combined — scheduled to run through May 2029, subject to termination rights.

SpaceX, for its part, said the arrangement would allow it to monetize some of its compute capacity while retaining enough to meet its own AI training and inference needs. As The New Stack reported in May, the deal also underscored just how central access to compute had become to competition between the leading AI labs. And it created an unusual commercial relationship: Anthropic was now buying a huge amount of compute from a company that also owned one of its direct AI rivals in xAI.

And that is what makes Anthropic’s response to the Cursor acquisition so notable. In response to Tom Brown’s post on X on Saturday, Replit founder and CEO Amjad Masad points to the contrast between Anthropic’s support for Cursor now and its treatment of Windsurf last year, suggesting the latter had been harsher than OpenAI’s decision to cut Cursor off.

“More likely answer is that you can’t do that here because you need the compute.”

“Maybe you changed your ways, but we all remember what you did to Windsurf, which was infinitely nastier,” Masad writes. “More likely answer is that you can’t do that here because you need the compute.”

Orosz essentially makes the same argument. If Anthropic was prepared to cut access when Windsurf merely looked likely to end up in OpenAI’s hands, why is it publicly promising MORE Claude capacity to Cursor after the company had actually been acquired by SpaceX?

“SpaceX basically in this enviable position where one of its biggest competitors depends on its compute infra!”

“Either SpaceX and Grok are not competitors to Anthropic (they are!); or, more likely, SpaceX leasing its Colossus 1 data center is more important to Anthropic than to stop offering Claude to SpaceX,” Orosz writes. “SpaceX basically in this enviable position where one of its biggest competitors depends on its compute infra!”

And so this effectively highlights how strong a position SpaceX finds itself in. It now owns xAI and Cursor, putting it in direct competition with Anthropic in foundation models through Grok and in AI coding tools through Cursor, while Anthropic is simultaneously paying it billions of dollars for compute capacity supporting Claude.

Whatever the stated rationale for treating Cursor as a “trusted partner,” that relationship leaves SpaceX with something neither Windsurf nor xAI had at the time Anthropic moved against them: a source of leverage over any decision on whether or not to cut access to Claude.

The post SpaceX is in an “enviable position”: why Anthropic is sticking with Cursor as OpenAI cuts access appeared first on The New Stack.

  •  

OpenAI wants to charge only when AI gets it right — here’s the catch

abstract yellow grade

AI companies have always charged customers for the tokens they use, whether the model gives them exactly what they need or completely misses the mark. Now, OpenAI is experimenting with a different approach by allowing customers to pay only when the AI gets the job done right.

First reported by The Information, the company has already started testing that approach with some enterprise customers, billing them based on successful outcomes rather than simply the amount of compute consumed along the way.

OpenAI has not publicly disclosed pricing or exactly how it determines when a particular task counts as a success. Still, the approach presents an interesting technical challenge for developers building AI agents. If a customer only pays when an agent completes a task, someone needs to determine exactly when that task is complete. As we know, determining if an AI agent succeeded is more complicated than counting tokens.

Counting tokens, not success

Some outcomes are easy for software to verify, such as a support ticket closing without human involvement, but other jobs leave room for interpretation. For example, take a coding agent asking to fix an authentication bug. It might rewrite the code and pass every test, only for the patch to cause another problem once it reaches production. In that case, the agent technically completed the task, but the customer probably wouldn’t consider it a successful outcome. 

With outcome pricing, getting through 90% of a task may not be enough for the run to count as billable.

The same issue arises with longer-running agents that interact with browsers, databases, APIs, and other systems. An agent may carry out nine of the ten steps before failing at the last one. With token pricing, all of that activity can still be charged for. With outcome pricing, completing 90% of a task may not be enough for the run to be billable.

Evals become billing infrastructure

Developers already use evals to catch problems with models and agents. OpenAI’s hosted tools, for example, can check responses against expected results and grade a model’s performance on a given task.

Braintrust adds visibility into what happens during an agent run. It records model calls, retrievals, and tool calls in a trace, then scores the run on factors such as task completion, factual accuracy, and correct tool use. Developers can also turn those traces into datasets for future testing.

Some results are easy: a unit test either passed or it didn’t, an API returned the expected response, or a database contains the record it was supposed to create. There isn’t much room for debate. 

It records model calls, retrievals, and tool calls within a trace, then scores the run for things like task completion, factual accuracy, and correct tool use.

Semantic evals are different because they require a judgment rather than a pass/fail. An LLM-as-a-judge can help developers compare two versions of an agent, but using that judgment to trigger a charge is another matter. A false positive could leave a customer paying for unfinished work, while a false negative could leave the vendor covering the cost of a successful run.

Grading your own work

When an AI company runs the agent and sets the criteria for success, it is effectively grading its own work and then billing the customer for the result.

That gets trickier with more subjective work. Asking an agent to generate a monthly sales report gives you something you can check. But asking it to generate a good one is different because someone still has to decide whether the result is actually any good. And if the goal is to improve conversion rates, figuring out how much of that improvement is attributable to the agent is even harder. 

There’s also the question of who gets blamed when something outside the agent fails. An agent might handle a support request correctly only to hit a timeout in the customer’s CRM. A coding agent could finish its work but fail because a separate service is unavailable. If those runs don’t count as successful, the vendor could end up paying for failures it didn’t cause.

Failed agents shrink margins

Outcome pricing changes who pays when an agent fails. Under token pricing, an agent can burn through tokens and retry failed steps without ever finishing the job. The customer still pays for that usage.

If it never finishes, the provider has spent money on compute without anything to bill.

If the customer pays only for successful outcomes, failed runs become the provider’s expense. An agent that completes a task on its first attempt is more profitable than one that requires 20 model calls and several retries. If it never finishes, the provider has spent money on compute resources without anything to bill for.

The post OpenAI wants to charge only when AI gets it right — here’s the catch appeared first on The New Stack.

  •  

AI agents are making retrieval engineering a core engineering discipline

An abstract digital image featuring an intricate network of glowing copper-orange and deep red strands swirling against a dark background, visualizing complex data pathways and retrieval engineering workflows for AI agents.

AI agents are changing retrieval requirements. As organizations move from chatbots to AI systems that investigate, reason, and act on users’ behalf, retrieval is becoming the foundation of application quality. Better retrieval doesn’t just produce better answers—it enables more capable assistants, more personalized experiences, and more trustworthy autonomous systems.

Traditional search and even many RAG applications could tolerate imperfect retrieval. If a user didn’t find exactly what they wanted, they refined the query or tried again. Agents don’t have that luxury.

“As organizations move from chatbots to AI systems that investigate, reason, and act on users’ behalf, retrieval is becoming the foundation of application quality.”

An AI agent plans, reasons, invokes tools, and increasingly makes decisions without a human reviewing every intermediate step. That raises the bar considerably. Retrieval is no longer about finding relevant information—it’s about consistently delivering the right evidence at the right time.

For engineers, this creates a familiar set of challenges:

  • Which signals matter most for this user?
  • How do fresh events change relevance?
  • How should structured, unstructured, and behavioral signals be combined?
  • When should a model influence ranking?
  • How do you optimize for business outcomes rather than similarity scores?

Those aren’t vector database problems. They’re Retrieval Engineering problems.

It’s no longer just about embeddings or vector search. It’s about engineering the entire retrieval workflow: combining hybrid retrieval, real-time signals, ranking, machine learning inference, and continuous experimentation to deliver the best possible decision at serving time.

A recent GigaOm Decision Brief argues that as retrieval becomes increasingly commoditized, competitive advantage shifts to decisioning—determining what an application or AI agent should see, and in what order, before it acts.

“As retrieval becomes increasingly commoditized, competitive advantage shifts to decisioning—determining what an application or AI agent should see, and in what order, before it acts.”

That aligns closely with the way we’ve been thinking about Retrieval Engineering:

Prompt engineering influences how a model reasons. Retrieval Engineering determines what it has to reason about.

As organizations move from copilots to production AI agents, I believe Retrieval Engineering will become a core engineering discipline alongside prompt engineering and model engineering.

If you’re interested in the engineering discipline itself, my earlier article explores Retrieval Engineering in more depth (https://thenewstack.io/ai-retrieval-engineering-bottleneck/). The new GigaOm paper complements that discussion by looking at why these engineering decisions increasingly influence product quality, customer experience, and ultimately business outcomes.

Read the GigaOm report. 

The post AI agents are making retrieval engineering a core engineering discipline appeared first on The New Stack.

  •  

Nvidia is paying $12.9 billion to keep open models on its chips 

I’m Matt Burns, Chief Content Officer at Insight Media Group. Each week, I round up the most important AI developments, explaining what they mean for people and organizations putting this technology to work. The thesis is simple: workers who learn to use AI will define the next era of their industries, and this newsletter is here to help you be one of them.


Nvidia agreed to pay $12.9 billion for Hugging Face this week, The Information reported. The purchase feels similar to when Microsoft bought GitHub. 

Nvidia followed developers. It bought the place they already go for AI models. Meanwhile, Anthropic and OpenAI are selling a premium into a market that’s increasingly treating free as the default.

It’s getting easier to run those open models, too. Earlier this week, Ollama shipped a release that lets Claude Desktop run Qwen, DeepSeek, and Kimi. You open Anthropic’s app, click the model picker, and pick Kimi K3 instead of Opus 5. 

Open models moved into the tools developers already use

Our own Paul Sawers covered the Ollama release for The New Stack. Version 0.33.0 launched on August 21 and uses a local proxy to get around Claude Desktop’s block on non-Anthropic model IDs. Flip “Use Ollama models” in the Mac menu bar, and whatever you’re running, locally or on Ollama Cloud, shows up in Anthropic’s model picker. “Open models should be easy to run, easy to build with, and available wherever people need them,” Ollama CEO Jeffrey Morgan said. The company raised a $65 million Series B in July.

The hardware is catching up, too. Apple launched new Mac Mini and Mac Studio configurations this week that seem purpose-built to run larger models locally. Frederic Lardinois wrote up Alibaba’s Qwen3.8-27B two weeks ago. The 4-bit build is 16.1GB and runs on a Mac with 32GB of unified memory; Alibaba’s own benchmarks put it at or past Opus 4.6 Max for coding and computer use. Those are Alibaba’s benchmarks, and early users say the model overthinks. It still fits on a laptop.

The cost of switching dropped with it. TNS writer Janakirm MSV started covering this in May, when OpenCode passed Claude Code on GitHub stars, from 157,000 to about 122,000. The real decision in front of most developers, he writes, is “whether their environment can tolerate a single-vendor harness at all.” Jason Calacanis put the commercial version on X on Wednesday: “Open source is solving 90% of startup use cases right now, so frugal founders aren’t paying for Fable.” Five hundred dollars per employee per month feels fine, he writes, but “when you start hitting five figures, CFOs start clenching.” Founders are the early signal. CFOs are the broader one.

Nvidia is paying for the place developers get their models

So what does $12.9 billion buy Nvidia? A website with about $150 million in revenue. Nvidia already builds open models of its own and gives them away, so it isn’t short on weights.

Microsoft paid $7.5 billion for GitHub in 2018 because that’s where code lived, and owning it made GitHub the surface for everything Microsoft wanted developers to do next. Hugging Face is that for weights. Every open model release and every fine-tune lands there. It’s the default place developers go to figure out how to run one.

Product analyst Aakash Gupta did the math to explain the price. Nvidia just reported $96.2 billion in quarterly revenue, so $12.9 billion is about 12 days of sales. And Nvidia’s largest customers are all building escape routes: OpenAI designing chips with Broadcom, Anthropic training on Amazon’s Trainium, Google a decade into its own TPUs. 

Open models are the counterweight, because a downloaded model gets fine-tuned and served on Nvidia CUDA by default. As long as developers keep choosing open models, they’re also choosing Nvidia’s software stack. “Nvidia spent 12 days of revenue to make sure the open-source rival to its own customers never dies,” Gupta writes.

It works because open models run on Nvidia hardware. Download Qwen, fine-tune it, and every step runs on CUDA unless you go out of your way. Broadcom and Amazon can build a chip. Neither of them can make ten thousand repos target it.

So the thing you did to get out from under one vendor’s pricing put you further under another’s. 

Hugging Face is also the leaderboard, the datasets, and the transformers library a good chunk of the industry uses. Nvidia will soon own all of it. The stalwart venture capitalist Bill Gurley posted the reason on Wednesday: “Open-models are the inevitable outcome of high stakes software competition,” he writes. “The more at stake, the more likely open wins. This is water running downhill.” 

He made the longer case in The Washington Post in July, under a headline that names the two companies preparing to go public: Open-model AI is good competition for Anthropic and OpenAI.

The obvious objection is that Hugging Face downloads aren’t production traffic, and that most real work still runs against a closed API. It’s fair. But Chinese open models have taken more than 30% of U.S. token usage on OpenRouter every week since February, peaking at 46%. Right now, that’s the corner of the market where switching is cheapest. 

The lesson for AI-native developers isn’t to replace Claude with Qwen or stop paying Anthropic. Closed models will still be the right answer for plenty of work. The lesson is to stop assuming today’s default will still be tomorrow’s.

Every AI-native application has defaults: the model, the registry, the harness, the API. Those become dependencies, and dependencies become leverage. Nvidia reportedly agreed to spend $12.9 billion on the place developers already go to find and fine-tune models.

That changes the job for AI-native developers. Five years ago, the best developers learned Kubernetes. Today, they spend their time comparing benchmark scores and arguing about which model is smartest. And that’s becoming the less important skill. The hard one is building systems that survive when the underlying defaults change. 

Because it will. 

The post Nvidia is paying $12.9 billion to keep open models on its chips  appeared first on The New Stack.

  •  

Google’s new legal AI exposes a bigger battle over the enterprise stack

Google Cloud launched Gemini Enterprise for Legal this week, a purpose-built agentic AI solution to automate legal workflows, including contract review, regulatory monitoring, document drafting, and data discovery. It signals that the next phase of enterprise AI competition will hinge on who can best specialize the stack — not just who can build the strongest foundation model.

The release comes alongside Gemini Enterprise for Financial Services, another agentic AI solution, this time geared towards financial professionals. Together, the two are the first offerings in what Google describes as “a series of specialized, packaged industry solutions built on top of the secure, fully governed Gemini Enterprise platform.”

Gemini Enterprise for Legal came about 24 hours after the debut of Thomson Reuters’ Thomson, its own AI model for legal, tax, and compliance work, which it spent $40 million developing. 

While Gemini Enterprise for Legal and Thomson look similar on the surface — both aim to make AI more useful for legal workflows — they are structurally different. Google’s new offering is an agentic system built around its existing Gemini models. Thomson, on the other hand, is a proprietary model that the company further trained on its own proprietary professional content and input from subject-matter experts. 

Still, in some important ways, the launches are two sides of the same coin. Both show companies are making a push to specialize AI for professional domains — but that specialization can come from different layers of the stack. 

Specialized AI doesn’t have to mean a specialized model

Google and Thomson Reuters are both making moves to build more specialized AI, but they’re attacking the beast from different angles. 

Thomson Reuters made a splash by taking an existing open-source foundation and spending millions to train the model with expert evaluation and decades of content from its own collection, including Westlaw, Practical Law, Checkpoint, and Reuters. The resulting Thomson even beat Gemini 3.1 Pro, Claude Opus 4.8, and GPT-5.5 in some benchmark evaluations. 

Google, meanwhile, is building much of its legal specialization around the model through agents, integrations, tools, and governance, though the company says its solution may also include model optimizations.

Both show companies are making a push to specialize AI for professional domains — but that specialization can come from different layers of the stack. 

Gemini Enterprise for Legal is built on Google’s own AI stack, which spans global infrastructure, custom silicon, foundation models, and an AI-ready data platform. While Thomson Reuters seems to be making the point that having proprietary data and using it to further model training can give it an edge, Google’s approach shows that specialized AI can also emerge by working above the foundation-model layer. 

The solution centers on four core components: 1) purpose-built specialized skills to guide AI agents through legal tasks; 2) secure Model Context Protocol (MCP) integrations to connect with legal industry platforms, like DocuSign; 3) access to a specialized network of third-party agents, legal tech providers, and consulting partners, like Accenture and Deloitte; 4) a control plane for risk management, audit logging, and governance.

With these pieces working together, Google says Gemini Enterprise for Legal can automate and speed up a range of legal operations, such as automating data discovery and Data Subject Access Requests (DSAR) responses, and autonomously tracking legislative updates, court dockets, and supervisory bodies to update policy drafts. It can also speed up contract review and negotiations, draft NDA documents, prepare court filings, redact legal documents, and build and update contracting playbooks.

Owning a model doesn’t necessarily mean going all in on it

A striking point in Thomson Reuters’ Thomson announcement was that the company’s independently trained model outperformed leading models on several professional and general-purpose evaluations. 

But that doesn’t mean the company’s forgoing frontier models altogether. Instead, it’s opting for a pick-and-choose strategy based on which model works best for the task at hand. 

Just look at CoCounsel Legal, Thomson Reuters’ AI assistant, built on Anthropic’s Claude Agent SDK, that helps legal professionals with research, analysis, and drafting. As its first deployment, Thomson will work within the Tabular Analysis product, the document-review tool that analyzes large volumes of documents. When Thomson gives a real advantage over other leading models, CoCounsel Legal will let it take the lead. But when other models are better suited for the task, the AI assistant will route work accordingly.

Model selection is now only one piece of the puzzle

Much of the conversation around adapting AI for domain-specific tasks has looked like a model problem: Take the best general-purpose model available and feed it the right data and context. But these two launches show the picture is getting more complex. 

Moving forward, the competitive advantage may increasingly belong to whoever can bring something uniquely valuable that others can’t easily get their hands on. 

Yes, the model is still a fundamental part of developing specialized AI, but it’s not the only way to differentiate. Moving forward, the competitive advantage may increasingly belong to whoever can bring something uniquely valuable that others can’t easily get their hands on. 

For Google, that’s its integrated AI and cloud stack, which Gemini Enterprise for Legal extends with specialized skills, MCP integrations, governance, and tools. For Thomson Reuters, that edge looks like something different: a proprietary, specialized model built on decades of content and expert knowledge that works inside a multi-model system.

Neither approach does away with the importance of the underlying model. But both suggest that a strong model alone isn’t enough to spearhead specialized AI. For developers, that begs the question: What part of the stack is worth owning? 

The post Google’s new legal AI exposes a bigger battle over the enterprise stack appeared first on The New Stack.

  •