❌

Vue normale

Reçu avant avant-hierThe New Stack

A third option is emerging in the fight over AI and your data

23 septembre 2026 à 21:31
Split-screen video interview with The New Stack host Alex Wilhelm and VAST Data cofounder Jeff Denworth.

Not your keys, not your coins. Not your model, not your data?

Over the summer, the tech industry was consumed by a debate about AI use in the enterprise and the need to protect IP. If an enterprise used proprietary models, was data leakage a necessary evil?

Companies seemed to have two options: They could use state-of-the-art, proprietary models and risk losing control of their data, or they could use open-weight models and never kiss the frontier.

Thankfully, a third option is emerging.

Consider the concern: Company A wants to use LLM B from AI Lab C, and they want to avoid training AI Lab C how to eat Company A’s lunch by building its capabilities into LLM B. A good way to resolve the tension would be to let Company A run LLM B on its own infrastructure, so there’s no risk of its information fleeing on the wind.

AI agents are “creating a whole different set of requirements at the data layer.”
–Vast Data co-founder Jeff Denworth

But that raises another problem: AI Lab C doesn’t want to allow Company A to run LLM B on its own GPUs because it doesn’t want to hand over its model weights. It’s the same IP issue the company ran into, in reverse. You have to solve the trust problem in both directions!

Enter VAST Data co-founder Jeff Denworth and a new product called DataEnclave, which aims to let AI labs and enterprise-scale companies deploy proprietary models in secure compute environments without risking data transfer in either direction. (DataEnclave uses Nvidia’s Confidential Computing technology to make the system tick; Vast Data’s core product is AI OS, infrastructure that fits beneath a company’s AI applications.) 

The New Stack had Denworth on the podcast to chat about the confidential computing market. I was curious about timing. Why did Vast build DataEnclave now? Nvidia began rolling out Confidential Computing in a serious way in 2024, after all. Denworth argues that the market needed the core technology, yes, but also demand.

And until late 2025, AI demand was modest compared to today’s token totals. Once agentic coding tools took off, corporate demand for AI products soared. This led to the pricing crisis we saw in early 2026, and the secure AI usage debate we endured over the summer. 

Performance drove demand, demand drove usage, and usage dug up fresh problems to solve. Now the question for the market is whether or not DataEnclave has solved enough concerns on both sides of the proprietary AI-proprietary data equation. The market will sort that out as it moves through early access and into general availability.

Our conversation goes deep into the arc of AI, where companies are in their AI journey today, and how much data remains to be unlocked inside the enterprise. If you want to feel the acceleration, it’s a fun one!

The post A third option is emerging in the fight over AI and your data appeared first on The New Stack.

AI spending can run negative. Qodo’s CEO built an ROI equation to fix it.

23 septembre 2026 à 15:00
Five stacks of mixed copper, silver, and brass coins arranged left to right in ascending height, like a bar chart, against a neutral beige background.

Flush with the proceeds of a $70 million Series B raised earlier this year, you might expect Qodo to spend freely on internal AI. After all, the startup uses artificial intelligence to ensure AI-generated code meets customer quality and governance requirements. An upstart technology company using AI to improve AI outputs is AI-pilled by definition.

Instead, the company has limits on AI consumption. Qodo CEO Itamar Friedman tells The New Stack that his engineers can access $10,000 worth of tokens per month, a cap that he described as “generous.” Most Qodo developers never reach it. The ceiling wasn’t enacted to “restrict usage,” Friedman says, but instead to drive “visibility and efficiency” at the startup so that it can “scale without runaway costs.” Put another way, the cap exists to make somebody answer this question: “Which path of automation or usage will be the best [use] of our money?”

Qodo’s AI footprint is larger than its developer token budget. The startup’s AI infrastructure spend — the cost of running the product for customers rather than the cost of its own engineers using AI — is growing at “roughly 5x year over year,” the company tells TNS in an email, reflecting both “increased user adoption” and its agents taking on more, and longer tasks as they mature. Qodo says it is pushing the other direction at the same time, driving down the cost of reviewed pull requests through routing and inference efficiency.

What Qodo runs on

The company also dogfoods heavily, running its pull requests through Qodo. Friedman said the product powers its entire software development life cycle (SDLC). Around that sits a stack most engineering organizations would recognize: Slack and Notion and their constituent “bots,” a centralized knowledge base built to be agent-readable, AI inside Google Workspace, and models from several providers including Google.

Qodo’s own product sits alongside Claude Code and other leading coding assistants rather than replacing them. Claude Code still holds the crown internally, but OpenAI’s Codex has been taking share, with staff “shifting quickly towards Codex.” Friedman tracks this two ways. He polls his 130-person staff, spread across offices in several countries, on the tools they prefer, and compares those answers against what the usage data shows.

Friedman reports that Qodo sees “roughly double” the number of PRs “every couple of months” alongside “a decreasing amount of bugs and incidents.” Hold on to those two numbers. They’re important here.

The bottleneck moved

More PRs and fewer bugs indicate that Qodo is onto something with its focus on software testing and governance. Its technology helps developers deal with an increasingly common issue: What do with all the code that AI agents generate? Companies that adopt AI coding tools often find that they create more code with machines than their humans can assess. As a result, the SDLC bottleneck simply shifts one step down the process.

“We solved the speed of writing code,” Friedman argues, “we didn’t solve the velocity of creating software.” The difference between accelerating one part of a task and its entire arc is the difference between AI hype and AI ROI. 

The software development example shows that when we consider AI costs and benefits, we need to think broadly. If we focus too much on a single metric, we might spend our entire budget on Claude Code credits while shipping no more software than before. Alongside a massive bill.

The AI ROI Equation

Friedman recommends an equation-based approach. The Qodo perspective on AI ROI is similar to a popular equation for happiness: Personal joy is the distance between your expectations and reality. The greater the expectations, the harder it is to be happy. The lower the expectations, the greater the chance of being content. 

This can be expressed as either simple subtraction or as a ratio:

  • Reality/expectations = Happiness, where larger results indicate greater joy

Take the same mathematical approach to AI ROI, per Friedman: Compare the positives against the negatives, add up all the good, and set it over all the bad.

  • AI benefits/AI costs = AI ROI, where larger results indicate greater return

Friedman found the shape of the equation in The Phoenix Project, the 2013 DevOps novel that contrasts types of software development work and sorts them into good and bad buckets. Plug those terms in:

  • (Features + Infrastructure)/(Incidents + Bugs) = Software development velocity

Now, those two numbers from earlier. Qodo has seen more PRs and fewer bugs thanks to AI. In DevOps terms, it’s shipping more and fixing less, so the equation returns a larger, better result. Feed the same terms into the AI ROI version, and it produces more benefits over fewer costs, and a larger final calculation. 

The fraction is not a thought experiment at Qodo. It’s the shape of what the company says is already happening to it.

Terms that have nothing to do with software development work too. Qodo runs AI inside Google Workspace, Slack, and Notion, and those benefits and costs go into the same calculation. 

The Qodo approach to measuring total AI ROI is less specific than The Phoenix Project’s DevOps equation, but the difference is acceptable. Friedman argues that you have to start somewhere: “I know [the equation is] a simplification,” the CEO tells TNS. “But what you can’t measure, you can’t improve.”

His argument is that imperfect beats absent. “Don’t think about it too much,” he says. “Try to put any number [in the AI ROI equation] and start tracking.” Being told not to overthink an equation is a great soundbite, but the benefit is real: A rough calculation on paper beats holding the same information in your head without form. In this case, the journey is a large part of the destination.

Friedman reckons that startups should pick no more than six or eight terms for their own calculations. That’s an afternoon’s work. A start on what will prove to be an ongoing exercise. 

Negative ROI

The fraction runs backward, too. 

Recall Friedman’s point about a company writing more code faster but not accelerating its software development speed. Stuff those terms in:

  • (Faster code generation + other AI benefits)/(Slower code review and approval + agentic coding costs + other AI costs) = Smaller AI ROI

That’s how a company spends a king’s ransom on AI credits and winds up nowhere or nonexistent. 

Which is not hypothetical at Qodo either. AI doesn’t excel everywhere, and Friedman named email automation as an example. The company went all in on automating it, then pulled back, “mov[ing] from AI automation to AI enhancement” after discovering that AI struggled to match writing tone and intelligently extract tasks from messages. The retreat is the interesting part: Qodo’s stated approach to any task is to “go all in on complete automation,” and then “take a step back to human judgment.” 

The CEO says that automation falls short today in two areas: When human judgment is required and when context is missing. The second cuts across everything from software development to personal productivity to answering customer questions. Without timely context, what can AI do other than filibuster? Qodo’s service helps answer the context issue for software development, but collecting a company’s data and making it accessible, timely, and well-governed for general agentic usage is a massive undertaking, and one that a host of startups want to help solve. If they can, everyone’s AI ROI math should improve.

No mandate, high expectations

Qodo doesn’t require its staff to use AI. As Friedman puts it, you won’t get fired simply because you’re “not AI all the way,” or “eating AI for breakfast.” The company expects staff to complete their work as efficiently as possible and leaves the method to them.

Employees make their own decisions and execute their own work. If they start to fall behind on assigned tasks, they’re expected to reach for automation. It’s a balanced approach with high expectations: An employee who isn’t as efficient as they could be with AI could find themselves at risk.

Friedman has been on the unpopular side of an AI argument before. When he was building Qodo in 2023 and talking up agents, “agents” was a “bad word,” dismissed as little more than “fluff.” Three years and a $70 million Series B, the bet has paid off.

His advice to founders starting now looks like his past. Predict “what’s going to happen two years from now,” he says, then solve for it immediately, because whatever looks like two years tends to arrive inside of twelve months. The future “is coming faster” than you think, he says.

It’s a lot to ask of anyone working from an incomplete picture. Predicting the future is hard, he admits, “but you have to.”

The post AI spending can run negative. Qodo’s CEO built an ROI equation to fix it. appeared first on The New Stack.

How to find failures without drowning in tracing data

3 septembre 2026 à 22:17
On The New Stack podcast, Sarah Hudspeth of Chronosphere, a Palo Alto Networks company, explains how teams can build a more effective tracing strategy.

A metrics dashboard can tell you a system’s health with ease. A log can help you understand a discrete failure. But if you want to understand where in a query’s journey things went awry, you need traces.

By tracking a request from its point of origin through data and microservices to the end user, traces offer unparalleled insight into how systems work and where failures occur. SREs offer the fastest path to remediation. That means less downtime, fewer burned-out developers, and happier customers.

Sadly, the promise of traces often doesn’t match the on-the-ground reality. 

Why? Simply collecting and holding onto all your company’s traces is an exercise in hoarding. Do you need to store terabytes of tracing data just to show when your systems worked? Not only is that much information expensive to hold onto, but collecting it can slow the very systems you are trying to monitor. And when you have all the stored tracing data, finding what you need in the ocean of information can take too long.

Is tracing cooked? Not at all.

Is tracing cooked? Not at all. There are several ways to beat back tracing data overload: Head sampling collects only a portion of tracing data, reducing storage concerns; tail sampling asks whether, after a trace is recorded, it is worth holding onto, making it easier to find what you’re looking for down the road. And dynamic sampling can automatically cull similar or highly repetitive traces, so you don’t accidentally flood your storage system with nearly identical data.

You can avoid the most common tracing pitfalls by building your observability system intelligently. That’s precisely what I was hoping to learn from Sarah Hudspeth of Chronosphere (a Palo Alto Networks company), who is my guest on the latest episode of The New Stack podcast.

Whether you are just starting your tracing journey or deep in the trenches looking for help, Hudspeth’s ability to turn abstract technical concepts into simple, digestible analogies is enviable. 

Hit play on the episode above, and let’s jump the chasm between the promise of tracing and getting it to work for you in a production setting.

The post How to find failures without drowning in tracing data appeared first on The New Stack.

Want to scale AI agents without breaking anything? Retrieval engineering is the answer.

3 septembre 2026 à 17:38
Abstract metallic circuit board with raised pathways and connection points illuminated in blue, cyan, and pink.

AI agents are multiplying as corporations adopt the technology in record numbers. Smarter underlying models, better tool use, and improved multi-agent collaboration have pushed agents to evolve beyond impressive demos into practical technology that companies marshal in production environments. But the job’s not finished. 

As companies deploy more agents, more often, and against longer tasks, the plumbing that provides their AI ephemera with the required information is buckling.

Here’s the problem: AI agents are sending waves of queries against company data, creating concurrency issues and exposing just how difficult it can be to ensure a company’s AI-legible information is fresh, served only when relevant, and quickly available.

Join the live conversation: On September 24 at 12 p.m. Eastern/9 a.m. Pacific, Whit Walters, Field CTO and Lead Analyst at GigaOm and author of Defeating the Integration Tax report, joins Bonnie Chase, Director of Product Marketing at Vespa.ai, to discuss what happens when retrieval architecture meets that workload.

And crucially, they will explore in this live conversation what changes when a team rebuilds it as a unified layer instead of a fragmented one.

Register for our free event on September 24

REGISTER NOW FOR THIS WEBINAR
By registering, you consent to The New Stack’s Privacy Policy, Terms of Use and to receiving email communication from The New Stack and our event partner. You may opt out at any time.

You might be asking yourself: How has this problem not been solved yet? Google famously handles tens of thousands of search queries every second; how difficult can it be to serve agents the information that they need when we’ve solved the human version of the same problem? It’s no small challenge, and it’s why retrieval engineering is a labor category you’ll hear more about in coming quarters.

So, why is the problem worse with AI? Agents don’t ask a single question. They may retrieve data, reason against it, and then go back for more context. That doesn’t sound too complicated, until we recall that companies often stitch multiple systems together to provide their agents with required information. In practice, that means fusing vector databases, ranking tools, and serving layers into a single hybrid retrieval system that serves ever more agentic queries.

Worse, when several agents ping the same cobbled-together architecture at once, relevance drift becomes a real issue. You might do all the work to get your company or team up and running with agents, only to see the effort fail because of stale data, generic answers, or even truncated results as retrieval plumbing stumbles.

Your AI agents can’t scale successfully if they get dumber the more agents you deploy. So join the conversation on September 24, where we’ll break down how you can solve your retrieval engineering woes.

What you’ll take away:

  • Why agent workloads create a fundamentally different retrieval challenge than added concurrency alone
  • The specific failure modes at agent scale — latency stacking, stale context, relevance drift
  • Why fragmented retrieval stacks amplify those failures
  • What a unified retrieval architecture looks like in practice

The post Want to scale AI agents without breaking anything? Retrieval engineering is the answer. appeared first on The New Stack.

❌