❌

Vue normale

Reçu avant avant-hier

Want to scale AI agents without breaking anything? Retrieval engineering is the answer.

3 septembre 2026 à 17:38
Abstract metallic circuit board with raised pathways and connection points illuminated in blue, cyan, and pink.

AI agents are multiplying as corporations adopt the technology in record numbers. Smarter underlying models, better tool use, and improved multi-agent collaboration have pushed agents to evolve beyond impressive demos into practical technology that companies marshal in production environments. But the job’s not finished. 

As companies deploy more agents, more often, and against longer tasks, the plumbing that provides their AI ephemera with the required information is buckling.

Here’s the problem: AI agents are sending waves of queries against company data, creating concurrency issues and exposing just how difficult it can be to ensure a company’s AI-legible information is fresh, served only when relevant, and quickly available.

Join the live conversation: On September 24 at 12 p.m. Eastern/9 a.m. Pacific, Whit Walters, Field CTO and Lead Analyst at GigaOm and author of Defeating the Integration Tax report, joins Bonnie Chase, Director of Product Marketing at Vespa.ai, to discuss what happens when retrieval architecture meets that workload.

And crucially, they will explore in this live conversation what changes when a team rebuilds it as a unified layer instead of a fragmented one.

Register for our free event on September 24

REGISTER NOW FOR THIS WEBINAR
By registering, you consent to The New Stack’s Privacy Policy, Terms of Use and to receiving email communication from The New Stack and our event partner. You may opt out at any time.

You might be asking yourself: How has this problem not been solved yet? Google famously handles tens of thousands of search queries every second; how difficult can it be to serve agents the information that they need when we’ve solved the human version of the same problem? It’s no small challenge, and it’s why retrieval engineering is a labor category you’ll hear more about in coming quarters.

So, why is the problem worse with AI? Agents don’t ask a single question. They may retrieve data, reason against it, and then go back for more context. That doesn’t sound too complicated, until we recall that companies often stitch multiple systems together to provide their agents with required information. In practice, that means fusing vector databases, ranking tools, and serving layers into a single hybrid retrieval system that serves ever more agentic queries.

Worse, when several agents ping the same cobbled-together architecture at once, relevance drift becomes a real issue. You might do all the work to get your company or team up and running with agents, only to see the effort fail because of stale data, generic answers, or even truncated results as retrieval plumbing stumbles.

Your AI agents can’t scale successfully if they get dumber the more agents you deploy. So join the conversation on September 24, where we’ll break down how you can solve your retrieval engineering woes.

What you’ll take away:

  • Why agent workloads create a fundamentally different retrieval challenge than added concurrency alone
  • The specific failure modes at agent scale — latency stacking, stale context, relevance drift
  • Why fragmented retrieval stacks amplify those failures
  • What a unified retrieval architecture looks like in practice

The post Want to scale AI agents without breaking anything? Retrieval engineering is the answer. appeared first on The New Stack.

AI agents are making retrieval engineering a core engineering discipline

30 août 2026 à 18:00
An abstract digital image featuring an intricate network of glowing copper-orange and deep red strands swirling against a dark background, visualizing complex data pathways and retrieval engineering workflows for AI agents.

AI agents are changing retrieval requirements. As organizations move from chatbots to AI systems that investigate, reason, and act on users’ behalf, retrieval is becoming the foundation of application quality. Better retrieval doesn’t just produce better answers—it enables more capable assistants, more personalized experiences, and more trustworthy autonomous systems.

Traditional search and even many RAG applications could tolerate imperfect retrieval. If a user didn’t find exactly what they wanted, they refined the query or tried again. Agents don’t have that luxury.

“As organizations move from chatbots to AI systems that investigate, reason, and act on users’ behalf, retrieval is becoming the foundation of application quality.”

An AI agent plans, reasons, invokes tools, and increasingly makes decisions without a human reviewing every intermediate step. That raises the bar considerably. Retrieval is no longer about finding relevant information—it’s about consistently delivering the right evidence at the right time.

For engineers, this creates a familiar set of challenges:

  • Which signals matter most for this user?
  • How do fresh events change relevance?
  • How should structured, unstructured, and behavioral signals be combined?
  • When should a model influence ranking?
  • How do you optimize for business outcomes rather than similarity scores?

Those aren’t vector database problems. They’re Retrieval Engineering problems.

It’s no longer just about embeddings or vector search. It’s about engineering the entire retrieval workflow: combining hybrid retrieval, real-time signals, ranking, machine learning inference, and continuous experimentation to deliver the best possible decision at serving time.

A recent GigaOm Decision Brief argues that as retrieval becomes increasingly commoditized, competitive advantage shifts to decisioning—determining what an application or AI agent should see, and in what order, before it acts.

“As retrieval becomes increasingly commoditized, competitive advantage shifts to decisioning—determining what an application or AI agent should see, and in what order, before it acts.”

That aligns closely with the way we’ve been thinking about Retrieval Engineering:

Prompt engineering influences how a model reasons. Retrieval Engineering determines what it has to reason about.

As organizations move from copilots to production AI agents, I believe Retrieval Engineering will become a core engineering discipline alongside prompt engineering and model engineering.

If you’re interested in the engineering discipline itself, my earlier article explores Retrieval Engineering in more depth (https://thenewstack.io/ai-retrieval-engineering-bottleneck/). The new GigaOm paper complements that discussion by looking at why these engineering decisions increasingly influence product quality, customer experience, and ultimately business outcomes.

Read the GigaOm report. 

The post AI agents are making retrieval engineering a core engineering discipline appeared first on The New Stack.

❌