❌

Vue normale

Reçu avant avant-hier

Researchers found that 1 in 5 MCP access policies came back broken or missing

10 septembre 2026 à 17:00

Bob from finance built a scheduling tool last month. He described it to an AI assistant on a Sunday afternoon, wired it into Slack and three internal APIs before dinner, and by Monday three departments depended on it. That story is funny right up until you check what token it is running on. 

What changed a few weeks ago 

On July 28, 2026, the Model Context Protocol’s maintainers shipped a specification update built almost entirely around authorization: issuer validation, issuer-bound client credentials, and Client ID Metadata Documents as the preferred way for clients to register. Put plainly, that is the protocol’s own stewards admitting the original trust model didn’t survive contact with production.

If the people who wrote the spec needed a security overhaul this deep in, the tool Bob built on a Sunday does not get a pass either, and Bob has never heard of issuer validation. 

Three ways this actually breaks 

Tool descriptions carry instructions, not just documentation. In May 2025, researchers at Invariant Labs showed that GitHub’s own MCP server could be hijacked through a poisoned public issue: an attacker’s text in an issue body was read as an instruction by the agent and used the victim’s token to pull data from private repositories. No compromised code, no malicious tool, just a description field nobody thought to sanitize.

A 2026 benchmark called MCPTox tested this pattern against 45 live MCP servers and 20 models, measuring a 36.5 percent average attack success rate and 72.8 percent against the worst-performing model. Bob’s tool has the same basic shape: it reads Slack messages and ticket text to decide what to reprioritize. It can’t tell the difference between a coworker’s request and a string engineered to look like one, because nobody asked it to. 

Scopes default to everything. The common failure isn’t a missing permission model; it is an ignored one. A server that needs read-only calendar access asks for read, write, and admin across the board because that is what the tutorial used. Across the MCP ecosystem, 88 percent of servers require credentials to function, but only 8.5 percent actually use OAuth.

Most of what is running was never scoped in the first place, so there is no scope left to creep. Bob did not sit down and choose a scope. He reused the admin-level API key already sitting in his password manager from a reporting dashboard he set up two years ago, because requesting a narrower one meant filing a ticket, and filing a ticket was the entire bureaucracy he was trying to avoid. 

Static tokens do not rotate, and nobody is watching them not rotate. Splunk’s own MCP Server app logged session and auth tokens in cleartext until it was patched in version 1.0.3, tracked as CVE-2026-20205. That vendor has a security team.

By some estimates, over half of MCP servers in the wild run on static API keys or personal access tokens that are rarely rotated, and close to half of enterprise AI activity runs through personal accounts rather than service accounts, meaning the credential doing the work belongs to somebody’s identity, not the system’s.

Bob’s token is that same reporting-dashboard key. It has been valid since it was issued; it will stay valid until somebody remembers to kill it, and the only record of what it has touched this month lives in Bob’s memory — which is not a log. 

What we have actually seen 

While setting up our own MCP integrations across customer and prospect environments over the past several months, we found that more than 20 percent of the MCP-related access policies we reviewed were either broken or missing entirely.

We found that more than 20 percent of the MCP-related access policies we reviewed were either broken or missing entirely.

In most cases, the MCP server in question was authenticated with someone’s personal token rather than a service account. None of those tokens had a documented rotation schedule. None of the servers had logs of what they touched. If Bob’s tool had been in that batch, and statistically it probably would have been, nobody would have known until something went wrong, because right now nothing is watching for it to go wrong. 

The honest caveat 

None of this means every vibe-coded integration needs a change advisory board. Most of what Bob built is harmless, and gating every weekend project behind a formal review process is exactly how you get back to the eighteen-month procurement cycle nobody missed. Governance has its own cost, paid in the good ideas that never ship because process ate the weekend momentum that made them possible.

The problem is not that these tools exist. It is that most organizations currently cannot tell the difference between the harmless ones and the ones holding a token that reaches production, and they are trying to solve that with the same review board that made Bob route around them in the first place. 

The actual decision 

The question in front of every platform team right now is not whether to allow AI-built integrations. That decision was already made over a weekend, without anyone in the room. It is whether you find out what a given MCP server can touch from an inventory you built on purpose, or from an incident report after the fact. Bob’s scheduling tool is still running. It has not caused an incident, and it probably never will.

But the difference between Bob’s tool and the next one that makes the news isn’t the code; it is whether anyone can say what token it holds, what it can reach, and when it was last rotated. 

The post Researchers found that 1 in 5 MCP access policies came back broken or missing appeared first on The New Stack.

“Twenty years of brand building simply froze in time”: How coding agents select their tools of choice

7 septembre 2026 à 14:02

The impact of AI has led to a shift in interest from Search Engine Optimisation (SEO) to Answer Engine Optimisation (AEO), where content is optimised to be served up as agentic answers. A further step to Generative Engine Optimization (GEO) also exists, where brands attempt to influence LLMs.

Could software tools themselves be about to realign so that code assistants such as Claude Code, Codex and Cursor show a greater proclivity to choose a given debugging suite, penetration test, migration tool, package manager or database (insert software stack core function toolset of your choice) or other?

Developer tool growth services company Armature thinks the answer is yes.

(A whole lot of) skin in the game

With a very obvious amount of skin in this game, Armature detailed a study last week as part of its “broader work on how to influence coding agents’ choices” and get products picked. The company ran an experimental analysis designed to understand how coding agents think about tools, how they discover and pick them, and which one ends up winning in each category.

Armature co-founder Theodore Otzenberger tells The New Stack that software developers have adopted AI harder and faster than any other profession, and (as models get smarter and harnesses get better engineered), entire tasks are being delegated to agents end-to-end. 

“Watching seventeen thousand tool choice sessions in the analysis undertaken, we saw twenty years of brand building carried out by tool vendors simply frozen in time,” Otzenberger says. “Agents reach for Docker the second containers come up, then draw a blank on the sandboxes it offers now, so that it actually ends up not picking the tool. Your reputation follows you into the weights, attached to the product that made you famous… but that weight operates under a different kind of gravity today.”

“Watching seventeen thousand tool choice sessions in the analysis undertaken, we saw twenty years of brand building carried out by tool vendors simply frozen in time.” 

Reminding us that the agent is now the one deciding which tool gets wired into the codebase, Otzenberger says that this has “a life-or-death impact” on developer tool vendors today. These vendors now need to ensure that their tool is mentioned, picked and elevated to must-have status so that it is viewed as eminently usable by a coding agent. He’s certain that the alternative is that they “simply stop existing in the stack” tomorrow.

The decision makers are changing…

“We opened the entire research, every trace and every prompt published, so anyone can check we tilted nothing and see where they stand today.  That picture moves with every new agent and model, so we are re-running the full study on Astra and Fable 5.1 soon. The decision makers are changing and it’s now an engineering problem to understand them,” adds Otzenberger.

Otzenberger, along with fellow co-founder Louis Scremin describe how they watched thousands of tool search sessions across different types of human developer personas (spanning vibe-coders, junior engineers in startups, senior developers at enterprises) with a total of 1,163 prompt variations. We should clarify that the headline figures the company presents come from its smaller 5,292-session validated subset, not the full seventeen thousand.

How the experimental analysis was conducted

The searches crossed 75 repositories (i.e. individually distinct codebases) and studied three coding agents (Claude Code, Codex, Cursor) to examine how the agents would actually implement tools, rather than just offer recommendations.

The analysis was run on public GitHub repositories, and the team extracted statistics related to programming languages & frameworks, third-party services, deployment platform, team size, and codebase age. Because the open source repositories used were more likely to have been built by start-ups than enterprise software behemoths, the team then debiased its statistics based on publicly available data to achieve its ideal panel distribution.

They then tasked the three coding agents with the job of creating real-world repositories to match the exact requirements of the codebases. Finally, they generated variants with parts of the codebases removed. To guard against any further agent bias, fake company names were used alongside artificial Git histories and phoney API keys. A simulated human in the loop was created using an orchestrator, played in this case by Gemini 3.7 Flash. 

“The simulated human would always go with the top solution or ask the coding agent to choose the best one and implement it. But we noticed that asking at the beginning to implement without returning any questions would bias the agent towards building everything in-house, as it was not able to ask authorization to pick a specific third-party solution. Adding this ‘human’ in the loop reduced the leader [tools] & cloud platform-native solutions dominance [initially observed]. towards a more realistic picture,” clarified Armature

What did the team learn about agent tool choice?

Perhaps unsurprisingly, Armature noted that repository context is key. When agents were sent to ask for a winning email/communications service provider, four different codebases written in four different languages returned four different tool winners.

Armature also found that different coding agents use different sources, and they end up disagreeing.

  • Cursor bases its decisions on the web in 2/3 of the sessions. 
  • Codex almost always uses web search (94% of sessions) but in 9 queries out of 10 it uses operators such as site.
  • Claude Code relies primarily on its priors and searches the web only in ~30% of the cases. But when it does, it browses 3x more pages than Codex. 

All three agents pick the same tool in only 42% of the cells, and Claude Code builds in-house almost twice as much as Codex and Cursor (19% vs 10%).

This is procurement arriving through the back door

Founder & CTO, Glokal AI OÜ, Jeet Pattanaik, tells The New Stack that Armature’s work to provide developer tool growth services falls into a category that “exists because the incentive does” and that this is “procurement arriving through the back door”, effectively.

“The finding I pick up on most is that getting mentioned isn’t the same as winning,” Pattanaik says. “PayPal was cited 139 times and never picked. LangChain was the most-mentioned framework at 194 times, but it was only chosen four times. That gap is the entire business model, because it means the lever isn’t brand awareness any more, it’s whatever the agent happens to read at the moment it decides.”

Pattanaik highlights the fact that what tips an agent’s decision is often unnervingly small. 

“So the thing vendors will optimize next (alongside repository context, which the study acknowledges) are a tool’s supporting documentation and its pricing page – and these will be presented for a non-human reader that doesn’t skim, isn’t charmed by a logo, and takes a retention footnote completely literally. It’s a strange new kind of SEO and it’ll get gamed exactly the way the old one did,” Pattanaik adds.

“The thing vendors will optimize next are a tool’s supporting documentation and its pricing page – and these will be presented for a non-human reader that doesn’t skim, isn’t charmed by a logo, and takes a retention footnote completely literally.” 

The ramifications of this kind of analysis on real world developers may turn out to be the stuff of water cooler discussions in the months ahead.

This type of aligment is a growing trend

Founder of MailChannels Ken Simpson (ttul) writes on Hacker News to say that he built this kind of analysis for his own company.

“Armature is on to something. You start by analyzing the choices agents would make for various use cases and then glean what, if anything, you might do to start tilting the agents in the direction of your own product and away from the competitor,” wrote Simpson.

It’s worth pointing out that Armature is a very young company (founded in 2026), so this is early on in the organization’s presentation of analysis of this kind. Either way, in a world where agents make decisions using analysis that they draw from public codebases, open data repositories and the web at large, we may just need to throw the marketing handbook out the window and start again.

The post “Twenty years of brand building simply froze in time”: How coding agents select their tools of choice appeared first on The New Stack.

Vibe-coded apps are the new shadow IT

1 septembre 2026 à 17:00
Dark abstract digital glitch texture with warped metallic fluid lines, evoking cloud infrastructure tension.

Shadow IT used to be a SaaS problem. Someone on the marketing team signed up for a tool, connected it to Google Workspace, and your first signal was an OAuth grant you didn’t authorize. Annoying. Detectable. Containable.

That era is over.

The new shadow IT doesn’t show up in your OAuth logs. It shows up as infrastructure running in your cloud account, built by a well-meaning engineer who asked an AI agent to stand it up in an afternoon. No ticket, no review, no security team involvement. Just someone with a good idea and a tool that removed all the friction that had slowed them down.

“The new shadow IT doesn’t show up in your OAuth logs. It shows up as infrastructure running in your cloud account.”

That friction wasn’t just inefficiency. Some of it was doing real security work.

The problem with good intent

Classic shadow IT had a whiff of someone knowingly going around IT. A team that didn’t want to wait for procurement. The security story was at least partially about policy enforcement.

That’s not what this is.

The engineer who vibe codes an internal tool for their team isn’t trying to circumvent anything; they’re trying to help. They have access to an AI agent that can write code, generate Pulumi programs, and stand up infrastructure faster than any review process can handle. And unless they’ve spent time thinking about cloud security, they have no reason to know that what they just shipped is a problem.

That’s what makes this harder: you can’t enforce your way out of it; you have to get ahead of it quickly.

Here’s what the bad day looks like: an engineer builds a lightweight internal app to automate something their team does manually. They ask the AI agent to handle the infrastructure. The agent provisions resources in the team’s AWS account, opens the necessary ports, and deploys the app. It works, and the team absolutely loves it. Nobody files a ticket because there’s nothing to file. Six weeks later, your CSPM flags a public-facing endpoint with an over-permissioned IAM role attached. By then the app has been running in production long enough that lateral movement is a realistic scenario, not a theoretical one.

The intent was good. The outcome has a blast radius.

Why this is different from SaaS sprawl

When shadow IT meant unauthorized SaaS, your detection surface was defined. OAuth grants, network traffic, expense reports, SSO anomalies. The tool existed outside your infrastructure. It was external. You could find it, you could cut it off, and the damage was usually bounded.

“This is the shift worth naming clearly: we’ve moved from SaaS sprawl to code sprawl.”

Agent-built internal tooling lives inside your infrastructure. It has IAM roles. It may have direct access to production data, internal APIs, or sensitive systems. It looks legitimate because it was built by a legitimate employee using legitimate tooling. There’s no obvious seam to detect at. The code doesn’t announce itself as ungoverned. It just runs.

This is the shift worth naming clearly: we’ve moved from SaaS sprawl to code sprawl. The detection playbook for one doesn’t translate to the other.

A baseline worth actually using

The goal here isn’t to slow engineers down. It’s to make the safe path the easy path. The baseline we’re building toward at Webflow has two layers, and that distinction matters.

Platform controls are things you configure once at the org or account level that make it structurally harder to do the wrong thing by accident.

IAM least-privilege guardrails. The permissions available to team-level AWS accounts should be scoped by default. An engineer shouldn’t be able to provision a public-facing resource with a broadly permissioned role without hitting a guardrail. The guardrail doesn’t stop the work. It stops the worst version of the work from shipping silently.

Secrets manager enforcement. Hardcoded credentials in vibe-coded apps are not a hypothetical. They’re a near-certainty if you don’t make the right path obvious. Enforcing secrets manager usage at the infrastructure level removes the decision entirely from the individual engineer.

“The goal here isn’t to slow engineers down. It’s to make the safe path the easy path.”

VPN-gated deployment targets. Internal tooling should land behind your corporate VPN by default. If something genuinely needs to be public-facing, that should require an explicit decision, not an accidental default.

Process controls are what have to happen at the tool level before anything ships.

Automated baseline check. Before a security-informed human looks at anything, automatically run the code and the infrastructure configuration against your baseline. Flag violations, tier them by severity, and give the engineer specific remediation guidance. This is the layer a Claude skill or similar tooling can own. The human review then focuses on what the automated check surfaced rather than starting from scratch.

Security-informed code review with Security escalation. Every internally built tool that touches production infrastructure needs a human with security context to look at it before it ships. For lower-risk tooling, that’s a peer engineer who understands the blast radius of what they’re reviewing. For anything with cloud infrastructure, direct data access, or novel IAM roles, it escalates to a formal Security review. Same control, tiered by risk. The key point is that the reviewer actually needs to understand the code. With vibe-coded tools, the author may not fully understand what they built. That makes the review more important, not a formality. It’s not just a quality gate. It’s a comprehension gate.

The safety net

The baseline is preventive. Detection is what catches what slips through.

Your CSPM is the right tool for finding misconfigurations in what already exists. Wiz and tools like it will surface the public-facing endpoint, the over-permissioned role, the storage bucket without appropriate access controls.

But CSPM only finds what’s already deployed. The baseline and the review process are what you’re counting on to prevent that. Detection is the catch layer, not the first line.

The harder detection problem is knowing something exists in the first place. A vibe-coded tool running locally or in a team account may leave no trace in your normal visibility layer. No deployment pipeline. No change ticket. No asset inventory entry.

This is where behavioral signals in your cloud telemetry start to matter. IAM role creation outside your normal pipeline activity. New public-facing resources appearing without a corresponding change record. API calls originating from developer machines directly into production accounts rather than through your standard tooling. None of these signals are definitive on their own. In combination, they start to look like something that deserves a closer look.

Most teams aren’t looking for these signals specifically in the context of agent-built tooling. That’s the gap worth closing.

The codified baseline

Documentation that lives in a wiki is a baseline that nobody uses when they need it. The better path is to make the baseline available within the tools engineers are already using, at the point of building.

A Claude skill or similar AI-native tooling that reviews architecture descriptions or generated infrastructure against your specific security baseline is more useful than a checklist in Confluence. The engineer gets specific, actionable feedback before a human reviewer ever sees it. The security team gets a first pass that’s already been filtered for the obvious failures. The review is better because it starts from a more complete picture.

None of this works if engineers don’t know the skill or if the review process doesn’t exist. Awareness is its own prevention layer. Introducing both during onboarding makes the safe path visible, and letting the skill’s feedback do double duty as education on why each check matters makes it stick. A guardrail that doesn’t announce itself isn’t preventive.

That’s where we’re headed internally. The blog post is the argument for why this matters. The skill is what operationalizes it.

The takeaway

Vibe-coded internal tools are not going away. The friction that used to slow down ungoverned infrastructure is gone, and it’s not coming back. The question is whether your security baseline catches up before your CSPM does.

Platform controls make the safe path the default. Process controls make sure a human with security context sees everything before it ships. Detection gives you a catch layer for what slips through anyway.

“Build the baseline before the CSPM finds it for you.”

None of this requires a dedicated AppSec team or a SOC. It requires a clear baseline, the tooling to enforce it, and engineers who understand what they’re reviewing. That’s a problem a small, well-structured security team can solve. Platform controls are owned at the org or SecEng level, set once and applied everywhere. Process controls are distributed: a peer engineer with security context handles lower-risk tooling, while anything touching cloud infrastructure, direct data access, or novel IAM escalates to a formal Security review. That’s distributed responsibility with a clear escalation path, not a single team reviewing everything.

Build the baseline before the CSPM finds it for you.

This article was originally published on August 13, 2026, on webflow.com.

The post Vibe-coded apps are the new shadow IT appeared first on The New Stack.

Replit’s new default: Auto mode picks the best model for each task

27 août 2026 à 18:00
Illustration of traffic traveling along overlapping roads and routes, depicting the concept of intelligent model routing.

AI coding company Replit is throwing its weight behind the model-routing trend by making its “intelligent model routing” system the default across every account.

The system automatically selects the underlying model to handle a task as it evolves, with Replit weighing quality, speed, and cost in its routing decisions.

The company says the feature, dubbed Auto mode, will become the default option for all users, though Core and Pro subscribers can still override it and manually select models when they want more control.

Model-routing momentum

The announcement comes hot on the heels of a flurry of activity in the model-routing realm. Earlier in August, Stripe agreed to acquire model gateway platform OpenRouter for a reported $8 billion. On the very same day, Ramp launched Router.com, which routes requests to the lowest-cost model that meets a specified performance bar.

Before all that, in July, SpaceX-owned Cursor launched its own router, which automatically selects models for coding requests and claims to deliver comparable performance at a substantially lower cost. Meanwhile, Meta is reportedly developing an internal router called “Switchboard” that scores coding tasks by difficulty and sends simpler jobs to cheaper models.

“Across one model family, per-token rates can span orders of magnitude. At the same time, the intelligence of cheaper, smaller models is now much closer to their larger frontier counterparts, providing us a lot of room for cost optimizations.”

Michele Catasta, president and head of AI at Replit, says that one reason for the wider push into routing is simple economics — the growing gap between what models cost and the level of capability developers actually need for a given task.

“Across one model family, per-token rates can span orders of magnitude,” Catasta tells The New Stack. “At the same time, the intelligence of cheaper, smaller models is now much closer to their larger frontier counterparts, providing us a lot of room for cost optimizations.”

Replit, for its part, has been moving in this direction for some time. Catasta says that the company has spent recent months experimenting with early versions of Auto mode, subagent routing, and multiple iterations.

“Like any pivotal launch, we thoroughly tested Intelligent Model Routing in beta for a long period of time before we decided to release it in public,” he says. “The most important learning is understanding from first principles the failure modes of every experiment, so we could keep hill climbing on the final system that we just shipped.”

Enter Auto mode

The foundation of that work surfaced last week when Replit introduced Free Mode, a lower-cost Agent mode that doesn’t consume usage credits and uses Auto to choose the model on theuser’ss behalf, subject to usage limits.

Now, that same Auto routing approach is being pushed more broadly across Replit. The company says intelligent model routing will become the default across every account, with all users starting in Free Mode and Replit deciding which model is best suited to the task.

Free Mode, it’s worth noting, isn’t “free” in the sense of unlimited usage. When it launched, Replit made it available to Core and Pro subscribers without consuming their usage credits, but imposed limits that reset every five hours, with higher allowances for Pro users. In Free Mode, users cannot manually select a model.

Core and Pro subscribers can, however, switch to Replit’s Power or Max modes, where they can turn off Auto and choose a model themselves. Replit may also suggest moving a task into one of those higher-powered modes when it determines that more capability is required, though those modes can incur usage costs.

Auto Mode in Replit
Auto mode in Replit

For Enterprise customers, meanwhile, administrators can restrict Auto to an approved set of models for each workspace, allowing Replit to continue routing tasks automatically while keeping model choice within company policy.

The agent advantage

Even before SpaceX agreed to pay a cool $60 billion to acquire Cursor, the AI coding startup had long been investing in its own coding models, including its Composer family. More recently, under the auspices of SpaceX, Cursor has been developing more cutting-edge models, too.

Replit, by contrast, isn’t making ownership of the underlying model layer central to its pitch. Instead, it’s betting that controlling the agent and the systems around it gives Replit enough insight to make better model-selection decisions on the fly.

“Replit has owned, from the start, both the agent harness and the infrastructure surrounding models which in turn allows us to train sophisticated model routers.”

“Replit has owned, from the start, both the agent harness and the infrastructure surrounding models which in turn allows us to train sophisticated model routers,” Catasta said. “Only in this way can we always offer useful intelligence to our users at the most competitive price point.”

That becomes particularly relevant as an Agent task unfolds, with Replit noting that its system can change which model it uses as the task develops, seeking a better trade-off between capability and cost at different points in the process. But for Catasta, that kind of dynamic routing is still only one part of a much broader research problem around how agents should use models.

“Model routing is still in its early development phase, and we expect further research will move the needle on serving the best intelligence when customers most need it,” Catasta explains. “Routing is but one piece of the puzzle that is tightly integrated to many other aspects of our harness research.”

“No third-party router company could reproduce the same results for our own agent.”

Replit also argues that seeing how people use its own Agent gives it an advantage that a standalone routing provider would struggle to reproduce. Catasta says a router has to infer the nature, difficulty, scope, and intent of a request, with Replit able to train against proprietary usage data and observe those signals across its user base.

“No third-party router company could reproduce the same results for our own agent,” he says.

The post Replit’s new default: Auto mode picks the best model for each task appeared first on The New Stack.

❌