❌

Vue normale

Reçu avant avant-hierThe New Stack

Amazon blocked Meta’s Muse. Then Shopify wired it into every store.

23 septembre 2026 à 18:39
Abstract image of a thin black frame shaped like an open doorway against a blurred gradient that runs from yellow and violet on the left to orange and red on the right.

Amazon started blocking Meta’s Muse from browsing and buying on Amazon.com on Sunday, roughly two weeks after the personal agent launched on September 8. Shoppers who ask Muse to buy something there now get a pop-up telling them that continued access by an unauthorized AI agent violates Amazon’s Conditions of Use.

Amazon’s objections have little to do with shopping itself. Meta never told Amazon that Muse would visit the store, the agent does not identify itself while it browses, and it appears to capture and store customer credentials. Those three properties describe almost every personal agent shipping this year. Grok Bot from xAI also drives signed-in browser sessions, and so does the open-source OpenClaw project that Muse is modeled on. The block is a category design problem rather than a disagreement between two companies.

What Amazon actually blocked

Muse runs on a dedicated virtual machine that Meta calls Muse Secure VM, and it reaches services in two ways: It uses built-in connectors for partners such as Gmail and OpenTable, and it drives an ordinary browser session for everything else. Shopping on Amazon.com used the second path.

That second path is what Amazon objects to. From the server’s side, a browser-driving agent looks like a signed-in customer with unusually fast reflexes, moving through search, product pages, account history, and checkout without ever declaring what it is. Amazon told GeekWire it asked Meta to exclude the store voluntarily, but Meta did not agree before the block went live.

The credential dispute is harder to settle from outside. Meta says Muse has no visibility into passwords or payment methods and that credentials sit in secure storage, while Amazon says the agent appears to capture and retain them. Both statements may be sincere, and the merchant can verify neither, since an unannounced session provides no evidence of which software holds the password.

The legal ground shifted seven weeks ago

Amazon reached for its Conditions of Use rather than the Computer Fraud and Abuse Act. Those terms, updated August 14, now require agents to identify themselves in user-agent strings and stop when asked. The likely reason sits in a ruling from early August. The Ninth Circuit vacated the preliminary injunction Amazon had won against Perplexity, and the panel held that a user directing the Comet assistant is the party accessing Amazon’s computers. Writing for the court, Judge Milan Smith described the assistant as a tool, not a person, for statutory purposes.

If you’re operating a public API or storefront, the ruling makes lawsuits a weaker tool for keeping agents out. Blocking them in your own infrastructure is now the more reliable option. A site cannot easily argue that an agent trespassed, so it has to decide for itself which automated clients it admits, publish that decision, and enforce it in its own infrastructure. Amazon’s pop-up is that enforcement, written in product rather than in a filing.

The identity layer already exists

Platform teams have solved a version of this problem before. Inside a service mesh, no workload is trusted by default because it looks like a normal client, and every call carries a verifiable identity that the receiving service checks before applying policy. Agent traffic on the public web faces the same requirement, and the specification is further along than most teams realize.

An IETF draft called Web Bot Auth builds on HTTP Message Signatures (RFC 9421). An agent signs its requests with a private key and publishes the matching public key at a well-known directory on its own domain. The verifier reads the Signature-Agent header, fetches the key set, and learns which operator is calling. Cloudflare validates these signatures at its edge for verified bots and agents. AWS WAF Bot Control added the same support for CloudFront distributions in November 2025.

The limits matter as much as the mechanism. A signature identifies the operator behind the agent, not the person it is acting for. The merchant learns that a request came from a named vendor, without learning whose account is in use or what the shopper approved. Amazon’s complaint about stored credentials sits in that gap. Signed identity settles the disclosure question and leaves authorization open.

Shopify took the other route within a day

While Amazon was blocking Muse, Shopify was wiring it in. On September 21, the two companies announced agentic checkout with Shop Pay across Shopify stores, extending the arrangement that made Meta an AI channel in Shopify Catalog on the day Muse launched. Muse reads structured product data and completes payment through a declared path, so the merchant knows an agent is transacting, and each purchase draws a single-use credential, so the card number never reaches Muse.

The plumbing for that path is public. Google and Shopify’s Universal Commerce Protocol covers discovery, cart, and checkout. The Agentic Commerce Protocol from OpenAI and Stripe covers checkout execution while the merchant stays the system of record. Google’s Agent Payments Protocol, donated to the FIDO Alliance in April, includes proof of the shopper’s authorization. A merchant that adopts it gets identity, scope, and an audit trail in the same transaction, which is what Amazon says it wanted and did not get.

Both routes follow from the business underneath them. Amazon runs its own storefront, recommendations, and assistant, so an outside agent that hides its identity takes the customer relationship and gives nothing measurable in return. Shopify sells infrastructure to merchants, so every new agent channel that reads its catalog and settles through Shop Pay reinforces the rails underneath. The key difference is who owns the demand surface, which explains why the same agent got a block from one company and a partnership from the other in the same 24 hours.

Choosing how to handle agent traffic

Most teams exposing an API or a storefront now have to make this call deliberately rather than by default. The decision depends on how much the business relies on the customer relationship at the point of contact and whether an agent can be identified when the customer arrives.

ScenarioRecommended optionRationale
Public content and catalog data, no account accessVerify signatures at the edge and allow named agentsWeb Bot Auth is checked by default on Cloudflare and AWS WAF, so the cost is policy configuration rather than engineering, though it tells you the operator and not the shopper
Agent transactions where you want the revenuePublish a declared channel using ACP, UCP, or an MCP serverStructured access gives scope and an audit trail, at the cost of building and maintaining a second interface alongside the site
Account access with stored credentialsRequire a scoped token, never a replayed passwordDelegated tokens can be revoked per agent, though few consumer agents support them yet, which pushes the burden back onto your login flow
Competitive surfaces you intend to keepState the rule in terms of service and enforce it at the edgeLegally durable after the Ninth Circuit ruling, though it invites the same public standoff Amazon is now in

Most real deployments will combine these rows rather than pick one. A retailer can verify signed agents on product pages, route purchases through a declared checkout, and still refuse an unannounced browser session inside a logged-in account. That combination is closer to Amazon’s position than its pop-up suggests.

What platform teams should do this quarter

Enterprise buyers and the teams running these systems face the same three questions, in a specific order.

Decide what an unidentified agent may do

The first question to settle is admission, and most sites have not settled it. They treat agent traffic as either a scraper to block or a browser to serve, and neither answer survives contact with a customer who wants an agent to act for them. Write policies for public pages, logged-in pages, and checkout separately, then publish them where an agent vendor can find them.

Give identified agents somewhere better to go

The second question is substitution, and it decides whether the first one holds. Blocking a browser-driving agent without offering a structured path leaves the demand intact and pushes it toward workarounds. Sabre reported that nearly 80 of its customers now pilot or run its MCP server for booking rather than let agents work through a booking screen. A catalog feed, an MCP server, or an ACP endpoint converts hostile traffic into a channel you can meter.

Fix credential handling before agents force it

The third question is authorization, and Amazon raised it loudest. An agent replaying a stored password is indistinguishable from credential stuffing at the network layer, regardless of any goodwill between the two companies. Scoped, revocable tokens tied to a named agent and a spending limit are the only version a risk owner can approve.

Where agent access is headed

Amazon and Meta will settle this commercially, because Amazon has an advertising arrangement that lets Facebook and Instagram users shop its products, and Meta buys compute from AWS, and neither gains from a long standoff over one shopping flow. The precedent is already set regardless of how they settle. Every site that matters to an agent now has to answer whether it admits anonymous automation, and it will enforce that answer through bot management rules and protocol endpoints rather than cease-and-desist letters.

Agent builders should read the block as an argument for declaring themselves. An agent that signs its requests, identifies its operator, and transacts via a published protocol can be allowed, rate-limited, and billed, while one that arrives disguised as a browser will keep encountering pop-ups. For developers building the services these agents reach, the signed identity layer arriving through Cloudflare, AWS, and the commerce protocols is the most useful infrastructure the open web has gained in years. It is worth adopting before the next agent shows up unannounced.

The post Amazon blocked Meta’s Muse. Then Shopify wired it into every store. appeared first on The New Stack.

Aider, Claude Code, and OpenClaw ran an identical model. Token use varied 70-fold.

27 août 2026 à 19:54
Collage of a woman climbing progressively taller stacks of coins, with arrows tracing her upward path.

When teams price an AI coding agent, they tend to scrutinize the model. But three recent benchmarking efforts suggest the harness — the software that steers it through tasks — may matter just as much.

The reason will be familiar to any web developer: Each inference request needs context. The relevant history must either be supplied again or reconstructed by the serving system. As a result, the provider processes large, overlapping blocks of text on every turn, including the harness’s system prompt and tool descriptions.

In June, an independent benchmark compared 12 configurations across two models on the same 12 Python tasks. In August, Composio compared eight harnesses using DeepSeek V4 Flash on 30 enterprise workflows. Artificial Analysis, meanwhile, continuously tracks harness-model pairings in its coding-agent index.

What the three benchmarks measured

Composio reported thirty workflows spanning Airtable, Gmail, Google Calendar, Google Sheets, GitHub, Slack, and PostHog. Each task ran under a 900-second ceiling. A programmatic verifier, rather than an LLM judge, graded the outcome, using isolated fixtures seeded with decoys and near-identical keys.

Composio reported 240 runs, of which 129 workflows completed successfully. Cost per successful task ranged from $0.028 for Pi Agent up to $0.195 for Claude Code. DeepAgents matched Claude Code’s pass rate exactly while costing a quarter as much per success.

The control was not perfect, and Composio disclosed that transparently. Pi ran a different reasoning setting across two model providers, and Prime Agent produced only 24 gradable runs out of 30. Those caveats rule out calling this a clean single-variable experiment.

The June benchmark measured tokens rather than dollars, and it stretched further. Its author reported runs of Aider, Claude Code, Codex, Goose, Hermes, Kilo, Kimi Code, Nanobot, OpenClaw, Opencode and Qwen Code, counting Aider’s architect mode separately. All twelve configurations ran through OpenRouter on the same tasks, so each harness used the same API and model. The suite ran on DeepSeek V4 Flash and then Nvidia’s Nemotron 3 Ultra. The second model offers a free tier on OpenRouter so that you can rerun it at no cost.

The post reported tokens per solved task ranging from roughly 3,500 for Aider in architect mode to 292,000 for OpenClaw. What makes that range usable is its stability, since the ordering barely moved between two unrelated models. That points to the harness software rather than model behavior.

Artificial Analysis approaches the same question with more statistical weight. Artificial Analysis reported an index combining DeepSWE, Terminal-Bench v2.1 from the Laude Institute, and Scale AI’s SWE-Atlas-QnA. That comes to 326 tasks, with pass rates averaged across three attempts each. It reports cost per task, token use, and wall time for each pairing. It also publishes a controlled comparison that holds Claude Opus 4.7 fixed while swapping between Claude Code, Cursor CLI, and Opencode.

The startup tax

At its core, the spread stems from a single measurement: the June benchmark, the startup tax. Before a prompt does any work, the harness ships its own baggage. That baggage is the system prompt, the tool descriptions, and the environment setup. The benchmark reported around 700 tokens for Aider in architect mode, against around 26,000 for OpenClaw.

A 40x overhead would be tolerable if it were paid once, but the resend pattern makes it otherwise. The author noted that a harness carrying a 26,000-token floor through fifteen turns spends roughly 390,000 input tokens on scaffolding alone.

The math behind that holds up under testing. The startup tax multiplied by the turn count predicts tokens per solved task, with an R-squared of 0.99 across both models. Developers looking to cut agent spend should look first at the prompt floor and the turn count before touching anything more sophisticated.

One caveat belongs on that regression figure. The June benchmark ran a single pass per harness, task, and model combination, so it carries no variance estimates, and agent runs are stochastic. Artificial Analysis carries more weight there, since it averages three attempts per task across 326 tasks. Consider the June overhead finding as evidence for the mechanism, and the larger index as the better instrument for comparing current production-scale pairings.

The expensive harnesses are not hoarding context; they are simply carrying a heavier floor.

What is surprising is what does not explain the spread. Every harness grew its context at a similar rate of a few hundred tokens per turn, so growth is not what separates them. The expensive harnesses are not hoarding context; they are simply carrying a heavier floor.

Cached tokens reorder the leaderboard

The two harness experiments converged on a second mechanism, and Artificial Analysis treats it as substantive in its methodology. It is the one platform teams are most likely to get wrong.

Composio disclosed that Claude Code drew only 1.5% of its input tokens from cache, against roughly 70% for Codex and 57% for OMP. Fresh input costs about five times as much as cached input. Claude Code’s token consumption, therefore, was comparable to its rivals’, while its invoice was not.

The June benchmark found a similar asymmetry in its own setup. On the DeepSeek run, Codex billed over a million tokens for the suite. Cache reads made up 77% of that, charged at roughly a tenth of the normal rate. Priced the way an invoice actually arrives, Codex came out cheaper per solved task than Claude Code, a harness that used half as many raw tokens.

The author attributed Claude Code’s near-zero cache share to the serving path rather than to its prompts. In that setup, Claude Code was the only harness communicating with OpenRouter via the Anthropic-style messages endpoint. The gateway’s translation of that dialect appeared to reduce the cache hits that identical traffic would have earned through the OpenAI-style endpoint. Gateway behavior changes, so treat that as one observed path.

Artificial Analysis builds that assumption into its cost model rather than discovering it. It warns that prompt cache hit rates vary widely with provider routing. Its cost model prices cached input and cache writes separately, rather than billing every prompt token at the uncached rate. A benchmark that has to price cache writes separately is telling developers how much the serving path matters.

Where the heavy harnesses earn their cost

Explaining the cost gap is not the same as crowning the cheapest harness. The agentic loop of look, act, check, and repeat is a cost multiplier. It helps most with unfamiliar code, failing tests, and changes across several files. On a small, well-described edit, the loop mostly re-confirms what a single call could have assumed.

But the hard-task test did not go the way the scaffolding argument predicts. Four harnesses spanning the cost range ran ten SWE-bench Lite tasks with generous limits, and all four resolved exactly one task, the same one. The author reported that Aider arrived at 0.8 million tokens, while Codex spent 15 million tokens. Ten tasks on one model is a thin basis for generalization, so consider it suggestive.

Quality is where the harness-sets-the-price framing needs qualifying, because the benchmarks argue against it. Composio reported pass rates across the eight harnesses, ranging from 46.7% for OpenCode to 66.7% for Pi Agent, a 20-point spread on one model. Artificial Analysis publishes the same kind of gap on a larger task set. Harness choice moves cost by multiples and moves task success by percentage points, which are different magnitudes rather than different directions.

What should platform teams measure?

Enterprises are already finding that owning the harness does not settle the cost question. Teams that built their own coding agents still pay for the underlying inference. Cost control has moved into the platform layer rather than the model contract.

QuestionWhat to measureWhy the obvious metric misleads
Which harness is cheaper?Cost per verified outcomeToken counts ignore pass rate and cache tier
Are we paying list price?Cached share on live trafficThe discount depends on the gateway and endpoint, not the harness alone
Will it survive a large prompt?Bytes transmitted versus bytes sentSilent truncation reads as success in the output

Cost per successful task, not cost per task

A pass rate and a token count, read independently, will incorrectly rank harnesses. Claude Code and DeepAgents finished the same number of Composio workflows, yet a successful Claude Code run cost more than four times as much. Procurement teams should ask for cost per verified outcome and refuse per-token comparisons.

Cached share on the actual serving path

Cache discounts are not a property of the harness alone. They depend on the endpoint dialect, the gateway, and the provider.

Cache discounts are not a property of the harness alone. They depend on the endpoint dialect, the gateway, and the provider.

Platform teams can verify the cached fraction on their own traffic in an afternoon, and that exercise is worth more than a model migration.

Prompt fidelity under load

The June benchmark injected 100,000 tokens of irrelevant log noise before each task and reported five distinct behaviors. Seven of the twelve configurations transmitted the prompt faithfully. Kilo and Opencode dropped 83%-89% of it while still reporting success.

Silent truncation is the dangerous one, because it looks exactly like success in the harness output.

OpenClaw refused to run, Kimi Code crashed, and Claude Code sent everything and then performed poorly. Silent truncation is the dangerous one, because it looks exactly like success in the harness output.

What comes next

Anthropic, OpenAI, Google, and Microsoft have already split on how to charge for the harness layer. DeepSeek then made every component of its own runtime swappable under an MIT license. Developers can now compare harness-model pairings in a publicly available, continuously updated index. That is the condition under which model pricing became contested in the first place.

For enterprises standardizing an agent platform this quarter, the harness deserves the same scrutiny as the model. In these experiments, harness choice produced cost spreads large enough to rival big differences in model pricing on workloads where the model held roughly steady. Developers, platform teams, and finance owners now have something reproducible to argue from. That is more than the harness conversation offered three months ago.

The post Aider, Claude Code, and OpenClaw ran an identical model. Token use varied 70-fold. appeared first on The New Stack.

Perplexity just separated reasoning from authority. Here’s why it matters for enterprises.

26 août 2026 à 18:18
Two stacked NVIDIA DGX Spark computers against a colorful, flowing abstract background.

Perplexity shipped Portable Computer this week, the local-first version of its Computer agent running on an Nvidia DGX Spark workstation, and the bill is steep: A DGX Spark starts at $4,700, and even an aging 24GB RTX 3090 sells well above $1,500.

The architectural choice beneath the surface deserves as much attention as the price tag: Most platforms building agents address reliability with greater intelligence. That typically includes a larger orchestrator model, a planner model, or a critic model reviewing the work.

Probabilistic reasoning proposes the next action, and deterministic software decides whether to execute it.

Perplexity separated the two jobs instead of stacking them. Probabilistic reasoning proposes the next action, and deterministic software decides whether to execute it.

The loop controller is code, and the decisions are still a model

Perplexity uses the term “orchestrator” to refer to the runtime controller rather than the planning model. That controller assembles context, enforces policy, and executes approved tool calls inside an OS-level sandbox. The company says it is deterministic code, not yet another model. The local model proposes the next action, including which tool to call and when to ask a cloud advisor for help. The split resembles a control plane architecture, where reasoning suggests, and inspectable software retains authority. Nate Kupp, Perplexity’s vice president of Computer Enterprise and Infrastructure, tells The New Stack that the harness accounted for most of the engineering work.

Same weights, better scores

Perplexity kept the base model and the silicon constant, putting Qwen3.8-27B on the same DGX Spark across three agent stacks. On its internal Local Knowledge Work Bench, a held-out set of 53 tasks, the company reported 82.6% for Computer. Pi scored 77.6% and Hermes 74%.

On ParseBench-100, a subset covering charts, layouts, tables, and formatting, the gap widened considerably. Perplexity reported 65.1% for Computer, while Hermes reported 34.6% and Pi reported 13.9%.

Because the base weights were constant, the gap measures the system around the model rather than a better model. That does not mean weights stopped mattering. Perplexity post-trained Qwen into PPLX 27B and reported 85.4%, above its own base-model number. Harness engineering and post-training are significant factors of this approach.

The word harness also covers a lot of ground, including prompts, tool schemas, context management, verification hooks, and document processing. Orchestration code is one part of that surface.

The security boundary lives outside the model

The sandbox is the boundary, not the determinism. Perplexity says the sandbox restricts processes, filesystem paths, and network access. If the sandbox is unavailable, the harness disables itself before making any tool call. Deterministic code is valuable here because it makes policy easy to inspect and helps the system fail safely. Deterministic code can still ship a vulnerability or faithfully execute a permitted mistake.

The key distinction is not between model orchestration and code orchestration. It is whether permission is enforced by something other than asking a model in plain English to comply.

The key distinction is not between model orchestration and code orchestration. It is whether permission is enforced by something other than asking a model in plain English to comply.

The same discipline shapes how context gets spent. Perplexity reports that Qwen3.8-27B advertises a 260,000-token context window but begins to struggle beyond 100,000 tokens. The harness therefore keeps the core prompt and toolset small and loads skills on demand. Commonly used connectors became command-line tools rather than full Model Context Protocol definitions sitting permanently in context.

Deterministic execution cannot rescue reasoning that exceeds the local model. On Terminal Bench 2.1, Perplexity reported 59.6% running locally. Letting the local agent consult Claude Opus 5 raised it to 73.0%, compared with 82.4% when Opus 5 worked alone. All of this remains vendor-reported evidence on a bench that the company has yet to open-source.

For enterprises evaluating local agents, the component to scrutinize is the layer that grants and denies authority, because that is where the platform’s engineering is most evident.

The post Perplexity just separated reasoning from authority. Here’s why it matters for enterprises. appeared first on The New Stack.

❌