❌

Vue lecture

[Launched] Generally Available: Instant Access for VM restore points

Instant Access for application-consistent restore points on virtual machines (that have Premium v2 or Ultra disks as data disks) is now generally available. With Instant Access, you can begin restoring a disk from a restore point as soon as the snapshot i
  •  

Amazon blocked Meta’s Muse. Then Shopify wired it into every store.

Abstract image of a thin black frame shaped like an open doorway against a blurred gradient that runs from yellow and violet on the left to orange and red on the right.

Amazon started blocking Meta’s Muse from browsing and buying on Amazon.com on Sunday, roughly two weeks after the personal agent launched on September 8. Shoppers who ask Muse to buy something there now get a pop-up telling them that continued access by an unauthorized AI agent violates Amazon’s Conditions of Use.

Amazon’s objections have little to do with shopping itself. Meta never told Amazon that Muse would visit the store, the agent does not identify itself while it browses, and it appears to capture and store customer credentials. Those three properties describe almost every personal agent shipping this year. Grok Bot from xAI also drives signed-in browser sessions, and so does the open-source OpenClaw project that Muse is modeled on. The block is a category design problem rather than a disagreement between two companies.

What Amazon actually blocked

Muse runs on a dedicated virtual machine that Meta calls Muse Secure VM, and it reaches services in two ways: It uses built-in connectors for partners such as Gmail and OpenTable, and it drives an ordinary browser session for everything else. Shopping on Amazon.com used the second path.

That second path is what Amazon objects to. From the server’s side, a browser-driving agent looks like a signed-in customer with unusually fast reflexes, moving through search, product pages, account history, and checkout without ever declaring what it is. Amazon told GeekWire it asked Meta to exclude the store voluntarily, but Meta did not agree before the block went live.

The credential dispute is harder to settle from outside. Meta says Muse has no visibility into passwords or payment methods and that credentials sit in secure storage, while Amazon says the agent appears to capture and retain them. Both statements may be sincere, and the merchant can verify neither, since an unannounced session provides no evidence of which software holds the password.

The legal ground shifted seven weeks ago

Amazon reached for its Conditions of Use rather than the Computer Fraud and Abuse Act. Those terms, updated August 14, now require agents to identify themselves in user-agent strings and stop when asked. The likely reason sits in a ruling from early August. The Ninth Circuit vacated the preliminary injunction Amazon had won against Perplexity, and the panel held that a user directing the Comet assistant is the party accessing Amazon’s computers. Writing for the court, Judge Milan Smith described the assistant as a tool, not a person, for statutory purposes.

If you’re operating a public API or storefront, the ruling makes lawsuits a weaker tool for keeping agents out. Blocking them in your own infrastructure is now the more reliable option. A site cannot easily argue that an agent trespassed, so it has to decide for itself which automated clients it admits, publish that decision, and enforce it in its own infrastructure. Amazon’s pop-up is that enforcement, written in product rather than in a filing.

The identity layer already exists

Platform teams have solved a version of this problem before. Inside a service mesh, no workload is trusted by default because it looks like a normal client, and every call carries a verifiable identity that the receiving service checks before applying policy. Agent traffic on the public web faces the same requirement, and the specification is further along than most teams realize.

An IETF draft called Web Bot Auth builds on HTTP Message Signatures (RFC 9421). An agent signs its requests with a private key and publishes the matching public key at a well-known directory on its own domain. The verifier reads the Signature-Agent header, fetches the key set, and learns which operator is calling. Cloudflare validates these signatures at its edge for verified bots and agents. AWS WAF Bot Control added the same support for CloudFront distributions in November 2025.

The limits matter as much as the mechanism. A signature identifies the operator behind the agent, not the person it is acting for. The merchant learns that a request came from a named vendor, without learning whose account is in use or what the shopper approved. Amazon’s complaint about stored credentials sits in that gap. Signed identity settles the disclosure question and leaves authorization open.

Shopify took the other route within a day

While Amazon was blocking Muse, Shopify was wiring it in. On September 21, the two companies announced agentic checkout with Shop Pay across Shopify stores, extending the arrangement that made Meta an AI channel in Shopify Catalog on the day Muse launched. Muse reads structured product data and completes payment through a declared path, so the merchant knows an agent is transacting, and each purchase draws a single-use credential, so the card number never reaches Muse.

The plumbing for that path is public. Google and Shopify’s Universal Commerce Protocol covers discovery, cart, and checkout. The Agentic Commerce Protocol from OpenAI and Stripe covers checkout execution while the merchant stays the system of record. Google’s Agent Payments Protocol, donated to the FIDO Alliance in April, includes proof of the shopper’s authorization. A merchant that adopts it gets identity, scope, and an audit trail in the same transaction, which is what Amazon says it wanted and did not get.

Both routes follow from the business underneath them. Amazon runs its own storefront, recommendations, and assistant, so an outside agent that hides its identity takes the customer relationship and gives nothing measurable in return. Shopify sells infrastructure to merchants, so every new agent channel that reads its catalog and settles through Shop Pay reinforces the rails underneath. The key difference is who owns the demand surface, which explains why the same agent got a block from one company and a partnership from the other in the same 24 hours.

Choosing how to handle agent traffic

Most teams exposing an API or a storefront now have to make this call deliberately rather than by default. The decision depends on how much the business relies on the customer relationship at the point of contact and whether an agent can be identified when the customer arrives.

ScenarioRecommended optionRationale
Public content and catalog data, no account accessVerify signatures at the edge and allow named agentsWeb Bot Auth is checked by default on Cloudflare and AWS WAF, so the cost is policy configuration rather than engineering, though it tells you the operator and not the shopper
Agent transactions where you want the revenuePublish a declared channel using ACP, UCP, or an MCP serverStructured access gives scope and an audit trail, at the cost of building and maintaining a second interface alongside the site
Account access with stored credentialsRequire a scoped token, never a replayed passwordDelegated tokens can be revoked per agent, though few consumer agents support them yet, which pushes the burden back onto your login flow
Competitive surfaces you intend to keepState the rule in terms of service and enforce it at the edgeLegally durable after the Ninth Circuit ruling, though it invites the same public standoff Amazon is now in

Most real deployments will combine these rows rather than pick one. A retailer can verify signed agents on product pages, route purchases through a declared checkout, and still refuse an unannounced browser session inside a logged-in account. That combination is closer to Amazon’s position than its pop-up suggests.

What platform teams should do this quarter

Enterprise buyers and the teams running these systems face the same three questions, in a specific order.

Decide what an unidentified agent may do

The first question to settle is admission, and most sites have not settled it. They treat agent traffic as either a scraper to block or a browser to serve, and neither answer survives contact with a customer who wants an agent to act for them. Write policies for public pages, logged-in pages, and checkout separately, then publish them where an agent vendor can find them.

Give identified agents somewhere better to go

The second question is substitution, and it decides whether the first one holds. Blocking a browser-driving agent without offering a structured path leaves the demand intact and pushes it toward workarounds. Sabre reported that nearly 80 of its customers now pilot or run its MCP server for booking rather than let agents work through a booking screen. A catalog feed, an MCP server, or an ACP endpoint converts hostile traffic into a channel you can meter.

Fix credential handling before agents force it

The third question is authorization, and Amazon raised it loudest. An agent replaying a stored password is indistinguishable from credential stuffing at the network layer, regardless of any goodwill between the two companies. Scoped, revocable tokens tied to a named agent and a spending limit are the only version a risk owner can approve.

Where agent access is headed

Amazon and Meta will settle this commercially, because Amazon has an advertising arrangement that lets Facebook and Instagram users shop its products, and Meta buys compute from AWS, and neither gains from a long standoff over one shopping flow. The precedent is already set regardless of how they settle. Every site that matters to an agent now has to answer whether it admits anonymous automation, and it will enforce that answer through bot management rules and protocol endpoints rather than cease-and-desist letters.

Agent builders should read the block as an argument for declaring themselves. An agent that signs its requests, identifies its operator, and transacts via a published protocol can be allowed, rate-limited, and billed, while one that arrives disguised as a browser will keep encountering pop-ups. For developers building the services these agents reach, the signed identity layer arriving through Cloudflare, AWS, and the commerce protocols is the most useful infrastructure the open web has gained in years. It is worth adopting before the next agent shows up unannounced.

The post Amazon blocked Meta’s Muse. Then Shopify wired it into every store. appeared first on The New Stack.

  •  

[Launched] Generally Available: High-scale mesh in Azure Virtual Network Manager

High-scale mesh using connected group in Azure Virtual Network Manager is now in general availability. In available regions, customers may connect up to 3,000 virtual networks in a single mesh connectivity configuration by default and higher scale IP conn
  •  

[Launched] Generally Available: Azure Virtual Network Manager IPAM in additional Azure regions

Azure Virtual Network Manager IP address management is now generally available in additional regions: US Gov Virginia, US Gov Texas and US Gov Arizona, and China North 3 and China East 3. In complex network environments, managing IP addresses effectively
  •  

[Launched] Generally Available: Azure VM Image Builder in sovereign and air-gapped clouds

Overview:Azure VM Image Builder is now generally available in Azure Government, China North 3, Azure Government Secret, and Azure Government Top Secret.You can now use the same managed image-building service across your sovereign and air-gapped environmen
  •  

Anthropic’s new Files API vs. pasting: It will save you time, but it won’t save you money.

Horizontal bands of distorted green digital numbers flicker across a black background.

Anthropic moved the Files API and computer-use toolset out of beta on August 19 and launched a browser-use toolset. It lets developers upload a document once and reference it by ID in subsequent requests, rather than sending its contents each time. The common alternative is what most developers do today: pasting the reference material into every prompt.

I wanted to measure how uploading a file once compares to pasting the document into each request, in terms of tokens, accuracy, and setup. I built a test with a verifiable answer key and ran the same workload through both, plus a third approach, prompt caching, that came out of the results.

The test

I wrote a fake API reference for an invoicing company called Ledgerline, about 1,200 words covering authentication, rate limits, idempotency, webhooks, bulk endpoints, and a sandbox. The scenario is a developer support bot that answers questions using only the Ledgerline API reference. Then I wrote five developer questions that the document can answer:

  • How to tell apart two different 403 errors, and the fix for each
  • How to safely retry payment creation after a timeout
  • How to verify webhook signatures and block replay attacks
  • The right way to create 2,000 invoices in one night
  • What to do after committing a secret API token to a public repo

Each question has a specific correct answer in the reference, including details that are easy to miss. One fix depends on knowing that token scopes can’t be edited after creation. Another requires a separate rate limit for the bulk endpoint.

I ran the same five questions three ways, in Python scripts against the API directly, on claude-sonnet-5 with identical instructions:

  • Arm 1 pasted the full reference into every request
  • Arm 2 uploaded the reference once through the Files API and referenced its file ID in every request
  • Arm 3 pasted the reference once into the system prompt with prompt caching enabled

The API reports token usage on every response, so each arm produced its own count.

Getting it running

Two things broke before the first run. The scripts originally set temperature to zero for reproducibility, and the API rejected it. Temperature is deprecated on claude-sonnet-5. Second, one response led with a thinking block instead of text, which caused my printing code to crash. Both fixes were one-liners.

The Files API upload itself was uneventful. One call, one file ID back, and the ID worked in every following request.

The results

All fifteen answers were correct against the answer key, across all three arms. Every arm caught the subtle details, the uneditable token scopes, the separate bulk rate limit, the constant-time signature comparison, the five-minute replay window. Whatever else changed between arms, answer quality did not.

Whatever else changed between arms, answer quality did not.

While the answer quality remained consistent, the token counts didn’t.

TestsRegular input tokensCache writeCache readOutput
Arm 1, paste every time15,246001,399
Arm 2, Files API15,371001,434
Arm 3, prompt caching2712,99011,9601,642

Arm 2 billed slightly more input than Arm 1, which was interesting. The marketing doesn’t say this, but I thought using the files from one API might require slightly fewer tokens. The uploaded file’s contents are still processed into every request, at roughly 3,050 input tokens per question either way, and referencing the file added a small amount of overhead on top, about 25 tokens per request. Across five requests, upload-once cost 125 more input tokens than pasting. There is no volume at which that flips.

Arm 3 is the one that behaved as I assumed the Files API would. The document was billed in full once, as a 2,990-token cache write on the first request. The four requests after that read it from cache, and cache reads bill at about one-tenth the rate of normal input tokens. The questions themselves cost between 48 and 63 regular input tokens each. Cache writes carry a 25 percent premium over normal input, so the first request is the most expensive, and subsequent requests are where the reduction occurs.

Two caveats on the caching numbers. The cache expires after five minutes of inactivity, so the reduction assumes requests keep coming in steadily. And caching required restructuring the request, moving the document into the system prompt with a cache marker.

When to use each approach

The interesting result is that the two features solve different problems. The Files API manages documents. Prompt caching lowers costs.

Files API

Use the Files API when the problem is the file itself. It gives you one uploaded copy referenced by ID instead of the document text living in your code; it handles formats you can’t paste, like PDFs and images; and stored files now support expiration settings. What it doesn’t change is cost. The document is processed for every request, and Anthropic’s announcement never claimed otherwise.

Prompt caching

Use prompt caching when the problem is paying for the same document on every request. In this workload, it cut billed input to roughly a third of pasting across five requests, and the gap keeps widening because every request after the first reads the document at about a tenth of the normal rate. The tradeoffs are the five-minute cache expiry, which assumes steady traffic, and restructuring your request to add the cache marker.

Use the files API and prompt caching together when both problems apply. Files API documents can be cache-marked the same way as pasted text. I tested them separately to isolate what each does on its own.

Paste the documents

Pasting still wins in some cases. Use pasting when prototyping and making one-off calls, where uploading first is just an extra step. This is also the best option for documents that change with every request, where nothing is reused so neither feature helps. Keep in mind that this is short reference text, since documents under 1,024 tokens can’t be cached on most models. Pasting also has the fewest moving parts: no upload step, no file IDs, no stored copies to manage.

I did this test expecting to find out whether the Files API beats pasting in terms of accuracy or cost. It turns out I asked the wrong question. They tie on cost, and the feature that wins on cost was something I tested at the last minute to try to force a different result.

The post Anthropic’s new Files API vs. pasting: It will save you time, but it won’t save you money. appeared first on The New Stack.

  •  

FinOps for the AI era: New flexible billing and cost controls for agents

Editor's note: A product image was updated after initial publication.


As AI takes on more complex work, business leaders face a new challenge: enabling rapid innovation using agents while protecting their margins and budgets. To get a real return on AI, financial operations (FinOps) and cost management must evolve alongside technology, giving you clear visibility, proactive cost controls, and flexible payment models that fit your needs. 

That’s why today we’re introducing expanded billing flexibility and new cost management tools for agent workloads across Gemini Enterprise and developer tools like Google Antigravity in Gemini Enterprise and Android Studio.

  • Flexible payment options: You can mix our existing, predictable per-user seat subscriptions with a new pay-as-you-go option in Gemini Enterprise app that lets you run agent workloads without hitting quota limits mid-task.

  • Developer access, one place to manage your AI: Google Antigravity and Android Studio AI use is now included in your Gemini Enterprise subscription (available for select customers and rolling out broadly soon), giving your developers more without giving you more to manage. Usage across Antigravity, the platform, and the app rolls up into a single view instead of separate licenses and billing silos.

  • Pay less as your usage grows: If your AI workloads are steady or climbing, Flexible Savings Plans let you commit to a monthly spend you're comfortable with and take 10–20% off your token costs — no minimums, no maximums, and no new billing silo to manage.

  • Consolidated spend guardrails: You can now set hard monthly caps on AI spend and projects, estimate agent runtime costs, and catch sudden budget spikes before they hit your invoice.

Give your teams flexibility without losing control over spend in Gemini Enterprise

Every organization operates differently. Even within the same business, no two teams consume AI in the same way. Your business users might rely on steady, everyday productivity tools. Meanwhile, your technical teams might run AI agent workloads in bursts. 

To help align costs with how work actually gets done, you can combine these payment and licensing choices and features across Gemini Enterprise:

Option

How it works

Why it helps optimize spend

Gemini Enterprise app per-user seat subscription

You pay a fixed monthly fee per user, which includes daily quota pools that are shared across your entire project.

Predictable budgeting. It provides finance teams with a clear, steady monthly baseline for teams with consistent daily productivity needs.

[New] Gemini Enterprise app pay-as-you-go consumption edition

*available for select customers and rolling out broadly soon

There is no upfront commitment or base subscription fee, meaning you pay strictly for the compute and tokens your teams consume at standard model API rates.

Only pay for what you use. Your spend scales up and down automatically with real usage, ensuring you never pay for empty seats when project demand dips.

[New for Antigravity in Gemini Enterprise] Consolidated pooled quotas

Daily usage allowances are pooled project-wide, letting business apps, developer tools, and custom agents draw from the same shared quota. Pooled quota is always exhausted first, and admins can control if overages are allowed, at which point it’s charged at pay-as-you-go rates. 

Maximized resource usage: Unused daily allowances from business users automatically absorb heavy developer or custom API agent demands, so no quota allowance goes to waste.

[Coming soon] Deferred execution pricing

*available for select workloads soon

Mark eligible agent workloads as deferred, and our intelligent scheduler in the Gemini Enterprise Agent Platform runs them during off-peak capacity windows.

Substantial discounts for work that can wait: AI workloads can run on separate, off-peak capacity, you pay up to half the inference cost and bypass standard quota limits entirely – letting you run substantially more agentic volume under the same budget.

Equip developers with advanced agentic tooling under a single Gemini Enterprise subscription

We’re rolling out access to Google Antigravity in Gemini Enterprise, an agent-first developer platform that brings powerful agentic coding and agent-building capabilities to technical teams, included with Gemini Enterprise subscriptions for eligible customers. In addition, Android developers can leverage the Google Antigravity quota included in their Gemini Enterprise subscriptions natively in Android Studio, the agentic IDE for professional Android development.

To be more efficient with agentic coding costs, we are pooling developer tools quota included in each Gemini Enterprise subscription and making it available across the whole Google Cloud project so your teams can benefit from the capacity you’re already purchasing. Your developers get access to advanced agentic tools, while you maintain centralized governance and control.

For a closer look into what’s new with Antigravity in Gemini Enterprise and how customers are putting it to work in production, take a look at our deep-dive.

Budget smarter with Gemini Enterprise Flexible Savings Plans (FSPs) 

If your organization has steady or growing AI workloads, Gemini Enterprise Flexible Savings Plans offer a simple, spend-based commitment model across Gemini Enterprise usage. FSPs are designed to lower token costs while keeping budgets flexible:

  • Programmatic savings: Receive 10% off for 1-year or 20% off for 3-year commitments for monthly spending across Gemini Enterprise.

  • Tailored to your pace: With no minimum or maximum spend requirements, you can determine a monthly commitment that fits your current traffic and make adjustments as your usage increases over time. 

  • Enterprise Agreement (EA) friendly: FSP spend seamlessly draws down against your existing Google Cloud EA, giving lines of business dedicated budget control without fragmenting your broader cloud commitments.

Gemini Enterprise Flexible Savings Plans are already available for self-serve customers and customers on enterprise agreements.

Give your teams the freedom to build while maintaining financial discipline

As a leader, your goal isn't to restrict the potential value of AI  – it's to remove the financial and operational risk that you face without managed AI costs. You should be able to give engineering, marketing, and operational teams the freedom to innovate with agents, but you should also have the visibility to trust what those agents are doing and the safety nets to protect your budget.

To bridge this gap, we've built robust, native governance tooling directly into the Google Cloud Billing Console around three simple goals:

1. Plan before you scale: The Google Cloud Pricing Calculator lets you estimate anticipated costs in Gemini Enterprise across per-user licenses, developer tools, and background agent runtimes. It gives you the numbers you need to build clear business cases upfront before project work begins.

2. Enforce boundaries without micromanaging spend: Instead of spending time tracking daily usage variations across project teams, let these tools do the monitoring for you:

  • Early anomaly detection: If a project’s AI spending trends higher than normal, the system flags the deviation with root cause analysis and pinpoints the top 3 SKUs driving the increase so you can see exactly what changed.

1 Jul22_Anomalies_Image1

Billing Console showing an Early Anomaly alert with the Root Cause Analysis (RCA) breakdown highlighting the driving SKUs

    • Project-level spend caps: When a project needs defined financial boundaries, you can set a firm monthly spend limit directly in the Google Cloud Billing Console. If a project hits its limit, the agent's API calls temporarily pause – protecting your budget without affecting the rest of your production infrastructure. Automated email alerts at 50%, 80% and 100% of the budget keep you informed of your progress against the spend limit. 
    • Overage controls: If a spend cap triggers, you can choose to resume work with a single click in the console. Alternatively, if your priority is continuous operation, you can turn on overages so excess usage smoothly transitions to consumption rates, which can draw directly against your FSP to keep overage unit costs heavily discounted.
3 PAYG Overage Enabled

Enabling overage pay-as-you-go for a project.

3. Get visibility into business value: Use centralized billing reports paired with the FinOps agent to generate natural-language cost insight summaries of where your budget went, making it simple to show ROI to leadership.

Cost overview FinOps

AI spending reporting in Google Cloud Console

Go deeper with AI cost optimization

To build a full-stack FinOps strategy that optimizes the cost, latency, and performance of your models and infrastructure, explore our detailed architecture specifications and frameworks:

  • How to outsmart infrastructure constraints with dynamic capacity management: Discover how to optimize your compute investments with capabilities in Google Kubernetes Engine and Google Compute Engine that automatically schedule and reallocate resources to avoid interruptions, over-provisioning, and over-reliance on any one hardware configuration.

  • Expanding Google Antigravity for Enterprise Customers: Read our developer tooling deep-dive to see how technical teams are accelerating software delivery with agent-first workflows.

  • What sports cars can teach us about optimizing AI spend: More tokens doesn't always mean better AI. Read our conversation with Mike Clark, Director of Product Management for Gemini Enterprise Agent Platform, on how to balance horsepower with efficiency and get the highest return out of every dollar you spend on AI. 

  • Protection during usage spikes: Your heavy workloads can surge during peak hours without forcing you to pay for expensive, dedicated infrastructure that sits idle the rest of the time. As your AI usage grows, Gemini models can automatically scale on demand without hitting artificial rate limits – processing up to 50 million tokens per minute.

  •  

[Launched] Generally Available: Summarized advertised gateway prefixes for route advertisement

Summarized advertised gateway prefixes for route advertisement is now generally available. You can specify aggregated (summarized) prefixes for an Azure gateway to advertise to your on-premises networks, rather than having every individual virtual network
  •  

How Deutsche Bank unlocked agility with an API-ready ecosystem

When people think about digital transformation in banking, they often focus on the visible results: mobile apps and new digital services. But there's an invisible infrastructure making all these services possible: APIs. At Deutsche Bank, we recognized that APIs aren't just technical plumbing; they're the nervous system of modern banking. 

A few years ago, our application landscape was dominated by monolithic systems. As we evaluated how to break them into modular, reusable APIs, one thing became clear: we couldn't just decompose our work into APIs — we needed a central API management platform (APIM) to manage what would emerge. We needed something where documentation, security policies, and governance all had to be built in from the start, not bolted on later. 

The question wasn't just how to modernize, but how to best serve our customers and position ourselves for tomorrow's opportunities, especially with emerging technological paradigm shifts. 

Needing a system that was adaptable, scalable, reliable, secure, and AI-ready for the demands of modern banking, we chose Google Cloud's Apigee as our APIM platform. 

Building the backbone: four key capabilities 

Today, Apigee manages our API ecosystem — from open banking APIs connecting us with fintech partners, to the internal microservices powering our various banking platforms, and even the client-facing applications that enable seamless digital experiences such as online banking. 

Here are four important capabilities the platform offers us:

1. Unified governance without sacrificing speed 

Apigee is the foundation of our API catalog. Every endpoint, version, and dependency is documented and discoverable. Development teams find and reuse existing APIs rather than rebuild functionality. We've moved from "Where's that customer data API?" — which took days — to a searchable, real-time catalog accessible to any developer. 

But governance isn't about bottlenecks, it's about guardrails, and with Apigee's policy framework, we automatically enforce standards. OpenAPI specifications, schema validation, and error handling are now baked into the platform. Teams move faster because they work within consistent frameworks. 

2. Security: the employee onboarding analogy 

When thinking about API security, imagine onboarding a new employee. You don't give them access to every system on day one. You follow the least privilege principle, so they get exactly the permissions needed for their role. If they switch departments, their access rights will be updated. If they leave the company, access is revoked immediately. Apigee works the same way for our services and applications. 

When connecting a new service — say, one that accesses customer accounts — we don't open the floodgates. Through OAuth2 scopes and API key management, we define precisely what that agent can access: 

  • Read account balances? Yes. 
  • Initiate wire transfers? No. 
  • Access 90-day transaction history? Yes. 
  • Full historical data? Only with elevated permissions. 

Like employee access, these permissions are centrally managed, regularly audited, and instantly revocable. Just as we track employee activity for compliance, Apigee logs every API call to see who accessed what data, when, and why. 

This becomes critical with high-volume automated systems. An automated service doesn't take breaks and can make thousands of calls per minute if misconfigured. Rate limiting and quota enforcement ensure that even when something goes wrong, the blast radius is contained. 

3. Resilience and performance at scale 

Banking doesn't have downtime. When customers check balances at 3 a.m. or markets surge with trading activity, our APIs must respond instantly and reliably. 

Apigee's load balancing and auto-scaling evenly distribute that traffic. Health checks and circuit breakers automatically route around struggling services, and for frequently accessed data, Apigee's caching delivers sub-millisecond responses without hitting backends. 

4. Observability: measuring everything 

Before Apigee, understanding API performance was like assembling a jigsaw puzzle with pieces from different boxes. Now we have unified dashboards showing real-time traffic, error rates by service, usage analytics by consumer, and compliance metrics. This visibility serves operations, product managers who track partner value, and security teams who identify anomalies.

DtBank_Apigee_1

Apigee provides a central suite of capabilities for managing the full API lifecycle

The path forward 

We built this infrastructure for the API economy, and in doing so, we have also built a strong foundation for the future of digital banking. As the industry evolves, this API-first approach will be critical for integrating next-generation services. 

As digital banking continues to advance, a shift toward intelligent services that can react, predict, and assist in real time is underway. Capabilities such as real-time pattern recognition, predictive insights, and AI-powered assistants are becoming part of everyday digital experiences, with their visibility and impact increasing as adoption accelerates. Each of these capabilities will consume APIs — and they will introduce new requirements: ultra-low latency, high-throughput data flows, and secure orchestration across multiple APIs. 

Because we invested in a flexible API platform with Apigee, we are well-positioned to adapt and optimize our infrastructure for these future needs, rather than having to rebuild it. 

Emerging standards: MCP, A2A, and the future 

The industry is exploring new integration standards. Protocols like Model Context Protocol (MCP) and Google's Agent2Agent (A2A) are interesting because they build on existing API infrastructure. 

Our Apigee-managed APIs are well-positioned to leverage these advancements. For instance, MCP could benefit from our OpenAPI specifications, and A2A could leverage our OAuth2 framework, with both relying on the governance we've built. 

We're also exploring patterns like placing new types of servers behind Apigee proxies to maintain security controls while enabling modern workflows. Our "always-API" pattern ensures that services benefit from centralized management, no matter how they are accessed.

DtBank_Apigee_2

MCP and A2A are complementary, MCP has a tools and resources focus, while A2A is focused on peer collaboration

The vision: APIs as universal interface 

Every banking capability will eventually be exposed as an API. That’s not because APIs are trendy, but because they're the most flexible, composable, and governable way to share functionality, whether consumed by mobile apps, partner fintechs, analytics platforms, or other automated agents. 

At Deutsche Bank, this shift is already taking shape. The same API foundation that powers our core platforms is now enabling our evolution toward more intelligent, AI-supported services across the bank. That foundation provides the consistency, governance, and scalability needed to bring these capabilities to life, ensuring that as new intelligent services emerge, they can be integrated seamlessly, securely, and at enterprise scale. 

Apigee makes this possible by providing governance that scales across all use cases. It's not about controlling innovation; it's about enabling it safely. 

Lessons learned 

  • Invest in excellent documentation. Semantic summaries and clear schemas aren't extras; they're foundational for both developers and AI. 

  • Treat security like employee onboarding. Least privilege and role-based access apply equally to APIs.

  • Observability is a competitive advantage. Unified analytics enable data-driven decisions. 

  • Plan for the future now. Your API management infrastructure becomes your advanced integration layer. 

  • Stay curious. Experiment with emerging standards. Flexibility wins. 

Conclusion 

We're at an inflection point. The API economy enabled fintech and open banking. Now, the same infrastructure can serve as the backbone for the next wave of innovation. Our investment in the API platform wasn't just about managing APIs better; it was about building a foundation for whatever comes next. 

As the industry transforms, we’re ready. The future belongs to organizations that move fast without breaking things. For us, that future is powered by Apigee.

  •  

Detect early and enforce firmly with Google Cloud's enhanced cost controls for AI spend

Generative AI can make cloud costs difficult to predict. A single five-word prompt can run complex operations and generate significant costs. Traditional metrics like requests per second no longer help you estimate your bill, increasing the risk of unexpected cost spikes.

While Google Cloud offers budget alerts, you need proactive guardrails that work out of the box - without the friction of writing custom JSON policies, configuring complex roles, or maintaining bespoke scripts.  

Today, we are announcing two native features in the Google Cloud Billing console: early anomalies on AI services and spend caps on Google Cloud Budgets. Together, these tools help you detect unusual spend early and enforce firm limits so you can build and experiment with confidence.

A two-pronged defense playbook

The inherent variability and potential for rapid scaling in AI usage demand a new level of responsiveness from cost management tools. To address this, we're providing tools that allow users to effortlessly set up a GCP defense playbook. 

1. Detect: early anomalies on AI services

This new feature, located within the Anomalies tool on the billing console, automatically alerts users of directional variances (anomalies) in daily AI costs within a project. While Google Cloud has offered standard Cost Anomaly Detection, this new capability represents a major architectural shift designed specifically for the speed of AI workloads.

How do early anomalies work?

  • Dynamic baseline modeling: The system automatically analyzes historical project data to build an expected seasonal baseline of daily service-level costs—no manual threshold configuration required.

  • Early cost monitoring: Instead of waiting for billing cycles to reconcile, the tool monitors early cost signals per service and alerts before actual costs are reported. 

  • Automated triage with RCA: If a daily cost signal trends abnormally, the system flags the deviation and generates a Root Cause Analysis (RCA) highlighting the top 3 SKUs driving the surge.

Jul22_Anomalies_Image1

Screenshot of the Billing Console showing an Early Anomaly alert with the Root Cause Analysis (RCA) breakdown highlighting the driving SKUs.

Early anomaly signals allow your team to investigate upward trending costs and intervene before they escalate dramatically. Monitoring these daily variances helps you easily identify the most anomalous high-velocity services, making them perfect candidates for a firm enforcement cap.

2. Enforce: Spend caps on Google Cloud Budgets

Once you identify a service that needs tight boundaries, the second step of the playbook comes in. Spend Caps is a new, native feature on Google Cloud Budgets that empowers you to set a monthly financial cap on specific services within a project. When accumulated spend reaches your defined cap, the system automatically restricts further cost-incurring usage for that specific service within that project. This provides a crucial safety net for your cloud finances without risking the rest of your infrastructure.

How do spend caps work?

During this Public Preview, you can apply Spend Caps to a single project and service for a fixed monthly timeframe.

Jul22_SpendCapsCreate_Image2

Screenshot of the Google Cloud Budgets interface highlighting the "Budget Type" selection with the new "Spend Cap" option enabled.

Once the cumulative costs reach your defined cap, Google Cloud automatically takes action to prevent further billable usage.

What makes Spend Caps so special? 

  • It is non-destructive i.e your data and resources are not deleted, and services outside the scope of the budget are entirely unaffected.

  • You stay informed i.e Billing Administrators and Project Owners receive automated email alerts at 50%, 80%, and 100% of the budget

  • One-click recovery i.e if a Spend Cap is triggered, the usage block remains in place until it is manually lifted from the Google Cloud Budgets UI with a single click.

Jul22_SpendCapsEdit_Image3

Screenshot of the "Lift spend cap" button in the Google Cloud Budgets UI showing the frictionless manual reversal flow

While spend caps successfully halt new on-demand charges by pausing usage, any underlying fixed commitment fees (such as Committed Use Discounts or Provisioned Throughput) will continue to bill at their flat contractual rate.

When do spend caps trigger?

Since AI and cloud costs can rack up at lightning speed, guardrails must act with equal urgency. Traditional billing data can sometimes take hours to reconcile, but Spend Caps for AI services trigger within minutes of hitting your defined threshold. This rapid, near real-time enforcement is specifically engineered to limit your AI specific financial exposure before a runaway model, an infinite loop, or a massive query can spike your bill. 

The defense playbook at a glance

By combining early anomaly signals with automated financial boundaries, Google Cloud provides a complete closed-loop defense for your AI spend.

Playbook step

Tool

What it does

Key benefit

Supported servicesPreview 

Step 1: Detect

Early Anomalies on Services

Analyzes early cost signals daily to alert you of directional spend variances before they hit your invoice.

Spot cost spikes early and pinpoint the exact SKU driving the trend.

Gemini API, Agent Platform, Cloud Run, Cloud Run Functions

Step 2: Enforce

Spend Caps on Budgets

Automatically halts usage for a specific project/service once your defined budget is breached.

Protection against runaway spend without deleting resources

Gemini API, Agent Platform, Cloud Run, Cloud Run Functions

By deploying this dual-pronged defense strategy that works out-of-the-box, your engineering teams can move quickly, experiment safely, and focus entirely on innovation.

Get started today

  • Explore early anomalies: Head to the "Anomalies" section of your Google Cloud Billing Console.

  • Set your first spend cap: Go to the Budgets & Alerts page in the Billing Console to apply native caps to your test or development environments.

  • For more details on configuration, service-specific behaviors, and permissions, read the Manage Anomalies and Spend Caps on Budgets documentation.
  •  

[In preview] Public Preview: Protect sensitive generative AI telemetry in Application Insights and Microsoft Foundry

Azure Monitor Application Insights now stores generative AI content in a dedicated GenAIContent table (AppGenAIContent in Log Analytics), making it possible to apply newly available access controls to sensitive AI telemetry. Because Application Insights p
  •  

[Launched] Generally Available: Support 5x churn in Azure Site Recovery

Azure Site Recovery now supports up to 5x churn (500 MB/s per VM). This major enhancement empowers customers to confidently run high IOPS workloads with Azure Site Recovery, ensuring robust disaster recovery for even the most demanding applications like d
  •  

[Launched] Generally Available: Azure Site Recovery support for Linux Azure VMs with NVMe disk controllers.

Azure Site Recovery now supports replication and disaster recovery for Linux Azure Virtual Machines running on NVMe-enabled Generation 2 VM families, such as the Da/Ea/Fa v6-series and Ebsv5/Ebdsv5 in the Azure-to-Azure scenario. This support is limited t
  •