Vue normale
-
Azure service updates
- [Launched] Generally Available: High-scale mesh in Azure Virtual Network Manager
[Launched] Generally Available: High-scale mesh in Azure Virtual Network Manager
-
Azure service updates
- [Launched] Generally Available: Azure Virtual Network Manager IPAM in additional Azure regions
[Launched] Generally Available: Azure Virtual Network Manager IPAM in additional Azure regions
-
Azure service updates
- [Launched] Generally Available: Azure VM Image Builder in sovereign and air-gapped clouds
[Launched] Generally Available: Azure VM Image Builder in sovereign and air-gapped clouds
FinOps for the AI era: New flexible billing and cost controls for agents
Editor's note: A product image was updated after initial publication.
As AI takes on more complex work, business leaders face a new challenge: enabling rapid innovation using agents while protecting their margins and budgets. To get a real return on AI, financial operations (FinOps) and cost management must evolve alongside technology, giving you clear visibility, proactive cost controls, and flexible payment models that fit your needs.
That’s why today we’re introducing expanded billing flexibility and new cost management tools for agent workloads across Gemini Enterprise and developer tools like Google Antigravity in Gemini Enterprise and Android Studio.
-
Flexible payment options: You can mix our existing, predictable per-user seat subscriptions with a new pay-as-you-go option in Gemini Enterprise app that lets you run agent workloads without hitting quota limits mid-task.
-
Developer access, one place to manage your AI: Google Antigravity and Android Studio AI use is now included in your Gemini Enterprise subscription (available for select customers and rolling out broadly soon), giving your developers more without giving you more to manage. Usage across Antigravity, the platform, and the app rolls up into a single view instead of separate licenses and billing silos.
-
Pay less as your usage grows: If your AI workloads are steady or climbing, Flexible Savings Plans let you commit to a monthly spend you're comfortable with and take 10–20% off your token costs — no minimums, no maximums, and no new billing silo to manage.
-
Consolidated spend guardrails: You can now set hard monthly caps on AI spend and projects, estimate agent runtime costs, and catch sudden budget spikes before they hit your invoice.
Give your teams flexibility without losing control over spend in Gemini Enterprise
Every organization operates differently. Even within the same business, no two teams consume AI in the same way. Your business users might rely on steady, everyday productivity tools. Meanwhile, your technical teams might run AI agent workloads in bursts.
To help align costs with how work actually gets done, you can combine these payment and licensing choices and features across Gemini Enterprise:
|
Option |
How it works |
Why it helps optimize spend |
|
Gemini Enterprise app per-user seat subscription |
You pay a fixed monthly fee per user, which includes daily quota pools that are shared across your entire project. |
Predictable budgeting. It provides finance teams with a clear, steady monthly baseline for teams with consistent daily productivity needs. |
|
[New] Gemini Enterprise app pay-as-you-go consumption edition *available for select customers and rolling out broadly soon |
There is no upfront commitment or base subscription fee, meaning you pay strictly for the compute and tokens your teams consume at standard model API rates. |
Only pay for what you use. Your spend scales up and down automatically with real usage, ensuring you never pay for empty seats when project demand dips. |
|
[New for Antigravity in Gemini Enterprise] Consolidated pooled quotas |
Daily usage allowances are pooled project-wide, letting business apps, developer tools, and custom agents draw from the same shared quota. Pooled quota is always exhausted first, and admins can control if overages are allowed, at which point it’s charged at pay-as-you-go rates. |
Maximized resource usage: Unused daily allowances from business users automatically absorb heavy developer or custom API agent demands, so no quota allowance goes to waste. |
|
[Coming soon] Deferred execution pricing *available for select workloads soon |
Mark eligible agent workloads as deferred, and our intelligent scheduler in the Gemini Enterprise Agent Platform runs them during off-peak capacity windows. |
Substantial discounts for work that can wait: AI workloads can run on separate, off-peak capacity, you pay up to half the inference cost and bypass standard quota limits entirely – letting you run substantially more agentic volume under the same budget. |
Equip developers with advanced agentic tooling under a single Gemini Enterprise subscription
We’re rolling out access to Google Antigravity in Gemini Enterprise, an agent-first developer platform that brings powerful agentic coding and agent-building capabilities to technical teams, included with Gemini Enterprise subscriptions for eligible customers. In addition, Android developers can leverage the Google Antigravity quota included in their Gemini Enterprise subscriptions natively in Android Studio, the agentic IDE for professional Android development.
To be more efficient with agentic coding costs, we are pooling developer tools quota included in each Gemini Enterprise subscription and making it available across the whole Google Cloud project so your teams can benefit from the capacity you’re already purchasing. Your developers get access to advanced agentic tools, while you maintain centralized governance and control.
For a closer look into what’s new with Antigravity in Gemini Enterprise and how customers are putting it to work in production, take a look at our deep-dive.
Budget smarter with Gemini Enterprise Flexible Savings Plans (FSPs)
If your organization has steady or growing AI workloads, Gemini Enterprise Flexible Savings Plans offer a simple, spend-based commitment model across Gemini Enterprise usage. FSPs are designed to lower token costs while keeping budgets flexible:
-
Programmatic savings: Receive 10% off for 1-year or 20% off for 3-year commitments for monthly spending across Gemini Enterprise.
-
Tailored to your pace: With no minimum or maximum spend requirements, you can determine a monthly commitment that fits your current traffic and make adjustments as your usage increases over time.
-
Enterprise Agreement (EA) friendly: FSP spend seamlessly draws down against your existing Google Cloud EA, giving lines of business dedicated budget control without fragmenting your broader cloud commitments.
Gemini Enterprise Flexible Savings Plans are already available for self-serve customers and customers on enterprise agreements.
Give your teams the freedom to build while maintaining financial discipline
As a leader, your goal isn't to restrict the potential value of AI – it's to remove the financial and operational risk that you face without managed AI costs. You should be able to give engineering, marketing, and operational teams the freedom to innovate with agents, but you should also have the visibility to trust what those agents are doing and the safety nets to protect your budget.
To bridge this gap, we've built robust, native governance tooling directly into the Google Cloud Billing Console around three simple goals:
1. Plan before you scale: The Google Cloud Pricing Calculator lets you estimate anticipated costs in Gemini Enterprise across per-user licenses, developer tools, and background agent runtimes. It gives you the numbers you need to build clear business cases upfront before project work begins.
2. Enforce boundaries without micromanaging spend: Instead of spending time tracking daily usage variations across project teams, let these tools do the monitoring for you:
-
Early anomaly detection: If a project’s AI spending trends higher than normal, the system flags the deviation with root cause analysis and pinpoints the top 3 SKUs driving the increase so you can see exactly what changed.
Billing Console showing an Early Anomaly alert with the Root Cause Analysis (RCA) breakdown highlighting the driving SKUs
-
- Project-level spend caps: When a project needs defined financial boundaries, you can set a firm monthly spend limit directly in the Google Cloud Billing Console. If a project hits its limit, the agent's API calls temporarily pause – protecting your budget without affecting the rest of your production infrastructure. Automated email alerts at 50%, 80% and 100% of the budget keep you informed of your progress against the spend limit.
-
- Overage controls: If a spend cap triggers, you can choose to resume work with a single click in the console. Alternatively, if your priority is continuous operation, you can turn on overages so excess usage smoothly transitions to consumption rates, which can draw directly against your FSP to keep overage unit costs heavily discounted.
Enabling overage pay-as-you-go for a project.
3. Get visibility into business value: Use centralized billing reports paired with the FinOps agent to generate natural-language cost insight summaries of where your budget went, making it simple to show ROI to leadership.
AI spending reporting in Google Cloud Console
Go deeper with AI cost optimization
To build a full-stack FinOps strategy that optimizes the cost, latency, and performance of your models and infrastructure, explore our detailed architecture specifications and frameworks:
-
How to outsmart infrastructure constraints with dynamic capacity management: Discover how to optimize your compute investments with capabilities in Google Kubernetes Engine and Google Compute Engine that automatically schedule and reallocate resources to avoid interruptions, over-provisioning, and over-reliance on any one hardware configuration.
-
Expanding Google Antigravity for Enterprise Customers: Read our developer tooling deep-dive to see how technical teams are accelerating software delivery with agent-first workflows.
-
What sports cars can teach us about optimizing AI spend: More tokens doesn't always mean better AI. Read our conversation with Mike Clark, Director of Product Management for Gemini Enterprise Agent Platform, on how to balance horsepower with efficiency and get the highest return out of every dollar you spend on AI.
-
Protection during usage spikes: Your heavy workloads can surge during peak hours without forcing you to pay for expensive, dedicated infrastructure that sits idle the rest of the time. As your AI usage grows, Gemini models can automatically scale on demand without hitting artificial rate limits – processing up to 50 million tokens per minute.

-
Azure service updates
- [Launched] Generally Available: Summarized advertised gateway prefixes for route advertisement
[Launched] Generally Available: Summarized advertised gateway prefixes for route advertisement
[Launched] Generally Available: Azure Virtual Network routing appliance
How Deutsche Bank unlocked agility with an API-ready ecosystem
When people think about digital transformation in banking, they often focus on the visible results: mobile apps and new digital services. But there's an invisible infrastructure making all these services possible: APIs. At Deutsche Bank, we recognized that APIs aren't just technical plumbing; they're the nervous system of modern banking.
A few years ago, our application landscape was dominated by monolithic systems. As we evaluated how to break them into modular, reusable APIs, one thing became clear: we couldn't just decompose our work into APIs — we needed a central API management platform (APIM) to manage what would emerge. We needed something where documentation, security policies, and governance all had to be built in from the start, not bolted on later.
The question wasn't just how to modernize, but how to best serve our customers and position ourselves for tomorrow's opportunities, especially with emerging technological paradigm shifts.
Needing a system that was adaptable, scalable, reliable, secure, and AI-ready for the demands of modern banking, we chose Google Cloud's Apigee as our APIM platform.
Building the backbone: four key capabilities
Today, Apigee manages our API ecosystem — from open banking APIs connecting us with fintech partners, to the internal microservices powering our various banking platforms, and even the client-facing applications that enable seamless digital experiences such as online banking.
Here are four important capabilities the platform offers us:
1. Unified governance without sacrificing speed
Apigee is the foundation of our API catalog. Every endpoint, version, and dependency is documented and discoverable. Development teams find and reuse existing APIs rather than rebuild functionality. We've moved from "Where's that customer data API?" — which took days — to a searchable, real-time catalog accessible to any developer.
But governance isn't about bottlenecks, it's about guardrails, and with Apigee's policy framework, we automatically enforce standards. OpenAPI specifications, schema validation, and error handling are now baked into the platform. Teams move faster because they work within consistent frameworks.
2. Security: the employee onboarding analogy
When thinking about API security, imagine onboarding a new employee. You don't give them access to every system on day one. You follow the least privilege principle, so they get exactly the permissions needed for their role. If they switch departments, their access rights will be updated. If they leave the company, access is revoked immediately. Apigee works the same way for our services and applications.
When connecting a new service — say, one that accesses customer accounts — we don't open the floodgates. Through OAuth2 scopes and API key management, we define precisely what that agent can access:
- Read account balances? Yes.
- Initiate wire transfers? No.
- Access 90-day transaction history? Yes.
- Full historical data? Only with elevated permissions.
Like employee access, these permissions are centrally managed, regularly audited, and instantly revocable. Just as we track employee activity for compliance, Apigee logs every API call to see who accessed what data, when, and why.
This becomes critical with high-volume automated systems. An automated service doesn't take breaks and can make thousands of calls per minute if misconfigured. Rate limiting and quota enforcement ensure that even when something goes wrong, the blast radius is contained.
3. Resilience and performance at scale
Banking doesn't have downtime. When customers check balances at 3 a.m. or markets surge with trading activity, our APIs must respond instantly and reliably.
Apigee's load balancing and auto-scaling evenly distribute that traffic. Health checks and circuit breakers automatically route around struggling services, and for frequently accessed data, Apigee's caching delivers sub-millisecond responses without hitting backends.
4. Observability: measuring everything
Before Apigee, understanding API performance was like assembling a jigsaw puzzle with pieces from different boxes. Now we have unified dashboards showing real-time traffic, error rates by service, usage analytics by consumer, and compliance metrics. This visibility serves operations, product managers who track partner value, and security teams who identify anomalies.
Apigee provides a central suite of capabilities for managing the full API lifecycle
The path forward
We built this infrastructure for the API economy, and in doing so, we have also built a strong foundation for the future of digital banking. As the industry evolves, this API-first approach will be critical for integrating next-generation services.
As digital banking continues to advance, a shift toward intelligent services that can react, predict, and assist in real time is underway. Capabilities such as real-time pattern recognition, predictive insights, and AI-powered assistants are becoming part of everyday digital experiences, with their visibility and impact increasing as adoption accelerates. Each of these capabilities will consume APIs — and they will introduce new requirements: ultra-low latency, high-throughput data flows, and secure orchestration across multiple APIs.
Because we invested in a flexible API platform with Apigee, we are well-positioned to adapt and optimize our infrastructure for these future needs, rather than having to rebuild it.
Emerging standards: MCP, A2A, and the future
The industry is exploring new integration standards. Protocols like Model Context Protocol (MCP) and Google's Agent2Agent (A2A) are interesting because they build on existing API infrastructure.
Our Apigee-managed APIs are well-positioned to leverage these advancements. For instance, MCP could benefit from our OpenAPI specifications, and A2A could leverage our OAuth2 framework, with both relying on the governance we've built.
We're also exploring patterns like placing new types of servers behind Apigee proxies to maintain security controls while enabling modern workflows. Our "always-API" pattern ensures that services benefit from centralized management, no matter how they are accessed.
MCP and A2A are complementary, MCP has a tools and resources focus, while A2A is focused on peer collaboration
The vision: APIs as universal interface
Every banking capability will eventually be exposed as an API. That’s not because APIs are trendy, but because they're the most flexible, composable, and governable way to share functionality, whether consumed by mobile apps, partner fintechs, analytics platforms, or other automated agents.
At Deutsche Bank, this shift is already taking shape. The same API foundation that powers our core platforms is now enabling our evolution toward more intelligent, AI-supported services across the bank. That foundation provides the consistency, governance, and scalability needed to bring these capabilities to life, ensuring that as new intelligent services emerge, they can be integrated seamlessly, securely, and at enterprise scale.
Apigee makes this possible by providing governance that scales across all use cases. It's not about controlling innovation; it's about enabling it safely.
Lessons learned
-
Invest in excellent documentation. Semantic summaries and clear schemas aren't extras; they're foundational for both developers and AI.
-
Treat security like employee onboarding. Least privilege and role-based access apply equally to APIs.
-
Observability is a competitive advantage. Unified analytics enable data-driven decisions.
-
Plan for the future now. Your API management infrastructure becomes your advanced integration layer.
-
Stay curious. Experiment with emerging standards. Flexibility wins.
Conclusion
We're at an inflection point. The API economy enabled fintech and open banking. Now, the same infrastructure can serve as the backbone for the next wave of innovation. Our investment in the API platform wasn't just about managing APIs better; it was about building a foundation for whatever comes next.
As the industry transforms, we’re ready. The future belongs to organizations that move fast without breaking things. For us, that future is powered by Apigee.

Detect early and enforce firmly with Google Cloud's enhanced cost controls for AI spend
Generative AI can make cloud costs difficult to predict. A single five-word prompt can run complex operations and generate significant costs. Traditional metrics like requests per second no longer help you estimate your bill, increasing the risk of unexpected cost spikes.
While Google Cloud offers budget alerts, you need proactive guardrails that work out of the box - without the friction of writing custom JSON policies, configuring complex roles, or maintaining bespoke scripts.
Today, we are announcing two native features in the Google Cloud Billing console: early anomalies on AI services and spend caps on Google Cloud Budgets. Together, these tools help you detect unusual spend early and enforce firm limits so you can build and experiment with confidence.
A two-pronged defense playbook
The inherent variability and potential for rapid scaling in AI usage demand a new level of responsiveness from cost management tools. To address this, we're providing tools that allow users to effortlessly set up a GCP defense playbook.
1. Detect: early anomalies on AI services
This new feature, located within the Anomalies tool on the billing console, automatically alerts users of directional variances (anomalies) in daily AI costs within a project. While Google Cloud has offered standard Cost Anomaly Detection, this new capability represents a major architectural shift designed specifically for the speed of AI workloads.
How do early anomalies work?
-
Dynamic baseline modeling: The system automatically analyzes historical project data to build an expected seasonal baseline of daily service-level costs—no manual threshold configuration required.
-
Early cost monitoring: Instead of waiting for billing cycles to reconcile, the tool monitors early cost signals per service and alerts before actual costs are reported.
-
Automated triage with RCA: If a daily cost signal trends abnormally, the system flags the deviation and generates a Root Cause Analysis (RCA) highlighting the top 3 SKUs driving the surge.
Screenshot of the Billing Console showing an Early Anomaly alert with the Root Cause Analysis (RCA) breakdown highlighting the driving SKUs.
Early anomaly signals allow your team to investigate upward trending costs and intervene before they escalate dramatically. Monitoring these daily variances helps you easily identify the most anomalous high-velocity services, making them perfect candidates for a firm enforcement cap.
2. Enforce: Spend caps on Google Cloud Budgets
Once you identify a service that needs tight boundaries, the second step of the playbook comes in. Spend Caps is a new, native feature on Google Cloud Budgets that empowers you to set a monthly financial cap on specific services within a project. When accumulated spend reaches your defined cap, the system automatically restricts further cost-incurring usage for that specific service within that project. This provides a crucial safety net for your cloud finances without risking the rest of your infrastructure.
How do spend caps work?
During this Public Preview, you can apply Spend Caps to a single project and service for a fixed monthly timeframe.
Screenshot of the Google Cloud Budgets interface highlighting the "Budget Type" selection with the new "Spend Cap" option enabled.
Once the cumulative costs reach your defined cap, Google Cloud automatically takes action to prevent further billable usage.
What makes Spend Caps so special?
-
It is non-destructive i.e your data and resources are not deleted, and services outside the scope of the budget are entirely unaffected.
-
You stay informed i.e Billing Administrators and Project Owners receive automated email alerts at 50%, 80%, and 100% of the budget
-
One-click recovery i.e if a Spend Cap is triggered, the usage block remains in place until it is manually lifted from the Google Cloud Budgets UI with a single click.
Screenshot of the "Lift spend cap" button in the Google Cloud Budgets UI showing the frictionless manual reversal flow
While spend caps successfully halt new on-demand charges by pausing usage, any underlying fixed commitment fees (such as Committed Use Discounts or Provisioned Throughput) will continue to bill at their flat contractual rate.
When do spend caps trigger?
Since AI and cloud costs can rack up at lightning speed, guardrails must act with equal urgency. Traditional billing data can sometimes take hours to reconcile, but Spend Caps for AI services trigger within minutes of hitting your defined threshold. This rapid, near real-time enforcement is specifically engineered to limit your AI specific financial exposure before a runaway model, an infinite loop, or a massive query can spike your bill.
The defense playbook at a glance
By combining early anomaly signals with automated financial boundaries, Google Cloud provides a complete closed-loop defense for your AI spend.
|
Playbook step |
Tool |
What it does |
Key benefit |
Supported servicesPreview |
|
Step 1: Detect |
Early Anomalies on Services |
Analyzes early cost signals daily to alert you of directional spend variances before they hit your invoice. |
Spot cost spikes early and pinpoint the exact SKU driving the trend. |
Gemini API, Agent Platform, Cloud Run, Cloud Run Functions |
|
Step 2: Enforce |
Spend Caps on Budgets |
Automatically halts usage for a specific project/service once your defined budget is breached. |
Protection against runaway spend without deleting resources |
Gemini API, Agent Platform, Cloud Run, Cloud Run Functions |
By deploying this dual-pronged defense strategy that works out-of-the-box, your engineering teams can move quickly, experiment safely, and focus entirely on innovation.
Get started today
-
Explore early anomalies: Head to the "Anomalies" section of your Google Cloud Billing Console.
-
Set your first spend cap: Go to the Budgets & Alerts page in the Billing Console to apply native caps to your test or development environments.
- For more details on configuration, service-specific behaviors, and permissions, read the Manage Anomalies and Spend Caps on Budgets documentation.
[In preview] Public Preview: AI Gateway in Azure API Management
-
Azure service updates
- [In preview] Public Preview: Protect sensitive generative AI telemetry in Application Insights and Microsoft Foundry
[In preview] Public Preview: Protect sensitive generative AI telemetry in Application Insights and Microsoft Foundry
[Launched] Generally Available: Support 5x churn in Azure Site Recovery
-
Azure service updates
- [Launched] Generally Available: Azure Site Recovery support for Linux Azure VMs with NVMe disk controllers.