❌

Vue normale

Reçu avant avant-hier

FinOps for the AI era: New flexible billing and cost controls for agents

26 août 2026 à 15:30

Editor's note: A product image was updated after initial publication.


As AI takes on more complex work, business leaders face a new challenge: enabling rapid innovation using agents while protecting their margins and budgets. To get a real return on AI, financial operations (FinOps) and cost management must evolve alongside technology, giving you clear visibility, proactive cost controls, and flexible payment models that fit your needs. 

That’s why today we’re introducing expanded billing flexibility and new cost management tools for agent workloads across Gemini Enterprise and developer tools like Google Antigravity in Gemini Enterprise and Android Studio.

  • Flexible payment options: You can mix our existing, predictable per-user seat subscriptions with a new pay-as-you-go option in Gemini Enterprise app that lets you run agent workloads without hitting quota limits mid-task.

  • Developer access, one place to manage your AI: Google Antigravity and Android Studio AI use is now included in your Gemini Enterprise subscription (available for select customers and rolling out broadly soon), giving your developers more without giving you more to manage. Usage across Antigravity, the platform, and the app rolls up into a single view instead of separate licenses and billing silos.

  • Pay less as your usage grows: If your AI workloads are steady or climbing, Flexible Savings Plans let you commit to a monthly spend you're comfortable with and take 10–20% off your token costs — no minimums, no maximums, and no new billing silo to manage.

  • Consolidated spend guardrails: You can now set hard monthly caps on AI spend and projects, estimate agent runtime costs, and catch sudden budget spikes before they hit your invoice.

Give your teams flexibility without losing control over spend in Gemini Enterprise

Every organization operates differently. Even within the same business, no two teams consume AI in the same way. Your business users might rely on steady, everyday productivity tools. Meanwhile, your technical teams might run AI agent workloads in bursts. 

To help align costs with how work actually gets done, you can combine these payment and licensing choices and features across Gemini Enterprise:

Option

How it works

Why it helps optimize spend

Gemini Enterprise app per-user seat subscription

You pay a fixed monthly fee per user, which includes daily quota pools that are shared across your entire project.

Predictable budgeting. It provides finance teams with a clear, steady monthly baseline for teams with consistent daily productivity needs.

[New] Gemini Enterprise app pay-as-you-go consumption edition

*available for select customers and rolling out broadly soon

There is no upfront commitment or base subscription fee, meaning you pay strictly for the compute and tokens your teams consume at standard model API rates.

Only pay for what you use. Your spend scales up and down automatically with real usage, ensuring you never pay for empty seats when project demand dips.

[New for Antigravity in Gemini Enterprise] Consolidated pooled quotas

Daily usage allowances are pooled project-wide, letting business apps, developer tools, and custom agents draw from the same shared quota. Pooled quota is always exhausted first, and admins can control if overages are allowed, at which point it’s charged at pay-as-you-go rates. 

Maximized resource usage: Unused daily allowances from business users automatically absorb heavy developer or custom API agent demands, so no quota allowance goes to waste.

[Coming soon] Deferred execution pricing

*available for select workloads soon

Mark eligible agent workloads as deferred, and our intelligent scheduler in the Gemini Enterprise Agent Platform runs them during off-peak capacity windows.

Substantial discounts for work that can wait: AI workloads can run on separate, off-peak capacity, you pay up to half the inference cost and bypass standard quota limits entirely – letting you run substantially more agentic volume under the same budget.

Equip developers with advanced agentic tooling under a single Gemini Enterprise subscription

We’re rolling out access to Google Antigravity in Gemini Enterprise, an agent-first developer platform that brings powerful agentic coding and agent-building capabilities to technical teams, included with Gemini Enterprise subscriptions for eligible customers. In addition, Android developers can leverage the Google Antigravity quota included in their Gemini Enterprise subscriptions natively in Android Studio, the agentic IDE for professional Android development.

To be more efficient with agentic coding costs, we are pooling developer tools quota included in each Gemini Enterprise subscription and making it available across the whole Google Cloud project so your teams can benefit from the capacity you’re already purchasing. Your developers get access to advanced agentic tools, while you maintain centralized governance and control.

For a closer look into what’s new with Antigravity in Gemini Enterprise and how customers are putting it to work in production, take a look at our deep-dive.

Budget smarter with Gemini Enterprise Flexible Savings Plans (FSPs) 

If your organization has steady or growing AI workloads, Gemini Enterprise Flexible Savings Plans offer a simple, spend-based commitment model across Gemini Enterprise usage. FSPs are designed to lower token costs while keeping budgets flexible:

  • Programmatic savings: Receive 10% off for 1-year or 20% off for 3-year commitments for monthly spending across Gemini Enterprise.

  • Tailored to your pace: With no minimum or maximum spend requirements, you can determine a monthly commitment that fits your current traffic and make adjustments as your usage increases over time. 

  • Enterprise Agreement (EA) friendly: FSP spend seamlessly draws down against your existing Google Cloud EA, giving lines of business dedicated budget control without fragmenting your broader cloud commitments.

Gemini Enterprise Flexible Savings Plans are already available for self-serve customers and customers on enterprise agreements.

Give your teams the freedom to build while maintaining financial discipline

As a leader, your goal isn't to restrict the potential value of AI  – it's to remove the financial and operational risk that you face without managed AI costs. You should be able to give engineering, marketing, and operational teams the freedom to innovate with agents, but you should also have the visibility to trust what those agents are doing and the safety nets to protect your budget.

To bridge this gap, we've built robust, native governance tooling directly into the Google Cloud Billing Console around three simple goals:

1. Plan before you scale: The Google Cloud Pricing Calculator lets you estimate anticipated costs in Gemini Enterprise across per-user licenses, developer tools, and background agent runtimes. It gives you the numbers you need to build clear business cases upfront before project work begins.

2. Enforce boundaries without micromanaging spend: Instead of spending time tracking daily usage variations across project teams, let these tools do the monitoring for you:

  • Early anomaly detection: If a project’s AI spending trends higher than normal, the system flags the deviation with root cause analysis and pinpoints the top 3 SKUs driving the increase so you can see exactly what changed.

1 Jul22_Anomalies_Image1

Billing Console showing an Early Anomaly alert with the Root Cause Analysis (RCA) breakdown highlighting the driving SKUs

    • Project-level spend caps: When a project needs defined financial boundaries, you can set a firm monthly spend limit directly in the Google Cloud Billing Console. If a project hits its limit, the agent's API calls temporarily pause – protecting your budget without affecting the rest of your production infrastructure. Automated email alerts at 50%, 80% and 100% of the budget keep you informed of your progress against the spend limit. 
    • Overage controls: If a spend cap triggers, you can choose to resume work with a single click in the console. Alternatively, if your priority is continuous operation, you can turn on overages so excess usage smoothly transitions to consumption rates, which can draw directly against your FSP to keep overage unit costs heavily discounted.
3 PAYG Overage Enabled

Enabling overage pay-as-you-go for a project.

3. Get visibility into business value: Use centralized billing reports paired with the FinOps agent to generate natural-language cost insight summaries of where your budget went, making it simple to show ROI to leadership.

Cost overview FinOps

AI spending reporting in Google Cloud Console

Go deeper with AI cost optimization

To build a full-stack FinOps strategy that optimizes the cost, latency, and performance of your models and infrastructure, explore our detailed architecture specifications and frameworks:

  • How to outsmart infrastructure constraints with dynamic capacity management: Discover how to optimize your compute investments with capabilities in Google Kubernetes Engine and Google Compute Engine that automatically schedule and reallocate resources to avoid interruptions, over-provisioning, and over-reliance on any one hardware configuration.

  • Expanding Google Antigravity for Enterprise Customers: Read our developer tooling deep-dive to see how technical teams are accelerating software delivery with agent-first workflows.

  • What sports cars can teach us about optimizing AI spend: More tokens doesn't always mean better AI. Read our conversation with Mike Clark, Director of Product Management for Gemini Enterprise Agent Platform, on how to balance horsepower with efficiency and get the highest return out of every dollar you spend on AI. 

  • Protection during usage spikes: Your heavy workloads can surge during peak hours without forcing you to pay for expensive, dedicated infrastructure that sits idle the rest of the time. As your AI usage grows, Gemini models can automatically scale on demand without hitting artificial rate limits – processing up to 50 million tokens per minute.

Detect early and enforce firmly with Google Cloud's enhanced cost controls for AI spend

28 juillet 2026 à 18:00

Generative AI can make cloud costs difficult to predict. A single five-word prompt can run complex operations and generate significant costs. Traditional metrics like requests per second no longer help you estimate your bill, increasing the risk of unexpected cost spikes.

While Google Cloud offers budget alerts, you need proactive guardrails that work out of the box - without the friction of writing custom JSON policies, configuring complex roles, or maintaining bespoke scripts.  

Today, we are announcing two native features in the Google Cloud Billing console: early anomalies on AI services and spend caps on Google Cloud Budgets. Together, these tools help you detect unusual spend early and enforce firm limits so you can build and experiment with confidence.

A two-pronged defense playbook

The inherent variability and potential for rapid scaling in AI usage demand a new level of responsiveness from cost management tools. To address this, we're providing tools that allow users to effortlessly set up a GCP defense playbook. 

1. Detect: early anomalies on AI services

This new feature, located within the Anomalies tool on the billing console, automatically alerts users of directional variances (anomalies) in daily AI costs within a project. While Google Cloud has offered standard Cost Anomaly Detection, this new capability represents a major architectural shift designed specifically for the speed of AI workloads.

How do early anomalies work?

  • Dynamic baseline modeling: The system automatically analyzes historical project data to build an expected seasonal baseline of daily service-level costs—no manual threshold configuration required.

  • Early cost monitoring: Instead of waiting for billing cycles to reconcile, the tool monitors early cost signals per service and alerts before actual costs are reported. 

  • Automated triage with RCA: If a daily cost signal trends abnormally, the system flags the deviation and generates a Root Cause Analysis (RCA) highlighting the top 3 SKUs driving the surge.

Jul22_Anomalies_Image1

Screenshot of the Billing Console showing an Early Anomaly alert with the Root Cause Analysis (RCA) breakdown highlighting the driving SKUs.

Early anomaly signals allow your team to investigate upward trending costs and intervene before they escalate dramatically. Monitoring these daily variances helps you easily identify the most anomalous high-velocity services, making them perfect candidates for a firm enforcement cap.

2. Enforce: Spend caps on Google Cloud Budgets

Once you identify a service that needs tight boundaries, the second step of the playbook comes in. Spend Caps is a new, native feature on Google Cloud Budgets that empowers you to set a monthly financial cap on specific services within a project. When accumulated spend reaches your defined cap, the system automatically restricts further cost-incurring usage for that specific service within that project. This provides a crucial safety net for your cloud finances without risking the rest of your infrastructure.

How do spend caps work?

During this Public Preview, you can apply Spend Caps to a single project and service for a fixed monthly timeframe.

Jul22_SpendCapsCreate_Image2

Screenshot of the Google Cloud Budgets interface highlighting the "Budget Type" selection with the new "Spend Cap" option enabled.

Once the cumulative costs reach your defined cap, Google Cloud automatically takes action to prevent further billable usage.

What makes Spend Caps so special? 

  • It is non-destructive i.e your data and resources are not deleted, and services outside the scope of the budget are entirely unaffected.

  • You stay informed i.e Billing Administrators and Project Owners receive automated email alerts at 50%, 80%, and 100% of the budget

  • One-click recovery i.e if a Spend Cap is triggered, the usage block remains in place until it is manually lifted from the Google Cloud Budgets UI with a single click.

Jul22_SpendCapsEdit_Image3

Screenshot of the "Lift spend cap" button in the Google Cloud Budgets UI showing the frictionless manual reversal flow

While spend caps successfully halt new on-demand charges by pausing usage, any underlying fixed commitment fees (such as Committed Use Discounts or Provisioned Throughput) will continue to bill at their flat contractual rate.

When do spend caps trigger?

Since AI and cloud costs can rack up at lightning speed, guardrails must act with equal urgency. Traditional billing data can sometimes take hours to reconcile, but Spend Caps for AI services trigger within minutes of hitting your defined threshold. This rapid, near real-time enforcement is specifically engineered to limit your AI specific financial exposure before a runaway model, an infinite loop, or a massive query can spike your bill. 

The defense playbook at a glance

By combining early anomaly signals with automated financial boundaries, Google Cloud provides a complete closed-loop defense for your AI spend.

Playbook step

Tool

What it does

Key benefit

Supported servicesPreview 

Step 1: Detect

Early Anomalies on Services

Analyzes early cost signals daily to alert you of directional spend variances before they hit your invoice.

Spot cost spikes early and pinpoint the exact SKU driving the trend.

Gemini API, Agent Platform, Cloud Run, Cloud Run Functions

Step 2: Enforce

Spend Caps on Budgets

Automatically halts usage for a specific project/service once your defined budget is breached.

Protection against runaway spend without deleting resources

Gemini API, Agent Platform, Cloud Run, Cloud Run Functions

By deploying this dual-pronged defense strategy that works out-of-the-box, your engineering teams can move quickly, experiment safely, and focus entirely on innovation.

Get started today

  • Explore early anomalies: Head to the "Anomalies" section of your Google Cloud Billing Console.

  • Set your first spend cap: Go to the Budgets & Alerts page in the Billing Console to apply native caps to your test or development environments.

  • For more details on configuration, service-specific behaviors, and permissions, read the Manage Anomalies and Spend Caps on Budgets documentation.
❌