❌

Vue normale

Reçu avant avant-hier

Storage Intelligence advisor: Know what changed in your storage estate and act on it

25 septembre 2026 à 18:00

The volume of data being generated today brings both opportunity and massive operational complexity. Most teams that operate at scale don't discover issues until they appear on an invoice — and by the time an unusual access pattern shows up as a line item, it has often been running for weeks. Understanding what happened means exporting inventory, joining it against access logs, and hoping someone still remembers which service account belongs to which job.

That workflow was manageable in the past, but today’s AI training and inference pipelines create data faster than governance systems can classify it, and read data in patterns that shift from week to week. 

Today we're announcing two new features for Google Cloud Storage: the general availability of Storage Intelligence advisor along with expanded capabilities in storage batch operations. Advisor tells you what changed in your storage estate and what to do about it. Batch operations can carry that decision out across millions of objects. These features are available now to all Storage Intelligence customers.

image1

Storage Intelligence advisor in cloud console.

Storage Intelligence advisor makes reporting easy

For the last decade, answering "what’s in my buckets?" has been a data engineering project. Export your inventory, load it somewhere queryable, join it against usage, build dashboards, and then maintain them. Storage Intelligence delivers visibility without the engineering overhead. Teams are voting with their workloads: the number of customers using Storage Intelligence to analyze datasets of over 1 billion objects has more than doubled this year. 

There are two ways to run a large storage estate. Teams can leverage daily activity data and metadata snapshots to build exactly the pipelines they need –Storage Intelligence still gives you that option – but most teams would prefer not to build pipelines if they don’t have to. They want to be told what changed in their storage environment and what to do about it. Storage Intelligence advisor is for them.

What Advisor gives you on day one

Storage Intelligence advisor brings visibility into your storage without having to perform any setup. Advisor starts from a curated set of findings. There's no schema to design, no pipeline to manage, and no dashboard to assemble. Enable Storage Intelligence on an organization, folder, or project, and charts and findings appear for the buckets in that scope. 

Shipt can now more quickly detect anomalies with Storage Intelligence advisor:

"Before Storage Intelligence advisor, tracking critical usage metrics and catching anomalies [in Google Cloud Storage] required heavy engineering and complex data pipelines. Now, with native, out-of-the-box dashboards, we can instantly identify usage spikes and drill down into the details. Having the visibility to immediately remediate unintended usage — without any configuration — has turned what used to be a major effort into a simple, self-service task." - Charley King, DataOps-DevOps Engineer, Shipt (a subsidiary of Target.com)

Once it’s installed, Advisor immediately starts analyzing the Cloud Storage estate, scanning for anomalies and optimization opportunities including:

  • A spike in Class A or B operations against Coldline or Archive data. Cold storage is cheap to use but expensive to access.

  • A spike in 429 errors. Where a request pattern is outrunning limits, timeouts follow.

  • A spike in cross-region egress.

  • Total consumption rising above a long-term trend.

Each finding is baselined from your project's own activity and metadata and works from daily snapshots of your storage usage, so a spike on one day is surfaced within 24 hours, not a line item you discover at the end of the month. In the last 30 days, over 6,000 findings have been generated across hundreds of customers. 

Take a runaway analytics job that issues millions of daily reads against Archive storage. Without Storage Intelligence advisor, this surfaces as a retrieval-fee weeks later on a bill. 

Advisor identifies the anomaly against your project’s baseline, attributes it to the responsible bucket, prefix, and service account, and points at the controls that apply: bulk-transition the affected objects to Cloud Storage Standard to stop retrieval charges, enable Autoclass so tiering follows real access patterns, or tighten access with Managed Folders so the job can’t reach data it was never meant to access. 

Act on findings with storage batch operations

Most storage recommendations go unactioned because carrying them out is a lot of work. Updating retention policies or storage classes across billions of objects means handling throttling, partial failures, and retries. Storage batch operations removes that work. Execution is fully managed and serverless, with progress tracking and automatic retries built in, so a recommendation becomes a policy-driven job rather than a project. 

Palo Alto Networks had this to say about batch operations:

"Object retention locks were essential for our security guardrails, but managing them across billions of objects was once a non-starter. Storage Intelligence changed that. Today, our team uses storage batch operations to seamlessly update retention policies on demand across our entire fleet." - Kurtis Nusbaum, Senior Principal Software Engineer, Palo Alto Networks

Because Storage Intelligence advisor and batch operations are part of the same Storage Intelligence subscription so customers can now quickly identify issues with Advisor and easily remediate those issues with batch operations.

Batch operations enables the following:

  • Remediating operational spikes: Bulk-transition high-traffic Archive or Coldline objects to Standard as soon as the pattern is detected, curbing retrieval and operation charges immediately.

  • Containing runaway growth. Mass-delete stale or temporary data across specific prefixes when the advisor flags above-trend storage growth.

  • Enforcing fleet-wide consistency. Apply metadata, tagging, retention, or encryption changes uniformly across massive object sets, with no dedicated compute to provision.

We also expanded and enhanced the existing capabilities of batch operations, making it easier to execute actions at scale:

  • Multi-bucket processing: Run a single job across up to a thousand buckets per project, rather than executing it bucket-by-bucket.

  • Dry-run validation. Simulate your transformations using dry-run mode before modifying live data. A dry run helps you safely preview a job's impact (including affected object counts, total size, and potential errors) before committing to permanent changes.

  • Advanced filters powered by Storage Insights datasets: Use Common Expression Language (CEL) expressions to select objects directly by specifying conditions that match fields in Insights datasets. For example, you can filter objects across your buckets by storage class, object size, creation date, or custom attributes.

Below is a CLI example demonstrating how to create a batch operations job using advanced filters. This job deletes all temporary objects belonging to the Standard storage class present in a user's "analytics" buckets.

code_block
<ListValue: [StructValue([('code', 'gcloud storage batch-operations jobs create bulk-delete-temp-objects \\\r\n --description="Bulk delete temporary objects in analytics buckets" \\\r\n --target-project="my-project-id" \\\r\n--insights-dataset-config="projects/my-project-id/locations/us-central1/datasetConfigs/my-dataset" \\\r\n --bucket-filters="name.startsWith(\'analytics-\')" \\\r\n --object-filters="storageClass == \'STANDARD\' && name.endsWith(\'.temp\')" \\\r\n --delete-object'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f522ded1990>)])]>

The evolution of storage management

Storage management shouldn’t be a reactive effort reserved for quarterly reviews and post-incident fire drills. It should be continuous, proactive, and contextual.

Storage Intelligence advisor and batch operations help to surface what changed and enable insights and action at scale. As Storage Intelligence gets better at recognizing which findings matter, Cloud Storage can carry more of the operating load for teams that need to manage storage at scale.

Storage Intelligence advisor and enhanced storage batch operations are generally available today.

To get started, enable Storage Intelligence on a project or org. If you haven't used Storage Intelligence before, a 30-day trial is available at no cost.

Scale your own way, using HPA with built-in support for PromQL metrics queries in GKE

23 septembre 2026 à 18:00

Earlier this year, we announced native support for Google Kubernetes Engine (GKE) custom metrics. This milestone allowed you to scrap external adapters and instead collect autoscaling metrics directly from your pods. By routing these metrics straight to the Horizontal Pod Autoscaler (HPA), we cut metrics reading latency down to 5 seconds.

Today, we are excited to introduce built-in support for processing Prometheus metrics, allowing you to use expressive PromQL queries to customize autoscaling triggers. With this update, HPA can now directly process autoscaling metrics present in Cloud Monitoring using Google Managed Service for Prometheus. Reading metrics from these backends will not require third-party adapters, leveraging the AutoscalingMetric integration used to support pod-level metrics. After the preview, we plan to support self-hosted Prometheus servers as we move to general availability. 

The challenge: Setting up Cloud Monitoring metrics

Support for custom pod-level metrics made autoscaling more straightforward, but production workloads often need to scale on multiple, complex infrastructure metrics. Common examples include scaling:

  • a worker pool based on the number of unacknowledged messages in a Pub/Sub topic

  • an inference service based on query-per-second (QPS) metrics stored in Cloud Monitoring / Prometheus

  • a webserver farm based on the 95th percentile of their measured response time

To achieve this, you used to need to deploy an external adapter like the Stackdriver Custom Metrics Adapter or the Prometheus adapter to retrieve the metrics from an external logging environment. While this sounds straightforward at first, these adapters introduce a lot of operational friction:

  • Management overhead: Platform teams have to install, configure, patch, and monitor these third-party components.

  • Reliability and inefficiency: Intermediate adapter pods reading from external systems introduce failure points in critical autoscaling loops. 

  • IAM complexity: Enabling secure cross-component communication requires setting up Kubernetes service account mappings to Cloud service accounts including their permissions.

And while setting up this system and maintaining it not impossible, it’s complex and features a complicated architecture:

1

How processing Prometheus Metrics in GKE can help

Extending the AutoscalingMetric object drastically simplifies this setup. Now you can read metrics from monitoring directly via PromQL and provide them to HPA via a high-performance, low-latency autoscaling pipeline, resulting in a simplified environment.

2

To prevent inefficiencies, we built this feature with minimal resource consumption in mind. The controller runs on the GKE control plane. It monitors your AutoscalingMetric custom resources and only deploys the system pod on your user nodes when a PromQL metric is actively requested. If no Prometheus metrics are configured, the controller is shut down, so there’s no resource overhead.

Configuring built-in Prometheus metrics

Configuring GKE to use PromQLl metrics is easy; here’s a sample configuration file providing PubSubs message queue depth as scaling metric:

code_block
<ListValue: [StructValue([('code', 'apiVersion: autoscaling.gke.io/v1beta1\r\nkind: AutoscalingMetric\r\nmetadata:\r\n name: gmp-metric\r\nspec:\r\n metrics:\r\n - promql:\r\n name: pubsub-queue-depth\r\n query: |\r\n {\r\n "pubsub.googleapis.com/subscription/num_undelivered_messages",\r\n subscription_id="my-subscription"\r\n }'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f522d455250>)])]>

Linking Prometheus metrics to your HPA

Once defined in your AutoscalingMetric resource, you can reference the metric in your standard HorizontalPodAutoscaler using the same intuitive format as raw custom metrics: autoscaling.gke.io|<custom-resource-name>|<metric-name>.

Scaling globally (Prometheus metric)

For global metrics like a queue size that returns a single aggregate value:

code_block
<ListValue: [StructValue([('code', 'apiVersion: autoscaling/v2\r\nkind: HorizontalPodAutoscaler\r\nmetadata:\r\n name: worker-hpa\r\nspec:\r\n scaleTargetRef:\r\n apiVersion: apps/v1\r\n kind: Deployment\r\n name: worker-deployment\r\n maxReplicas: 10\r\n metrics:\r\n - type: External\r\n external:\r\n metric:\r\n name: autoscaling.gke.io|gmp-metric|pubsub-queue-depth\r\n target:\r\n type: AverageValue\r\n averageValue: 100 # maintain queue size at ~100 per pod'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f522d485390>)])]>

Scaling on Cloud Monitoring per-Pod metrics

GKE natively supports scale based on the most recent gauge metric values, but PromQL offers greater flexibility, allowing you to scale across time windows and calculate rates or histogram percentiles.

To use this capability, configure your PromQL metric to include a label for the pod name, then assign type: Pods within your AutoscalingMetric manifest. Below is an example that calculates a Pod's average memory usage over a five-minute rolling window.

code_block
<ListValue: [StructValue([('code', 'apiVersion: autoscaling.gke.io/v1beta1\r\nkind: AutoscalingMetric\r\nmetadata:\r\n name: per-pod-stored-metric\r\nspec:\r\n metrics:\r\n - promql:\r\n name: container-memory-metric\r\n query: |\r\n sum by ("pod")\r\n (avg_over_time({"container_memory_working_set_bytes"}[5m]))\r\n type: Pods # The promql query returns per-pod metrics'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f522d486c10>)])]>

Key benefits

  • No adapter maintenance: No pods to install, configure, or upgrade. The entire lifecycle is fully managed within GKE.

  • Streamlined security: Out of the box, the Kubernetes Default Node Service Agent has read permissions to Cloud Monitoring and Google Managed Prometheus in the same project. No extra IAM service accounts, keys, or federation parameters are required.

  • Low latency and fast scalability: The new Autoscaling Metric system polls the backend every 15 seconds, helping ensure fast scaling reactions.

  • Rich query capabilities: Leverage the full power of PromQL (including rate calculations, averages, and percentiles) to translate high-level business and user-experience objectives directly into scaling.

  • Support for the new HPA scale-to-zero capability: Utilize it for scaling workloads to zero replicas when demand hits zero (e.g., Pub/Sub queue size) and, more crucially, back up from zero replicas quickly using CapacityBuffers API.

Try it today 

By natively supporting both custom container metrics and Prometheus metrics, GKE now  offers a more robust, performant, and low-friction autoscaling experience. Built-in support for Prometheus Metrics is in preview now. To learn more about setting up your first AutoscalingMetric resource, check out the latest GKE autoscaling documentation.

Introducing Amazon CloudWatch Omni: collaborative AI-powered observability for your applications

23 septembre 2026 à 00:23

Amazon CloudWatch now offers CloudWatch Omni, an AI-powered observability experience for the applications and AI agents you run together. You reach Omni through a dedicated URL for your organization and sign in with the identities you already manage, so working in Omni does not require access to the AWS Management Console. Omni is built on OpenTelemetry: the telemetry you already send to CloudWatch appears in Omni with nothing to reconfigure, and any other workload you instrument with OpenTelemetry sends its telemetry to an OpenTelemetry Protocol (OTLP) endpoint.

CloudWatch Omni offers both agent observability and application observability in a single experience. In our companion post, we introduced the agent observability capabilities of Omni for generative AI and agentic workloads. In this post, we present the application observability experience.

Engineering teams spend a significant portion of their observability time maintaining dashboards, tuning thresholds, and switching between tools to piece together what happened during an incident. When an issue crosses team boundaries, context gets lost in Slack threads and screenshots rather than flowing naturally to the next engineer. CloudWatch Omni changes this by organizing observability around your applications rather than individual signals, and bringing your whole team into the same workspace.

What CloudWatch Omni brings
CloudWatch Omni addresses three problems that engineering teams told us they face today.

One collaborative experience for your whole team. Every engineer accesses CloudWatch Omni through a single URL with enterprise SSO (via IAM Identity Center, supporting Okta, EntraID, and other providers). No AWS Console access is required. SREs, developers, database engineers, and managers share the same data and investigation context. When an investigation escalates, the next person joins the same session with full context already in front of them.

The system adapts as your applications evolve. CloudWatch Omni discovers your services, maps dependencies, and adjusts alarms automatically. Instead of manually curating dashboards and tuning thresholds, you declare what matters (availability targets, latency budgets, error rate thresholds) and Omni adapts as your system changes. When you deploy new services, Omni updates the application topology automatically.

AI-powered investigation with Amazon DevOps Agent. Amazon DevOps Agent participates alongside your team in investigation sessions, correlating signals and suggesting next steps. The agent works from the same telemetry your engineers see, so its suggestions are grounded in the actual state of your application. It identifies correlated events across services, traces root cause paths through your dependency graph, and maintains investigation history for post-incident review.

How an investigation works
When something breaks, CloudWatch Omni opens an investigation session pre-loaded with context. Here is a typical incident workflow:

An alarm fires on elevated error rates in your checkout service. Omni opens a session showing the service topology, correlated signals (a deployment 10 minutes earlier, increased latency from a downstream payment API), and DevOps Agent’s initial analysis.

Your on-call SRE confirms the deployment correlation, pulls in the trace view to identify failing endpoints, and checks if the payment API latency correlates with a capacity limit.

The SRE escalates to the payments team. The payments engineer joins the same session and sees everything found so far, plus DevOps Agent’s correlation with a configuration change in the payment provider’s API gateway. They identify the root cause and roll back.

The entire investigation history is captured automatically. No separate incident report needed.

Walkthrough: setting up your first Space
To set up CloudWatch Omni for your team, open the CloudWatch console and click “Try CloudWatch Omni.”


Figure 1. CloudWatch console — Omni setup page

Next, connect your identity provider through IAM Identity Center (supporting Okta, Azure AD, and other SAML 2.0 providers). Once connected, your team members access Omni directly at your dedicated URL without needing AWS Console credentials.

Create a Space for your team. A Space groups the applications your team owns and the telemetry associated with them.


Figure 2. CloudWatch Omni Home — your team’s workspace with application monitoring, analytics, and agent observability

Once created, Omni discovers your services automatically and maps the dependencies between them. You see your application topology immediately.


Figure 3. Application topology — services and dependencies mapped automatically

You can ask CloudWatch Omni any question about your applications in plain English, and Omni will analyze your telemetry data and surface insights.

Figure 4. Interact with your telemetry in natural language

You can also set up service health alerts, configure what matters to your team, and trigger an AWS DevOps agent investigation to identify the root cause and develop a mitigation plan.


Figure 5. Investigation session — DevOps Agent identifies root causes and suggests next steps

Application-centric organization
CloudWatch Omni organizes telemetry by application rather than by infrastructure component. The system automatically discovers services from the telemetry data and AWS Config resource discovery, maps dependencies, and lets you see your application as a connected system rather than a collection of isolated resources.

Each team gets a Space that contains the applications they own. A Space points at existing CloudWatch data (logs, metrics, traces, and alarms) with no additional data movement required. Dynamic views replace the maintenance burden of static dashboards, providing ongoing visibility into SLOs and application health.

Getting started
Getting started takes minutes and doesn’t require reconfiguration of your existing CloudWatch setup.

If you’re an existing CloudWatch customer: Click “Try CloudWatch Omni” in the CloudWatch console. All your existing telemetry (logs, metrics, traces, and alarms) is immediately available. Workloads are discovered automatically, and you can start an investigation or browse your application topology right away.

For organization-wide deployment: An administrator configures a domain, connects your identity provider via IAM Identity Center, defines Spaces for teams and environments, and invites users. Each Space points at existing CloudWatch data with no additional data movement required.

For applications in other environments: CloudWatch Omni provides connectors that make it easy to bring in telemetry from additional environments. All ingested telemetry appears alongside your AWS data in the same Spaces and investigation sessions.

For generative AI and agentic workloads: The same CloudWatch Omni experience delivers purpose-built observability for AI agents, including trace exploration, evaluation frameworks, and real-time monitoring. In our companion post, we introduced the agent observability capabilities of Omni; for that walkthrough, see Introducing Amazon CloudWatch Omni: AI-powered observability for generative AI and agentic workloads.

Things to know

  • CloudWatch Omni extends CloudWatch. Existing alarms, dashboards, APIs, and console workflows continue unchanged.
  • Access is through a dedicated web application with enterprise SSO. Engineers don’t need AWS Console access to use it.
  • Once you setup, DevOps Agent is enabled by default in every Omni investigation session.

Pricing and availability
Amazon CloudWatch Omni is now available. Existing CloudWatch customers can try it directly from the CloudWatch console. For pricing details, visit the Amazon CloudWatch pricing page.

To get started, visit Amazon CloudWatch Omni or click “Try CloudWatch Omni” in the Amazon CloudWatch console.

If you want to call APIs, search documentation, find regional availability, and check troubleshooting about this feature, try using the AWS MCP Server and plugins with your preferred AI tool. Share your feedback on AWS re:Post or reach out through your usual AWS Support contacts.

9/24/2026 – Editor’s note: Azure AD has been renamed to Microsoft Entra ID.

— Daniel Abib

Introducing Amazon CloudWatch Omni: AI-powered observability for generative AI and agentic workloads

23 septembre 2026 à 00:21

Today, Amazon CloudWatch introduces CloudWatch Omni, a unified observability experience for application and AI workloads that is app-centric, AI-powered, built on open standards, and delivered off-console. CloudWatch Omni is a purpose-built observability, evaluation, and experimentation solution for AI agents. It helps teams design, evaluate, and operate AI agents across any model provider, framework, or runtime, with an eval-driven workflow, support for the tools you already use, and observability delivered where you work: directly in your IDE and through a standalone web experience, separate from the AWS Management Console.

Organizations deploying agentic AI systems face observability challenges that traditional monitoring can’t address. Agent behavior is non-deterministic: a prompt change can degrade response quality even when standard metrics show no errors. Teams spend hours manually reviewing logs across multiple systems, unable to pinpoint what changed or why. Existing tools force teams to choose between siloed generative AI monitoring or fragmented solutions requiring constant context-switching between their coding environment and browser-based dashboards.

CloudWatch Omni captures every trace and includes built-in evaluators for correctness, coherence, retrieval quality, and tool selection, among others. You can compare prompt versions side by side in the playground, build test datasets from production traffic, run experiments across different configurations, and detect regressions automatically.

Two surfaces for development and operations
CloudWatch Omni delivers observability through two complementary surfaces. Developers get a native extension inside VS Code and Kiro (the currently supported IDEs), where traces appear as you run your agent with a playground and evaluators a click away. Operators get a standalone web experience, separate from the AWS Management Console to monitor the fleet, accessible through SSO with no AWS console needed. Both share the same data: the trace a developer debugs is the trace an operator investigates.

The Cloud Login feature connects your local IDE environment to your AWS account, enabling you to send telemetry data to Amazon CloudWatch for persistent storage, share traces with your team, and access production dashboards. This connection is optional. You can use CloudWatch Omni entirely locally during development, then connect to the cloud when you are ready to monitor agents in production.

Getting started
CloudWatch Omni offers two ways to get started: through the IDE extension (for VS Code and Kiro) or directly through the cloud experience, where you can start sending telemetry data to CloudWatch without installing any IDE extension. In this walkthrough, I install the extension, create an agent, run it, and explore the traces and evaluation tools from my IDE.

After installing the CloudWatch Omni extension from the VS Code Marketplace, the CloudWatch Omni icon appears in the Activity Bar. From the welcome screen, I selected Get started with Sample Project to load a pre-configured agent with sample trace data or use shortcut to Command Palette using Command + Shift + P (on macOS) or Ctrl + Shift + P (on Windows/Linux) and select Omni: Create a new Project

CloudWatch Omni welcome screen and create new project in VS Code
Figure 1. CloudWatch Omni welcome screen & create new project in VS Code

The sample project comes with an agent implementation and example datasets. Part of the getting-started experience is adding OpenTelemetry instrumentation, and CloudWatch Omni guides you through each step. You can also create a new agent from scratch. CloudWatch Omni walks you through the process using an interactive chat where you define the agent’s purpose, select a model provider, and configure tools. All data is stored locally by default. You can optionally connect to AWS to send data to Amazon CloudWatch.

After verifying the configuration, I started the local dev server and sent a question to the agent. What makes this different from a typical chatbot interface is what happens next: selecting View Trace shows exactly how the agent processed the request.

CloudWatch Omni guides your AI code assistant to configure the development environment
Figure 2. CloudWatch Omni guides your AI code assistant to configure the local development environment for testing

CloudWatch Omni integrates with AI code assistants such as Kiro, Claude Code, and Codex to streamline the setup process. These assistants can configure the Dev Server, install dependencies, and set up instrumentation on your behalf, so you can go from installation to running your first traced agent session in minutes without manual configuration.

Interacting with the agent and viewing traces
Figure 3. Interacting with the agent and viewing traces

Traces are essential for understanding AI agent behavior. Unlike traditional request-response systems, agents make multiple decisions per invocation: choosing tools, composing prompts, and chaining sub-calls. Without full trace visibility, diagnosing why an agent produced an incorrect answer or took an unexpected path becomes guesswork. CloudWatch Omni records every step in a structured timeline so you can pinpoint exactly where behavior diverged.

The Trace Explorer shows a detailed breakdown of every step the agent took (LLM calls, tool invocations, and reasoning steps) in a structured, hierarchical timeline. I could drill into any span to inspect inputs, outputs, token usage, and latency.

Trace Explorer showing the agent execution timeline
Figure 4. Trace Explorer showing the agent’s execution timeline

The Trace Explorer also supports Compare mode, which places two traces side by side to see how different prompts or configurations affect behavior. Compare mode is especially helpful when debugging regressions. And with Ask Assistant, an AI agent analyzes your traces to surface patterns and anomalies, answering questions like “Why did the agent call this tool twice?”

Comparing two traces side by side
Figure 5. Comparing two traces side by side

Evaluation is what turns observability into actionable quality improvement for generative AI. Traditional metrics like latency and error rate cannot tell you whether an agent’s response was helpful, coherent, or factually correct. Evaluators score each response against quality dimensions, letting you measure what users actually experience and catch regressions that standard monitoring misses entirely.

CloudWatch Omni includes 17 built-in evaluators for metrics like coherence, helpfulness, faithfulness, and routing correctness. I selected traces from the Trace Explorer, chose evaluators, and ran an evaluation, getting per-example scores and aggregate metrics without building any custom evaluation framework.

Running evaluations on traces
Figure 6. Running evaluations on traces

From there, I used the Playground to test different system prompts side by side, comparing multiple model and prompt configurations in real time to see how each variation affects output quality before committing changes. With the Experiments view, I could run the same dataset against two agent variants and compare their evaluation scores, latency, and token usage side by side to pick the best-performing configuration.

Figure 7. Comparing evaluations across agent variants in the Omni Experiments console

With Prompt Management, you can version and track prompt configurations over time, making it easy to roll back when a new version underperforms.

CloudWatch Omni also provides a Session Explorer to review full conversation histories and understand how agents handle multi-turn interactions, along with an Agent Topology view that visualizes the architecture of your agent system, including sub-agents, tools, and their interconnections. You can drill into any node to inspect performance and identify bottlenecks.

CloudWatch Omni also offers a dedicated web experience accessible from any browser without an IDE. Teams can access all capabilities collaboratively, including application monitoring, analytics, agent observability, and AI-powered investigations.

CloudWatch Omni web experience with application monitoring, analytics, and agent observability
Figure 8. CloudWatch Omni web experience with application monitoring, analytics, and agent observability

I curated traces into golden datasets for structured experimentation. The Experiment function runs the agent against a dataset and automatically scores results, creating benchmarks for regression testing whenever prompts or agent logic change.

If you already have an agent built with a supported framework, CloudWatch Omni provides two paths to add instrumentation: Auto-instrument with Kiro, which detects your framework and configures tracing automatically, or manual instrumentation with ready-to-use code snippets for Python and TypeScript. For detailed instrumentation guides, see the CloudWatch Omni documentation.

Supported frameworks and open standards
The walkthrough above uses the sample project, but CloudWatch Omni works with the agent frameworks teams are already using: LangChain, LangGraph, CrewAI, OpenAI SDK, Strands, Vercel AI SDK, and more, in both Python and TypeScript. It also provides native observability for agents built with Amazon Bedrock AgentCore, and uses AgentCore’s evaluation capabilities to assess agent quality directly within the Omni workflow.

Instrumentation uses open standards (OpenInference and ADOT), whether your agents run on Lambda, ECS, EKS, or other clouds. For evaluation, Omni integrates with third-party evaluators including AutoEval and DeepEval alongside built-in datasets, a playground, and batch experiments. No re-platforming required.

CloudWatch Omni brings agent observability and application observability together in a single experience. For the application observability experience, read the companion post Introducing Amazon CloudWatch Omni: collaborative AI-powered observability for your applications.

Pricing and availability
Amazon CloudWatch Omni is now generally available. The IDE extension is free to use. You don’t need an AWS account to get started. You only need AWS credentials for Amazon Bedrock models, or API keys for other providers like OpenAI or Anthropic. Get started today by installing the extension from the VS Code Marketplace.

To explore all capabilities and get started quickly, visit CloudWatch on AWS Builder Center.

If you want to call APIs, search documentation, find regional availability, and check troubleshooting about this feature, try using the AWS MCP Server and plugins with your preferred AI tool. Share your feedback on AWS re:Post or reach out through your usual AWS Support contacts.

Happy building!

— Daniel Abib

Best practices for handling cloud reliability incidents

15 septembre 2026 à 18:00

Cloud outages can range from global service disruptions to issues isolated to a specific region, zone, or even just your project, workload or application. If you suspect a Google Cloud Platform outage is impacting your services, we recommend you follow a structured “Verify→ Investigate→Report→Resolve→Review" workflow to resolve it. And before that outage occurs, you should also have prepared your environment for an eventual disruption by designing for failure, and actively practicing the steps you need to take to restore service. 

In this blog, we summarize the key reliability incident handling best practices to help you design and practice your reliability incident response capabilities and minimize impact. Rather than an exhaustive guide, this is meant as a primer on only the most important practices for advisory purposes. Please note that we do not cover additional practices specific to security incidents here. 

Beyond the base steps covered here, you may want to also explore how AI agents and tools are starting to transform incident handling. Check out this episode of the Prodcast, where Googlers explore the latest trends of leveraging agentic AI in Site Reliability Engineering (SRE) to detect issues early and prevent disruptions. Try Cloud Assist investigations, or explore Agent Skills and remote managed MCP servers to give you another set of tools for quickly pinpointing an issue. Before getting into these advanced techniques, we focus below on the foundational steps to good incident handling.

1. Prepare

Long before things start to go sideways, you should have spent significant time preparing for an outage along at least four dimensions: design, data, playbooks and training.

  • Design: Think ahead and mitigate future incidents by designing automated response actions, like a load balancer shifting traffic away from slow or unresponsive instances, or by automating as much of your incident response playbook as possible. Review designs of all critical applications to automate as many actions as possible to accelerate response and recovery.

  • Data: When a disruption occurs, having meaningful data at your fingertips vastly improves response capabilities. Use Cloud Logging, Cloud Trace and Cloud Monitoring, or other third-party observability tools, and replicate that data to a redundant stack in a separate location from the systems being observed. Make sure, in advance of any incident, that time stamps are synced across your observability streams for easy correlation, or know how to do that on-demand during an outage, when time is of the essence.

  • Playbook: A well-thought-out playbook documenting your incident response processes, including crystal clear role and responsibility definitions for all personas, is paramount to efficient incident response. Who is responsible to do what? Who needs to be notified or mobilized for each type of disruption? How can they be reached? What tools and data are available? How are results communicated? How do teams hand over to the next shift during long running incidents? etc. Conduct a simulated incident response and critically review every step to find where your playbook needs clarification. Without clear responsibilities, mitigation inevitably takes longer.

  • Training: Hopefully, service disruptions are rare events. To ensure your staff knows and remembers how to react, they need to retrain on the process several times per year by running simulated cross-team incident response drills. A retrospective on the simulated exercise will help identify warranted improvements.

2. Verify

Despite your best efforts, sooner or later, a service disruption will occur, which you can detect via any number of mechanisms:

Now, you need to determine what broke and who should ultimately fix the problem:

  • Google, e.g., a bug, code roll-out, hardware failure, etc.

  • You, e.g., a configuration change, elevated load, quota ceiling, etc.

  • Third party, e.g., a directory hosted by a different cloud provider

If Google has declared an incident and started working to fix the problem, estimate whether you can possibly reestablish service sooner, for example by failing over to a secondary stack (see the ‘Typical Causes’ table below). You can determine whether Google has declared an incident and will provide a fix by consulting:

  • Personalized Service Health: Check this first. Personalized Service Health shows incidents specifically relevant to your projects and regions, distinguishing between incident types:. 

    • Emerging Incidents: Google has received an alert, on-callers are investigating, impact is yet unknown

    • Confirmed Incidents: Google has investigated and found customers are impacted

Located within the Google Cloud console, Personalized Service Health often displays limited-scope incidents that don't appear on the public dashboard. Personalized Service Health also offers a mobile client for Android and iOS smartphones, assuming you can use your work ID and credentials on the phone.

  • Gemini Cloud Assist, which is integrated with Personalized Service Health, so you can use it to query that information in natural language.

  • Cloud Service Health dashboard: This is the public-facing non-authenticated web page for broad, severe incidents affecting many customers. Limited blast radius disruptions are not externalized to the public. All its content is available in Personalized Service Health as well. If ever Personalized Service Health goes down, Cloud Service Health serves as an alternative channel built on a separate infrastructure.

  • Known Issues: In the console, navigate to Support > Cases, view a case, and use the resource selector on the console toolbar to find the specific cloud resource you’re interested in. Then click Known issues. If your issue matches one listed here, you can link a support case to it, so you will receive automatic updates in your case record. If you don’t find a match, open a new support case. Google will automatically match the case to a related incident, as soon as one is declared.

  • Google declared incidents are updated as new information becomes available, so check back regularly, or set up a Personalized Service Health alert policy to be notified each time new information becomes available.

If you host cloud resources in multiple clouds, a good practice is to check early on whether the problem occurs for multiple cloud providers. If so, the problem is likely external to the providers and caused either by you or by a third-party service that your application interacts with.

3. Investigate

To determine the blast radius within your cloud footprint of Google-declared reliability incidents, first check Personalized Service Health updates for a description of the technical problem. Knowing what to look for will allow you to map your blast radius and decide on suitable contingency actions quicker.

If Google hasn’t declared an incident, try to rule out configuration errors or issues within your environment by checking:

  • Cloud Monitoring: Look for spikes in error rates (e.g. 5xx errors), increased latency, or drops in traffic in your dashboards.

  • Cloud Logs: Use Log Explorer to look for specific error messages like DEADLINE_EXCEEDED, SERVICE_UNAVAILABLE, or specific API errors.

  • Quotas: Ensure you haven't hit a project quota (e.g., CPU, API rate limits), which can often mimic the behavior of an outage.

  • Change history: Check your log of recently applied changes. Not all problems manifest immediately, but proximity on a timeline can be a powerful indicator of causality, even if it’s not proof. Also check whether Google rolled out any updates just before the symptoms started. See the Unified Maintenance Management interface in Cloud Hub.

Absent a clear culprit, such as a traffic spike or a DDOS attack, and if symptoms manifested immediately after rolling out a change, a good strategy is to back out that change and attempt to return to a last known good configuration. 

4. Report

If the Cloud Service Health and Personalized Service Health dashboards are green but your metrics show a failure, you must report it to Google. 

  • Determine priority:

  • File a case: Go to Support > Cases > Create Case in the console.

    • Explain quantifiable business impact to rationalize the submitted priority and prevent it from being reset when Cloud Support prioritizes cases. A clear and accurate rationale helps!

  • Essential information to include:

    • Project ID and affected region/zone

    • Timestamps (when it started and if it's ongoing) with a clearly labeled timezone

    • Specific error messages or log snippets

    • Scope: Is it affecting all users/systems, or a specific subset/location?

Escalation for Premium/Enhanced support

If you have a Premium or Enhanced support plan and a P1 case is not receiving the attention it requires, use the Escalate button within the support case in the console. This alerts a support manager to investigate and rectify the situation.

5. Resolve

By taking these steps, you are well on your way to resolving the outage. In the meantime, here are some ways to mitigate the impact of the outage and communicate with impacted stakeholders.

While waiting for a resolution:

  • Communicate: Notify your stakeholders and customers. Transparency helps manage expectations and reduces duplicate internal reports.

  • Fail over: If you have a multi-regional architecture, consider shifting traffic to a healthy region. As a best practice, first ensure that the disruption is at the infrastructure level and not at your workload level. 

  • Check for workarounds: While working on a permanent fix, Google often posts temporary workarounds in the Service Health Dashboard updates, or in Personalized Service Health updates.

  • Consider your regulatory reporting requirements: Know whether your organization is subject to regulatory reporting requirements, and what the required deadlines are for both initial and follow-up reporting. Google Cloud prepares Incident Reports for incidents that meet certain criteria — see details here for how to get those reports. Premium Support customers can also request an Incident Summary, which is an Incident Report customized to your account’s specific hosting location, time stamps, etc.

De-escalation and closure

Once systems are stable, Google downgrades the severity levels and deactivates the active on-call escalation chain. Google only closes an incident in Personalized Service Health when it has taken all the mitigation steps covering all impacted customers. Your specific services might be restored sooner than the incident closure time, if other customers are restored later than you. The incident is officially closed on the Google Cloud Status Dashboard when systems have run stably for a designated auto-close duration. Verify that your services are operating normally at this point. And if your incident responders aren’t compensated for extra time spent on the incident, find a way to thank them.

6. Review

After the problem has been fixed and operations have returned to a normal, steady state, it’s time to conduct a post-mortem analysis to identify how your team can respond better in future service disruptions. A “blameless” approach is essential to surfacing meaningful and impactful improvements that can be made to your incident response process. Ask questions like:

  • What went well?

  • What could we have done better?

  • Where did we get lucky?

  • Where did we get unlucky?

Then decide what changes can be made to improve your playbook, tools and training.

At Google, we often publish a post-mortem or Incident Report for major outages, available via Personalized Service Health. Review this to understand the root cause and adjust your own disaster recovery plans to prevent or reduce future impact. Customers with a Premium Support plan can request an Incident Summary for a Google-caused incident they were impacted by and for which they opened a P1 case. An Incident Summary is an Incident Report customized for your environment (e.g., start and end times of impact).

Typical causes, comms and prevention strategies

To help you prepare and plan ahead, here’s an overview of some typical incidents based on the symptoms reported in Cloud Service Health and Personalized Service Health along with guidance on what Google communications to expect, and some generic mitigation or prevention strategies you can build into your playbooks.

Blast radius

Typical cause

Comms

Strategy

Single zone or region.Subset of products.

Typical of a software problem triggered by a rollout. Learning points:

- Understand the location scope (zones and regions) of your workload

- Products can depend on other products

Major incidents are communicated via Cloud Service Health.Major and Minor (by number of customers, not severity) incidents are communicated via Personalized Service Health.

Highly localized incidents are not communicated via Cloud Service Health or Personalized Service Health.

Fail over, if so configured, but verify the health of the secondary stack first.

Single zone.Most or all products.

Typical of a power or cooling issue.

Check Cloud Service Health and Personalized Service Health.

Fail over to a different zone, if so configured.

Single region.

Most or all products.

Typical of a backbone networking infrastructure issue 

Check Cloud Service Health and Personalized Service Health.

Fail over to a different region, if so configured.

Control plane issue for a product

Typical of a late detected issue

Communicated via Personalized Service Health if significant customer impact is verified.

Look for workarounds. Wait for Google to fix. Fail over, if so configured.

Multi-regional issue with a global product

Rare but possible, typically detected quickly. Learnings: Mitigation options can be limited. Try regional variants, alternative products with similar functionality

Check Cloud Service Health and Personalized Service Health.

Wait for Google to fix. In the meantime, verify via Google Comms and your own investigation that this is truly Google’s problem to fix.

Capacity / Stockout issue

System-level demand exceeding capacity in the product/location/model. (Cloud is designed to scale, but limits always exist, so proper planning is advised)

Error message. No incident will be declared.

Place reservations for predicted capacity needs (if cost is acceptable). Flexibility in zone placement can also help.

Quota exhaustion

Difficult / inaccurate prediction of traffic

Error message. No incident will be declared.

Review consumption trends against ceiling regularly.

Go deeper

This document offers only a condensed summary of key points. If you have an active Premium Support contract with Google Cloud, reach out to your account team for a deeper review of your response plans. For a comprehensive treatise on how to build reliable services and how to respond to incidents, we strongly recommend Google’s SRE Book, which is available as a free download. A new version of the SRE book is releasing ~Oct 2026 and will be available for purchase on O’Reilly Media. We’re also working on a future primer that explores AI-supported incident handling in-depth — stay tuned!

New AI-powered quick assessments in Migration Center turbocharge modernization

24 août 2026 à 18:00

Technology leaders are under mounting pressure to modernize infrastructure, control multi-cloud operational spend, and build data foundations for generative AI. However, the discovery required for that level of transformation can entail weeks of manual spreadsheet analysis, mapping in-house infrastructure, and reconciling siloed, piecemeal cost estimates across disparate teams and sources. To help, we’re announcing AI-powered Quick Assessments in Migration Center, which delivers near-instant total cost of ownership (TCO) modeling and automated service mapping.

Compare this to legacy assessment processes, which can stall digital transformation initiatives before they even launch. Manual discovery can delay migration timelines by months, increase engineering overhead, and often miscalculates complex financial models. By replacing manual discovery with AI-assisted automation, IT gains instant, actionable visibility into the TCO and return on investment (ROI) for a given migration initiative. 

Now, organizations can generate comprehensive migration financial models in minutes rather than months. Teams ingest raw infrastructure data or cloud billing reports and quickly receive an optimized target bill of materials (BOM), service mapping coverage, and projected savings. Decision makers interact with an agentic assistant that explains underlying financial assumptions, recommends technical cost optimizations, and exports ready-to-share executive reports.

Inside the AI-powered Migration Center

This is made possible with AI-assisted Quick Assessments alongside enhanced Cloud Billing assessment capabilities, both integrated into the new AI-powered Migration Center.

AI-assisted Quick Assessment for on-premises workloads

Designed for enterprise customers and partners, AI-assisted Quick Assessment automates on-premises infrastructure evaluation to provide rapid financial modeling. Let’s walk through these new capabilities: 

  • Instant Compute Engine TCO estimates convert VMware inventory exports (such as RVTools) or aggregated infrastructure inputs into precise Compute Engine cost targets:

1

Migration Center’s Quick TCO Estimator

2

Migration Center’s Quick TCO Estimator results page

  • Advanced architecture modeling supports latest-generation Gen4 compute instances and high-performance Hyperdisk storage pools:
3

Migration Center’s Quick TCO Estimator detailed results page

  • Customizable financial controls allow teams to adjust on-premises baseline cost assumptions to match internal accounting standards:
4

Migration Center’s Quick TCO Estimator detailed results page (continued)

  • Context-aware agentic chat recommends tailored technical cost optimizations aligned with your specific business constraints (such as regional location or compliance needs), clearly explaining the underlying logic and financial assumptions.
5

Migration Center’s agentic chat capabilities

6

Migration Center’s agentic chat capabilities (continued)

7

Migration Center’s agentic chat capabilities (continued)

  • Automated Business Case and Google Sheets export generates ready-to-use reports capturing the complete recommended BOM, TCO comparison, and ROI analysis:
8

Migration Center’s business case

9

The path forward

Modernizing your infrastructure starts with fast and accurate data. Migration Center’s Gemini-powered features simplify cloud evaluation, empowering IT decision makers to build defensible business cases generated by machine-learning.

Try Migration Center directly in the console today, or take a free migration and modernization assessment to evaluate your workloads and accelerate your strategic cloud journey with Google Cloud.

Introducing Database Operations Agents: The future of autonomous database management

4 août 2026 à 18:00

As part of the Agentic Data Cloud launch at Google Cloud Next ‘26, we announced two AI-powered database agents to simplify database management. These include the Database Onboarding Agent for Day 0 operations — setup, configuration, and initial deployment — as well as the Database Observability Agent for Day 1 and 2 operations, including monitoring, troubleshooting, and ongoing maintenance. 

These agents are always on, informed by Google’s years of experience, and integrated across Google surfaces such as Chat, CLI, the Google Cloud console, Managed Context Protocol (MCP) servers, and third-party tools — including your preferred integrated development environment (IDE), so you get help where and when you need it. 

Traditionally, managing and creating databases has involved a combination of manual architecture planning, custom scripts, and distinct tools. Teams handle database provisioning, schema design, index configuration, and query tuning, alongside performance monitoring—often cycling through repeated testing and optimization cycles as application demands change. Although this method is functional, it demands substantial technical skill and continuous attention throughout the entire database lifecycle. For example, developers often fear making an update that may limit their ability to scale the system later. Similarly, when an application slows down, finding the exact query or resource constraint causing the issue can take hours of manual investigation and troubleshooting.

Intelligent AI-powered agents can simplify database lifecycle management by automating many of these tasks such as recommending the right database type for the workload, detecting anomalies, recommending the right configurations, optimizing queries, and providing actionable insights to improve operational efficiency. By embedding these capabilities directly into workflows where you need them, agents help organizations build, operate, and optimize databases more efficiently while reducing operational overhead.

Let’s take a closer look at these new database agents.

Database Observability Agent: From diagnosis to remediation 

The Observability Agent empowers Site Reliability Engineers (SREs), DevOps pros, DBAs and developers to diagnose complex issues and remediate them using simple natural language prompts.

As your operations scale, identifying subtle issues like query hotspots or lock contention becomes an expensive burden. The database observability agent uses Google’s operational expertise and the reasoning capabilities of Gemini to solve these challenges. By automatically connecting telemetry across multiple sources including Database Insights, Cloud Monitoring, Cloud Logging, and Cloud Trace the agent provides a clear root cause analysis in minutes.

Beyond just identifying the "why," the agent suggests recommended actions to fix the issues found, and can execute validated actions with your approval. For example, if it detects a bottleneck, it might suggest you "Enable connection pooling for Cloud SQL instance," providing the rationale and expected impact before you commit to the change. Some capabilities include:

  • Fleet-level troubleshooting: The Observability Agent is integrated with Database Center so you can use Gemini Chat to ask complex fleet-wide questions like, "Which databases in my fleet consumed the most CPU in the last 7 days?" to receive a summarized analysis across your entire fleet.

  • In-product investigations: The agent correlates complex telemetry across Database telemetry, Cloud Monitoring, Cloud Logging, Cloud Trace, and multiple other data sources to pinpoint issues like latency spikes or lock contention. (In preview with select customers)

  • Validated remediations: Instead of just identifying problems, the agent provides crisp recommendations and can execute validated actions with your approval, such as adding indexes  for a Cloud SQL instance. (In preview with select customers)

  • MCP tools: The Observability Agent derives insights with the help of tools such as system metrics, query metrics, fleet inventory, and issues, which are also available as MCP tools via the Database Insights MCP Server and Database Center MCP Server. 

Integration that fits your workflow

You can access these Database Observability Agent capabilities directly within your existing database management processes. The agent powers several experiences, including:

  • Cloud Assist chat: Ask questions in natural language, for example, "What is the CPU utilization trend for my top Cloud SQL instances?" to get a summarized analysis complete with charts. Then, within the Chat window, you can start an investigation for any issues found,and get a root-cause analysis and remediations.

1
  • In-product investigations: Use Gemini Cloud Assist to investigate and remediate issues in-context on relevant database pages from the console.
2
  • Developer tools: Consume the agent’s capabilities through Antigravity or an IDE of your choice. This is augmented by the rich set of observability MCP tools that Google provides. All of these tools are available on Google Remote MCP servers. Combining them together is like giving developers a virtual DBA to optimize their databases, but all within their IDEs.

Supports multiple managed databases  

You can use Observability Agent to get answers to your database queries, to access any database metric instantaneously, or to leverage AI-powered diagnosis to resolve complex problems. The agent covers a broad set of issues across a variety of Google Cloud databases, including:

  • Cloud SQL: Troubleshoot and optimize your database instance load, query performance or connectivity issues for all Cloud SQL database engines. For Cloud SQL for PostgreSQL, leverage the agent to troubleshoot common database issues.

4

Similarly, the agent helps you identify issues, find their root cause, and take remediation actions for other supported databases and issue types.

  • Spanner: Here, the most common troubleshooting scenario involves optimizing read and write latencies. The agent helps you do that in minutes, covering a broad set of scenarios ranging from hotspots to lock contentions.
  • AlloyDB: Troubleshoot and optimize your database instance load, query performance or replica lag issues.
  • Bigtable: Diagnose and optimize your read and write latencies, complete with crisp, actionable recommendations.

Database Onboarding Agent

The new Database Onboarding Agent is your active partner during the database selection process. Instead of spending hours reading documentation, you can describe your application requirements to the agent in natural language. The agent understands technical metrics like IOPS, latency limits, and replication lag, so it can provide a sound recommendation. You can access the Database Onboarding Agent’s capabilities directly within the Gemini chat interface. With the Database Onboarding Agent, you get:

  • Recommends database solutions: Analyzes user requirements regarding workload performance, scale, data type, and reliability to suggest optimal Google Cloud Managed Database services (e.g., Cloud SQL, Spanner, AlloyDB).
  • Smart recommendations: The agent reflects your requirements back to you, such as recommending AlloyDB for a high availability configuration, helping you have confidence in its selections.
  • Streamlined configuration: Once you choose a service, the agent generates the required commands. You can then use these commands to provision your database instance, configure the correct features, and deploy it.

Get started

The Database Observability and Onboarding Agent’s capabilities are available for a wide range of services, including AlloyDB, Bigtable, Cloud SQL (PostgreSQL, MySQL, SQL Server), Firestore, Memorystore, and Spanner. These agents are currently available via Gemini Cloud Assist. Explore AI assisted troubleshooting and Gemini Chat for AlloyDB, Cloud SQL, Spanner, and Visit Gemini Cloud Assist page to learn more.

Accelerate your infrastructure deployments by up to 4x with AWS CloudFormation Express mode

30 juin 2026 à 23:30

Today, we’re announcing AWS CloudFormation Express mode, a new deployment mode that accelerates deployments for developers and AI tools iterating on infrastructure. Express mode accelerates deployments by completing when CloudFormation confirms resource configuration is applied, rather than waiting for extended stabilization checks. This reduces deployment time by up to 4 times for iterative development workflows and production scenarios.

How it works
Every CloudFormation deployment performs stabilization checks after resource configuration is applied. These checks serve an important purpose when you need to confirm resources can serve traffic before shifting load.

However, many workflows do not require full stabilization to proceed. Express mode benefits two primary use cases: iterative development workflows and production scenarios where you are comfortable with eventual stabilization. These use cases include iterating on infrastructure configurations during development, testing individual components of your application, and AI-assisted infrastructure development that benefits from sub-minute feedback loops.

With Express mode, CloudFormation completes deployments when resource configuration is applied, without waiting for stabilization checks. Resources continue becoming operational in the background. CloudFormation automatically retries dependent resources that encounter transient failures during provisioning within the same stack, without requiring any customer intervention. This built-in resilience handles timing issues between resources as they stabilize. Express mode changes when the deployment completes, not how resources are provisioned.

For example, when I create an Amazon Simple Queue Service (SQS) queue with a dead letter queue (DLQ), Standard mode takes 64 seconds, but Express mode completes in up to 10 seconds. In the case of deleting an AWS Lambda function with network interface attachment, Standard mode takes 20–30 minutes, but Express mode completes in up to 10 seconds based on my benchmarking test.

Get started with CloudFormation Express mode
When you create a CloudFormation stack in the AWS Management Console, choose Enable in the Express mode under Stack deployment options.

You can also use AWS Command Line Interface (AWS CLI), AWS SDKs, or IaC tools like AWS Cloud Development Kit (CDK), and AI tools such as Kiro.

Activate Express mode by setting the --deployment-config parameter to EXPRESS when creating, updating, or deleting stacks. No template changes are required. Express mode disables rollback by default for the fastest iteration experience. To re-enable rollback, set disableRollback to false in the deployment-config for production environments, or implement monitoring/cleanup mechanisms for failed deployments.

aws cloudformation create-stack \ 
   --stack-name my-app \ 
   --template-body file://template.yaml \ 
   --deployment-config '{"mode": "EXPRESS", "disableRollback": true}' \

For example, use the Express mode when you build infrastructure incrementally, adding resources one at a time. Ensure your IAM role templates follow the principle of least privilege.

# Iteration 1: Deploy IAM role
aws cloudformation create-stack \
--stack-name my-microservice \
--template-body file://iteration1-iam.yaml \
--deployment-config '{"mode": "EXPRESS"}' \
--capabilities CAPABILITY_IAM
--role-arn arn:aws:iam::123456789012:role/CloudFormationDeployRole

# Iteration 2: Add Lambda function
aws cloudformation update-stack \
--stack-name my-microservice \
--template-body file://iteration2-lambda.yaml \
--deployment-config '{"mode": "EXPRESS"}' \
--capabilities CAPABILITY_IAM
--role-arn arn:aws:iam::123456789012:role/CloudFormationDeployRole

# Iteration 3: Add SQS queue and event source mapping
aws cloudformation update-stack \
--stack-name my-microservice \
--template-body file://iteration3-sqs.yaml \
--deployment-config '{"mode": "EXPRESS"}' \
--capabilities CAPABILITY_IAM
--role-arn arn:aws:iam::123456789012:role/CloudFormationDeployRole

For AWS CDK, activate Express mode with the cdk deploy --express command when you deploy your CDK stack. This command retrieves your generated CloudFormation template and deploys it through the CloudFormation Express mode, which provisions your resources as part of a CloudFormation stack.

Express mode works with all existing CloudFormation templates and supports all CloudFormation features including change sets and nested stacks. When you enable Express mode on a parent stack, all nested stacks also use Express mode. If you need resources to be fully operational before proceeding with traffic or testing, continue using the default deployment behavior, which performs stabilization checks before completing.

Now available
AWS CloudFormation Express mode is available today in all AWS commercial Regions at no additional cost. For Regional availability and a future roadmap, visit the AWS Capabilities by Region. If you want to call APIs, search documentation, find regional availability, and check troubleshooting about this new feature, try using the AWS MCP Server and plugins with your preferred AI tool. To learn more, visit the CloudFormation documentation.

Start accelerating your deployments today, and send feedback to AWS re:Post for AWS CloudFormation or through your usual AWS Support contacts.

— Channy

❌