❌

Vue normale

Reçu avant avant-hier

How to upskill enterprise AI builders by using daily micro habits

18 septembre 2026 à 18:00

As enterprises invest in generative AI, tech leaders keep seeing the same pattern: Developers test AI tools for a week, hit setup problems, and then drift back to the backlog. Nothing ships.

The real gap is enablement. In this landmark Harvard Business Review article, Josh Bersin and Marc Zao-Sanders noted that knowledge workers carve out just five minutes a day for formal learning. Most enterprise training programs still lean on week-long classroom bootcamps, multi-week certification tracks, and passive video lectures, none of which fit into the time developers actually have. 

With the Build with Gemini event series underway, Google Cloud Consulting is seeing more leaders rethink AI enablement by building quick, daily practice into their teams' routines. In this post, we'll walk through a four-pillar approach and the lessons from our global developer challenges to share what micro-habit upskilling looks like.

Moving from workshops to daily practice

The traditional method…

…now becomes

Multi-week, semi-annual classroom bootcamps

Five-minute hands-on exercises

Local workstation configuration and credential setup

Pre-configured browser-based sandboxes

Mandatory attendance and compliance checks

Daily streaks, badges, and team challenges

Multiple-choice quiz completion

Deployable agent tools and reusable code 

Rolling out a model like this comes down to keeping each task small and manageable. Here's how we structure that work across engineering teams:

  1. Make micro-learning a habit. Offer short objectives that each cover one skill, like connecting a model to a database schema or validating structured output, in place of full-day training blocks.

  2. Give teams browser-based sandboxes. Setup is where most training stalls, so remove it. With a pre-configured, managed cloud environment, developers open a tab and are writing code within minutes, with no credentials to request and nothing to install or maintain on their own machines.

  3. Build in daily streaks. Milestones, shared wins, and teammates comparing solutions turn practice into a normal part of the workday.

  4. End every session with something that runs. Each exercise should leave behind a working component, and over time those components accumulate into a shared library of code and prompts the whole team can pull from.

Lessons from the Advent of Agents program

image

When Google Cloud launched Advent of Agents, a daily agent-building program for developers, we wanted to test one question: what happens when you remove setup and scheduling from technical enablement?

Each day, developers got one short, real-world agent exercise they could run right in the browser, with no half-day to block off and no setup guide to read first. 

  • 150,000+ developers participated across global teams.

  • 859,000+ hands-on code executions in browser-based environments.

  • 31% of participants returned daily, more than triple the 10% industry average for self-paced tech, and significantly exceeding the standard 5%–15% MOOC benchmark

  • 32,000+ participants built working agent components.

The above data was accessed via Advent of Agents Google Analytics metrics.

Keeping each exercise under five minutes and pre-wiring the sandboxes removed the two things that usually stall workplace training: setup time and scheduling. The numbers suggest developers will make time to learn when the exercise fits into the day they already have.

Putting micro-enablement into practice

image1

AI enablement doesn't have to pause your sprints. It takes a consistent habit of practice and the tools that let teams build alongside their regular work.

  • Experience live building. Bring your engineering teams to a Build with Gemini workshop. The events are complimentary and run different tracks according to technical depth, from no-code for business leaders to code-first for developers, with live hands-on labs supported by Google Cloud experts.

  • Build skills with GEAR. Enroll your technical and business teams in the Gemini Enterprise Agent Ready (GEAR) program. Membership is free and includes monthly learning credits on Google Skills, hands-on labs, and skill badges, with learning paths for developers, line-of-business leaders, and IT decision-makers.

Start small, build often

Developing AI skills starts with a change in routine. Short, daily, hands-on exercises let developers learn by doing, and the working code they produce along the way becomes the team's starting library for production work.

Give your developers a few minutes a day and a sandbox that's ready when they are. Start with one exercise this week and see how small, daily habits can build AI capability across your organization.

How Orange built FinOps accountability, and why agents are next

16 septembre 2026 à 18:00

At Orange, the leading France-based multinational telecom provider, there are days when engineering teams set aside their delivery backlogs and spend the day cleaning up cloud spend together. There's a leaderboard. There are goodies on the line. Experienced practitioners guide the newcomers, so people learn the work while doing it. By the end of the day, sponsors can see the results.

Orange calls these FinOps Clean Days. Together with gamified hackathons, they've earned the company's 100-plus person FinOps community a Net Promoter Score within the organization that’s above 70.

Those numbers point at something the wider industry is wrestling with. Recent State of FinOps reports identify getting engineers to take action as one of the top challenges organizations face. Moving from awareness to action means finding ways to build FinOps accountability, and to get teams to genuinely care.

That makes FinOps a business change problem. And business change problems have known solutions. We spoke with Camille Marini, the FinOps lead at Orange, to get a deeper understanding of how the company overcame these hurdles to accelerate AI adoption and ROI, and how your organization might follow the same course.

Why the Clean Days work

Orange has held two principles since it set up its FinOps team. First, Cloud FinOps is a shared responsibility, with every stakeholder in a project involved in their own way. And the only path to that shared responsibility runs through communication and a deliberate change effort. 

“We insisted on the concept of shared responsibility across the organization for our FinOps practices,” Marini told us. “It’s very similar to how we approach cloud security. We needed to make teams understand that every single stakeholder in a project is involved in FinOps, each in their own way, if we are going to achieve responsible and impactful AI spending and usage.”

Those principles led Orange to create a FinOps Community of Practice, with support from Google Cloud Consulting. The team ran it on standardized communication channels so the methodology reached well beyond the central group, and kept the meetings actionable, sharing optimizations and billing updates so every session provided value.

The Clean Days came from a clear-eyed reading of how agile teams actually operate. In agile environments with deployment running constantly, optimization work rarely wins against the sprint. Delivery priorities, backlogs, and daily operations take the available time first. So Orange created protected time, made it collaborative, and made it fun.

McKinsey's four building blocks of change explain why this approach lands. Any large organizational change, the framework holds, requires action across four areas:

  • Conviction and understanding: "I know what is expected of me and I agree with it."

  • Formal mechanisms: "The structures, processes, and systems reinforce the change."

  • Role modeling: "I see my leaders and colleagues behaving differently."

  • Talent and skills: "I have the skills and opportunities to behave in a new way."

Map Orange's practice onto those blocks and the pattern is visible. Gamification and rewards give engineers colleagues to emulate: The leaderboard makes different behavior visible, and sponsors see the quick wins for themselves. Experienced practitioners guiding novices builds talent and skills through the community itself. The regular sessions, sharing optimizations and billing updates, build the conviction that comes from knowing where the money goes.

image2

FinOps activities mapped to the four building blocks of change, with the points where AI agents can reinforce them.

What happens beyond 100 people

A community of 100 engaged people is an achievement. But in an organization with thousands of engineers, no central FinOps team can reach everyone directly. The question for leaders is how to extend what a community like Orange's creates — the awareness, the shared ownership, the habit of acting — to people the FinOps team will never meet.

This is where AI agents extend the capabilities of a FinOps team with two core benefits. They take on complex, time-intensive activities that previously needed a human, and they reduce friction around FinOps for individuals across the business.

Getting teams to adopt them takes a strategy aimed at your own organization's pain points, which often come from high cognitive load, unclear accountability, or competing priorities. Start by finding where engagement drops off in your FinOps lifecycle:

  • An awareness gap: If teams are unsure of their spend impact, an insight agent can push real-time cost data into their daily tools.

  • A bandwidth gap: If engineers are too busy with backlogs, a remediation agent can identify quick wins and present them as ready-to-merge code changes.

  • A complexity gap: If reporting feels like a manual chore, an orchestration agent can gather the data and simplify the process.

Start with trust, then add autonomy

The sensible path runs in sequence. Establish the community practice, the way Orange did. Then introduce read-only agents that inform and suggest. Only once those are established across the community should you build agents that execute changes. Direct action carries operational risk, so manage it carefully. It's also where significant wins often sit.

How you build depends on who's building. For teams that want to deploy quickly with minimal code, the Gemini Enterprise App provides a no-code environment for creating agents. For developers who need granular control, the Gemini Enterprise Agent Platform (formerly Vertex AI) offers advanced tools for launching and governing agents built with frameworks like the Agent Development Kit (ADK).

Cloud FinOps is moving beyond centralized reporting toward action that happens where the work does. The organizations getting there start with the culture, then use agents to carry it further than any one team could reach. 

Orange's numbers came out of the community work. Building that foundation is the part worth copying first. When you're ready to extend it, Google Cloud Consulting can help you shape the community practice, and the Gemini Enterprise App is a low-lift way to put your first read-only agent in front of your teams.

Best practices for handling cloud reliability incidents

15 septembre 2026 à 18:00

Cloud outages can range from global service disruptions to issues isolated to a specific region, zone, or even just your project, workload or application. If you suspect a Google Cloud Platform outage is impacting your services, we recommend you follow a structured “Verify→ Investigate→Report→Resolve→Review" workflow to resolve it. And before that outage occurs, you should also have prepared your environment for an eventual disruption by designing for failure, and actively practicing the steps you need to take to restore service. 

In this blog, we summarize the key reliability incident handling best practices to help you design and practice your reliability incident response capabilities and minimize impact. Rather than an exhaustive guide, this is meant as a primer on only the most important practices for advisory purposes. Please note that we do not cover additional practices specific to security incidents here. 

Beyond the base steps covered here, you may want to also explore how AI agents and tools are starting to transform incident handling. Check out this episode of the Prodcast, where Googlers explore the latest trends of leveraging agentic AI in Site Reliability Engineering (SRE) to detect issues early and prevent disruptions. Try Cloud Assist investigations, or explore Agent Skills and remote managed MCP servers to give you another set of tools for quickly pinpointing an issue. Before getting into these advanced techniques, we focus below on the foundational steps to good incident handling.

1. Prepare

Long before things start to go sideways, you should have spent significant time preparing for an outage along at least four dimensions: design, data, playbooks and training.

  • Design: Think ahead and mitigate future incidents by designing automated response actions, like a load balancer shifting traffic away from slow or unresponsive instances, or by automating as much of your incident response playbook as possible. Review designs of all critical applications to automate as many actions as possible to accelerate response and recovery.

  • Data: When a disruption occurs, having meaningful data at your fingertips vastly improves response capabilities. Use Cloud Logging, Cloud Trace and Cloud Monitoring, or other third-party observability tools, and replicate that data to a redundant stack in a separate location from the systems being observed. Make sure, in advance of any incident, that time stamps are synced across your observability streams for easy correlation, or know how to do that on-demand during an outage, when time is of the essence.

  • Playbook: A well-thought-out playbook documenting your incident response processes, including crystal clear role and responsibility definitions for all personas, is paramount to efficient incident response. Who is responsible to do what? Who needs to be notified or mobilized for each type of disruption? How can they be reached? What tools and data are available? How are results communicated? How do teams hand over to the next shift during long running incidents? etc. Conduct a simulated incident response and critically review every step to find where your playbook needs clarification. Without clear responsibilities, mitigation inevitably takes longer.

  • Training: Hopefully, service disruptions are rare events. To ensure your staff knows and remembers how to react, they need to retrain on the process several times per year by running simulated cross-team incident response drills. A retrospective on the simulated exercise will help identify warranted improvements.

2. Verify

Despite your best efforts, sooner or later, a service disruption will occur, which you can detect via any number of mechanisms:

Now, you need to determine what broke and who should ultimately fix the problem:

  • Google, e.g., a bug, code roll-out, hardware failure, etc.

  • You, e.g., a configuration change, elevated load, quota ceiling, etc.

  • Third party, e.g., a directory hosted by a different cloud provider

If Google has declared an incident and started working to fix the problem, estimate whether you can possibly reestablish service sooner, for example by failing over to a secondary stack (see the ‘Typical Causes’ table below). You can determine whether Google has declared an incident and will provide a fix by consulting:

  • Personalized Service Health: Check this first. Personalized Service Health shows incidents specifically relevant to your projects and regions, distinguishing between incident types:. 

    • Emerging Incidents: Google has received an alert, on-callers are investigating, impact is yet unknown

    • Confirmed Incidents: Google has investigated and found customers are impacted

Located within the Google Cloud console, Personalized Service Health often displays limited-scope incidents that don't appear on the public dashboard. Personalized Service Health also offers a mobile client for Android and iOS smartphones, assuming you can use your work ID and credentials on the phone.

  • Gemini Cloud Assist, which is integrated with Personalized Service Health, so you can use it to query that information in natural language.

  • Cloud Service Health dashboard: This is the public-facing non-authenticated web page for broad, severe incidents affecting many customers. Limited blast radius disruptions are not externalized to the public. All its content is available in Personalized Service Health as well. If ever Personalized Service Health goes down, Cloud Service Health serves as an alternative channel built on a separate infrastructure.

  • Known Issues: In the console, navigate to Support > Cases, view a case, and use the resource selector on the console toolbar to find the specific cloud resource you’re interested in. Then click Known issues. If your issue matches one listed here, you can link a support case to it, so you will receive automatic updates in your case record. If you don’t find a match, open a new support case. Google will automatically match the case to a related incident, as soon as one is declared.

  • Google declared incidents are updated as new information becomes available, so check back regularly, or set up a Personalized Service Health alert policy to be notified each time new information becomes available.

If you host cloud resources in multiple clouds, a good practice is to check early on whether the problem occurs for multiple cloud providers. If so, the problem is likely external to the providers and caused either by you or by a third-party service that your application interacts with.

3. Investigate

To determine the blast radius within your cloud footprint of Google-declared reliability incidents, first check Personalized Service Health updates for a description of the technical problem. Knowing what to look for will allow you to map your blast radius and decide on suitable contingency actions quicker.

If Google hasn’t declared an incident, try to rule out configuration errors or issues within your environment by checking:

  • Cloud Monitoring: Look for spikes in error rates (e.g. 5xx errors), increased latency, or drops in traffic in your dashboards.

  • Cloud Logs: Use Log Explorer to look for specific error messages like DEADLINE_EXCEEDED, SERVICE_UNAVAILABLE, or specific API errors.

  • Quotas: Ensure you haven't hit a project quota (e.g., CPU, API rate limits), which can often mimic the behavior of an outage.

  • Change history: Check your log of recently applied changes. Not all problems manifest immediately, but proximity on a timeline can be a powerful indicator of causality, even if it’s not proof. Also check whether Google rolled out any updates just before the symptoms started. See the Unified Maintenance Management interface in Cloud Hub.

Absent a clear culprit, such as a traffic spike or a DDOS attack, and if symptoms manifested immediately after rolling out a change, a good strategy is to back out that change and attempt to return to a last known good configuration. 

4. Report

If the Cloud Service Health and Personalized Service Health dashboards are green but your metrics show a failure, you must report it to Google. 

  • Determine priority:

  • File a case: Go to Support > Cases > Create Case in the console.

    • Explain quantifiable business impact to rationalize the submitted priority and prevent it from being reset when Cloud Support prioritizes cases. A clear and accurate rationale helps!

  • Essential information to include:

    • Project ID and affected region/zone

    • Timestamps (when it started and if it's ongoing) with a clearly labeled timezone

    • Specific error messages or log snippets

    • Scope: Is it affecting all users/systems, or a specific subset/location?

Escalation for Premium/Enhanced support

If you have a Premium or Enhanced support plan and a P1 case is not receiving the attention it requires, use the Escalate button within the support case in the console. This alerts a support manager to investigate and rectify the situation.

5. Resolve

By taking these steps, you are well on your way to resolving the outage. In the meantime, here are some ways to mitigate the impact of the outage and communicate with impacted stakeholders.

While waiting for a resolution:

  • Communicate: Notify your stakeholders and customers. Transparency helps manage expectations and reduces duplicate internal reports.

  • Fail over: If you have a multi-regional architecture, consider shifting traffic to a healthy region. As a best practice, first ensure that the disruption is at the infrastructure level and not at your workload level. 

  • Check for workarounds: While working on a permanent fix, Google often posts temporary workarounds in the Service Health Dashboard updates, or in Personalized Service Health updates.

  • Consider your regulatory reporting requirements: Know whether your organization is subject to regulatory reporting requirements, and what the required deadlines are for both initial and follow-up reporting. Google Cloud prepares Incident Reports for incidents that meet certain criteria — see details here for how to get those reports. Premium Support customers can also request an Incident Summary, which is an Incident Report customized to your account’s specific hosting location, time stamps, etc.

De-escalation and closure

Once systems are stable, Google downgrades the severity levels and deactivates the active on-call escalation chain. Google only closes an incident in Personalized Service Health when it has taken all the mitigation steps covering all impacted customers. Your specific services might be restored sooner than the incident closure time, if other customers are restored later than you. The incident is officially closed on the Google Cloud Status Dashboard when systems have run stably for a designated auto-close duration. Verify that your services are operating normally at this point. And if your incident responders aren’t compensated for extra time spent on the incident, find a way to thank them.

6. Review

After the problem has been fixed and operations have returned to a normal, steady state, it’s time to conduct a post-mortem analysis to identify how your team can respond better in future service disruptions. A “blameless” approach is essential to surfacing meaningful and impactful improvements that can be made to your incident response process. Ask questions like:

  • What went well?

  • What could we have done better?

  • Where did we get lucky?

  • Where did we get unlucky?

Then decide what changes can be made to improve your playbook, tools and training.

At Google, we often publish a post-mortem or Incident Report for major outages, available via Personalized Service Health. Review this to understand the root cause and adjust your own disaster recovery plans to prevent or reduce future impact. Customers with a Premium Support plan can request an Incident Summary for a Google-caused incident they were impacted by and for which they opened a P1 case. An Incident Summary is an Incident Report customized for your environment (e.g., start and end times of impact).

Typical causes, comms and prevention strategies

To help you prepare and plan ahead, here’s an overview of some typical incidents based on the symptoms reported in Cloud Service Health and Personalized Service Health along with guidance on what Google communications to expect, and some generic mitigation or prevention strategies you can build into your playbooks.

Blast radius

Typical cause

Comms

Strategy

Single zone or region.Subset of products.

Typical of a software problem triggered by a rollout. Learning points:

- Understand the location scope (zones and regions) of your workload

- Products can depend on other products

Major incidents are communicated via Cloud Service Health.Major and Minor (by number of customers, not severity) incidents are communicated via Personalized Service Health.

Highly localized incidents are not communicated via Cloud Service Health or Personalized Service Health.

Fail over, if so configured, but verify the health of the secondary stack first.

Single zone.Most or all products.

Typical of a power or cooling issue.

Check Cloud Service Health and Personalized Service Health.

Fail over to a different zone, if so configured.

Single region.

Most or all products.

Typical of a backbone networking infrastructure issue 

Check Cloud Service Health and Personalized Service Health.

Fail over to a different region, if so configured.

Control plane issue for a product

Typical of a late detected issue

Communicated via Personalized Service Health if significant customer impact is verified.

Look for workarounds. Wait for Google to fix. Fail over, if so configured.

Multi-regional issue with a global product

Rare but possible, typically detected quickly. Learnings: Mitigation options can be limited. Try regional variants, alternative products with similar functionality

Check Cloud Service Health and Personalized Service Health.

Wait for Google to fix. In the meantime, verify via Google Comms and your own investigation that this is truly Google’s problem to fix.

Capacity / Stockout issue

System-level demand exceeding capacity in the product/location/model. (Cloud is designed to scale, but limits always exist, so proper planning is advised)

Error message. No incident will be declared.

Place reservations for predicted capacity needs (if cost is acceptable). Flexibility in zone placement can also help.

Quota exhaustion

Difficult / inaccurate prediction of traffic

Error message. No incident will be declared.

Review consumption trends against ceiling regularly.

Go deeper

This document offers only a condensed summary of key points. If you have an active Premium Support contract with Google Cloud, reach out to your account team for a deeper review of your response plans. For a comprehensive treatise on how to build reliable services and how to respond to incidents, we strongly recommend Google’s SRE Book, which is available as a free download. A new version of the SRE book is releasing ~Oct 2026 and will be available for purchase on O’Reilly Media. We’re also working on a future primer that explores AI-supported incident handling in-depth — stay tuned!

New AI-powered quick assessments in Migration Center turbocharge modernization

24 août 2026 à 18:00

Technology leaders are under mounting pressure to modernize infrastructure, control multi-cloud operational spend, and build data foundations for generative AI. However, the discovery required for that level of transformation can entail weeks of manual spreadsheet analysis, mapping in-house infrastructure, and reconciling siloed, piecemeal cost estimates across disparate teams and sources. To help, we’re announcing AI-powered Quick Assessments in Migration Center, which delivers near-instant total cost of ownership (TCO) modeling and automated service mapping.

Compare this to legacy assessment processes, which can stall digital transformation initiatives before they even launch. Manual discovery can delay migration timelines by months, increase engineering overhead, and often miscalculates complex financial models. By replacing manual discovery with AI-assisted automation, IT gains instant, actionable visibility into the TCO and return on investment (ROI) for a given migration initiative. 

Now, organizations can generate comprehensive migration financial models in minutes rather than months. Teams ingest raw infrastructure data or cloud billing reports and quickly receive an optimized target bill of materials (BOM), service mapping coverage, and projected savings. Decision makers interact with an agentic assistant that explains underlying financial assumptions, recommends technical cost optimizations, and exports ready-to-share executive reports.

Inside the AI-powered Migration Center

This is made possible with AI-assisted Quick Assessments alongside enhanced Cloud Billing assessment capabilities, both integrated into the new AI-powered Migration Center.

AI-assisted Quick Assessment for on-premises workloads

Designed for enterprise customers and partners, AI-assisted Quick Assessment automates on-premises infrastructure evaluation to provide rapid financial modeling. Let’s walk through these new capabilities: 

  • Instant Compute Engine TCO estimates convert VMware inventory exports (such as RVTools) or aggregated infrastructure inputs into precise Compute Engine cost targets:

1

Migration Center’s Quick TCO Estimator

2

Migration Center’s Quick TCO Estimator results page

  • Advanced architecture modeling supports latest-generation Gen4 compute instances and high-performance Hyperdisk storage pools:
3

Migration Center’s Quick TCO Estimator detailed results page

  • Customizable financial controls allow teams to adjust on-premises baseline cost assumptions to match internal accounting standards:
4

Migration Center’s Quick TCO Estimator detailed results page (continued)

  • Context-aware agentic chat recommends tailored technical cost optimizations aligned with your specific business constraints (such as regional location or compliance needs), clearly explaining the underlying logic and financial assumptions.
5

Migration Center’s agentic chat capabilities

6

Migration Center’s agentic chat capabilities (continued)

7

Migration Center’s agentic chat capabilities (continued)

  • Automated Business Case and Google Sheets export generates ready-to-use reports capturing the complete recommended BOM, TCO comparison, and ROI analysis:
8

Migration Center’s business case

9

The path forward

Modernizing your infrastructure starts with fast and accurate data. Migration Center’s Gemini-powered features simplify cloud evaluation, empowering IT decision makers to build defensible business cases generated by machine-learning.

Try Migration Center directly in the console today, or take a free migration and modernization assessment to evaluate your workloads and accelerate your strategic cloud journey with Google Cloud.

❌