❌

Vue normale

Reçu avant avant-hier

Building operational resilience with agentic AI in financial services

18 août 2026 à 16:00

For financial institutions, operational resilience has long been embedded in regulatory and supervisory expectations — to say nothing of the high expectations of consumers. With the implementation of the European Union’s Digital Operational Resiliency Act (DORA), those expectations have become even more stringent, with more explicit, harmonized, and evidence-driven requirements. Firms must now demonstrate that their critical business services and supporting digital infrastructures can withstand disruption, support coordinated response, and recover with control.

To meet these conditions, Deutsche Bank developed an AI-powered agentic resilience platform that modernized its regulatory tabletop resilience exercises at scale and turned manual preparation into context-aware and evidence-ready simulations grounded in actual operational data. The platform builds enterprise context from architecture, data flows, logs, incident history, alerting signals, and operational telemetry to generate scenarios, simulated operational evidence, structured session records, and regulator-ready artifacts.

At many large banks with operations that span interdependent applications, data flows, and third-party services, this is a critical and even existential shift. Across financial services, supervisory expectations are evolving and as they do, banks’ tabletop exercises must reflect their production dependencies, real operating conditions, and compliance with consistent evidence standards more directly.

As Deutsche Bank considered how to successfully and efficiently make this shift at scale, it looked to its long-time partner, Google Cloud, and its growing suite of agentic AI tools.

From tabletop exercises to resilience intelligence

With its agentic resilience platform, DB has been able to transform its tabletop exercises from manual preparation to a continuous intelligence model. And it’s been able to extend the same agentic layer to root-cause analysis when real operational context is needed.

This means that every scenario it runs is based on real enterprise signals. The platform can then reflect true system dependencies, failure patterns, and business impact instead of relying on static inputs that are more likely to return assumptions than real-time insights.

By using Gemini Enterprise Agent Platform, DB has been able to migrate this operational context into structured scenarios with clear timelines, decision points, and expected responses. This has ensured that each exercise is grounded in real system behavior that produces consistent, audit-ready evidence that meets regulatory expectations.

Dual orchestration for control and flexibility

In order to deliver both regulator-grade control and operational flexibility, DB’s platform introduced a dual-orchestration architecture that separates workflows into two complementary execution models.

First, for regulator-aligned execution, the bank is using LangGraph to ensure that it generates every scenario through a traceable, deterministic process — with clear lineage from input context to output — that supports the auditability required for supervisory review.

Next, for its adaptive and investigative scenarios, DB is using Google Agent Development Kit (ADK) to enable agent-driven coordination. This approach allows the bank’s platform to dynamically analyze conditions and generate responses without predefined execution paths.

image1

Figure 1. Architecture for context assembly, orchestration, and scenario generation.

With this architectural separation, the platform can combine governed execution with adaptive investigation while preserving a common intelligence layer. The same agents and tools can reason over architecture, data-flow diagrams, logs, and code artifacts across tabletop scenario generation and related incident-analysis workflows. Importantly, this supports a consistent resilience model across both planned exercises and real operational events.

Deutsche Bank’s objective with this platform was to engineer a resilience model for critical financial systems that meets regulatory expectations — even within highly complex, distributed environments. By linking dynamically generated scenarios to real business context and combining governed orchestration with adaptive analysis, the platform has given us an intelligent, continuously adaptive model for operational resilience.” – Sanjay Tripathi, Managing Director, Global Head of Surveillance Technology & Compliance Cloud & AI Transformation Lead, Deutsche Bank

Powering generation and governance with Google Cloud

Google Cloud’s suite of agentic tools is providing the foundation for scaling Deutsche Bank’s platform across its many governed, enterprise-grade resilience workflows. Here’s how:

  • Cloud Run supports elastic execution of scenario and evidence-generation services. 

  • Gemini Enterprise Agent Platform transforms operational context into structured resilience scenarios.

  • Google ADK enables adaptive agent coordination.

  • Cloud SQL provides durable persistence for scenarios, session artifacts, and review records.

Collectively, these services give DB support for the traceable generation, controlled execution, and persistent evidence record required for compliance review and continuous improvement.

Scalable, evidence-ready resilience testing

Every scenario generated by Deutsche Bank’s platform drives a structured tabletop session for the teams that run response, escalation, and recovery. Because these exercises are grounded in real enterprise context, they reflect operational reality while also strengthening consistency across teams and creating audit-ready evidence that meets regulatory expectations. For institutions that operate under DORA or similar frameworks, this makes it easier to demonstrate controlled, coordinated, and disciplined response at scale. 

This model is now being applied across multiple DB portfolios, which is helping the bank establish more consistent and scalable resilience paradigms and a replicable blueprint for the broader financial sector.

In this model, root-cause analysis acts as the feedback loop between real incidents and future resilience testing. The resulting insights from production events can inform future tabletop scenarios, while exercise outcomes can strengthen response playbooks, escalation paths, and recovery readiness.

All of this extends the platform’s value from planned resilience exercises to real operational events while keeping scenario-based resilience testing as the primary use case.

As adoption expands, this platform brings consistency by embedding Google Cloud’s methodology for context-aware resilience. It eliminates fragmented manual approaches and establishes a cross-functional, AI-informed operating model across the bank.

Toward resilience intelligence

The bank’s next step is to extend this approach into a broader resilience intelligence layer, which is possible because it can deploy the same patterns to support playbook refinement, recovery-readiness assessments, and continuous validation of controls against evolving system conditions.

For financial institutions, this is a strategic shift. As systems become more distributed and regulatory expectations more demanding, banks must move from periodic resilience testing to continuous, intelligence-driven capabilities. At Deutsche Bank, Google Cloud is making that transition simple across the organization.

Learn more about Google Cloud’s methodology for context-aware resilience in this article.

How Deutsche Bank unlocked agility with an API-ready ecosystem

4 août 2026 à 18:00

When people think about digital transformation in banking, they often focus on the visible results: mobile apps and new digital services. But there's an invisible infrastructure making all these services possible: APIs. At Deutsche Bank, we recognized that APIs aren't just technical plumbing; they're the nervous system of modern banking. 

A few years ago, our application landscape was dominated by monolithic systems. As we evaluated how to break them into modular, reusable APIs, one thing became clear: we couldn't just decompose our work into APIs — we needed a central API management platform (APIM) to manage what would emerge. We needed something where documentation, security policies, and governance all had to be built in from the start, not bolted on later. 

The question wasn't just how to modernize, but how to best serve our customers and position ourselves for tomorrow's opportunities, especially with emerging technological paradigm shifts. 

Needing a system that was adaptable, scalable, reliable, secure, and AI-ready for the demands of modern banking, we chose Google Cloud's Apigee as our APIM platform. 

Building the backbone: four key capabilities 

Today, Apigee manages our API ecosystem — from open banking APIs connecting us with fintech partners, to the internal microservices powering our various banking platforms, and even the client-facing applications that enable seamless digital experiences such as online banking. 

Here are four important capabilities the platform offers us:

1. Unified governance without sacrificing speed 

Apigee is the foundation of our API catalog. Every endpoint, version, and dependency is documented and discoverable. Development teams find and reuse existing APIs rather than rebuild functionality. We've moved from "Where's that customer data API?" — which took days — to a searchable, real-time catalog accessible to any developer. 

But governance isn't about bottlenecks, it's about guardrails, and with Apigee's policy framework, we automatically enforce standards. OpenAPI specifications, schema validation, and error handling are now baked into the platform. Teams move faster because they work within consistent frameworks. 

2. Security: the employee onboarding analogy 

When thinking about API security, imagine onboarding a new employee. You don't give them access to every system on day one. You follow the least privilege principle, so they get exactly the permissions needed for their role. If they switch departments, their access rights will be updated. If they leave the company, access is revoked immediately. Apigee works the same way for our services and applications. 

When connecting a new service — say, one that accesses customer accounts — we don't open the floodgates. Through OAuth2 scopes and API key management, we define precisely what that agent can access: 

  • Read account balances? Yes. 
  • Initiate wire transfers? No. 
  • Access 90-day transaction history? Yes. 
  • Full historical data? Only with elevated permissions. 

Like employee access, these permissions are centrally managed, regularly audited, and instantly revocable. Just as we track employee activity for compliance, Apigee logs every API call to see who accessed what data, when, and why. 

This becomes critical with high-volume automated systems. An automated service doesn't take breaks and can make thousands of calls per minute if misconfigured. Rate limiting and quota enforcement ensure that even when something goes wrong, the blast radius is contained. 

3. Resilience and performance at scale 

Banking doesn't have downtime. When customers check balances at 3 a.m. or markets surge with trading activity, our APIs must respond instantly and reliably. 

Apigee's load balancing and auto-scaling evenly distribute that traffic. Health checks and circuit breakers automatically route around struggling services, and for frequently accessed data, Apigee's caching delivers sub-millisecond responses without hitting backends. 

4. Observability: measuring everything 

Before Apigee, understanding API performance was like assembling a jigsaw puzzle with pieces from different boxes. Now we have unified dashboards showing real-time traffic, error rates by service, usage analytics by consumer, and compliance metrics. This visibility serves operations, product managers who track partner value, and security teams who identify anomalies.

DtBank_Apigee_1

Apigee provides a central suite of capabilities for managing the full API lifecycle

The path forward 

We built this infrastructure for the API economy, and in doing so, we have also built a strong foundation for the future of digital banking. As the industry evolves, this API-first approach will be critical for integrating next-generation services. 

As digital banking continues to advance, a shift toward intelligent services that can react, predict, and assist in real time is underway. Capabilities such as real-time pattern recognition, predictive insights, and AI-powered assistants are becoming part of everyday digital experiences, with their visibility and impact increasing as adoption accelerates. Each of these capabilities will consume APIs — and they will introduce new requirements: ultra-low latency, high-throughput data flows, and secure orchestration across multiple APIs. 

Because we invested in a flexible API platform with Apigee, we are well-positioned to adapt and optimize our infrastructure for these future needs, rather than having to rebuild it. 

Emerging standards: MCP, A2A, and the future 

The industry is exploring new integration standards. Protocols like Model Context Protocol (MCP) and Google's Agent2Agent (A2A) are interesting because they build on existing API infrastructure. 

Our Apigee-managed APIs are well-positioned to leverage these advancements. For instance, MCP could benefit from our OpenAPI specifications, and A2A could leverage our OAuth2 framework, with both relying on the governance we've built. 

We're also exploring patterns like placing new types of servers behind Apigee proxies to maintain security controls while enabling modern workflows. Our "always-API" pattern ensures that services benefit from centralized management, no matter how they are accessed.

DtBank_Apigee_2

MCP and A2A are complementary, MCP has a tools and resources focus, while A2A is focused on peer collaboration

The vision: APIs as universal interface 

Every banking capability will eventually be exposed as an API. That’s not because APIs are trendy, but because they're the most flexible, composable, and governable way to share functionality, whether consumed by mobile apps, partner fintechs, analytics platforms, or other automated agents. 

At Deutsche Bank, this shift is already taking shape. The same API foundation that powers our core platforms is now enabling our evolution toward more intelligent, AI-supported services across the bank. That foundation provides the consistency, governance, and scalability needed to bring these capabilities to life, ensuring that as new intelligent services emerge, they can be integrated seamlessly, securely, and at enterprise scale. 

Apigee makes this possible by providing governance that scales across all use cases. It's not about controlling innovation; it's about enabling it safely. 

Lessons learned 

  • Invest in excellent documentation. Semantic summaries and clear schemas aren't extras; they're foundational for both developers and AI. 

  • Treat security like employee onboarding. Least privilege and role-based access apply equally to APIs.

  • Observability is a competitive advantage. Unified analytics enable data-driven decisions. 

  • Plan for the future now. Your API management infrastructure becomes your advanced integration layer. 

  • Stay curious. Experiment with emerging standards. Flexibility wins. 

Conclusion 

We're at an inflection point. The API economy enabled fintech and open banking. Now, the same infrastructure can serve as the backbone for the next wave of innovation. Our investment in the API platform wasn't just about managing APIs better; it was about building a foundation for whatever comes next. 

As the industry transforms, we’re ready. The future belongs to organizations that move fast without breaking things. For us, that future is powered by Apigee.

From maintenance to innovation: Checkout's migration to Managed Service for Apache Airflow

22 juillet 2026 à 16:00

Data engineering teams often face a “Day 2” operational reality after building a data platform: the ongoing work of maintaining the orchestrator itself.

For the Data Platform team at Checkout.com, managing a self-hosted Apache Airflow environment on another hyperscaler was consuming time the team wanted to spend elsewhere as server management, patching, and incident response were pulling focus from building pipelines.

By migrating to Managed Service for Apache Airflow (Gen 3), Google Cloud’s fully managed Airflow service, Checkout.com transformed its reliability and cost structure. Here’s how they built a more scalable, cost-efficient, and robust data foundation.

The starting point: self-managed Airflow

Before the migration, Checkout.com ran Airflow on self-managed infrastructure. While functional, maintaining the underlying resources required significant attention. Patching, upgrades, and server management created regular interruptions.

Operational data from the past year illustrates some of the challenges the company was navigating:

  • Reducing operational friction: In its self-managed environment, Checkout.com faced stability challenges, particularly during high-load periods. 

  • Complex dependency management: Upgrading packages and ensuring compatibility was a constant, manual struggle. With Managed Airflow (Gen 3), the company was able to simplify this by handling dependencies at the image level, ensuring seamless compatibility out-of-the-box during routine environment upgrades.

  • DAG sync time: Syncing DAGs to the scheduler took approximately six minutes after deployment to S3, which affected iteration speed.

  • Manual processes: Scaling required manual intervention, and onboarding new teams meant manually creating secrets and variables for dbt.

The solution: Managed Service for Apache Airflow (Gen 3)

Checkout.com’s team migrated to Managed Airflow to offload infrastructure responsibility and take advantage of Google Cloud's managed scalability. The results were immediate and measurable across three areas: reliability, cost, and developer velocity.

checkout-managed-airflow-chart

Dynamic scaling in action

In the previous elastic container service setup, the team allocated the maximum number of workers required for peak loads. This meant paying for peak capacity around the clock, regardless of actual usage.

Managed Airflow provides built-in dynamic scaling, eliminating the need for manual resource management. The environment automatically adjusts the number of workers based specifically on the workload demands. When tasks spike, the system scales up; when they drop, it scales down to save resources. Similarly, moving from fixed provisioning to dynamic scaling reduced monthly costs by an estimated 30%.

Reliability and DAG isolation

Achieving increased stability was a primary driver for Checkout.com’s migration since in the past, a single problematic DAG could affect its entire environment. 

Managed Airflow introduced a number of critical architecture improvements:

  • DAG isolation: Each DAG runs in its own execution environment. If one DAG fails or consumes excessive resources, it doesn’t affect the entire environment.

  • Managed operations: Google Cloud handles patching and upgrades during scheduled windows, removing the need for manual upgrade management.

  • Improved visibility: Integration with Google Cloud Monitoring and Cloud Logging provides clear visibility into task execution. Engineers can now debug issues independently without escalating to the platform team.

Faster developer workflows

The migration also improved day-to-day workflows for Checkout.com’s data engineers.

  • Faster deployments: Using Cloud Storage for DAGs enabled near-instant syncing.
  • Simpler onboarding: Teams no longer needed platform support to create variables before onboarding.
  • Modernizing dbt execution: One of the company’s most significant wins was changing how it runs dbt. Previously, its engineers had to manually install and manage complex virtual environments for every supported dbt version. By leveraging containerized dbt runs, Managed Airflow (Gen 3) eliminates dependency bottlenecks. This ensures complete dependency isolation, allowing teams to run any required dbt model with minimal setup and no manual infrastructure overhead.
  • Environment updates: The company no longer needs to redeploy the entire Airflow environment to add new roles or update Python packages.
  • AI-powered troubleshooting with Gemini Cloud Assist: In a self-managed environment, a failed task often triggered a frantic hunt through fragmented logs and metrics. With Managed Airflow, Checkout.com can initiate a Gemini investigation directly from its Airflow DAG UI in the Google Cloud console. 

Gemini doesn't just provide generic error messages; it generates a scorecard that evaluates different hypotheses with both supporting and contradictory evidence, which can drastically reduce mean time to recovery.

Conclusion

For Checkout.com, the move to Managed Airflow (Gen 3) marked a strategic shift, one that freed its engineers to focus on delivering value.

"With Managed Service for Apache Airflow, we’ve achieved significant improvements in efficiency, scalability, and reliability. Managed infrastructure, automated scaling, faster deployments, and isolated execution environments have transformed how we operate." — Keisi Mancellari, Data Platform Engineer, Checkout.com

With a stable, scalable, and cost-efficient platform in place, Checkout.com is now able to  focus on the future of its data pipelines, confident that its orchestration layer is ready for whatever comes next.

Learn more about Managed Service for Apache Airflow and how it can support your data platform.


Special thanks to the following contributors to this post: Serge Bouschet and Keisi Mancellari

❌