❌

Vue lecture

Maximizing Apache Spark availability: Mitigating compute stockouts with flexible VMs and other best practices

The surge in AI development has created unprecedented demand for compute capacity around the globe. This can have negative implications for data processing and pipelines with Apache Spark. Whether you are managing your own Spark infrastructure or using a managed service, you can face availability constraints. However, a significant advantage of using Google’s Managed Service for Apache Spark is the availability of flexible VMs, which provide a targeted mechanism to adopt a dynamic, resource-agnostic philosophy and ensure your pipelines remain operational, even during regional or zonal capacity stockouts.

Understanding capacity stockouts

Capacity stockouts occur when demand for a specific machine family (such as N2 or N2D) exceeds available capacity in a target zone or region. For time-sensitive analytics pipelines, rigid single-VM requirements transform standard provisioning into a single point of failure which can result in cluster creation delays, failed executions, and potentially compromised business SLAs.

Flexible VMs

Flexible VMs fundamentally overhaul how a Managed Spark cluster requests compute resources. Rather than binding a cluster to a rigid instance type, flexible VMs allow teams to establish an ordered list of acceptable machine families for master, primary worker, and secondary worker nodes.

Key features

  • Multi-family blending: Mix nodes across diverse machine types and generations, combining Gen2 families (e.g., N2, N2D) with Gen4 families (e.g., N4, C4) in a single configuration.

  • Mixed storage support: Broaden available capacity pools by allowing storage options to dynamically adapt to the underlying host family's supported disk types.

  • Comprehensive cluster coverage: Apply flexible rules to primary workers, secondary (preemptible/spot) workers, and master nodes to guarantee cluster provisioning end-to-end.

Ranked configuration: A strategy for success

A successful flexible VM implementation relies on intentional ranking. By defining a clear hierarchy of options, Managed Spark clusters automatically attempt provisioning, systematically mitigating stockout risks without requiring manual intervention. To improve the availability of  suitable VMs, we recommend specifying at least two machine families in the highest priority (Rank 0) flexible VM list.

As an example, for production pipelines standardizing on n2d-standard-16 shapes, the following tiering strategy provides robust resilience against capacity constraints:

Rank

Machine family examples

Storage recommendation

Rank 0 (Primary)

n2d-standard-16, n2-standard-16

Standard Local SSD or PD

Rank 1

n4-standard-16, n4d-standard-16

Hyperdisk Balanced

Rank 2

c4-standard-16, c3-standard-22

Hyperdisk Balanced

Rank 3 

e2-standard-16

Standard PD

code_block
<ListValue: [StructValue([('code', 'gcloud dataproc clusters create $CLUSTER_NAME \\\r\n--num-workers=10 \\\r\n--zone="" \\\r\n--region=us-east1 \\\r\n--worker-instance-selection=\'{"machineTypes":["n2d-standard-16","n2-standard-16"],"rank":0,"diskConfig":{"bootDiskType":"pd-standard","bootDiskSizeGb":400}}\' \\\r\n--worker-instance-selection=\'{"machineTypes":["n4-standard-16","n4d-standard-16"],"rank":1,"diskConfig":{"bootDiskType":"hyperdisk-balanced","bootDiskSizeGb":400}}\' \\\r\n--worker-instance-selection=\'{"machineTypes":["c4-standard-16","c3-standard-22"],"rank":2,"diskConfig":{"bootDiskType":"hyperdisk-balanced","bootDiskSizeGb":400}}\' \\\r\n--worker-instance-selection=\'{"machineTypes":["e2-standard-16"],"rank":3, "diskConfig":{"bootDiskType":"pd-ssd","bootDiskSizeGb":400}}\' \\\r\n--master-instance-selection=\'{"machineTypes":["n4-standard-16","n4d-standard-16"],"rank":0,"diskConfig":{"bootDiskType":"hyperdisk-balanced","bootDiskSizeGb":400}}\''), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb3fbbd5a90>)])]>

For pipelines standardizing on legacy n1-standard-16 shapes, the following tiering strategy helps transition workloads toward newer, more available architectures while preserving operational stability:

Rank

Machine family examples

Storage recommendation

Rank 0 (Primary)

n1-standard-16

n2-standard-16

Standard Local SSD or PD

Rank 1

n2d-standard-16

Standard Local SSD or PD

Rank 2

n4-standard-16

n4d-standard-16

Hyperdisk Balanced

Rank 3

e2-standard-16

Standard PD

Leveraging Hyperdisk Balanced

Unlocking maximum availability with flexible VMs often requires adopting modern storage architectures like Hyperdisk Balanced. Newer instance families (including N4 and C4) rely on Hyperdisk to deliver predictable performance across variable VM sizes. Starting with default IOPS and throughput settings typically provides a reliable baseline for the majority of distributed Spark jobs.

Trade-offs and key considerations

While flexible VMs  dramatically improve cluster provisioning success, aligning them with enterprise requirements involves evaluating several architectural and financial factors:

1. Resource quotas

It is no longer enough to have one specific machine (e.g., N2) quota. You need to ensure you have sufficient compute and disk quotas allocated for all specific machine types and disks (including Hyperdisk) defined in their flexible VM lists.

2. Compute flexible Committed Use Discounts (CUDs)

Traditional, resource-based CUDs are tied to specific machine families, which limits flexibility. Adopt Compute flexible Committed Use Discounts (CUDs) to apply savings across multiple VM families and regions.

3. Performance Characteristics

Performance can vary between machine generations, as well as between Local SSD and Hyperdisk. While the Managed Spark team maintains internal benchmarks for these comparisons, actual outcomes are workload-dependent. Testing your specific Spark jobs across these families is essential for understanding SLA impacts.

Additional recommendations

In addition to implementing flexible VMs, there are several other key architectural and scheduling strategies to improve resource availability and workload stability:

  • AutoZone: Implement AutoZone routing to allow Managed Spark to automatically select the zone best suited to execute the job based on current capacity.

  • Smaller machine shapes: Avoid high in demand, large-core shapes. Design workloads and YARN containers to utilize smaller machine shapes (such as 4, 8, or 16 cores). These smaller shapes are much easier to fulfill from the available GCE on-demand pool.

  • Autoscaling: Deploy cluster autoscaling with reasonable maxInstances to manage capacity effectively for bursty or unpredictable workloads without relying on rigid, massive upfront provisioning.

  • Partial cluster creation: Configure a minimum acceptable number of primary workers. This allows clusters to spin up under resource constraints and begin executing, while autoscaling can dynamically add remaining workers as resources become available.

  • Establish regional fallbacks: Some regions, such as us-central1, can experience  high demand. Setting up fallbacks to other regions reduces capacity stockout risks.

Keep your Spark jobs running with flexible VMs

Managing your own Apache Spark infrastructure can be complex, especially when capacity stockouts disrupt your data processing. Utilizing a managed service like Managed Service for Apache Spark provides unique advantages — including built-in platform resilience and access to flexible VMs. By adopting a prioritized fallback strategy with flexible VMs, you can protect your workloads from regional hardware shortages and keep your critical pipelines running.

Ready to improve your Spark workload resilience? Start configuring flexible VMs for your Managed Spark clusters today.

  •  

Google Cloud partners with CIQ to provide an enterprise-grade experience for Rocky Linux

At Google Cloud, we strive to offer a great customer experience for enterprises by building a robust and supported platform for running all Linux-based workloads.

This mission is why we were one of the first cloud providers to offer purpose-built Rocky Linux images when Rocky Linux debuted last year as a replacement option for CentOS. We were also one of the first hyperscalers to sponsor the Rocky Enterprise Software Foundation (RESF) to support the open source community behind this Linux distribution. With these efforts, we’re pleased that many customers are already running Rocky Linux in Google Cloud today.

Today, we’re excited to announce that we’re taking another step in furthering the support we provide for Rocky Linux. We’re partnering with CIQ—the company started by CentOS co-founder and Rocky Linux founder Gregory Kurtzer featuring core expertise across Linux, cloud, HPC, containers and security— so we can provide customers a new and improved experience for Rocky Linux on Google Cloud. 

Starting today, customers can leverage Google’s support offerings to file support cases for Rocky Linux. Google support teams and the Rocky Linux experts at CIQ are working together to address customer issues to help ensure they get enterprise-grade support. If you already have a paid support plan with Google, you will be able to open a case for an issue related to Rocky Linux. Google teams can expediently help resolve the issues, backed by CIQ expertise, giving you an integrated experience of using Rocky Linux on Google Cloud. 

"We asked ourselves, how do we bring the best value to everyone? Through this partnership, anytime you use our Rocky Linux on Google Cloud, both Google and CIQ jointly have your back! From the cloud platform itself, all the way through the enterprise operating system, every aspect of using Google Cloud is supported by a single call to Google, and together, we are your escalation team.”—Gregory Kurtzer, CEO of CIQ and Founder/Director of Rocky Linux and the RESF

In addition to CIQ-backed support for Rocky Linux, Google is also working with CIQ to provide a streamlined product experience - with plans to include performance-tuned Rocky Linux images, out-of-the-box support for specialized Google infrastructure, tools to help support easy migration, and more. We’re doing these updates in a community-friendly way. Together with CIQ, Google is helping to create a Rocky Linux Cloud SIG that aims to provide optimized, standardized, and simplified Rocky Linux experience. 

If you’re currently looking for alternatives to CentOS as it reaches end of life, Rocky Linux on Google Cloud can have you covered both from a product and support perspective. So, take Rocky for a spin if you haven’t already, and if you have questions or suggestions on how we can help you, please don’t hesitate to reach out to us. To learn more, please also join us for a webinar discussion on April 6th 2022 at 11.00am PT.

  •  

[Launched] Generally Available: Azure Developer CLI (azd) Extension Framework

The Azure Developer CLI (azd) Extension Framework is now generally available. The framework enables developers, teams, and partners to extend Azure Developer CLI with custom capabilities that support their preferred application development workflows.With
  •  

AWS Weekly Roundup: EC2 application status checks, IAM role manager, OpenAI Daybreak on Bedrock, and more (August 17, 2026)

Last week, AWS contributors joined the OpenSearch and Valkey communities at Open Source Summit Korea 2026 and MCP DevSummit Seoul 2026 to meet open source developers and contributors. At the four-day event, community leaders and users of these Linux Foundation open source projects gathered to share knowledge, collaborate on solutions, and push the projects forward.

Leaders of the Korean OpenSearch communities volunteered to participate in the booth, and also had time to network and interact in the user group meetup.

OpenSearch is an open source, enterprise-grade search and observability suite that brings order to unstructured data at scale. On June 9, 2026, OpenSearch 3.7 introduced new tools designed to query, alert, and track SLOs across logs, traces, and metrics through a single interface and retrieve vectors up to 5.5x faster for improved search performance. Since July 30, 2026, you can run OpenSearch version 3.7 on Amazon OpenSearch Service for improvements in vector search performance, search relevance, and Query Insights.

Valkey is an open source high-performance key/value datastore that supports a variety of workloads such as caching, message queues, and it can act as a primary database. On May 19, 2026, Valkey 9.1 introduced a redesigned I/O threading model that improves throughput by up to 17% and reduces memory usage for strings under 128 bytes by up to 20%. Since June 23, 2026, you can run Valkey 9.1 in Amazon ElastiCache for node-based clusters, delivering higher throughput, improved memory efficiency, and stronger access control for multi-tenant workloads.

You can connect with AWS contributors at upcoming OpenSearch and Valkey community events.

Last week’s launches
Here are some launches that got my attention:

For a full list of AWS announcements, be sure to keep an eye on the What’s New with AWS page.

Other AWS news
Here are some additional projects and news items you may find interesting:

  • The deprecation of email validation in AWS Certificate Manager: ACM will discontinue support for email-validated public certificates by September 30, 2027. If you use email validation for your ACM public certificates, you need to migrate to DNS validation before that date. For Amazon CloudFront distributions, HTTP validation is also available.
  • The next-generation AWS VPN Client with CLI support and admin controls: You can use a new AWS VPN Client built on OpenVPN3. With the new client, you get full backward compatibility with existing AWS Client VPN endpoints while delivering the automation capabilities and security posture that enterprise networking teams have been asking for.
  • Oracle Exadata on Exascale for Oracle AI Database@AWS: ExaDB-XS brings Exadata-class performance and availability through a consumption-based model. With ExaDB-XS, you can scale compute and storage independently in small increments and pay only for what you consume.

For a full list of AWS blog posts, be sure to keep an eye on the AWS Blogs page.

Learn more about AWS, browse and join upcoming AWS-led in-person and virtual events, startup events, and developer-focused events including AWS Summits and AWS Community Days. Join the AWS Builder Center to connect with builders, share solutions, and access content that supports your development.

That is all for this week. Check back next Monday for another Weekly Roundup!

— Channy

  •  

AWS Weekly Roundup: AWS Heroes Summit, Web Search on Amazon Bedrock, Dogwood, Kiro Crew, and more (August 10, 2026)

Last week, we brought together AWS Heroes from around the world to connect, collaborate, and celebrate the builders who go above and beyond for the AWS community.

The AWS Heroes Summit, an invite-only annual gathering, brings global experts specializing in fields like AI, serverless, and containers together for direct collaboration, technical deep-dives, and feedback sessions with internal AWS product and service teams.

Day 1 started with an inspiring fireside chat from AWS CEO Matt Garman. From an insightful AMA with James Hamilton on Day 2 to breakout sessions from various product teams that sparked new ideas, our AWS Heroes excelled at sharing knowledge, lifting each other up, and turning conversations into collaborations. To learn more, read the attendee feedback on LinkedIn.

Last week’s launches
Here are some launches that got my attention:

  • Web Search on Amazon Bedrock: Amazon Bedrock now enables OpenAI models (GPT-5.4, GPT-5.5, and GPT-5.6 Sol/Terra/Luna) to browse and retrieve information from the internet, allowing AI applications to access up-to-date information beyond their training data. This capability opens new possibilities for building AI agents and applications that can answer questions using real-time web content while maintaining data residency within your secured AWS environment with zero data egress. To get started, visit the AI blog post and the Amazon Bedrock User Guide.
  • Runtime Instances on Amazon Bedrock AgentCore: You can now deploy and run AI agents on dedicated runtime instances through Amazon Bedrock AgentCore, providing more control over agent execution environments with predictable performance and cost. To get started, visit Sébastien’s blog post and AgentCore documentation.
  • Vector search for Amazon DynamoDB: You can store and query vector embeddings alongside your existing data in DynamoDB without managing a separate vector database. DynamoDB already supports storing memory for AI agents, and with vector search you can now add semantic retrieval over that memory for agentic grounding, with predictable performance. To learn more, visit Esra’s blog post and Amazon DynamoDB Developer Guide.
  • AWS Transform continuous modernization now generally available: This capability helps engineering teams analyze and remediate technical debt across source code repositories at scale. You can modernize mainframe and legacy workloads with an ongoing, automated approach rather than a one-time migration event. To learn more, visit Micah’s preview blog post. You can also try the AWS Transform Kiro Power and agent plugins.
  • Up to 3,000 Mbps for AWS Lambda function bandwidth: AWS Lambda functions now support increased network bandwidth, enabling data-intensive workloads and faster communication between Lambda functions and other AWS services. This feature enables functions outside a VPC that are configured with 2 GB of memory or more to access network bandwidth that scales proportionally, from 625 Mbps at 2 GB up to 3,000 Mbps at 10 GB.

For a full list of AWS announcements, be sure to keep an eye on the What’s New with AWS page.

Other AWS news
Here are some additional projects and news items that you may find interesting:

  • Introducing Dogwood: Runtime Verification for AI Agents: AWS open-sourced Dogwood, a purpose-built governance language for AI agents to support Cedar policies and add temporal conditions. Powering Dogwood, Amazon Bedrock AgentCore introduced temporal policies whose decisions depend on the history of an agent’s actions within a session, not on the current request alone.
  • AWS supports Agent Plugins: An Open Standard for Portable Agent Extensions: AWS announced support for Agent Plugins, an open source, vendor-neutral specification that gives AI agent extensions a common packaging format so you can package an extension once and ship it to any client, including Kiro, VS Code, Cursor, or any tool that implements the spec.
  • Introducing Kiro Crew: Kiro Crew is a persistent, self-evolving workspace that keeps work moving, online or off, enabling collaborative multi-agent development workflows. It’s built for engineering work that goes beyond a single chat session, and spans repos, tools, and days. You can run several efforts in parallel or hand work to subagents that report back, so nothing waits in line.

For a full list of AWS blog posts, be sure to keep an eye on the AWS Blogs page.

Learn more about AWS, browse and join upcoming AWS-led in-person and virtual events, startup events, and developer-focused events including AWS Summits and AWS Community Days. Join the AWS Builder Center to connect with builders, share solutions, and access content that supports your development.

That is all for this week. Check back next Monday for another Weekly Roundup!

— Channy

  •  

[Launched] Generally Available: Azure Functions support for Python 3.14

You can now develop functions using Python 3.14 locally and deploy them to Azure Functions plans on Linux. Upgrade your apps to Python 3.14 to take advantage of enhanced security, a longer support window, and continued compatibility with the Azure Functio
  •  

[Launched] Generally Available: Ingest OTLP signals into Azure Monitor with the OpenTelemetry Collector

Azure Monitor's support for native ingestion of OpenTelemetry Protocol (OTLP) signals is now generally available. This enables you to send telemetry data directly from OpenTelemetry-instrumented applications and platforms to Azure Monitor. You can configu
  •