❌

Vue lecture

Why your startup needs open models alongside frontier APIs

Every week, I talk with founders who are building at an unbelievable pace. Teams are moving from inception to product-market fit faster than ever, with foundation models wired deeply into their core product workflows.

Yet as startup architectures mature, a clear divide has emerged between teams struggling with margins and those scaling sustainably. The most effective engineering teams have abandoned the one-size-fits-all model strategy.

In the early days of LLMs the default architecture was simple: send every interaction to the largest model available. But as applications move into production, serving millions of people and running autonomous multi-agent workflows, relying on a single frontier model starts to strain in three places:

  • Latency penalties: Relying entirely on cloud round trips makes it difficult to deliver the sub-second responsiveness that interactive mobile and desktop apps require.

  • Infrastructure overhead: Self-hosting large open models with more than 70 billion parameters forces early-stage teams to act like infrastructure providers, pulling senior engineers on cluster provisioning and multi-GPU orchestration.

  • Margin erosion: Sending high-frequency, structured tasks (like intent routing, JSON extraction, or status validation) to general-purpose frontier endpoints spends capital that could be funding product differentiation.

Great engineering teams pick the right tool for each job. Most production requests don’t require a frontier generalist, and routing every call to one can actually slow your product down. Instead, the winning pattern is a compound AI stack: pairing frontier models for complex synthesis with compact, open-weight models that you can tune, control, and run anywhere. 

It’s for these reasons that an open model like Gemma belongs in your model lineup. With more than one billion downloads across the developer community, Gemma 4 is our most capable open model family to date, using the same foundational research and technology behind the Gemini models.

Built under one roof

Gemma is built by Google DeepMind using the same foundational research and architecture advances behind the Gemini models. Because they share common DNA and developer tooling, your team can prototype in Google AI Studio and design hybrid architectures where Gemini and Gemma work together.

Released under a commercially permissive Apache 2.0 license, Gemma 4 is engineered for parameter and token efficiency. Rather than forcing a single model architecture onto every hardware target, Gemma 4 spans five sizes across four specialized architectures: compact E2B and E4B models with native audio and vision for mobile and edge devices; an encoder-free 12B Unified multimodal model; a 26B A4B Mixture-of-Experts (MoE) model that activates only 4B parameters per token for high-throughput serving; and a dense 31B model that fits on a single GPU for maximum reasoning quality and fine-tuning. Every model includes configurable thinking modes, native function calling, up to 256K context, and built-in Multi-Token Prediction (MTP) draft models for speculative decoding. 

Real proof: How startups are winning with Gemma

Founders are using Gemma to solve urgent problems around unit economics, output accuracy, and responsiveness.

  • Flipping the architecture: Cue is a voice-activated desktop assistant that runs natively on a user's machine to automate everyday tasks. They integrated Gemma 4 E4B via Ollama on local hardware to handle real-time transcript formatting. While they originally planned for Gemma to be a weak offline fallback, benchmarking proved it was so fast and precise that they made it their default engine—driving a 44% latency drop (from 876 ms to 488 ms).

  • True edge independence: Mobile development studio HubX built BetterSpeak, a voice-based interactive mobile English-learning tutor that simulates immersive, real-time voice conversations. To bypass cellular network lag and avoid charging users expensive subscription fees to cover cloud hosting, they packaged a 4-bit quantized Gemma 4 E2B model (~2.9 GB) natively on-device. The result is an offline, speech-to-speech mobile tutor that costs them $0 in server bills.

  • Scientific discovery and air-gapped security: K-Dense has built Faraday, an AI-powered scientific collaborator optimized end-to-end across hardware, software, and sensor suites, powered by Gemma 4 together with K-Dense's Scientific Agent Skills. Faraday runs fully air-gapped, making it suitable for secure, proprietary scientific work in pharma and biotech. Deployed on an NVIDIA DGX Spark, Gemma 4 can also be fine-tuned locally on a user's own proprietary datasets.

  • Unlocking infinite gameplay and retention: Gaming company Latitude integrated Gemma across their AI-native game products. By swapping in Gemma for AI Dungeon, they significantly improved player retention, while their new AI RPG platform Voyage leverages Gemma to deliver high intelligence at a cost that enables unlimited user gameplay with ultra-fast latency.

Four workloads where Gemma wins for startups

If you’re evaluating where Gemma fits into your stack today, start with these four jobs:

1. Edge and local execution (low latency, true privacy)

If you’re building mobile apps, developer desktop tools, robotics, or offline-first experiences, every cloud round-trip adds latency that people can feel. Gemma can run directly on laptops (including Apple silicon), smartphones, and local appliances. Your users get immediate feedback, and sensitive data never has to leave their device.

You can handle many local interactions on-device for zero incremental cost, and keep a bridge to frontier models in the cloud for the requests that need it. When a local workflow calls for large-scale multimodal reasoning, long-context data synthesis, or complex planning, your application can route that specific request to Gemini.

2. High-throughput triage and agent routing

In multi-agent architectures, agents spend a surprising amount of tokens on simple tasks like checking statuses, classifying intent, and routing tickets. With Gemma as your front-line gatekeeper, those high-volume background tasks run on a compact model and your team can save frontier reasoning for the requests where it creates product value.

3. Task-specific fine-tuning for real moats

Adapting a model to your proprietary data is one way to build a competitive moat. Because Gemma gives you full access to model weights and has a compact memory footprint, your team can run parameter-efficient fine-tuning (LoRA or QLoRA) on a single GPU in hours rather than days.

4. Turnkey vertical starting lines

DeepMind releases domain-specific variants of Gemma, so you don’t have to start from scratch. One example is MedGemma. MedGemma scores 87.7% on the MedQA benchmark, matching the clinical accuracy of frontier models at roughly one-tenth the inference cost. In a blind clinical study, board-certified radiologists judged that 81% of chest X-ray reports generated by the lightweight MedGemma 1.5 4B were accurate enough to result in equivalent patient management compared to reports written by human experts.

Beyond healthcare, biotech startups use C2S Scale to model virtual cellular responses and accelerate oncology research. Meanwhile, DataGemma cross-references more than 240 billion public data points to help reduce numerical hallucinations. If you’re operating in a specialized market, starting with a model that already speaks your industry's language can save you engineering time and compute budget.

Deploy wherever your business lives

Gemma is designed to fit into your existing engineering stack without lock-in:

  • Apache 2.0 licensing: Gemma 4 ships under the Apache 2.0 license, giving startups the freedom to fine-tune, quantize, redistribute, and deploy commercial products on-premises or at the edge with full ownership of their custom weights.

  • Day-zero open tooling: Run and fine-tune Gemma with the tools your engineers already use, including vLLM, Ollama, llama.cpp, LM Studio, MLX, Unsloth, Hugging Face, Kaggle, Keras, PyTorch, JAX, and LiteRT-LM.

  • Serverless and managed cloud deployment: Prototype immediately in Google AI Studio, scale to zero on serverless GPUs with Cloud Run, or deploy dedicated endpoints from Model Garden on Gemini Enterprise Agent Platform when traffic surges and you don’t want to manage GPU clusters.

  • Enterprise-ready safety: Gemma undergoes rigorous pre-release safety evaluations, data filtering, and red-teaming, and pairs with ShieldGemma 2 to help you meet enterprise compliance requirements.

Build with Gemma: What to do this week

Great technical architecture isn't about finding one model to do everything. It’s about assembling the right tool for each job so you can move faster, protect your runway, and ship a superior product.

Here’s my challenge to your engineering team this week:

  1. Audit your model calls: Look at your logging dashboard and identify three high-volume, deterministic tasks (such as intent classification, JSON validation, or summarization) currently running on your most expensive models.

  2. Benchmark Gemma: Run a quick test with a compact Gemma model locally or on a single endpoint. Measure the latency and calculate what happens to your gross margins when that workload runs with lower inference cost.

  3. Redirect your runway: Take the capital and engineering hours you save on compute and invest them back into your core differentiators.

You can download the Gemma weights directly or deploy them through Model Garden. If you need compute credits and technical architecture reviews to get up and running, the Google for Startups team is ready to help you build - learn more.

  •  

Maximizing Apache Spark availability: Mitigating compute stockouts with flexible VMs and other best practices

The surge in AI development has created unprecedented demand for compute capacity around the globe. This can have negative implications for data processing and pipelines with Apache Spark. Whether you are managing your own Spark infrastructure or using a managed service, you can face availability constraints. However, a significant advantage of using Google’s Managed Service for Apache Spark is the availability of flexible VMs, which provide a targeted mechanism to adopt a dynamic, resource-agnostic philosophy and ensure your pipelines remain operational, even during regional or zonal capacity stockouts.

Understanding capacity stockouts

Capacity stockouts occur when demand for a specific machine family (such as N2 or N2D) exceeds available capacity in a target zone or region. For time-sensitive analytics pipelines, rigid single-VM requirements transform standard provisioning into a single point of failure which can result in cluster creation delays, failed executions, and potentially compromised business SLAs.

Flexible VMs

Flexible VMs fundamentally overhaul how a Managed Spark cluster requests compute resources. Rather than binding a cluster to a rigid instance type, flexible VMs allow teams to establish an ordered list of acceptable machine families for master, primary worker, and secondary worker nodes.

Key features

  • Multi-family blending: Mix nodes across diverse machine types and generations, combining Gen2 families (e.g., N2, N2D) with Gen4 families (e.g., N4, C4) in a single configuration.

  • Mixed storage support: Broaden available capacity pools by allowing storage options to dynamically adapt to the underlying host family's supported disk types.

  • Comprehensive cluster coverage: Apply flexible rules to primary workers, secondary (preemptible/spot) workers, and master nodes to guarantee cluster provisioning end-to-end.

Ranked configuration: A strategy for success

A successful flexible VM implementation relies on intentional ranking. By defining a clear hierarchy of options, Managed Spark clusters automatically attempt provisioning, systematically mitigating stockout risks without requiring manual intervention. To improve the availability of  suitable VMs, we recommend specifying at least two machine families in the highest priority (Rank 0) flexible VM list.

As an example, for production pipelines standardizing on n2d-standard-16 shapes, the following tiering strategy provides robust resilience against capacity constraints:

Rank

Machine family examples

Storage recommendation

Rank 0 (Primary)

n2d-standard-16, n2-standard-16

Standard Local SSD or PD

Rank 1

n4-standard-16, n4d-standard-16

Hyperdisk Balanced

Rank 2

c4-standard-16, c3-standard-22

Hyperdisk Balanced

Rank 3 

e2-standard-16

Standard PD

code_block
<ListValue: [StructValue([('code', 'gcloud dataproc clusters create $CLUSTER_NAME \\\r\n--num-workers=10 \\\r\n--zone="" \\\r\n--region=us-east1 \\\r\n--worker-instance-selection=\'{"machineTypes":["n2d-standard-16","n2-standard-16"],"rank":0,"diskConfig":{"bootDiskType":"pd-standard","bootDiskSizeGb":400}}\' \\\r\n--worker-instance-selection=\'{"machineTypes":["n4-standard-16","n4d-standard-16"],"rank":1,"diskConfig":{"bootDiskType":"hyperdisk-balanced","bootDiskSizeGb":400}}\' \\\r\n--worker-instance-selection=\'{"machineTypes":["c4-standard-16","c3-standard-22"],"rank":2,"diskConfig":{"bootDiskType":"hyperdisk-balanced","bootDiskSizeGb":400}}\' \\\r\n--worker-instance-selection=\'{"machineTypes":["e2-standard-16"],"rank":3, "diskConfig":{"bootDiskType":"pd-ssd","bootDiskSizeGb":400}}\' \\\r\n--master-instance-selection=\'{"machineTypes":["n4-standard-16","n4d-standard-16"],"rank":0,"diskConfig":{"bootDiskType":"hyperdisk-balanced","bootDiskSizeGb":400}}\''), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb3fbbd5a90>)])]>

For pipelines standardizing on legacy n1-standard-16 shapes, the following tiering strategy helps transition workloads toward newer, more available architectures while preserving operational stability:

Rank

Machine family examples

Storage recommendation

Rank 0 (Primary)

n1-standard-16

n2-standard-16

Standard Local SSD or PD

Rank 1

n2d-standard-16

Standard Local SSD or PD

Rank 2

n4-standard-16

n4d-standard-16

Hyperdisk Balanced

Rank 3

e2-standard-16

Standard PD

Leveraging Hyperdisk Balanced

Unlocking maximum availability with flexible VMs often requires adopting modern storage architectures like Hyperdisk Balanced. Newer instance families (including N4 and C4) rely on Hyperdisk to deliver predictable performance across variable VM sizes. Starting with default IOPS and throughput settings typically provides a reliable baseline for the majority of distributed Spark jobs.

Trade-offs and key considerations

While flexible VMs  dramatically improve cluster provisioning success, aligning them with enterprise requirements involves evaluating several architectural and financial factors:

1. Resource quotas

It is no longer enough to have one specific machine (e.g., N2) quota. You need to ensure you have sufficient compute and disk quotas allocated for all specific machine types and disks (including Hyperdisk) defined in their flexible VM lists.

2. Compute flexible Committed Use Discounts (CUDs)

Traditional, resource-based CUDs are tied to specific machine families, which limits flexibility. Adopt Compute flexible Committed Use Discounts (CUDs) to apply savings across multiple VM families and regions.

3. Performance Characteristics

Performance can vary between machine generations, as well as between Local SSD and Hyperdisk. While the Managed Spark team maintains internal benchmarks for these comparisons, actual outcomes are workload-dependent. Testing your specific Spark jobs across these families is essential for understanding SLA impacts.

Additional recommendations

In addition to implementing flexible VMs, there are several other key architectural and scheduling strategies to improve resource availability and workload stability:

  • AutoZone: Implement AutoZone routing to allow Managed Spark to automatically select the zone best suited to execute the job based on current capacity.

  • Smaller machine shapes: Avoid high in demand, large-core shapes. Design workloads and YARN containers to utilize smaller machine shapes (such as 4, 8, or 16 cores). These smaller shapes are much easier to fulfill from the available GCE on-demand pool.

  • Autoscaling: Deploy cluster autoscaling with reasonable maxInstances to manage capacity effectively for bursty or unpredictable workloads without relying on rigid, massive upfront provisioning.

  • Partial cluster creation: Configure a minimum acceptable number of primary workers. This allows clusters to spin up under resource constraints and begin executing, while autoscaling can dynamically add remaining workers as resources become available.

  • Establish regional fallbacks: Some regions, such as us-central1, can experience  high demand. Setting up fallbacks to other regions reduces capacity stockout risks.

Keep your Spark jobs running with flexible VMs

Managing your own Apache Spark infrastructure can be complex, especially when capacity stockouts disrupt your data processing. Utilizing a managed service like Managed Service for Apache Spark provides unique advantages — including built-in platform resilience and access to flexible VMs. By adopting a prioritized fallback strategy with flexible VMs, you can protect your workloads from regional hardware shortages and keep your critical pipelines running.

Ready to improve your Spark workload resilience? Start configuring flexible VMs for your Managed Spark clusters today.

  •  

Google Cloud partners with CIQ to provide an enterprise-grade experience for Rocky Linux

At Google Cloud, we strive to offer a great customer experience for enterprises by building a robust and supported platform for running all Linux-based workloads.

This mission is why we were one of the first cloud providers to offer purpose-built Rocky Linux images when Rocky Linux debuted last year as a replacement option for CentOS. We were also one of the first hyperscalers to sponsor the Rocky Enterprise Software Foundation (RESF) to support the open source community behind this Linux distribution. With these efforts, we’re pleased that many customers are already running Rocky Linux in Google Cloud today.

Today, we’re excited to announce that we’re taking another step in furthering the support we provide for Rocky Linux. We’re partnering with CIQ—the company started by CentOS co-founder and Rocky Linux founder Gregory Kurtzer featuring core expertise across Linux, cloud, HPC, containers and security— so we can provide customers a new and improved experience for Rocky Linux on Google Cloud. 

Starting today, customers can leverage Google’s support offerings to file support cases for Rocky Linux. Google support teams and the Rocky Linux experts at CIQ are working together to address customer issues to help ensure they get enterprise-grade support. If you already have a paid support plan with Google, you will be able to open a case for an issue related to Rocky Linux. Google teams can expediently help resolve the issues, backed by CIQ expertise, giving you an integrated experience of using Rocky Linux on Google Cloud. 

"We asked ourselves, how do we bring the best value to everyone? Through this partnership, anytime you use our Rocky Linux on Google Cloud, both Google and CIQ jointly have your back! From the cloud platform itself, all the way through the enterprise operating system, every aspect of using Google Cloud is supported by a single call to Google, and together, we are your escalation team.”—Gregory Kurtzer, CEO of CIQ and Founder/Director of Rocky Linux and the RESF

In addition to CIQ-backed support for Rocky Linux, Google is also working with CIQ to provide a streamlined product experience - with plans to include performance-tuned Rocky Linux images, out-of-the-box support for specialized Google infrastructure, tools to help support easy migration, and more. We’re doing these updates in a community-friendly way. Together with CIQ, Google is helping to create a Rocky Linux Cloud SIG that aims to provide optimized, standardized, and simplified Rocky Linux experience. 

If you’re currently looking for alternatives to CentOS as it reaches end of life, Rocky Linux on Google Cloud can have you covered both from a product and support perspective. So, take Rocky for a spin if you haven’t already, and if you have questions or suggestions on how we can help you, please don’t hesitate to reach out to us. To learn more, please also join us for a webinar discussion on April 6th 2022 at 11.00am PT.

  •